Score
Designs and implements instrumentation, metrics, and reporting to measure and attribute token usage and associated costs (monetary, compute, energy, latency) for model requests. Builds token-accounting systems to profile tokens per request, compare cost across methods, and surface trade-offs and savings including on-device prompt usage.
This study addresses the limitations of large language model (LLM)-based multi-agent systems in software engineering, particularly the lack of transparency in resource consumption, unpredictability of costs, and unclear environmental impact. To this end, it introduces the first standardized token consumption evaluation framework tailored for agent-based software engineering. By analyzing execution trajectories from the ChatDev framework across 30 development tasks, the work maps internal agent interactions to standard software engineering phases—design, coding, completion, code review, testing, and documentation—and quantifies the distribution of input, output, and reasoning tokens across these stages. The analysis reveals that the code review phase alone accounts for 59.4% of total token usage, with input tokens comprising 53.9% of the total, indicating that cost is primarily driven by automated refinement and validation rather than initial code generation. These findings provide empirical foundations for optimizing workflows, forecasting costs, and designing efficient agent collaboration protocols.
This study addresses the challenge of precisely quantifying the environmental impact of large model inference in cloud environments by proposing a systematic evaluation method that separately measures the energy consumption, water usage, and carbon emissions associated with input and output tokens. By integrating GPU benchmarking, Bayesian linear regression, and Monte Carlo simulation, this work pioneers a reliable estimation of energy consumption uncertainty for both closed-source and open-source models with unknown configurations. The proposed framework provides actionable quantitative evidence to support the reduction of cloud service costs, electricity consumption, and carbon footprints.
This study addresses the unpredictable token consumption caused by context inflation during the execution of large language model (LLM) agents. To tackle this issue, we propose a segmented and composable cost modeling approach that constructs composable cost representations by integrating context growth records with a dynamic evidence fusion algorithm. Through cumulative estimation and dynamic updating mechanisms, the method achieves precise real-time prediction without requiring additional LLM invocations. Experimental results demonstrate that our approach reduces the average prediction error by 14.5% and saves 21.3% in token consumption under budget constraints, while maintaining a single-prediction latency of only 32.8 milliseconds. These findings indicate that the proposed method effectively balances computational efficiency with predictive accuracy for resource-aware agent deployment.
This study addresses the lack of a systematic overview of open-source software energy measurement tools, which hinders energy-aware software design and tool selection. From a mining software repositories (MSR) perspective, the authors employ qualitative content analysis to screen and categorize 585 GitHub projects, identifying 24 high-quality open-source energy measurement tools. The work systematically characterizes these tools in terms of architectural design, measurement granularity—spanning from CPU-level to process, container, and AI workload levels—and their capabilities for carbon emission estimation. By elucidating evolutionary trends in tool development, this research provides software architects with a structured foundation and practical guidance for informed tool selection in energy-efficient software engineering.
This study addresses the absence of a systematic, reusable, and empirically grounded end-to-end approach that integrates incentive mechanisms, governance structures, and tokenomics in current token economic designs. To bridge this gap, the paper proposes the Token Economic Design Method (TEDM), which, for the first time, unifies these three dimensions into a structured and actionable design framework, with explicit emphasis on sociotechnical context and early-stage design considerations. Developed through the design science research paradigm and informed by qualitative synthesis, co-design case studies, and expert interviews, TEDM was empirically validated through its application to the Currynomics stablecoin ecosystem and subsequent expert evaluation. The results demonstrate that TEDM effectively supports the analysis and construction of tokenized ecosystems, offering practical and reusable design guidance.
研究通过改变任务说明来影响AI编码任务中的令牌消耗,使用Kimi K3模型测试不同任务说明对令牌消耗的影响,并提出了一种预测令牌消耗的方法。
This study addresses the sustainability challenges faced by operators due to cost-benefit mismatches and price volatility in AI services. To manage such uncertainty, it proposes a structured contract infrastructure that decouples service consumption rights from payment collection rights. Methodologically, this work pioneers the introduction of financial derivatives logic into the AI computing market, enabling standardized commitments and risk hedging across heterogeneous services. Furthermore, it leverages smart contracts to formalize participant roles, ownership structures, and settlement rules, validating system logic through API replay. Experimental results demonstrate that forward contracts effectively reduce average expenditures, while financing mechanisms significantly enhance user contributions under low-capital conditions.
本文研究了GPU上大语言模型推理工作负载的请求和令牌能耗问题,通过分解能耗模型并评估不同类型模型在不同条件下的能耗,提出应同时优化请求和令牌能耗。
This study addresses the lack of systematic modeling approaches for token economies and quantitative analysis of event impacts. Building upon the DeTEcT framework, it proposes the first integrated methodology that combines formal token economy simulation with significance-based measurement of event effects. By introducing an event impact analysis framework augmented with numerical simulation techniques and wealth distribution metrics, the work enables quantitative assessment of wealth redistribution effects triggered by endogenous policy changes—such as Bitcoin Improvement Proposals (BIPs). Using Bitcoin as a case study, the approach demonstrates its effectiveness and practicality in capturing economic dynamics and evaluating the consequences of significant protocol-level events.
This study addresses the insufficient accuracy of compliance reporting in Ethereum node energy estimation caused by overlooking hardware attributes and hosting environment heterogeneity. We propose a fine-grained power quantification method leveraging publicly available properties of the peer-to-peer network. By collecting node characteristics via web crawling, this work is the first to utilize such features to differentiate client types, system architectures, and cloud service providers, thereby overcoming the limitations of conventional single typical-power assumptions. Furthermore, we integrate a rule engine with a random forest model to map empirically measured energy consumption data, enabling precise per-node power estimation. Experimental results demonstrate that the estimated total network power consumption is 415 kW, representing a 3.9% reduction compared to traditional methods. This approach effectively bridges the assessment gap in the Cambridge Centre for Alternative Finance (CCAF) framework, significantly enhancing the accuracy of energy consumption measurements.