On this page
The short answer
Measure agent cost against successful task completion, including retries, tool calls, failed attempts, and required review. Compare candidates under the same operating limits, and keep reliability and latency visible rather than collapsing every tradeoff into one score.
What to take away
- Cheap attempts can become expensive completed tasks.
- Averages can conceal slow or costly failure paths.
- Compare candidates under equivalent resource constraints.
Sources checked . Research synthesis and Rubrex recommendations; examples are illustrative, not client results.
Behavior and resource use need more than one metric
A June 2026 preprint proposes entropy-based descriptions of agent behavior alongside familiar outcome measures. Reliability research separately examines consistency and robustness. These sources suggest looking at repeated or unproductive behavior, while leaving the actual cost model dependent on the deployed system.
Evidence: Entropy-Based Observability for AI Agent Behavior [1]Towards a Science of AI Agent Reliability [2]
Define the accounting boundary
Decide whether the unit is one attempt, one user request, or a completed business task. Include generation, retrieval, tools, orchestration, retries, and any human review required for the workflow. Keep billable provider charges separate from internal labor estimates so assumptions remain inspectable.
Do not compare a candidate with unrestricted retries against a baseline allowed one attempt without disclosing the difference. Extra attempts may be a legitimate product choice, but the comparison then concerns two operating policies as well as two models. Record timeout and concurrency limits because they can affect both cost and completion.
Report distributions and failure paths
Rubrex recommends showing task success, median and tail latency, total resource use, and cost per completed task under a specified policy. Retain failed and abandoned attempts in the cost denominator. If no tasks succeed, report that fact instead of presenting an undefined ratio as zero.
Segment tasks by complexity and tool dependence. A routing change may lower the average by moving easy tasks to a cheaper path while making difficult tasks worse. Inspect both slices and the routing errors. A favorable overall number should not erase a critical subgroup regression.
Illustrative arithmetic: the cheaper attempt loses
Consider two hypothetical systems using abstract cost units. System A spends 1 unit per attempt and completes 50 of 100 independent tasks. System B spends 1.5 units per attempt and completes 90. With one attempt per task, cost per success is 2 units for A and about 1.67 for B. These invented numbers illustrate the calculation, not provider pricing or measured performance.
The result changes if failures require manual repair, retries are allowed, or the successful tasks differ in difficulty. Document those policies before using the ratio to choose a system.
Choose an operating point, not a universal winner
Set minimum quality and permission requirements first. Among candidates that satisfy them, compare cost and latency under the expected workload. Preserve several operating points when the tradeoff is meaningful. A model that suits quick classification may not suit a long tool workflow.
- Include failed runs in total resource use.
- Record retry and timeout policies.
- Separate task classes and routing mistakes.
- Recheck current provider prices when converting units to money.
Limits of the evidence
The example uses invented cost units and assumes one attempt per independent task. It excludes pricing claims. Research on behavioral metrics does not establish that a particular entropy measure predicts cost or reliability in your application.
Common questions
Is cost per token enough?
No. Token price omits how many tokens, retries, tool operations, and review steps are needed to complete a task.
Should the fastest agent always win?
Only after it meets the required quality and operating constraints. Fast incorrect completion is not useful throughput.
Sources & further reading
Primary sources behind this briefing. A source’s findings apply to its own study conditions; publication on arXiv does not establish peer review.
- Entropy-Based Observability for AI Agent Behavior arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.
- Towards a Science of AI Agent Reliability arXiv · 2026 · PreprintReviewed: abstract and publication record. Accessed September 25, 2026.
Questions or corrections? Write to Rubrex. Read our editorial approach.