Method
How these numbers are produced, and what they exclude.
- History avoidedthe ratio
- Tokens of record delivered to the agent, divided by the total size of the record, measured across the benchmark packs. It is a measure of retrieval, not of spend.
- Task contextexcluded, deliberately
- The code, the files, the prompt and the tool output are identical either way, so they cancel out of the comparison. The ratio measures only the variable being tested.
- The 50,000-token thresholda standing ruling
- The measured point below which the dollar difference stops being material. It scopes where the cost argument applies; the currency and accuracy arguments apply at any size.
SourcesThe forecast that AI coding costs will surpass average developer salaries, and bloated context windows as a driver, are Gartner's. The finding that committed context files tend to reduce task success while adding measurable inference cost comes from published research on instruction adherence. Both are linked on the benchmark page.