What it costs an agent to translate a resource file
Run 20260812T160551Z-92885b, captured 2026-08-12
Every figure on this page comes from the artifacts listed at the bottom. Token figures are Claude's own accounting from the stream and print mode result objects. Currency figures are approximations reported by the Claude Code CLI, not billing figures.
The question
A coding agent asked to translate an i18n resource file can do the work itself, or it can hand the file to a translator running on the same machine. The first spends API tokens on every value. The second spends none on the values, and BetterTranslator can do it with no agent in the loop at all. The benchmark measures what each way costs, on one file, in tokens, in approximate dollars, in requests and in wall clock.
Three arms were measured, named as the artifacts name them.
| Arm | What happens | Who translates |
|---|---|---|
| direct | Claude Code reads the file and writes the translation itself. No MCP server is configured. | The agent |
| mcp-path | Claude Code calls translate_file on the local BetterTranslator MCP server and polls job_status. The file content never enters the conversation. | The local model |
| mcp-value | Claude Code sends the values through translate_text and writes the file itself. | The local model |
A fourth path, BetterTranslator's own window or bt translate with no agent involved, spends zero API tokens by construction and is carried in the projection as the zero row.
How the test was built
| Held constant | Value |
|---|---|
| Agent | Claude Code CLI 2.1.228 |
| Model | claude-opus-5 |
| Prompt | identical across arms apart from one sentence |
| Session | fresh session per run |
| Repetitions | three per cell, the first excluded as cache warm-up |
| Language pair | English to Czech |
| Local engine | TranslateGemma, translategemma-4b-it-Q4_K_M.gguf, Q4_K_M |
| MCP transport | stdio, server bettertranslator 1.0.0, protocol 2024-11-05, 10 tools |
| Host | Windows 11, .NET SDK 10.0.302 |
| Corpus | 4,666 keys in 627,920 bytes, sha256 a8fdce5a, of which 1,242 were eligible under the selection rule |
A value qualifies for a slice when it is a string of 20 to 180 characters carrying no pluralisation pipe, HTML tag or line break, and still holding at least 4 words and 20 ASCII letters once placeholder tokens are removed. Every selected value is prose a translator has to change. Each slice is a prefix of the next larger one. The full source hash is a8fdce5aba678743b2aac57a9b8c06405e41b836cd174051fc09357da648a633.
| Slice | Keys | Bytes | sha256 |
|---|---|---|---|
| slice-10 | 10 | 1,205 | be639f19 |
| slice-50 | 50 | 5,109 | f0214d54 |
| slice-200 | 200 | 21,295 | 6a58de05 |
| slice-500 | 500 | 55,642 | 57e3b4d6 |
Accounting note: this Claude Code build emits no OpenTelemetry metrics or events, so per-request accounting is read from Claude's stream-json usage blocks and the print mode result object. 4 exporter configurations were tried and none produced output. As a cross-check the harness compares both sources per run: 22 of 33 runs agree within 2 percent, and the 11 that disagree carry both numbers in summary.csv.
Per slice
| Slice and arm | Requests | cacheCreation | cacheRead | Total tokens | Approx USD | Wall clock s | Values delivered |
|---|---|---|---|---|---|---|---|
| slice-10 direct | 3.0 | 17,546 | 103,780 | 121,337 | 0.2602 | 23.4 | 10 of 10 |
| slice-10 mcp-path | 7.0 | 18,289 | 269,247 | 287,669 | 0.3362 | 17.3 | 10 of 10 |
| slice-50 direct | 3.0 | 22,704 | 105,376 | 128,091 | 0.4016 | 64.0 | 50 of 50 |
| slice-50 mcp-path | 13.5 | 18,523 | 541,119 | 560,011 | 0.4827 | 27.3 | 49 of 50 |
| slice-200 direct | 4.0 | 50,665 | 144,893 | 195,574 | 0.9510 | 143.4 | 198 of 200 |
| slice-200 mcp-path | 26.5 | 13,656 | 1,125,718 | 1,140,077 | 0.7553 | 78.6 | 195 of 200 |
| slice-500 direct | 3.5 | 92,294 | 140,597 | 232,903 | 2.2062 | 539.3 | 498 of 500 |
| slice-500 mcp-path | 24.0 | 24,068 | 1,032,698 | 1,057,235 | 0.8214 | 57.2 | none |
The third arm, for completeness. It was not run at 500 keys by instruction.
| Slice and repetition | Requests | Total tokens | Approx USD | Wall clock s | Values delivered | Gate |
|---|---|---|---|---|---|---|
| slice-10 | 5.0 | 212,361 | 0.3693 | 35.3 | 10 of 10 | pass, spread 32.7 percent |
| slice-50 | 15.5 | 740,057 | 0.8774 | 162.0 | 50 of 50 | pass, spread 48.0 percent |
| slice-200 rep 2 | 15 | 1,000,865 | 2.6973 | 639.2 | 198 of 200 | fail, model fallback |
| slice-200 rep 3 | 9 | 478,525 | 3.0136 | 206.1 | 0 of 200 | fail, model fallback and session limit |
What delegating changes
The delegated arm against the agent translating the file itself. The 500 key row is not comparable, because the delegated arm produced no file.
| Slice | Delta tokens | Delta tokens percent | Delta USD approximate | Delta USD percent | Delta requests | Delta wall clock s |
|---|---|---|---|---|---|---|
| slice-10 | +166,332 | +137.1% | +0.0760 | +29.2% | +4.0 | -6.1 |
| slice-50 | +431,920 | +337.2% | +0.0811 | +20.2% | +10.5 | -36.8 |
| slice-200 | +944,503 | +482.9% | -0.1957 | -20.6% | +22.5 | -64.8 |
| slice-500 | not comparable |
Tokens and money do not rank these arms the same way. In tokens there is no crossover anywhere between 10 and 500 keys: the delegated path costs more at every measured slice. In approximate dollars the two lines cross at about 87 keys, because the delegated path swaps fresh input for cache reads and cache reads bill far below it. At 200 keys the delegated path is the cheaper of the two in money while spending nearly six times the tokens.
Where the tokens go: the payload really does stay out of the context. On the smallest slice the whole MCP exchange is 306 bytes of tool arguments and 707 bytes of tool results against a 1,205 byte input file, and the entire tool surface adds 743 tokens of cache creation over the agent path. The cost is request count. Every job_status poll is a full API request that replays the whole conversation, so the arm's bill is set by how often it asks whether the local model has finished: 7.0 requests at 10 keys, 26.5 at 200. The fitted floor with an empty file is 244,502 tokens. Hold everything else constant and give the arm the agent path's request count, as a blocking call or a wait tool would, and the same 200 key work lands near 183,575 tokens against 195,574. That last figure is arithmetic on measured per-request cost, not a measured run.
The whole file
Projected, not measured: the 4,666 key file does not fit one context window, so the projection repeats the largest measured slice, 200 keys, as 24 chunks. It is arithmetic over measured runs.
| Path | Tokens | Approx USD | Claude requests | Wall clock |
|---|---|---|---|---|
| Claude Code alone (direct) | 4,693,764 | 22.82 | 96 | 0h57m |
| Claude Code plus local BetterTranslator (mcp-path) | 27,361,836 | 18.13 | 636 | 0h31m |
| Claude Code plus local BetterTranslator (mcp-value) | 24,020,760 | 64.74 | not published | 2h49m |
| BetterTranslator alone, no agent | 0 | 0.00 | 0 | not published |
Two cells are withheld. The mcp-value request count has no gate-passing basis at 200 keys, since both repetitions fell back to another model. The local wall clock is withheld because the artifacts disagree with each other.
The mcp-path row is cheaper in money than the agent path only above the 87 key break-even, and it reaches that price with 27,361,836 tokens against 4,693,764. The mcp-value projection rests on runs the gate excluded and should be read as an upper bound rather than a result.
The fairness gate
A run passes only when its output file is a JSON object whose key set and key ordering match the slice exactly, whose values are all non-empty strings, whose placeholder multiset per value matches the source, and none of whose values equal their source. The last check has one allowance: a value may equal its source when every alphabetic word in it is an allowlisted brand, protocol or locale token or begins with a capital letter, which covers product names and title-case proper nouns, or when the value is a command line carrying an executable name or a long flag.
Across all 33 recorded runs, summary.csv reports 0 key-set failures, 0 key-order failures and 0 placeholder failures. 29 of the 33 runs delivered a file at all. Every exclusion below is an untranslated value, a model fallback, a session limit or a missing file.
| Arm | Slice | Rep | Total tokens | Requests | Values | Reason |
|---|---|---|---|---|---|---|
| mcp-path | slice-50 | 2 | 581,184 | 14 | 49 of 50 | 1 value equal to its source outside the allowlist |
| mcp-path | slice-50 | 3 | 538,837 | 13 | 49 of 50 | 1 value equal to its source outside the allowlist |
| mcp-path | slice-200 | 2 | 1,549,789 | 36 | 195 of 200 | 4 values equal to their source outside the allowlist |
| mcp-path | slice-200 | 3 | 730,364 | 17 | 195 of 200 | 4 values equal to their source outside the allowlist |
| mcp-value | slice-200 | 2 | 1,000,865 | 15 | 198 of 200 | model fallback across claude-opus-5 and claude-opus-4-8 |
| mcp-value | slice-200 | 3 | 478,525 | 9 | 0 of 200 | model fallback, then API error 429 on the account session limit |
| mcp-path | slice-500 | 2 | 513,111 | 12 | 0 of 500 | output file was never written |
| mcp-path | slice-500 | 3 | 1,601,358 | 36 | 0 of 500 | output file was never written |
Of 33 runs, 33 billed a request, 1 ended on the account session limit and 3 fell back from the requested model, which makes those uncomparable.
Limitations
The MCP server costs more tokens, not fewer. At every measured slice it spent more than the agent translating the file itself: 137.1 percent more at 10 keys, 337.2 percent at 50, 482.9 percent at 200. Its supported advantage is money above the 87 key break-even, and that advantage exists because cache reads bill below fresh input. It is not a token saving and does not become one at any size measured here.
No quality judgement exists. Nothing in this run scores adequacy, fluency or meaning. The only value-level measurement is whether an output value differs from its source. On the slices it was measured on, the local engine left 2.3 percent of values in English. The agent path passed the gate at every slice. That comparison covers untranslated values and nothing else. Any claim about how well either side translates is outside these artifacts.
The delegated path produced no file at 500 keys. All 3 repetitions billed requests and delivered nothing. The agent stopped waiting while the local job was still running, and the MCP server dies with the session that started it, so the finished work was thrown away. Those runs are a delivery failure, not a cheap result, which is why they are excluded from the difference table and why the largest slice has no head to head figure.
The whole file numbers are projections. No run has ever translated all 4,666 keys under measurement. The projection multiplies the 200 key slice by 24. Chunking differently, or hitting a context or session limit part way through, changes the result.
Thin repetitions. Two comparable repetitions per cell, because repetition 1 is excluded as cache warm-up. The report marks the cells whose spread is wide: the agent path at 500 keys (15.0 percent), mcp-value at 10 keys (32.7 percent) and at 50 keys (48.0 percent).
One file, one pair, one machine. A single en.json translated English to Czech, on one Windows host, with one local model at one quantization, against one CLI version and one Claude model. Both sides of this comparison move over time.
Reproducing this
btbench prepare --source <path to en.json> --model claude-opus-5 --sizes 10,50,200,500
btbench run --run <run-id> --arms direct,mcp-path --reps 3
btbench collect --run <run-id>
btbench compare --run <run-id>
prepare probes the MCP server, hashes the source and cuts the slices. run executes every arm and repetition in a fresh session. collect re-derives summary.csv from the transcripts. compare writes compare.md. The orchestrator is a .NET project that builds to btbench.
The artifacts
Every figure above resolves to one of these files, at the section, table row and column the registry records. The four the figures cite travel with the site so the check runs offline.
| File | Contents |
|---|---|
| report.md | The full report, all arms and slices |
| compare.md | The two-arm head to head |
| summary.csv | One row per run, every measured field |
| environment.json | Versions, git commit, MCP probe, local model, corpus hashes |
| raw.jsonl | One row per API request |
| runs.jsonl | One row per run as recorded by the orchestrator |
| logs | Per-run stream-json transcripts and stderr |