How to cut your AI agent's bill, make cost visible and keep the cache warm
Every agent message re-reads the whole conversation, so cost climbs with each one. How to see it per session, and the 3 cache rules that keep it down.
· Jeremy

Give an agent a 40-message task and it won't cost four times what a 10-message task costs. It'll cost more than that, for a simple reason: every reply re-reads the whole conversation so far. Message 40 doesn't just answer your latest request, it reads messages 1 through 39 again first. The longer a conversation runs, the more expensive each new message gets.
There are two things you can do about that: see the cost instead of guessing at it, and keep the part of the system that controls most of the bill, the cache, working the way it's supposed to.
See the cost, don't guess
Three numbers are worth watching on any agent session: what it costs at message 10, what it costs at message 40 (or wherever the task ends), and what share of that spend is cache reads versus fresh, uncached input. Together they tell you whether a long task is behaving normally or quietly getting expensive.
agentproto's own sessions already carry this. Every message writes a usage_snapshot event with the fields that make up those three numbers: costUsd, tokensIn, tokensOut, cacheReadTokens, cacheWriteTokens. You can read them for one session right now:
agentproto sessions show <id-or-name> --jsonHere's real output, from the very session that wrote this paragraph:
"costUsd": 0.7263239,
"tokensIn": 50,
"tokensOut": 11226,
"cacheReadTokens": 2024257,
"cacheWriteTokens": 83645tokensIn is what this one message sent fresh: 50 tokens. cacheReadTokens is everything from earlier in the conversation it read again instead: just over two million by this point. That gap between fresh input and cache reads is, in a single session, the whole story this post is about.
For a wider view, those same events sit locally at ~/.agentproto/sessions/<id>/events.jsonl, one per message, read-only, nothing sent anywhere. That's what let us compute the three numbers above across hundreds of sessions at once, not just one (details in the method section below).
Keep the cache warm
Reading a long conversation isn't free, but if a part of it was already read on an earlier message, it can be served from a cache instead, at a fraction of the price. That's what keeps most agent conversations affordable at all. It only works under three conditions.
- Keep the start of the conversation stable. Put instructions, tool definitions, and files first, and don't edit them mid-task. The cache works like a prefix match: change anything near the top of the conversation and every message after it stops matching, so the whole thing gets read fresh again, at full price.
- One model per session. A cache is built for one model's own read of the conversation. Switch models mid-task and the new model starts from zero, caching nothing from before. If a task genuinely needs a different model partway through, that's a new session, not a model switch inside the old one.
- Don't leave a session sitting idle for long. A cache doesn't last forever, it goes cold after a few quiet minutes. Come back to a conversation after that and the next message pays full price to rebuild it, same as if nothing had ever been cached. Keeping sessions focused, one task each, rather than one long session left open across many unrelated ones, is what keeps this from happening by accident.

Get those three rules right and the cache does most of the work for you, which is the next section, in numbers.
What we found when we checked
We looked at our own telemetry to see whether this actually holds up: 160 real agentproto sessions long enough to measure directly, out of 68,523 total.
A 40-message session costs about 4x a 10-message one, not more. (A widely shared model for agent costs, from a Microsoft talk on TokenOps, describes a no-caching scenario reaching 16x at that point; ours don't, because caching is doing its job.)
Where the money actually goes: across 293 sessions and $1,241 of real spend, 72% of the total was cache reads. Caching cut that part of the bill by roughly 9x, about $865 actually spent versus an estimated $8,650 if none of it had been cached.
Three things we looked for but couldn't measure, because nobody had used them yet in this data: compacting a long conversation down, a hard cost ceiling on a session, and restarting a task fresh from a handoff checkpoint. We're not claiming those don't help, only that we don't have a real before-and-after number for them yet.
Method
The numbers above come from reading each session's local events.jsonl under ~/.agentproto/sessions/, specifically every usage_snapshot event, and treating a session's position in that list as its message number. Read-only, nothing leaves the machine, so the same method reproduces against any agentproto install with different numbers.
- Scanned 68,523 session directories. 293 had enough usage data to analyze at all (most sessions are short, under 20 messages, or report usage through a mechanism that lacks the token-type breakdown). 160 of those 293 were long enough to compute a real
cost(message 40) / cost(message 10)ratio. - Median ratio across those 160 sessions: 4.3x (interquartile range about 3.9x to 4.8x, minimum 2.7x, maximum 12.9x). The marginal cost of message 40 alone versus message 10 alone, not the running total, has a median ratio of just 1.08x: it's the accumulation of many messages that adds up, not any single message exploding in price.
- Token-type breakdown across 293 sessions and $1,241 of metered spend: cache reads 72.1%, model output 15.5%, cache writes 12.4%, fresh uncached input roughly 0%. A cache read prices at about a tenth of fresh input (confirmed by regressing cost deltas against token deltas: $0.20 per million tokens for cache reads versus $2.00 for input, matching the published rate card exactly).
- This is one daemon's own dogfooding and development sessions (building agentproto itself, product QA, assorted coding tasks), not a controlled benchmark, and the task mix isn't constant across messages. Turn 1 typically carries a session's one-time setup cost (reading project instructions, initial cache writes), which inflates the message-10 baseline and is a plausible reason the measured ratio lands below what a naive doubling-of-doubling model would predict.
- On the three "not measured" items:
session_compactwas called zero times across all 68,523 sessions. The only session-ending reasons observed anywhere in the data arecompleted,exited,cancelled,error,aborted, andwatchdog-timeout, none tied to a cost cap. Two handoff-checkpoint attempts exist in the data (a deliberate smoke test), and the resulting session's context came out larger, not smaller, right after the handoff, so that one test doesn't show a saving either way.