Why your agent subscription runs out so fast
Your coding agent re-reads the whole conversation at every step. Where your allowance really goes, and four rules to pay less.
· Jeremy
A coding agent re-reads everything you have said to it each time it takes a step. Even with a cache, a saved copy that makes each read a tenth of the price or less, the re-reading is most of the cost. When the cache vanishes, the agent pays more than full price to read it all again.
One evening in August, a developer sent five messages to Claude Code, Anthropic's coding agent. The developer goes by thepiesco on GitHub, the site where developers file bug reports. By the end, roughly half of a five-hour usage window was gone. A subscription hands out its allowance in windows like that one.
That evening the tool made 339 requests to the model, across several sessions running at once. A single message from you can set off dozens of steps, each one a request, and at every step the whole conversation goes back to the model with your newest line on top. That re-reading is what eats the allowance. If you want the fix first, the four rules are under "What to do."

At every step the model gets the whole stack again. The new message is the thin green slice.
Paying full price for all that re-reading would be unaffordable. So the provider keeps a saved copy of everything before your newest message, called the cache. Reading it costs a tenth of the full price, or less.
I run coding agents all day and my subscriptions kept running out, so I used agentproto, a tool that keeps a record of my sessions, to price them. Developers often write that long sessions get steeply more expensive. Gareth Bland, chief data scientist at Microsoft, describes in a video an agent that pays full price for the re-reading at every step: four times as many steps cost sixteen times as much.
My long sessions, with the cache working, say something milder. Stretched four times longer, they cost 4.3 times as much, close to a straight line.

My long sessions sit almost on the straight line, far below the uncached curve; only the end points are measured.
Across 293 of my sessions, about 72 cents of every dollar went to the agent reading the conversation back. I pay a flat subscription, so those dollars are a yardstick: the list price of the same usage on the API, the route developers use when they buy the model directly. They pay per token, and a token is the word-sized piece of text a model reads and charges by.

Counted in money, cache reads are most of the bill; counted in tokens, they are nearly all of it.
A read from the cache is cheap, and it is still where the money goes, because the reads never stop.
Losing the cache is what multiplies that cost, and it is what happened to thepiesco. In the logs the cache kept vanishing, and each time the whole conversation was written back in. "This is not user behaviour," thepiesco wrote in the bug report. Each of those is a cold start: with no cache, the model reads everything as if it were new. On a subscription Claude Code keeps the cache for an hour, and writing a cache meant to last that long costs twice the full price. That is how a lost cache ends up costing more than full price.

Sixteen of the 339 requests carried half the evening's cost, as thepiesco counted it in the bug report.
Most people never see logs like these. They see the allowance gone and look for something to change.
What to do
- Pick the model when the task starts and keep it. Each model keeps its own cache, so a switch starts a new one. (Documented.)
- Stay in one session until the task is done, then end it. Staying is cheaper than starting over, because each step reads the cheap cache. The last rule is the one exception. (The discount is documented; the cost at four times the length is measured on my long sessions.)
- Keep /clear, the command that empties the conversation, for the end of a task. A clear in the middle throws the cache away. (Reasoned, untested.)
- Before a break of over an hour, or when the conversation is full, write four lines yourself and start fresh. Full means close to the most the model can hold. What you decided and ruled out, what is finished, what comes next, which files matter. You choose what survives; the summary the agent writes on its own when the conversation is full chooses for you. On the API, read "an hour" as five minutes. (The timing is measured and documented; the four lines are reasoned.)
The last two rules look like they disagree, and the difference is whether the cache is still worth keeping. A clear in the middle of a task throws away a cache that is working. After an hour away the cache has expired on its own, so starting fresh from your own four lines is a cheap restart. A full conversation gets summarized anyway, so you may as well write the summary yourself. On a subscription, a short break needs none of it: in my long sessions, pauses of up to nearly an hour caused no cold start.
To see whether your cache is holding, type /usage in a recent version of Claude Code. The line headed Prompt cache counts cold starts, which it calls misses. It also says whether the cache is warm, meaning still saved, or cold, meaning gone. If it reads cold when you come back to a task, my advice is to write your four lines and start fresh.
On a subscription, that screen flags the misses once they account for 10 percent or more of your recent usage. If your version shows no Prompt cache line, a cold start on a long conversation shows up as one ordinary message that takes a visible bite out of the allowance.
How long the cache lasts
On the API, where you pay per token, the cache lasts five minutes. Every message that uses it restarts the clock. On a subscription, Claude Code's documentation sets the limit at an hour.

A coffee break finds the cache cold on the API and warm on a subscription.
My own logs agree with the longer figure, when the cache behaves. In thepiesco's logs it had been written to last an hour and still vanished, with 33 to 61 seconds between requests. The developer suspected it was pushed out while several sessions competed for space.
Was it a bug, and was it fixed? The thread does not say. An automated note later closed the report for inactivity. No person replied and no fix is recorded. So it may still happen to you, and the Prompt cache line in /usage is where you would see it.
Switching, compacting, clearing
Three habits feel like thrift and can cost more. I did not measure them; the case rests on the documentation and on reasoning.
The first is moving a task to a different model. The documentation says it plainly: "Each model has its own cache." The new model's first step is a cold start.
Say you switch from Sonnet 5 to the cheaper Haiku 4.5, two of the models Claude Code can run. Haiku has no cache of your conversation, so its first step writes all of it in as new. Staying on Sonnet would have meant reading its cache instead. Take the list prices for those two versions. Writing the conversation into Haiku's cache costs at least six times as much per token as reading Sonnet's cache would have.
The second is compacting, the automatic summary: the agent summarizes the conversation when it is full. A summary is shorter and so cheaper to carry, but the specifics tend to go first, the file paths and the reasons something was ruled out. Then the agent has to rediscover them or I have to explain them again, and every extra step reads the whole conversation back.
The third, clearing, is the blunt version: /clear empties the conversation outright. A developer who writes as recca0120 built a cost model and concluded that "frequent '/clear' can cost more than keeping a long session alive."
The one rule
The four rules come down to one: keep the cache while it is working. Let it go when the task is done, when it has expired, or when the conversation is full.

When the cache has expired, Claude Code says so in a plain sentence like this one, as quoted in a bug report.
Someone had left a conversation alone for nearly eleven days, and when they came back, the first thing the agent had to do was read it all again, from the top.