Prompt caching in the Claude API lets a request skip reprocessing the start of its prompt when an earlier request has already written that same start to the cache. You turn it on with a cache_control field, either once at the top level of the request, where the API chooses the spot for you, or on a specific content block, which Anthropic calls an explicit cache breakpoint.
The prompt caching docs list one prompt as a common mistake. Its large system context is split over five blocks that are the same on every request. A sixth block holds a timestamp and the user’s message, so it is different every time. cache_control sits on block 6, the last block in the prompt.
On the first request nothing is cached yet. The API processes the whole prompt and writes one entry, at block 6. The docs describe that entry as “a hash of the prefix ending at that block,” and the hash is cumulative, so it covers blocks 1 through 6 with the timestamp inside it. The five static blocks are part of that hash, but none of them gets an entry of its own, because the API writes nothing at positions before the breakpoint.
The second request arrives with the same five blocks and a new timestamp. The API computes the hash at block 6, and since the timestamp changed, the hash changed, and there is no entry to match. It then walks backward a block at a time, from block 5 down to block 1, checking the prefix hash at each position against the cache. Those five blocks are word for word what the first request sent. But the API is looking for an entry that some earlier request wrote at that position, and no request has ever written one at block 5 or anywhere before it. The walk comes back empty. The second request writes its own entry at block 6, with its own timestamp in the hash, and the third request will miss that one in the same way, because its timestamp will be different again.
So with the breakpoint on block 6, every request is a cache write and none is a read. That shows up on the bill, because a write is not priced like ordinary input. On the pricing page, a write to the five-minute cache costs 1.25 times the base input price, and a read costs a tenth of it on most models (a few, Opus 5.5 among them, read for less). The same page says the five-minute cache “pays off after one cache read.” Since that read never comes, every request pays the write price on everything up to block 6, which costs more than sending the same prompt with no cache_control at all.
Switching to automatic caching doesn’t get around this, since automatic caching puts the breakpoint on the last cacheable block, and in this prompt that is block 6 again. The fix is an explicit breakpoint one block up, on block 5, the last block that stays the same from one request to the next:
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=1024,
system=[
{"type": "text", "text": context_1},
# context_2 through context_4, unchanged
{
"type": "text",
"text": context_5,
"cache_control": {"type": "ephemeral"},
},
],
messages=[
{"role": "user", "content": f"[{timestamp}] {user_message}"},
],
)
Now the first request writes its entry at block 5, and that hash ends before the timestamp, which means a new timestamp no longer changes it. When the second request computes the hash at block 5, it gets the same value and finds the entry the first request left right there, so it never needs the walk back described above. It reads blocks 1 through 5 from the cache. Only the timestamp and the message go through as ordinary input. The docs put the rule this way: place cache_control “on the last block whose prefix is identical across the requests you want to share a cache.” If anything in blocks 1 through 5 changes between two requests, even slightly, the second request writes instead of reading.
Caching has no effect on Claude’s reply, so the place to see whether a request hit the cache is the response’s usage object. It splits the input tokens three ways: cache_creation_input_tokens for tokens written to the cache on this request, cache_read_input_tokens for tokens read from it, and input_tokens for the tokens after the last breakpoint. That last one is worth keeping an eye on, because it is not the size of the prompt.
With the breakpoint on block 6, every response looks alike: nearly the whole prompt appears under cache_creation_input_tokens, and cache_read_input_tokens stays at 0. With the breakpoint on block 5, the first response looks much the same, since that request still has to write the entry, but the timestamp and message now show up under input_tokens.
The second response is the one that shows whether the move worked. When you test it, remember to send that request within five minutes of the first one, since by default an entry lasts five minutes counted from the start of the request that used it. It also can’t go out alongside the first, because the entry only becomes available once the first response begins. The static context should then appear under cache_read_input_tokens, cache_creation_input_tokens should be 0, and input_tokens should be just the last block. If both cache fields read 0, nothing was cached at all. That usually means the prefix is shorter than the minimum the model will cache (512 tokens on Opus 5.5, more on some other models), and the API returns no error when that happens.
This is just one small idea from prompt caching: where to put the breakpoint, and how to check that it worked. If you’ve seen the cache behave differently from what’s described here, I’d like to hear about it.


