<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>Claude API on KbWen Blog</title>
    <link>https://www.kbwen.com/tags/claude-api/</link>
    <description>KbWen is a practical technology blog about AI systems, machine learning, Python, data engineering, and software development.</description>
    <generator>Hugo</generator>
    <language>zh-tw</language>
    <image>
      <url>https://www.kbwen.com/images/og-default.png</url>
      <title>KbWen Blog</title>
      <link>https://www.kbwen.com/</link>
    </image>
    
    <lastBuildDate>Mon, 05 Oct 2026 09:50:00 +0800</lastBuildDate><atom:link href="https://www.kbwen.com/tags/claude-api/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Claude API 串流時的工具呼叫與不完整的 JSON</title>
      <link>https://www.kbwen.com/claude-api-streaming-tool-call-partial-json/</link>
      <pubDate>Mon, 05 Oct 2026 09:50:00 +0800</pubDate><dc:creator>KbWen</dc:creator>
      <guid>https://www.kbwen.com/claude-api-streaming-tool-call-partial-json/</guid>
      <description>用 Claude API 開串流的時候，參數會是一段段字串。這篇簡單的看這些片段長什麼樣子，以及程式該怎麼接手處理。</description>
      <content:encoded><![CDATA[<p>用 Claude API 的時候把串流打開（<code>&quot;stream&quot;: true</code>），回覆會變成一連串 server-sent events，文字會一小段一小段送過來，可以邊收取並同時印出文字。Claude 決定呼叫工具的時候，工具的參數也是用同樣的方式分段送來，只是這些片段是 JSON 的一部分，處理方式不太一樣。</p>
<p><a href="https://platform.claude.com/docs/en/build-with-claude/streaming">官方的串流文件</a>裡有個查天氣的例子：使用者問舊金山的天氣，Claude 先回了一句話（index 0 的文字區塊），接著開了 index 1 的工具區塊，呼叫 <code>get_weather</code>。下面是事件：</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-text" data-lang="text"><span class="line"><span class="cl">event: content_block_start
</span></span><span class="line"><span class="cl">data: {&#34;type&#34;:&#34;content_block_start&#34;,&#34;index&#34;:1,&#34;content_block&#34;:{&#34;type&#34;:&#34;tool_use&#34;,&#34;id&#34;:&#34;toolu_01T1x1fJ34qAmk2tNTrN7Up6&#34;,&#34;name&#34;:&#34;get_weather&#34;,&#34;input&#34;:{}}}
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">event: content_block_delta
</span></span><span class="line"><span class="cl">data: {&#34;type&#34;:&#34;content_block_delta&#34;,&#34;index&#34;:1,&#34;delta&#34;:{&#34;type&#34;:&#34;input_json_delta&#34;,&#34;partial_json&#34;:&#34;&#34;}}
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">event: content_block_delta
</span></span><span class="line"><span class="cl">data: {&#34;type&#34;:&#34;content_block_delta&#34;,&#34;index&#34;:1,&#34;delta&#34;:{&#34;type&#34;:&#34;input_json_delta&#34;,&#34;partial_json&#34;:&#34;{\&#34;location\&#34;:&#34;}}
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">event: content_block_delta
</span></span><span class="line"><span class="cl">data: {&#34;type&#34;:&#34;content_block_delta&#34;,&#34;index&#34;:1,&#34;delta&#34;:{&#34;type&#34;:&#34;input_json_delta&#34;,&#34;partial_json&#34;:&#34; \&#34;San&#34;}}
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">event: content_block_delta
</span></span><span class="line"><span class="cl">data: {&#34;type&#34;:&#34;content_block_delta&#34;,&#34;index&#34;:1,&#34;delta&#34;:{&#34;type&#34;:&#34;input_json_delta&#34;,&#34;partial_json&#34;:&#34; Francisc&#34;}}
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">event: content_block_delta
</span></span><span class="line"><span class="cl">data: {&#34;type&#34;:&#34;content_block_delta&#34;,&#34;index&#34;:1,&#34;delta&#34;:{&#34;type&#34;:&#34;input_json_delta&#34;,&#34;partial_json&#34;:&#34;o,&#34;}}
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">event: content_block_delta
</span></span><span class="line"><span class="cl">data: {&#34;type&#34;:&#34;content_block_delta&#34;,&#34;index&#34;:1,&#34;delta&#34;:{&#34;type&#34;:&#34;input_json_delta&#34;,&#34;partial_json&#34;:&#34; CA\&#34;}&#34;}}
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl">event: content_block_stop
</span></span><span class="line"><span class="cl">data: {&#34;type&#34;:&#34;content_block_stop&#34;,&#34;index&#34;:1}
</span></span></code></pre></div><p>仔細看這六段字串，第一段是空的，第二段是 <code>{&quot;location&quot;:</code>，只有左大括號和一個 key；<code> &quot;San</code> 開了引號沒有關，<code> Francisc</code> 則斷在單字的中間。每段都不能單獨拿去給 JSON 解析器，那肯定會錯，因為它們只是同一串文字切開後的段落。要等到 <code> CA&quot;}</code> 送來，頭尾符號都有了，接著 <code>content_block_stop</code> 表示這個區塊結束；這時把六段照順序接起來，才會是 <code>{&quot;location&quot;: &quot;San Francisco, CA&quot;}</code>。</p>
<p>所以程式在收參數的時候，要做的事情其實很簡單。<a href="https://platform.claude.com/docs/en/agents-and-tools/tool-use/fine-grained-tool-streaming">fine-grained tool streaming 文件</a>把它寫成三步：收到 tool_use 的 <code>content_block_start</code> 時準備一個空字串，每個 <code>input_json_delta</code> 來就把 <code>partial_json</code> 接到後面，等到 <code>content_block_stop</code> 再解析，解析要包在 try 裡。如果用的是 Python、TypeScript 或 Go 的 SDK，裡面已經有 helper 會把片段接好；但如果是直接處理事件，或是想自己決定怎麼處理的時候，才需要照這三步驟。</p>
<h2 id="eager_input_streaming">eager_input_streaming</h2>
<p>查天氣的參數只有一個城市名，很短也很快就送完了。如果參數很長，像是整份文件或整段程式碼，預設的做法就會讓人等比較久：API 會先把每個參數值緩衝起來、驗證過之後才送出。</p>
<p>在自己定義的工具上把 <code>eager_input_streaming</code> 設成 <code>true</code>，請求本身也記得要設成串流，這個參數就不經過伺服器端的緩衝跟 JSON 驗證，Claude 開始的同時，片段也就跟著跑出來。fine-grained tool streaming 文件的範例就打開了這個設定，片段一到就印出來，讓人看到參數寫到哪裡；不過要知道此時印出來的是還沒接完的字串，還不能當參數用。接的方式跟前面一樣，但是伺服器沒有先驗證，所以到了 <code>content_block_stop</code>，接完的字串不保證是合法的 JSON。</p>
<p>因此，我們在使用時，要不要替某個工具打開 <code>eager_input_streaming</code>，可以確認情境，有時候很長再空等待那也許是個可以考慮的方向。
總之這篇提供個 Claude 的不同小用法，有任何討論和想法歡迎提供。</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Where the cache_control breakpoint goes in Claude prompt caching</title>
      <link>https://www.kbwen.com/claude-prompt-caching-breakpoint-placement/</link>
      <pubDate>Mon, 05 Oct 2026 09:45:00 +0800</pubDate><dc:creator>KbWen</dc:creator>
      <guid>https://www.kbwen.com/claude-prompt-caching-breakpoint-placement/</guid>
      <description>A walk through how the Claude API writes and looks up prompt cache entries, using the docs&amp;#39; example of a breakpoint that never gets a cache hit, and how to check the result in the usage fields.</description>
      <content:encoded><![CDATA[<p>Prompt caching in the Claude API lets a request skip reprocessing the start of its prompt when an earlier request has already written that same start to the cache. You turn it on with a <code>cache_control</code> field, either once at the top level of the request, where the API chooses the spot for you, or on a specific content block, which Anthropic calls an explicit cache breakpoint.</p>
<p>The <a href="https://platform.claude.com/docs/en/build-with-claude/prompt-caching">prompt caching docs</a> list one prompt as a common mistake. Its large system context is split over five blocks that are the same on every request. A sixth block holds a timestamp and the user&rsquo;s message, so it is different every time. <code>cache_control</code> sits on block 6, the last block in the prompt.</p>
<p>On the first request nothing is cached yet. The API processes the whole prompt and writes one entry, at block 6. The docs describe that entry as &ldquo;a hash of the prefix ending at that block,&rdquo; and the hash is cumulative, so it covers blocks 1 through 6 with the timestamp inside it. The five static blocks are part of that hash, but none of them gets an entry of its own, because the API writes nothing at positions before the breakpoint.</p>
<p>The second request arrives with the same five blocks and a new timestamp. The API computes the hash at block 6, and since the timestamp changed, the hash changed, and there is no entry to match. It then walks backward a block at a time, from block 5 down to block 1, checking the prefix hash at each position against the cache. Those five blocks are word for word what the first request sent. But the API is looking for an entry that some earlier request wrote at that position, and no request has ever written one at block 5 or anywhere before it. The walk comes back empty. The second request writes its own entry at block 6, with its own timestamp in the hash, and the third request will miss that one in the same way, because its timestamp will be different again.</p>
<p>So with the breakpoint on block 6, every request is a cache write and none is a read. That shows up on the bill, because a write is not priced like ordinary input. On the <a href="https://platform.claude.com/docs/en/about-claude/pricing">pricing page</a>, a write to the five-minute cache costs 1.25 times the base input price, and a read costs a tenth of it on most models (a few, Opus 5.5 among them, read for less). The same page says the five-minute cache &ldquo;pays off after one cache read.&rdquo; Since that read never comes, every request pays the write price on everything up to block 6, which costs more than sending the same prompt with no <code>cache_control</code> at all.</p>
<p>Switching to automatic caching doesn&rsquo;t get around this, since automatic caching puts the breakpoint on the last cacheable block, and in this prompt that is block 6 again. The fix is an explicit breakpoint one block up, on block 5, the last block that stays the same from one request to the next:</p>
<div class="highlight"><pre tabindex="0" class="chroma"><code class="language-python" data-lang="python"><span class="line"><span class="cl"><span class="n">client</span> <span class="o">=</span> <span class="n">anthropic</span><span class="o">.</span><span class="n">Anthropic</span><span class="p">()</span>
</span></span><span class="line"><span class="cl">
</span></span><span class="line"><span class="cl"><span class="n">response</span> <span class="o">=</span> <span class="n">client</span><span class="o">.</span><span class="n">messages</span><span class="o">.</span><span class="n">create</span><span class="p">(</span>
</span></span><span class="line"><span class="cl">    <span class="n">model</span><span class="o">=</span><span class="s2">&#34;claude-opus-5-5&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">max_tokens</span><span class="o">=</span><span class="mi">1024</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">    <span class="n">system</span><span class="o">=</span><span class="p">[</span>
</span></span><span class="line"><span class="cl">        <span class="p">{</span><span class="s2">&#34;type&#34;</span><span class="p">:</span> <span class="s2">&#34;text&#34;</span><span class="p">,</span> <span class="s2">&#34;text&#34;</span><span class="p">:</span> <span class="n">context_1</span><span class="p">},</span>
</span></span><span class="line"><span class="cl">        <span class="c1"># context_2 through context_4, unchanged</span>
</span></span><span class="line"><span class="cl">        <span class="p">{</span>
</span></span><span class="line"><span class="cl">            <span class="s2">&#34;type&#34;</span><span class="p">:</span> <span class="s2">&#34;text&#34;</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">            <span class="s2">&#34;text&#34;</span><span class="p">:</span> <span class="n">context_5</span><span class="p">,</span>
</span></span><span class="line"><span class="cl">            <span class="s2">&#34;cache_control&#34;</span><span class="p">:</span> <span class="p">{</span><span class="s2">&#34;type&#34;</span><span class="p">:</span> <span class="s2">&#34;ephemeral&#34;</span><span class="p">},</span>
</span></span><span class="line"><span class="cl">        <span class="p">},</span>
</span></span><span class="line"><span class="cl">    <span class="p">],</span>
</span></span><span class="line"><span class="cl">    <span class="n">messages</span><span class="o">=</span><span class="p">[</span>
</span></span><span class="line"><span class="cl">        <span class="p">{</span><span class="s2">&#34;role&#34;</span><span class="p">:</span> <span class="s2">&#34;user&#34;</span><span class="p">,</span> <span class="s2">&#34;content&#34;</span><span class="p">:</span> <span class="sa">f</span><span class="s2">&#34;[</span><span class="si">{</span><span class="n">timestamp</span><span class="si">}</span><span class="s2">] </span><span class="si">{</span><span class="n">user_message</span><span class="si">}</span><span class="s2">&#34;</span><span class="p">},</span>
</span></span><span class="line"><span class="cl">    <span class="p">],</span>
</span></span><span class="line"><span class="cl"><span class="p">)</span>
</span></span></code></pre></div><p>Now the first request writes its entry at block 5, and that hash ends before the timestamp, which means a new timestamp no longer changes it. When the second request computes the hash at block 5, it gets the same value and finds the entry the first request left right there, so it never needs the walk back described above. It reads blocks 1 through 5 from the cache. Only the timestamp and the message go through as ordinary input. The docs put the rule this way: place <code>cache_control</code> &ldquo;on the last block whose prefix is identical across the requests you want to share a cache.&rdquo; If anything in blocks 1 through 5 changes between two requests, even slightly, the second request writes instead of reading.</p>
<p>Caching has no effect on Claude&rsquo;s reply, so the place to see whether a request hit the cache is the response&rsquo;s <code>usage</code> object. It splits the input tokens three ways: <code>cache_creation_input_tokens</code> for tokens written to the cache on this request, <code>cache_read_input_tokens</code> for tokens read from it, and <code>input_tokens</code> for the tokens after the last breakpoint. That last one is worth keeping an eye on, because it is not the size of the prompt.</p>
<p>With the breakpoint on block 6, every response looks alike: nearly the whole prompt appears under <code>cache_creation_input_tokens</code>, and <code>cache_read_input_tokens</code> stays at 0. With the breakpoint on block 5, the first response looks much the same, since that request still has to write the entry, but the timestamp and message now show up under <code>input_tokens</code>.</p>
<p>The second response is the one that shows whether the move worked. When you test it, remember to send that request within five minutes of the first one, since by default an entry lasts five minutes counted from the start of the request that used it. It also can&rsquo;t go out alongside the first, because the entry only becomes available once the first response begins. The static context should then appear under <code>cache_read_input_tokens</code>, <code>cache_creation_input_tokens</code> should be 0, and <code>input_tokens</code> should be just the last block. If both cache fields read 0, nothing was cached at all. That usually means the prefix is shorter than the minimum the model will cache (512 tokens on Opus 5.5, more on some other models), and the API returns no error when that happens.</p>
<p>This is just one small idea from prompt caching: where to put the breakpoint, and how to check that it worked. If you&rsquo;ve seen the cache behave differently from what&rsquo;s described here, I&rsquo;d like to hear about it.</p>
]]></content:encoded>
    </item>
    
    <item>
      <title>Claude API 的 prompt caching：cache_control 放在哪裡</title>
      <link>https://www.kbwen.com/claude-prompt-caching-breakpoint-placement-zh/</link>
      <pubDate>Mon, 05 Oct 2026 09:40:00 +0800</pubDate><dc:creator>KbWen</dc:creator>
      <guid>https://www.kbwen.com/claude-prompt-caching-breakpoint-placement-zh/</guid>
      <description>來看看 Claude API 的 prompt caching 是怎麼找快取的，以及 cache_control 放在哪裡會影響有沒有正確讀到。</description>
      <content:encoded><![CDATA[<p>Claude API 有提示快取（prompt caching）的功能，讓每次都相同的文字內容不必重新處理。做法是在請求裡加上 <code>cache_control</code>，第一次送出時，系統會把那段內容寫進快取，之後的請求如果一樣，就直接從快取讀，所花費的時間和費用都會比較少。</p>
<p>標上 <code>cache_control</code> 的區塊，<a href="https://platform.claude.com/docs/en/build-with-claude/prompt-caching">官方文件</a>稱為斷點（breakpoint）。快取只會在斷點寫入一筆資料，內容是整段前綴的雜湊。雜湊是累積的，所以斷點本身或它之前任何一個區塊有改動，下一次出來的值就會是另一個雜湊。不過讀取的時候，系統會先拿斷點位置的雜湊去找，沒找到就往回退一步，換位置的雜湊再找一次，這樣一步一步找。</p>
<p>規則放到實際的請求裡，有些要注意的地方。文件裡有個例子，請求前面五個區塊是固定的系統內容，每次送出去都完全相同；第六個區塊放的是這次的時間戳記，以及使用者剛打的訊息。如果把 <code>cache_control</code> 標在最後面，也就是第六個區塊上，第一次請求確實會寫進快取，只是寫進去的雜湊內容，是從第一個區塊一路到第六個區塊，也就是說，時間戳記也算在裡面。</p>
<p>但是到了第二次請求，時間戳記換了，第六個區塊的雜湊也當然會不同，因為時間不一樣了，系統找不到對應的紀錄，於是像剛剛說的開始往回找，第五、第四，一路到第一個區塊。但因為這幾個位置上一次都沒有寫過任何東西，因為寫入只發生在斷點，所以一直找到最前面，也不會讀到雜湊。前面五個區塊明明完全沒變，但是往回找只是查之前的請求已經寫過的紀錄，不會因為內容沒變就另外補存一筆。結果導致每一次請求都會再寫進一筆新的，到了下一次請求，又因為時間戳記不同而對不上。</p>
<p>如果把斷點往前移到第五個區塊，也就是最後一個每次都相同的區塊，第一次請求寫入的就會在這裡，而不是包含會改變的時間。到了第二次請求時，我們可以在第五個區塊上對應到雜湊，前面五個區塊就從快取讀取出來，會變的第六個區塊照一般的輸入處理，也就不會影響快取。</p>
<p>如果用的是自動快取，也就是把 <code>cache_control</code> 放在請求的最外層、讓系統自己決定斷點，在這個例子裡也會碰到相同的問題。自動快取會把斷點放在最後一個可以快取的區塊，而這裡的最後一個區塊剛好就是每次都會變的那一個，所以建議還是改用明確的斷點。</p>
<p>那放錯位置的影響是什麼呢？請求照樣會成功，回答也跟沒有開快取時一樣，但差別會出現在你的花費上，快取能有效地降低花費。5 分鐘快取的寫入價格是一般輸入的 1.25 倍，讀取則只要一般輸入的一點點。斷點放錯地方，每一次反而要多付一點快取的價格，讀取卻一次也沒發生過，實際比較下來比完全不開快取還多付一些。那關於寫入和讀取的價格差別怎麼影響整體的設計，可以參考之前寫的相關文章：<a href="/token-cost-and-budget-tiers/">Token 成本的真相：分級，但別分太細</a>。</p>
<p>因為快取不會有錯誤訊息，要確認有沒有讀到，就要看回應裡的 <code>usage</code>。<code>cache_creation_input_tokens</code> 是這次寫進快取的 token，<code>cache_read_input_tokens</code> 是從快取讀出來的 token。第一次請求本來就只會有寫入，所以主要看第二次：同樣的開頭內容再送一次（兩次之間要在快取的有效時間內，預設是 5 分鐘），<code>cache_read_input_tokens</code> 應該會大於 0。如果第二次還是只有寫入、讀取是 0，就可以回頭看看斷點是不是標在每次都會變的區塊上，或是其他可能錯誤。另外如果沒有被快取，這種情況同樣不會回傳錯誤，原因可能是內容沒有達到該模型的最低長度。</p>
<p>最後說明，這篇是照 Anthropic 目前的 prompt caching 文件和價格頁整理的。各個模型的最低長度、快取設計以及讀取價格也不太一樣，而且會隨新模型更新，實際的數字以文件上的表格為準，這裡還是提供一個小概念和想法讓大家參考。</p>
]]></content:encoded>
    </item>
    
  </channel>
</rss>
