TL;DR: Keep one string per tool-use block, append each partial_json to it, and parse it at content_block_stop, ready for that parse to fail, because a max_tokens stop can cut a parameter off in either mode. With eager_input_streaming on, the input also arrives unchecked, so it may not be valid JSON.

When a streamed response from the Claude API calls a tool, the tool’s arguments arrive as a run of input_json_delta events, each carrying a piece of a JSON string. The streaming documentation shows one such call in full, for a get_weather tool asked about San Francisco. This is the tool-use block from that example, start to stop:

event: content_block_start
data: {"type":"content_block_start","index":1,"content_block":{"type":"tool_use","id":"toolu_01T1x1fJ34qAmk2tNTrN7Up6","name":"get_weather","input":{}}}

event: content_block_delta
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":"{\"location\":"}}

event: content_block_delta
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":" \"San"}}

event: content_block_delta
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":" Francisc"}}

event: content_block_delta
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":"o,"}}

event: content_block_delta
data: {"type":"content_block_delta","index":1,"delta":{"type":"input_json_delta","partial_json":" CA\"}"}}

event: content_block_stop
data: {"type":"content_block_stop","index":1}

Appended in order, the pieces in that block make {"location": after the second delta and {"location": "San Francisc after the fourth, and nothing parses as JSON until the last fragment closes the quote and the brace. That is why the parse belongs at content_block_stop, for this call and for every other tool call.

Current models, the streaming page says, emit one complete key and value from the input at a time, so there can be delays between events while the model works. Once a key and value are complete, they go out as several deltas of chunked partial JSON, “so that the format can automatically support finer granularity in future models.” So for a tool that doesn’t set the eager_input_streaming flag, a fragment arriving is not a sign of how far the model has got through the value. A client that prints each fragment the moment it arrives is printing pieces of a value that is already done.

Turning on eager_input_streaming

Without the flag, the fine-grained tool streaming page says, the API buffers and validates each parameter value before streaming it back, so nothing prints for a large parameter until Claude has finished generating it. Its example is a make_file tool asked to write a long poem to poem.txt, whose lines_of_text parameter is an array of lines.

Setting "eager_input_streaming": true on a user-defined tool turns that buffering off. One thing to remember is that the request itself needs streaming turned on as well. Fragments start arriving as soon as Claude begins the parameter, which is how the poem can fill a terminal while it is being written. The page says these fragments are typically longer, too, with fewer breaks in the middle of a word. The get_weather trace above comes from a tool that sets no flag, and it has one of those breaks: Francisc and o, arrive as two separate fragments.

Turning the flag on also means the API no longer checks a tool’s input before sending it, so the accumulated string may be partial or invalid JSON. The events themselves don’t change, though: they are the same input_json_delta events, and a client appends them the same way, so the code that collects them can stay as it is.

When the input doesn’t parse

A response can stop at max_tokens in the middle of a parameter, with the flag or without it. The string then ends partway through, much like the halfway points in the trace above, and it won’t parse, so the parse at the stop event should be ready to fail in both modes. The stop reason shows when that has happened, so it’s worth checking before deciding what to do next: retry the request with a higher max_tokens, or repair the partial input. With the flag on, the input can also fail to parse because nothing checked it. The tool can’t run on that, and the fine-grained page describes how to report it back to Claude as an error result.

The flag is set per tool, so it can be on for a tool like make_file that writes out long text and off for a tool like get_weather that takes one short string. This is only one small part of the API, the moment when a tool’s arguments come in as pieces, and if you have handled those pieces differently in your own client, I’d be glad to hear how.