Anthropic shipped anthropic-sdk-python v1.9.0 on September 28, and the release notes read less like a maintenance bump than like a wiring diagram for how the company wants agents built. Three items in the changelog matter more than the rest: a new between_tools thinking type, the ability to optionally run tool calls while the reply streams, and cache diagnostics reaching general availability on Message and MessageCreateParams. The same release adds a model id, claude-sonnet-5-5, and reorders the Model types so known ids list first.
Version numbers in vendor SDKs are usually noise. This one is not, because the two behavioral features point at the same architectural claim: that a model’s reasoning should be able to interleave with tool execution rather than bracket it.
between_tools is the interesting one
Most agent loops today look the same. The model emits a tool call, the harness executes it, the result goes back as a new turn, and the model thinks again. Thinking, when it is exposed at all, happens before the tool call or after the result lands. The between_tools thinking type implies a third position: reasoning that occurs in the gap, while tools are in flight.
That is a different control flow. If a model can think between tool invocations, the harness no longer has to serialize every step into a strict request-response ladder. It can let the model reason about a partially completed tool batch, or about a tool that is still running, without burning a full round trip to do it.
The bug fix attached to the same feature is the tell that this is not trivial plumbing. Commit a3834d4, tracking issue #952, makes the helpers “degrade between_tools thinking to disabled on fallback hops.” So the SDK has fallback logic: if a request lands somewhere that cannot honor the new thinking type, the client quietly turns it off rather than failing. That is the behavior of a feature being rolled out across a heterogeneous fleet, not a feature that is universally available. Anthropic is shipping the client-side contract before the server-side guarantee is uniform.
Tool calls that run while the reply streams
The second feature, commit 26d0812, lets tool calls optionally execute while the reply streams. Combine it with between_tools and the shape becomes clear. The reply is no longer a thing that finishes and then triggers work. The reply is a channel, and tool execution is a parallel activity on that channel.
There is a real cost here that the release notes do not address. Streaming tool execution means side effects can fire before the model has finished speaking. If a tool writes a file, sends a message, or hits a paid endpoint, the caller has to decide whether a half-finished stream is enough authorization to act. The SDK makes it optional, which is the right default, but it also means every agent framework built on this will need its own policy for when a streamed tool call is committed versus speculative.
Cache diagnostics go GA
Cache diagnostics reaching general availability is the least glamorous item and probably the most immediately useful. Prompt caching is one of the few levers that moves both latency and cost for long-context agent workloads, and it has historically been opaque: you either got a cache hit or you did not, and figuring out why was guesswork. Diagnostics on Message and MessageCreateParams means the response now carries structured information about what was cached and what was not.
The related fix, fad840c for issue #929, makes stream() and parse() accept diagnostics. That matters because streaming is how most production agents talk to the API. A diagnostics feature that only worked on the non-streaming path would have been useless to the people who need it most.
Anthropic is shipping the client-side contract for interleaved reasoning before the server-side guarantee is uniform.
The smaller items are not filler
claude-sonnet-5-5 appears as a new model id, and the chore e5a082a reorders the Model types to list known ids first. Neither tells us anything about the model’s capabilities, context window, or price. The release notes do not say. Anyone reading a benchmark claim into a changelog line is inventing it.
Two other additions are worth noting for operators. include_inherited and source land on workspace rate limits, which suggests rate limits now have a hierarchy where a workspace can inherit a parent’s limits, and the API will tell you where a given limit came from. And the Managed Agents events list filter gains typed event type values, tightening a previously loose string field.
There is also a small correctness fix with outsized implications for anyone piping files into the API: commit 9e9709d stops sending a placeholder filename for unnamed file uploads. If your pipeline was relying on that placeholder to identify uploads, it is now gone.
What it means for builders
The direction is legible. Anthropic is pushing reasoning and tool use toward a single interleaved stream, and it is doing so inside the client SDK before the server surface is fully uniform, using fallback degradation to paper over the gaps. That is a bet that agent harnesses will get more concurrent, not less.
The practical consequence for anyone building on this: your agent loop is probably more serial than it needs to be. If between_tools thinking and streamed tool execution both work as described, the round-trip-per-tool pattern that most frameworks inherited from chat completions is leaving latency and reasoning quality on the table.
The unresolved question is the one the changelog cannot answer. Degrading between_tools to disabled on fallback hops means the same request can produce different reasoning behavior depending on where it lands. For an eval suite, that is a confound. For a production agent, it is a silent behavior change. Watch whether Anthropic publishes which endpoints and model ids actually honor the new thinking type, because until that list exists, the feature is real in the client and conditional on the server.