llama.cpp v0.6.0 landed on October 5, and the headline is not a speedup or a new quantization format. It is an API. The release introduces llama_batch_ext, an extended batch type built around a new llama_process() call that accepts mixed token and embedding inputs in the same batch, along with per-token “state” embeddings for MTP and deepstack models (release notes, PR #24669). That is a structural change to how the runtime thinks about a forward pass. For most of its life, llama.cpp has assumed a batch is a sequence of token IDs. v0.6.0 assumes a batch can be tokens, embeddings, or both, and that individual positions can carry auxiliary state.

The reason is the models. GLM-5.3-Flash, which the release calls GLM5-Next, is a 320B text-and-vision hybrid built on a KDA/DSA attention scheme with mHC and MoE layers (#27773). Qwen4Exp gets MTP speculative decoding at roughly 1.5x decode speedup on an Nvidia DGX Spark, plus a set of correctness fixes (#29761, #29751). These are not vanilla transformer decoders. Multi-token prediction, deepstack embeddings, and hybrid attention all need to move more than one kind of tensor through the graph per step, and the old batch abstraction was the bottleneck.

The decision-model endpoint is the tell

The /v1/systemone server API is the most interesting thing in the changelog. It adds a dedicated decision pipeline for five models: laya, julia-1, lev, openjev (with vision), and kev (#29818), later extended to a sixth, nimble (#29844). These are not chat models. They are classifiers or policy models that take a context and emit a decision.

llama.cpp has spent years positioning itself as the local inference layer for open-weight chat models. Shipping a first-class endpoint for decision models, with its own pipeline and its own model registry, says the project sees a second workload arriving. Agent loops, routers, guardrails, and tool-selection policies all want a small, fast model that returns a label rather than a paragraph. That workload has different batching, different memory, and different latency requirements than chat. The /v1/systemone endpoint is llama.cpp admitting it.

The Clef model, supported for both text and vision (#29831, #29969), fits the same pattern. So does the addition of classifier_pooling support for rerankers (#29627) and the change to /v1/embeddings to accept typed vision, audio, and video content (#29556). The server is widening from a chat completions box into a general inference router.

ggml 0.26.0 and the hardware story

The runtime sits on ggml, which bumps to v0.26.0 in this release. The backend work is where the numbers live. On Metal, a new tensor-API flash attention kernel for F16 KV (#29570) and few-row MMA mat-mul kernels for speculative and batched decoding (#29869) deliver up to roughly 3x faster mat-mul on Apple GPUs. Vulkan gets sparse flash attention for quantized K/V (#29639). CUDA gets a model-driven W4A4 path using NVFP4 and MXFP4 (#24364) and shared-expert fusion into MMVQ (#29184). The Hexagon backend adds a sampler and more quant types. Windows ARM64 MSVC builds are now enabled.

The batch API and the decision-model endpoint point the same direction: llama.cpp is being rebuilt for models that do more than emit the next token.

Two smaller items matter for anyone running long contexts. llama_prefetch_rows() uses MADVISE-based prefetching of PLE tensors in Qwen4Exp and Gemma4 (#29599), and Qwen4Exp halves its indexer score memory (#29825). A separate fix saves exact rotation metadata so a mismatched KV cache rotation can be restored correctly (#28498). These are the unglamorous changes that decide whether a 320B hybrid actually runs on a workstation.

What to watch

The session and state formats both bumped: LLAMA_SESSION_VERSION to 11 and LLAMA_STATE_SEQ_VERSION to 4. Any tool that persists sessions across versions needs to handle that. The Web UI also got a real overhaul, with a Hugging Face Hub data layer (#27947), a model download pipeline (#27959), and memory-fit estimation (#27957). That is llama.cpp moving toward being a self-contained local app rather than a library that expects you to bring your own front end.

The opinion: llama.cpp’s center of gravity is shifting from “run Llama locally” to “run whatever the open-weight ecosystem ships, including the small models that make agents work.” The llama_batch_ext API and the /v1/systemone endpoint point the same direction. The project is being rebuilt for models that do more than emit the next token, and the release cadence suggests the maintainers expect that shift to keep accelerating. For anyone building on llama.cpp, the practical question is whether your code assumes a batch is tokens. In v0.6.0, that assumption is now outdated.