Magnitude, a Y Combinator S25 company founded by Anders and Tom, posted its Launch HN this week with a claim worth reading closely: an inference engine that compiles and tunes its GPU kernels on your actual device before a model runs, and beats llama.cpp by 92% on decode for a Qwen 3.6 35B A3B model at 4-bit quantization and 64k context on an M4 Pro with 48 GB of unified memory. On a DGX Spark over CUDA, the same setup shows 19% faster decode and 23% faster prefill. The engine is Apache 2.0, written in Rust, and ships as a desktop app that connects to Pi, OpenCode, Hermes, Codex, Claude Code, Cline, and anything else that speaks the OpenAI-compatible API.

The headline number is the Metal result, and it is the one that should make people pay attention, because the delta between the two benchmarks tells you where Magnitude’s bet actually lives.

Why the Metal number is so much bigger

On CUDA, Magnitude gains 19% decode over llama.cpp. On Apple Silicon, it gains 92%. That asymmetry is not an accident of benchmarking. llama.cpp’s CUDA backend has had years of NVIDIA-focused tuning from contributors with access to the same hardware classes Magnitude is targeting. Its Metal backend has not received the same intensity of per-chip optimization, partly because Apple’s GPU toolchain is less mature and partly because the community of people who write hand-tuned Metal kernels for LLM decode is small.

Magnitude’s on-device autotuner sidesteps that problem. Rather than shipping kernels precompiled for broad hardware classes, the engine writes kernels with flexible parameters and then tunes them against the actual chip it finds. On a Mac, that means the autotuner is doing work that llama.cpp’s Metal path mostly does not: fitting tile sizes, memory layouts, and fusion decisions to the specific M4 Pro rather than to a generic Apple Silicon target. The 30 tok/s to 57 tok/s jump is what that gap looks like when you close it.

This is a real technical story, but the framing around it is the more interesting one.

The agent-first framing is the actual claim

Magnitude’s pitch is not “faster llama.cpp.” It is that every existing engine was designed for the wrong workload. vLLM and SGLang optimize for batched datacenter inference, which trades single-session latency for throughput. llama.cpp and Ollama optimize for broad hardware compatibility. oMLX and ds4 optimize for specific hardware or models at the cost of engine completeness. None of them, the founders argue, were built for the shape of a local agent session: long-lived, often several running concurrently, with the user still wanting to do other work on the same machine.

That is a defensible critique. Agent sessions do behave differently from chat. A coding agent holds a large prefix cache across many turns, spawns sub-sessions, and sits idle between tool calls. A chat session is a single short stream. The design choices Magnitude makes, dynamic memory allocation that grows with session count and frees when agents stop, and hybrid paged attention that shares prefix caches across concurrent sessions while preserving memory adjacency for single-session speed, are direct responses to that shape. The 27–28% per-agent memory reduction is the number that matters most for anyone trying to run three agents on a 48 GB machine.

The honest caveat: these benchmarks are the founders’ own, on a single model (Qwen 3.6 35B A3B), with no speculative decoding, against one competitor. Independent reproduction is the next thing to watch. “Up to 2x” is doing load-bearing work in the marketing, and the CUDA result, at 19%, is the more representative number for anyone on NVIDIA hardware.

What the roadmap says about where this goes

The three items Magnitude lists as next are worth reading as a statement of intent. Expert streaming, storing MoE experts on RAM or disk and loading them just-in-time, is the feature that would let a 48 GB Mac run models that do not fit in memory. That is the same problem Apple’s own unified-memory architecture was supposed to solve and mostly has not, because the software stack has not kept up. A kernel compiler that fuses operations automatically is the natural extension of the current autotuner, which only tunes a few parameters. Multi-device utilization, the last item, is the one that would put Magnitude in conversation with the distributed inference crowd.

Magnitude’s real bet is that local agents are a distinct workload, and that the engine that wins it will not be the one built for datacenter batches or for maximum hardware compatibility.

What this means for builders

If you are running local agents today, Magnitude is worth a download, with the caveat that the benchmark numbers come from the vendor and the CUDA gains are modest. The more consequential question is what happens to the local inference stack over the next year. llama.cpp has been the default for so long that its limitations have become invisible to the people building on it. A competitor that tunes kernels on-device, ships an actual desktop app with agent connections built in, and treats memory pressure as a first-class problem is a useful forcing function regardless of whether Magnitude itself wins.

The thing to watch is not the 92% number. It is whether the autotuner holds up across the long tail of hardware Magnitude claims to support, AMD GPUs, older Apple Silicon, CPU-only machines, and whether the kernel compiler on the roadmap actually ships. On-device compilation is a strong idea with a real cost: it means every user pays a tuning tax before their first token, and it means Magnitude owns a compiler problem that llama.cpp’s precompiled-kernel approach avoids. If the tuning is fast enough to be invisible, the approach generalizes. If it is not, the 92% becomes a benchmark artifact rather than a product.

The Qwen 3.6 35B A3B benchmark is a single data point from the people selling the engine. The next data point will come from someone with no stake in the result.