The BerriAI/litellm repository now describes itself as “the fastest, litest AI Gateway,” and the phrase that matters is not “fastest.” It is “Rust core with Python SDK.” For a project that spent its first two years as a thin Python wrapper over a hundred provider SDKs, that is a structural change, not a marketing line. LiteLLM is repositioning from library to infrastructure, and the pieces it is adding tell you which parts of the AI stack its maintainers think are permanent.

The headline numbers are modest and specific. LiteLLM claims 8ms P95 latency at 1,000 requests per second, with Netflix listed among its open-source adopters. The gateway routes to more than 100 providers, from OpenAI and Anthropic to Bedrock, Azure, Vertex AI, vLLM, and Nvidia NIM. None of that is new. What is new is the packaging: an independent litellm-core distribution that ships the same import litellm API with no optional extras, no CLI entry points, and no bundled dashboard. The docs are blunt about the constraint: install one SDK distribution per environment, because litellm and litellm-core own overlapping Python files.

The split is the story

That warning is a tell. BerriAI is carving the Python surface into two products: a slim library for people who just want completion() to work, and a full distribution for people running the proxy, the CLI, and the admin dashboard. The build path is still manual. You run python scripts/build_core_distribution.py --out-dir dist/core, then pip install dist/core/litellm_core-*.whl, and the builder needs Git, uv, and the Rust toolchain. The docs say this is “while its release integration is pending.” So the Rust core is real enough to build from a checkout and not yet real enough to install from PyPI as a first-class artifact.

Read that as a statement of intent. Python is where LiteLLM’s users live. Rust is where its latency budget lives. A gateway that sits in front of every model call in a company’s stack cannot afford to be the slowest hop in the chain, and a Python process holding a thousand concurrent connections is exactly that. Moving the hot path to Rust while keeping the Python API is the standard playbook now. It is how Pydantic, Ruff, and a dozen other tools bought themselves a second act.

MCP and A2A are the real land grab

The more consequential additions are protocol-shaped. LiteLLM now bridges MCP servers into any model through an experimental_mcp_client that loads tools in OpenAI format, and it exposes an MCP Gateway so you can call tools through /chat/completions with a server_url and server_label. It also speaks A2A, both as a client through litellm.a2a_protocol.A2AClient and as a gateway that fronts registered agents behind a master key or virtual key.

This is where the project stops being a convenience wrapper and starts being a control plane. If your MCP servers and your A2A agents both terminate at the same gateway that already handles your model routing, then the gateway sees every tool call, every agent hop, and every token spend in one place. That is a genuinely defensible position, and it is why the “guardrails, load balancing, and logging” line in the repo description matters more than the provider count.

The docs also flag a real-world sharp edge that most gateway vendors paper over. For MCP OAuth, an upstream provider may advertise dynamic client registration and then refuse requests with HTTP 401 or 403. LiteLLM’s answer is to let you configure a pre-registered client_id and client_secret, skipping dynamic registration. The docs add a warning worth quoting in spirit: reaching a provider’s authorization page does not establish that login or tool calls will succeed. That is an unusually honest sentence for a vendor README, and it points at the mess underneath the MCP standardization story.

The harness move is the interesting one

The feature I would watch hardest is agent harness support. LiteLLM now exposes a litellm.agent() call that runs Claude Code, Codex, OpenCode, or Deep Agents against any model, in a local sandbox, and returns text, cost, and touched files. Set LITELLM_PROXY_API_BASE and LITELLM_PROXY_API_KEY, and every call the agent makes routes through your gateway, tagged harness,claude_code.

If the gateway sees every tool call, every agent hop, and every token spend, it stops being a wrapper and becomes the control plane.

That is a direct shot at the assumption that coding agents are welded to one vendor’s models. It also creates an obvious tension. Anthropic and OpenAI have little incentive to make their harnesses portable, and the tagging mechanism is the kind of thing that works until a provider changes an endpoint. But the demand is real, because teams want cost visibility and policy enforcement over agent traffic, and today that traffic mostly bypasses whatever governance they built for plain chat completions.

What this means for builders

The provider table is the part that should make you pause. It lists Abliteration, AI/ML API, AI21, Aleph Alpha, Amazon Nova, Anthropic, Anyscale, AssemblyAI, Auto Router, Bedrock, Sagemaker, Azure, Baseten, Bytez, Cerebras, Clarifai, Cloudflare AI Workers, Codestral, Cognition, Cohere, CometAPI, CompactifAI, Dashscope, Databricks, Deepgram, DeepInfra, Deepseek, ElevenLabs, Fireworks AI, FriendliAI, Galadriel, GitHub Copilot, Groq, Heroku, Huggingface, Hyperbolic, IBM Watsonx, Jina AI, Lambda, LM Studio, Maritalk, Meta Llama API, Mistral, Moonshot, Nebius, Novita, Nscale, Nvidia NIM, Ollama, OpenRouter, Perplexity, Predibase, Qwen, Replicate, Sagemaker, Sambanova, Snowflake, Together AI, Triton, Vercel AI Gateway, vLLM, VoyageAI, WandB, xAI, and Xinference. That is not a curated list. It is a land grab, and it is the strongest argument that LiteLLM is trying to be the neutral layer precisely because no single lab can be.

The bet underneath all of this is that model providers will keep churning and the routing layer will not. If that bet is right, LiteLLM’s Rust core, its MCP and A2A gateways, and its harness support are all the same move: own the seam where models, tools, and agents meet, and make that seam fast enough and boring enough that nobody rips it out. The open question is whether BerriAI can finish the litellm-core release integration before a hyperscaler ships the same thing as a managed product. The build script in the README suggests they know the clock is running.