What it is
LiteLLM is an open source AI Gateway that gives a single, unified interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure and more) using the OpenAI format. It can be used as a Python SDK for direct library integration, or deployed as a Proxy Server (AI Gateway) that serves as a centralized service for a team or organization. It also covers A2A agents, MCP tools and agent harnesses.
Who it's for
- Developers who want one OpenAI-style interface across many LLM providers
- Teams or organizations that want a self-hosted central gateway with virtual keys and spend tracking
- Developers connecting MCP servers or A2A agents to LLMs
Requirements
Requirements
- Python environment with uv (used in the documented install commands)
- API keys for the LLM providers you call (e.g. OPENAI_API_KEY, ANTHROPIC_API_KEY)
- Building litellm-core from the checkout requires Git, uv and the Rust build toolchain
- A2A gateway calls require a2a-sdk>=1.1.0
Setup
Install the Python SDK
Install litellm with uv.
shelluv add litellmInstall and start the AI Gateway (Proxy Server)
Install the proxy extra as a uv tool and start it with a model.
shelluv tool install 'litellm[proxy]' litellm --model gpt-4oOptional: build litellm-core from the checkout
An independent litellm-core distribution provides the SDK with the same
import litellmAPI, with no extras, CLI entry points or dashboard. Install only one SDK distribution per environment, in a fresh environment without litellm.shellpython scripts/build_core_distribution.py --out-dir dist/core python -m pip install dist/core/litellm_core-*.whl
Examples
Call multiple providers with the SDK
pythonfrom litellm import completion
import os
os.environ["OPENAI_API_KEY"] = "your-openai-key"
os.environ["ANTHROPIC_API_KEY"] = "your-anthropic-key"
# OpenAI
response = completion(model="openai/gpt-4o", messages=[{"role": "user", "content": "Hello!"}])
# Anthropic
response = completion(model="anthropic/claude-sonnet-4-20250514", messages=[{"role": "user", "content": "Hello!"}])What it does: Switch providers by changing the model string while keeping the same completion call.
Use the gateway via the OpenAI client
pythonimport openai
client = openai.OpenAI(api_key="anything", base_url="http://0.0.0.0:4000")
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Hello!"}]
)What it does: Points the standard OpenAI client at the running LiteLLM proxy.
Call MCP tools through the gateway
bashcurl -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
-H 'Authorization: Bearer <your-master-key>' \
-H 'Content-Type: application/json' \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Summarize the latest open PR"}],
"tools": [{
"type": "mcp",
"server_url": "litellm_proxy/mcp/github",
"server_label": "github_mcp",
"require_approval": "never"
}]
}'What it does: After adding an MCP server to the gateway, its tools are invoked through /chat/completions.
Run an agent harness on any model
pythonimport litellm
from litellm import Harness, sandbox
result = litellm.agent(
Harness.CLAUDE_CODE, # or Harness.CODEX, Harness.OPENCODE, Harness.DEEPAGENTS
"Find why tests/test_router.py is flaky and fix it.",
sandbox=sandbox.local("./repo"),
model="litellm_proxy/claude-sonnet-4-5", # a model group on your AI Gateway
)
print(result.text, result.cost, [f.path for f in result.files])What it does: Runs a coding agent through the AI Gateway when LITELLM_PROXY_API_BASE and LITELLM_PROXY_API_KEY are set.
Pros & cons
Pros
- Pro:Unified OpenAI-format API across 100+ providers, so switching providers needs no code rewrite
- Pro:Gateway ships with virtual keys, spend tracking, guardrails, load balancing and an admin dashboard
- Pro:Documented 8ms P95 latency at 1k RPS
- Pro:Covers endpoints beyond chat, including responses, embeddings, images, audio, batches, rerank, A2A and MCP
Cons
- Con:Provider support differs by endpoint, so not every provider supports every endpoint
- Con:Install variants (litellm vs litellm-core) overlap and only one may be installed per environment
- Con:Core distribution build from checkout requires Git, uv and a Rust toolchain, and its release integration is noted as pending
Images
