Repo Trending

BerriAI/litellm

Open source AI Gateway and Python SDK to call 100+ LLM providers in OpenAI format, with virtual keys, spend tracking, guardrails, load balancing and an admin dashboard.

  • 60.6k GitHub stars
  • Python
uv add litellm
BerriAI/litellm preview image

What it is

LiteLLM is an open source AI Gateway that gives a single, unified interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure and more) using the OpenAI format. It can be used as a Python SDK for direct library integration, or deployed as a Proxy Server (AI Gateway) that serves as a centralized service for a team or organization. It also covers A2A agents, MCP tools and agent harnesses.

Who it's for

  • Developers who want one OpenAI-style interface across many LLM providers
  • Teams or organizations that want a self-hosted central gateway with virtual keys and spend tracking
  • Developers connecting MCP servers or A2A agents to LLMs

Requirements

Requirements

  • Python environment with uv (used in the documented install commands)
  • API keys for the LLM providers you call (e.g. OPENAI_API_KEY, ANTHROPIC_API_KEY)
  • Building litellm-core from the checkout requires Git, uv and the Rust build toolchain
  • A2A gateway calls require a2a-sdk>=1.1.0

Setup

  1. Install the Python SDK

    Install litellm with uv.

    shell
    uv add litellm
  2. Install and start the AI Gateway (Proxy Server)

    Install the proxy extra as a uv tool and start it with a model.

    shell
    uv tool install 'litellm[proxy]'
    litellm --model gpt-4o
  3. Optional: build litellm-core from the checkout

    An independent litellm-core distribution provides the SDK with the same import litellm API, with no extras, CLI entry points or dashboard. Install only one SDK distribution per environment, in a fresh environment without litellm.

    shell
    python scripts/build_core_distribution.py --out-dir dist/core
    python -m pip install dist/core/litellm_core-*.whl

Examples

Call multiple providers with the SDK

python
python
from litellm import completion
import os

os.environ["OPENAI_API_KEY"] = "your-openai-key"
os.environ["ANTHROPIC_API_KEY"] = "your-anthropic-key"

# OpenAI
response = completion(model="openai/gpt-4o", messages=[{"role": "user", "content": "Hello!"}])

# Anthropic
response = completion(model="anthropic/claude-sonnet-4-20250514", messages=[{"role": "user", "content": "Hello!"}])

What it does: Switch providers by changing the model string while keeping the same completion call.

Use the gateway via the OpenAI client

python
python
import openai

client = openai.OpenAI(api_key="anything", base_url="http://0.0.0.0:4000")
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello!"}]
)

What it does: Points the standard OpenAI client at the running LiteLLM proxy.

Call MCP tools through the gateway

bash
bash
curl -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
  -H 'Authorization: Bearer <your-master-key>' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Summarize the latest open PR"}],
    "tools": [{
      "type": "mcp",
      "server_url": "litellm_proxy/mcp/github",
      "server_label": "github_mcp",
      "require_approval": "never"
    }]
  }'

What it does: After adding an MCP server to the gateway, its tools are invoked through /chat/completions.

Run an agent harness on any model

python
python
import litellm
from litellm import Harness, sandbox

result = litellm.agent(
    Harness.CLAUDE_CODE,  # or Harness.CODEX, Harness.OPENCODE, Harness.DEEPAGENTS
    "Find why tests/test_router.py is flaky and fix it.",
    sandbox=sandbox.local("./repo"),
    model="litellm_proxy/claude-sonnet-4-5",  # a model group on your AI Gateway
)

print(result.text, result.cost, [f.path for f in result.files])

What it does: Runs a coding agent through the AI Gateway when LITELLM_PROXY_API_BASE and LITELLM_PROXY_API_KEY are set.

Pros & cons

Pros

  • Pro:Unified OpenAI-format API across 100+ providers, so switching providers needs no code rewrite
  • Pro:Gateway ships with virtual keys, spend tracking, guardrails, load balancing and an admin dashboard
  • Pro:Documented 8ms P95 latency at 1k RPS
  • Pro:Covers endpoints beyond chat, including responses, embeddings, images, audio, batches, rerank, A2A and MCP

Cons

  • Con:Provider support differs by endpoint, so not every provider supports every endpoint
  • Con:Install variants (litellm vs litellm-core) overlap and only one may be installed per environment
  • Con:Core distribution build from checkout requires Git, uv and a Rust toolchain, and its release integration is noted as pending

Images