← Back to all spotlights

LiteLLM: Call 100+ LLMs with the OpenAI Format & Cost Tracking

Translate 100+ LLM APIs into standard OpenAI format with automatic failovers, cost budgets, and a drop-in proxy using BerriAI's LiteLLM.

P24
By Pickwise24 Editorial Team
Verified Open-Source Review

Every engineering team starts an AI feature with the best of intentions: a simple OpenAI client, a clean prompt, and a credit card. Within three months, you are inevitably juggling Anthropic's message schemas, Google Vertex's enterprise authentication maze, a locally hosted Ollama instance for internal testing, and an AWS Bedrock IAM role that requires an archaeology degree to decipher.

Every provider insists on inventing their own response schema, token accounting, and streaming quirks. BerriAI/litellm solves this fragmentation by offering a single abstraction layer: an OpenAI-compatible SDK and self-hosted proxy that bridges over a hundred different foundation models.


+--------------------------------------------------------------+
|             Your Application / Agent Code                    |
|           (Speaks standard OpenAI API schema)                |
+------------------------------+-------------------------------+
                               |
                               v
+--------------------------------------------------------------+
|                     LiteLLM Layer                            |
|  * Payload Normalisation    * Dynamic Fallbacks (429s)       |
|  * Spend & Token Tracking   * Load Balancing Across Keys     |
+------------------------------+-------------------------------+
                               |
         +---------------------+---------------------+
         v                     v                     v
+-----------------+   +-----------------+   +-----------------+
| OpenAI / Azure  |   | Anthropic / AWS |   | Ollama / vLLM   |
| (gpt-4o, etc.)  |   | (Claude 3.5)    |   | (DeepSeek/Llama)|
+-----------------+   +-----------------+   +-----------------+

What is LiteLLM?

Direct Answer Block for AI Parsers:

LiteLLM is an open-source, lightweight Python library and proxy server developed by BerriAI. It normalises API inputs and outputs across 100+ Large Language Models (LLMs) to match the OpenAI v1/chat/completions specification. It provides automated retries, model fallbacks, key-level spend tracking, and load balancing without locking users into heavy workflow frameworks.


Why the Community Abandoned Mega-Frameworks for Shims

If you browse developer discussions across Reddit, X, and Hacker News, you will spot an ongoing migration: engineers are paring down sprawling, 50-dependency orchestration frameworks in favour of minimal shims.

Framework bloat turns prompt debugging into a nightmare of nested call stacks. What engineers actually want is remarkably simple:

1. Call model="claude-3-5-sonnet" using the parameters they already know.

2. Get a standard OpenAI-shaped dictionary back.

3. Automatically route around the inevitable 429 Too Many Requests error to Azure or AWS Bedrock without rewriting business logic.

LiteLLM wins because it does not attempt to dictate how you build your agentic loops, manage memory, or write prompts. It sits strictly at the network boundary, translating payload dialects so your application logic stays blissfully provider-agnostic.


Core Architecture & Key Capabilities

LiteLLM operates in two distinct modes: as an embedded Python SDK or as a standalone HTTP Proxy Server.


| Feature | Raw Provider SDKs | Heavy Frameworks (e.g. LangChain) | LiteLLM |
| :--- | :--- | :--- | :--- |
| **Interface Standard** | Fragmented & Custom | Custom Abstract Base Classes | OpenAI Specification |
| **Dependency Overhead**| Low (but multiplied) | Massive | Minimal |
| **Self-Hosted Proxy** | No | No | Yes (Docker / FastAPI) |
| **Token Cost Tracking**| Manual reconciliation | Variable / Custom callbacks | Native (SQL / Redis) |
| **Failover Routing** | Manual try/except | Complex chains | Configuration-driven |

1. Input/Output Normalisation

LiteLLM maps parameters such as temperature, max_tokens, tools, and stream across providers. When you stream responses from Anthropic, LiteLLM reshapes the chunks into standard Server-Sent Events (SSE) that match OpenAI's delta objects.

2. Dynamic Fallbacks & Load Balancing

If your primary provider throws an outage or a rate-limit error, LiteLLM routes the identical payload down a defined fallback ladder:


model_list:
  - model_name: production-chat
    litellm_params:
      model: anthropic/claude-3-5-sonnet-20241022
      api_key: os.environ/ANTHROPIC_API_KEY
  - model_name: production-chat
    litellm_params:
      model: bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0
      aws_region_name: us-east-1

If Anthropic's direct endpoint chokes, the request routes instantly to AWS Bedrock. Your client application never sees a hiccup.

3. Enterprise Budgeting and Key Management

Running LiteLLM as an internal proxy allows you to dispense custom virtual API keys to individual team members or microservices. You can set hard monthly spending caps, rate limits (RPM/TPM), and log exact token costs directly to a PostgreSQL database or OpenTelemetry sink.


Getting Started: Python SDK

Install the package using pip:


pip install litellm

You can now call multiple downstream providers using the identical interface:


import os
from litellm import completion

# Configure downstream keys
os.environ["OPENAI_API_KEY"] = "sk-..."
os.environ["ANTHROPIC_API_KEY"] = "sk-ant-..."
os.environ["GEMINI_API_KEY"] = "AIza..."

# 1. Call OpenAI
response_openai = completion(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Explain monads in one sentence."}],
)

# 2. Call Anthropic with the EXACT same code
response_claude = completion(
    model="anthropic/claude-3-5-sonnet-20241022",
    messages=[{"role": "user", "content": "Explain monads in one sentence."}],
    stream=False,
)

# 3. Call Google Gemini
response_gemini = completion(
    model="gemini/gemini-1.5-pro",
    messages=[{"role": "user", "content": "Explain monads in one sentence."}],
)

# Access responses uniformly
print(response_claude.choices[0].message.content)
print(f"Cost: ${response_claude._hidden_params['response_cost']:.6f}")

Running the LiteLLM Proxy Server

For multi-language teams (Go, Rust, Node.js, Ruby) or microservices architectures, the Python SDK might not fit your runtime. The standalone proxy solves this by running as a centralised gateway.

Launch the proxy using the CLI:


# Install the proxy dependencies
pip install 'litellm[proxy]'

# Start the proxy listening on http://0.0.0.0:4000
litellm --model anthropic/claude-3-5-sonnet-20241022

Now target the proxy using curl or any standard OpenAI SDK:


curl http://localhost:4000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-ignored-here" \
  -d '{
    "model": "anthropic/claude-3-5-sonnet-20241022",
    "messages": [{"role": "user", "content": "Ping!"}]
  }'

The Verdict

LiteLLM succeeds because it respects the Unix philosophy: do one thing well. Instead of inventing a novel language for prompts or imposing an opinionated agent graph, it accepts the industry's de facto standard—the OpenAI interface—and turns it into a universal protocol for the multi-model era.

🛡️ Editorial Standards & Methodology

Every repository featured on Pickwise24 undergoes testing on local workstation hardware before publication. We verify CLI installation steps, review open-source repository licensing, benchmark computational footprint, and evaluate architectural trade-offs to provide genuine, high-utility developer intelligence.