LiteLLM Gateway on a VPS: Unified AI API Proxy for OpenAI, Anthropic, Ollama, and 100+ Models
LiteLLM is an open-source AI gateway that provides a single OpenAI-compatible endpoint that routes to 100+ LLM providers — OpenAI (GPT-4o), Anthropic (Claude), Google (Gemini), Cohere, Azure OpenAI, local Ollama, and many more. Instead of hardcoding API clients for each provider, your applications talk to one LiteLLM endpoint. Features include load balancing across providers, automatic fallbacks, per-team rate limits, cost tracking, and centralized API key management.
Why LiteLLM Gateway
- Unified endpoint: One OpenAI-compatible URL for all LLM providers — no SDK changes when switching models
- Cost control: Track spend per team/project, set budget limits, auto-fallback to cheaper models
- Load balancing: Distribute requests across multiple API keys or model providers
- Observability: Centralized logging of all LLM calls, latency, tokens, cost
- API key management: Issue virtual keys to teams; rotate upstream keys without application changes
Step 1: Docker Compose Setup
<code">mkdir -p /opt/litellm && cd /opt/litellm nano docker-compose.yml
<code">version: '3.8'
services:
litellm:
image: ghcr.io/berriai/litellm:main-latest
container_name: litellm
restart: always
ports:
- "127.0.0.1:4000:4000"
volumes:
- ./config.yaml:/app/config.yaml:ro
command: ["--config", "/app/config.yaml", "--port", "4000", "--num_workers", "2"]
environment:
LITELLM_MASTER_KEY: ${LITELLM_MASTER_KEY}
OPENAI_API_KEY: ${OPENAI_API_KEY}
ANTHROPIC_API_KEY: ${ANTHROPIC_API_KEY}
GEMINI_API_KEY: ${GEMINI_API_KEY}
extra_hosts:
- "host.docker.internal:host-gateway"
<code">cat > .env << 'EOF' LITELLM_MASTER_KEY=sk-master-your-strong-key-here OPENAI_API_KEY=sk-your-openai-key ANTHROPIC_API_KEY=sk-ant-your-anthropic-key GEMINI_API_KEY=your-gemini-key EOF chmod 600 .env
Step 2: LiteLLM Configuration
<code">nano config.yaml
<code">model_list:
# OpenAI models
- model_name: gpt-4o
litellm_params:
model: openai/gpt-4o
api_key: os.environ/OPENAI_API_KEY
- model_name: gpt-4o-mini
litellm_params:
model: openai/gpt-4o-mini
api_key: os.environ/OPENAI_API_KEY
# Anthropic Claude
- model_name: claude-sonnet
litellm_params:
model: anthropic/claude-3-5-sonnet-20241022
api_key: os.environ/ANTHROPIC_API_KEY
- model_name: claude-haiku
litellm_params:
model: anthropic/claude-3-5-haiku-20241022
api_key: os.environ/ANTHROPIC_API_KEY
# Google Gemini
- model_name: gemini-flash
litellm_params:
model: gemini/gemini-2.0-flash
api_key: os.environ/GEMINI_API_KEY
# Local Ollama models (no cost)
- model_name: mistral-local
litellm_params:
model: ollama/mistral
api_base: http://host.docker.internal:11434
- model_name: llama3-local
litellm_params:
model: ollama/llama3.2
api_base: http://host.docker.internal:11434
# Load balanced: route across multiple OpenAI keys for rate limit handling
- model_name: gpt-4o-balanced
litellm_params:
model: openai/gpt-4o
api_key: os.environ/OPENAI_API_KEY
- model_name: gpt-4o-balanced
litellm_params:
model: openai/gpt-4o
api_key: os.environ/OPENAI_API_KEY_2 # Second key for load balancing
router_settings:
routing_strategy: least-busy # Route to least busy model
litellm_settings:
set_verbose: false
drop_params: true # Ignore unsupported parameters silently
fallbacks:
- gpt-4o:
- claude-sonnet # If OpenAI is down, fallback to Anthropic
- claude-sonnet:
- gpt-4o
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
database_url: "sqlite:///./litellm.db" # Store usage data locally
<code">docker compose up -d
Step 3: Nginx Reverse Proxy
<code">sudo nano /etc/nginx/sites-available/litellm
<code">server {
listen 443 ssl http2;
server_name llm-gateway.yourdomain.com;
ssl_certificate /etc/letsencrypt/live/llm-gateway.yourdomain.com/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/llm-gateway.yourdomain.com/privkey.pem;
location / {
proxy_pass http://127.0.0.1:4000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_read_timeout 300s;
proxy_buffering off;
}
}
<code">sudo certbot --nginx -d llm-gateway.yourdomain.com sudo systemctl reload nginx
Step 4: Use with OpenAI SDK
<code">from openai import OpenAI
# All applications use the same LiteLLM endpoint
client = OpenAI(
api_key="sk-master-your-strong-key-here", # LiteLLM master key
base_url="https://llm-gateway.yourdomain.com/",
)
# Switch between any model by changing model name only:
for model in ["gpt-4o", "claude-sonnet", "mistral-local", "gemini-flash"]:
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "What is 2+2?"}],
max_tokens=50,
)
print(f"{model}: {response.choices[0].message.content}")
Step 5: Virtual Keys for Teams
<code"># Create virtual API keys for different teams with budget limits:
curl -X POST https://llm-gateway.yourdomain.com/key/generate \
-H "Authorization: Bearer sk-master-your-strong-key-here" \
-H "Content-Type: application/json" \
-d '{
"key_alias": "engineering-team",
"max_budget": 100.00, # $100 budget limit
"budget_duration": "monthly",
"tpm_limit": 100000, # 100K tokens/minute
"allowed_models": ["gpt-4o-mini", "claude-haiku", "mistral-local"],
"metadata": {"team": "engineering"}
}'
# Returns: {"key": "sk-engineering-xxxxx", "max_budget": 100.0}
# Engineering team uses sk-engineering-xxxxx
# Finance team gets their own key with different model access and limits
Step 6: Cost Tracking Dashboard
<code"># View usage and costs via API: curl https://llm-gateway.yourdomain.com/spend/logs \ -H "Authorization: Bearer sk-master-key" | python3 -m json.tool # Or enable built-in UI: # Add to config.yaml: litellm_settings: ui_username: admin ui_password: StrongUIPassword! # Access dashboard: https://llm-gateway.yourdomain.com/ui # Shows: total spend, spend by model, spend by API key, request volume
Getting Started
LiteLLM uses 200–400 MB RAM. A 2 GB Ubuntu VPS at VPS.DO runs LiteLLM as a lightweight proxy. Since LiteLLM forwards API calls to external providers, the VPS handles routing logic only — compute-intensive inference happens at OpenAI/Anthropic/etc. For organizations with multiple teams using AI, LiteLLM provides centralized cost control that a scattered direct-API-key approach cannot.
Conclusion
LiteLLM gateway centralizes AI API management across 100+ providers behind one OpenAI-compatible endpoint — enabling seamless model switching, automatic fallbacks, per-team budget limits, and centralized cost tracking. Application code never changes when switching from GPT-4o to Claude to local Ollama; only the model name parameter changes. For teams running multiple AI applications or experimenting across providers, LiteLLM is the infrastructure layer that makes model diversity manageable.
web site
October 10, 2026Porn on wap net eroticaPersonal erotic massageKim kardashian 2nd sex videoProstatis from excessive masturbationStacy dash assFree sexy flash
videos sexyTeachers real hidden cameras nudeJenna hustler dvdDragonball sailormoon hentai videos xnxx Lick
you up and down lyricGloryhole addresses glory holeKaren loves kate xxxEbony nude
feetLed watches vintageTeen attitudes stdsNaked body huffington postFreebig butt pornBest female anal lubeCell
phone tutorials for adults https://amp.xnxxtubexxx.cc/ Friends pornoBreast areoleThailand’s sex industryDried
pussy willow cheapFuck buddy in bowie txOblivion boundless pleasureKick ass endangered speciesBlack legs fetish xnxx Same sex
commitment ceremony sampleTeen young amateurLori douglas nude dark cavern picsGeisha lounge hollywoodAmateur gratuit analMore teen girlMistress juliya sexy picsBreast enlargement inexpensiveJapan av
matureFacial pierceingTeen allure galleryBecome an ebony porn starErin andrews nude picsBig titty amateur
sex videosHelp for troubled teens in nevadaBasket ball dick