Gateways

LLM gateways and proxies for routing, caching, rate limits and cost control across providers.

OmniRoute leads with 86, ahead of LiteLLM (84) and freellmapi (76). 14 projects ranked by score.

The ranking

Gateways: full ranking
RankProjectAdoptionFreshnessMaintenanceEasy to runAgent-readyScore
1OmniRouteOpenAI-compatible gateway that routes requests across hundreds of AI providers75.1k stars, MIT, last commit Oct 202694100100677086 out of 100
2LiteLLMOpenAI-format gateway and Python SDK for calling 100+ LLM providers61k stars, custom license, last commit Oct 20268910083833084 out of 100
3freellmapiOpenAI-compatible router that fails over across free LLM provider tiers33.1k stars, MIT, last commit Oct 2026761009867076 out of 100
#4ContextForge MCP GatewayRegistry and proxy federating MCP, A2A, REST and gRPC behind one endpoint4.6k stars, Apache-2.0, last commit Oct 20262410090677068 out of 100
#5HigressEnvoy-based API gateway with LLM proxy plugins and MCP server hosting Live demo ↗ (opens in a new tab)9.5k stars, Apache-2.0, last commit Oct 20265210078334561 out of 100
#6BifrostGo AI gateway with web UI, fallbacks, budgets and semantic caching8.7k stars, Apache-2.0, last commit Oct 20264310082334559 out of 100
#7PlanoEnvoy-based data plane that routes, traces and guards agent traffic7.1k stars, Apache-2.0, last commit Oct 20263610076335558 out of 100
#8GoModelGo AI gateway with OpenAI and Anthropic APIs, caching and budgets Live demo ↗ (opens in a new tab)1.2k stars, MIT, last commit Oct 2026110087507057 out of 100
#9agentgatewayOne proxy for LLM, MCP and A2A traffic with auth and RBAC5.3k stars, Apache-2.0, last commit Oct 2026301008833054 out of 100
#10optillmOpenAI-compatible proxy applying inference-time reasoning techniques Live demo ↗ (opens in a new tab)4.3k stars, Apache-2.0, last commit Oct 20261510077334052 out of 100
#11Portkey GatewayNode.js LLM gateway with fallbacks, load balancing and guardrails13.2k stars, MIT, last commit May 202660803334046 out of 100
#12MetaMCPAggregates MCP servers into namespaced endpoints with auth and middleware2.7k stars, MIT, last commit Jun 2026785550037 out of 100
#13CoAIMulti-user chat site plus OpenAI-compatible proxy with billing9.3k stars, Apache-2.0, last commit Mar 20264854033034 out of 100
#14mcpoExposes any MCP server as an OpenAPI HTTP endpoint4.4k stars, MIT, last commit Feb 20262062033029 out of 100

Momentum, Verified build, Docs and Privacy are not measured yet; their weight goes to the signals shown. A dash means the signal is not scored for that kind of project. Hover a number for its rating in words.

Reviews

186 out of 100

OmniRoute

OpenAI-compatible gateway that routes requests across hundreds of AI providers

75.1k stars, MIT, last commit Oct 2026

OmniRoute exposes one OpenAI-compatible endpoint at localhost:20128/v1 and routes requests to a catalog of 357+ providers, including many free tiers, with automatic fallback between them. It also accepts Claude, Gemini and Responses API formats, supports MCP and A2A, and has a dashboard for keys, quotas and free-tier usage. Providers are connected with your own accounts or API keys.

Strengths

  • Single endpoint with automatic fallback across many providers and model IDs
  • Dashboard page tracks free-tier pools and remaining quota
  • Install via npm, Docker or Electron; MIT license
  • Documents the free-tier token math and flags providers with risky terms

Weaknesses

  • Provider count in the README varies (290, 357, 370) across sections
  • The 'auto' model needs at least one eligible connected provider to route
  • Providers marked tos:avoid, such as Kiro, are excluded from auto routing by default
  • Free-tier token estimate depends on third-party limits that change
  • no GPU
  • Docker + Compose
  • Compose runs Redis, Qdrant
  • Models: OpenAI API, Claude API, Gemini API, Responses API
  • port 20128
284 out of 100

LiteLLM

OpenAI-format gateway and Python SDK for calling 100+ LLM providers

61k stars, custom license, last commit Oct 2026

LiteLLM translates calls to 100+ providers (OpenAI, Anthropic, Gemini, Bedrock, Azure and others) into the OpenAI format, either as a Python SDK or as a proxy server. The proxy adds virtual keys, spend tracking, guardrails, load balancing and an admin dashboard, and it also exposes A2A agent and MCP server gateways. The README reports 8ms P95 latency at 1k RPS.

Strengths

  • One OpenAI-style API across 100+ providers and many endpoint types
  • Proxy includes virtual keys, spend tracking, load balancing and admin dashboard
  • Also gateways A2A agents and MCP servers
  • Use as a Python library or as a standalone proxy

Weaknesses

  • License reported as NOASSERTION; an enterprise tier exists, feature split unclear from README
  • Proxy listens on port 4000 and its database requirements are not stated in the README excerpt
  • Provider coverage varies by endpoint; many providers support only chat-style endpoints
  • Python-based, so latency figures depend on the benchmark setup
  • no GPU
  • Docker + Compose
  • Compose runs PostgreSQL
  • Models: OpenAI, Anthropic, Gemini, AWS Bedrock, Azure
  • port 4000
376 out of 100

freellmapi

OpenAI-compatible router that fails over across free LLM provider tiers

33.1k stars, MIT, last commit Oct 2026

FreeLLMAPI exposes one /v1 endpoint (chat, responses, completions, embeddings, images, video, audio) and routes requests across free tiers from 34 providers, plus custom OpenAI-compatible endpoints. Provider keys are AES-256-GCM encrypted in SQLite, per-key RPM/RPD/TPM/TPD counters keep requests under quotas, and a 429 or 5xx triggers fallover to the next model. It also serves Anthropic Messages, Gemini and opt-in Ollama surfaces, and ships a React dashboard and desktop apps.

Strengths

  • Also speaks Anthropic, Gemini and Ollama formats, so Claude Code and Codex CLI connect
  • Per-key rate counters and automatic fallover on 429/5xx across providers
  • Keys AES-256-GCM encrypted in SQLite; apps only see one unified token
  • Runs on Node 20+ at about 40 MB idle RSS, or via Docker

Weaknesses

  • Free installs get new models 30 days after premium; same-day catalog costs $19/yr
  • Single-user by design; no multi-user setup described
  • Depends on free tiers that providers can change or retire without notice
  • Catalog sync pulls a signed feed from freellmapi.co
  • no GPU
  • Docker + Compose
  • Needs SQLite, Node 20+, provider API keys
  • Models: OpenAI-compatible API, Anthropic Messages API, Gemini API, Ollama API
  • port 3001
#468 out of 100

ContextForge MCP Gateway

Registry and proxy federating MCP, A2A, REST and gRPC behind one endpoint

4.6k stars, Apache-2.0, last commit Oct 2026

ContextForge is IBM's Python registry and proxy that federates MCP servers, A2A agents and REST or gRPC APIs into one MCP-compliant endpoint with auth, rate limiting, retries, an Admin UI and OpenTelemetry tracing. It installs from PyPI (mcpgateway on port 4444), as a GHCR container, via Docker Compose with PostgreSQL, Redis and nginx, or with a Helm chart, and virtualizes legacy REST services as MCP tools.

Strengths

  • Federates MCP, A2A, REST and gRPC (via reflection) behind one MCP endpoint
  • Transports: HTTP, JSON-RPC, WebSocket, SSE, Streamable HTTP, stdio
  • Admin UI with live log viewer; OpenTelemetry to Phoenix, Jaeger, Zipkin
  • Helm chart with HPA, Redis clustering and Grafana dashboards

Weaknesses

  • arm64 containers unsupported in production; Apple Silicon needs Rosetta or PyPI
  • Local Docker builds fail without the CI-only wheel closure; pull the GHCR image
  • Will not start without generated JWT_SECRET_KEY and AUTH_ENCRYPTION_SECRET
  • Large surface: 55+ tables, 40+ plugins, nginx and pgAdmin in the Compose stack
  • no GPU
  • Docker + Compose
  • Needs PostgreSQL (production; SQLite for dev), Redis (caching and federation)
  • Models: A2A agents: OpenAI, Anthropic, custom
  • port 4444
#561 out of 100

Higress

Envoy-based API gateway with LLM proxy plugins and MCP server hosting

Live demo ↗ (opens in a new tab)9.5k stars, Apache-2.0, last commit Oct 2026

Higress is a CNCF sandbox API gateway on Istio and Envoy, extended with Wasm plugins in Go, Rust or JS. Its AI plugins proxy mainstream LLM providers with token rate limiting, load balancing, caching and observability, and host remote MCP servers, including ones generated from OpenAPI specs. A Docker all-in-one image exposes the console on 8001 and the gateway on 8080/8443; Helm covers Kubernetes.

Strengths

  • Envoy-based with millisecond config reloads and no connection drops
  • Hosts MCP servers with auth, rate limits and audit; OpenAPI-to-MCP converter
  • Token rate limiting, multi-model load balancing and caching for LLM routes
  • Also a Kubernetes ingress controller and Gateway API implementation

Weaknesses

  • Images only on Alibaba Cloud registries; pulls can time out outside the mirrors
  • Istio and Envoy underneath; heavier than single-binary LLM proxies
  • AI features are Wasm plugins on a general API gateway
  • Docs split across higress.ai and higress.cn
  • no GPU
  • Docker
  • Models: mainstream LLM providers, domestic and international, via the ai-proxy plugin
  • port 8001
#659 out of 100

Bifrost

Go AI gateway with web UI, fallbacks, budgets and semantic caching

8.7k stars, Apache-2.0, last commit Oct 2026

Bifrost is a Go AI gateway that fronts 23+ providers (OpenAI, Anthropic, Bedrock, Vertex and more) with one OpenAI-compatible API and drop-in paths for the OpenAI, Anthropic and GenAI SDKs. It starts with npx or Docker on port 8080 with a web UI, and adds fallbacks, load balancing, semantic caching, MCP tool access, virtual keys, budgets and Prometheus metrics. Clustering, guardrails and the MCP gateway are enterprise features.

Strengths

  • Single Go binary via npx or Docker with zero-config web UI on 8080
  • Drop-in base URLs for OpenAI, Anthropic and Google GenAI SDKs
  • Virtual keys, team budgets, OIDC provisioning and Prometheus metrics
  • 11 microsecond added latency at 5k RPS in its own benchmark

Weaknesses

  • Guardrails, clustering, adaptive load balancing and MCP gateway are enterprise-only
  • Benchmarks are self-reported on t3 instances
  • 23+ providers, fewer than LiteLLM or Portkey
  • Semantic caching needs a vector store backend
  • no GPU
  • Models: OpenAI, Anthropic, AWS Bedrock, Google Vertex, Azure, Cerebras, Cohere, Mistral, Ollama, Groq and more
  • port 8080
#758 out of 100

Plano

Envoy-based data plane that routes, traces and guards agent traffic

7.1k stars, Apache-2.0, last commit Oct 2026

Plano is an Envoy-based proxy for agentic apps: a YAML file declares agents (HTTP servers with an OpenAI chat endpoint), model providers and listeners, and Plano routes each turn to the right agent with its 4B orchestrator model (hosted free or run locally). It also routes LLM calls by model name, alias or preference, captures OpenTelemetry traces with no instrumentation and applies guardrails via filter chains.

Strengths

  • Declarative multi-agent orchestration; agents are plain OpenAI-compatible HTTP servers
  • Zero-code OpenTelemetry traces and agentic signals for every request
  • Filter chains add moderation, jailbreak checks and memory out of process
  • Model routing by name, alias or preference across providers

Weaknesses

  • Agent routing depends on Plano's own orchestrator model; hosted by default
  • Install prerequisites live in external docs; README shows only YAML and curl
  • Envoy underneath; heavier than a single-binary proxy
  • No port or resource guidance beyond example listeners
  • no GPU
  • Docker
  • Needs Plano-Orchestrator routing model (hosted or local)
  • Models: OpenAI, Anthropic and other providers configured as model_providers
#857 out of 100

GoModel

Go AI gateway with OpenAI and Anthropic APIs, caching and budgets

Live demo ↗ (opens in a new tab)1.2k stars, MIT, last commit Oct 2026

GoModel is a Go AI gateway (install script or container on port 8080) exposing OpenAI-compatible /v1 and Anthropic /v1/messages endpoints in front of OpenAI, Anthropic, Gemini, Bedrock, Azure, Ollama, vLLM, SGLang and more. It adds exact and semantic caching, cost tracking, budgets, rate limits, failover, an MCP gateway, guardrails and a dashboard with playground. Compose adds Redis, PostgreSQL, MongoDB and Prometheus.

Strengths

  • Single Go binary or container; official OpenAI and Anthropic SDKs work unchanged
  • Budgets, rate limits, cost tracking and a usage API per user, team or key
  • Exact and semantic caching, failover with circuit breakers, provider key rotation
  • Dashboard with playground, live request stream, Prometheus and OpenTelemetry

Weaknesses

  • Prompt compression, intelligent routing and OIDC SSO are in the paid Pro build
  • Pre-1.0; roadmap points to an upcoming 0.2.0 release
  • Full Compose stack pulls in Redis, PostgreSQL, MongoDB and Prometheus
  • Benchmarks against LiteLLM and Portkey are self-run
  • no GPU
  • Docker + Compose
  • Needs Redis, PostgreSQL, MongoDB (Compose infrastructure)
  • Models: OpenAI, Anthropic, xAI, Gemini, Vertex AI, Cohere, DeepSeek, Groq, Fireworks, OpenRouter, Azure OpenAI, Bedrock, Ollama, SGLang, vLLM, llm-d, ElevenLabs and any OpenAI-compatible provider
  • port 8080
#954 out of 100

agentgateway

One proxy for LLM, MCP and A2A traffic with auth and RBAC

5.3k stars, Apache-2.0, last commit Oct 2026

Agentgateway is a Linux Foundation proxy for agent traffic: an LLM gateway (OpenAI-compatible API, budgets, failover), an MCP gateway federating tools over stdio, HTTP, SSE and Streamable HTTP, and an A2A gateway. It adds JWT, API key and OAuth auth, CEL RBAC, rate limits, guardrails and OpenTelemetry, and runs standalone from YAML or as a Kubernetes controller with Gateway API.

Strengths

  • One proxy for LLM, MCP and A2A traffic with an OpenAI-compatible API
  • MCP federation over stdio, HTTP, SSE and Streamable HTTP plus OpenAPI tools
  • JWT, API key and OAuth auth with CEL-based RBAC and rate limits
  • Standalone YAML mode or Kubernetes controller with Gateway API

Weaknesses

  • README has no install command, ports or resource needs; quickstart is external
  • Inference routing assumes Kubernetes Inference Gateway extensions
  • Marked in active development; roadmap is the issue tracker
  • Guardrail backends beyond regex are cloud services (OpenAI, Bedrock, Model Armor)
  • no GPU
  • Docker
  • Models: OpenAI, Anthropic, Gemini, Bedrock and other providers; self-hosted models via inference routing
#1052 out of 100

optillm

OpenAI-compatible proxy applying inference-time reasoning techniques

Live demo ↗ (opens in a new tab)4.3k stars, Apache-2.0, last commit Oct 2026

OptiLLM is an OpenAI-compatible proxy (pip or Docker, port 8000) that applies inference-time techniques such as mixture of agents, N-sample selection, self-consistency, MCTS, CePO and MARS to any upstream model, selected by a model-name prefix like moa-gpt-4o-mini. Plugins add an MCP client, memory, PII anonymization, code execution, JSON outputs and provider failover; upstreams are OpenAI, Cerebras, Azure or anything LiteLLM supports.

Strengths

  • 20+ techniques selected by model-name prefix, e.g. moa-gpt-4o-mini
  • Per-technique benchmarks listed (MARS +30 points on AIME 2025 with Gemini 2.5 Flash Lite)
  • Plugins for MCP client, memory, PII anonymization, code execution, JSON output
  • Works with any OpenAI-compatible endpoint; LiteLLM covers other providers

Weaknesses

  • Techniques multiply upstream calls (bon, MoA, MCTS), raising cost and latency
  • Runs on Flask's development server by default
  • Decoding techniques (cot_decoding, AutoThink) need the local inference path
  • Web search plugin drives Chrome through Selenium
  • GPU optional
  • Docker + Compose
  • Models: OpenAI, Cerebras, Azure OpenAI, any OpenAI-compatible endpoint, LiteLLM providers, local models via the built-in inference server
  • port 8000
#1146 out of 100

Portkey Gateway

Node.js LLM gateway with fallbacks, load balancing and guardrails

13.2k stars, MIT, last commit May 2026

Portkey Gateway is a Node.js proxy that routes requests to 250+ LLM providers through an OpenAI-style API on port 8787, runnable with npx, Docker or Cloudflare Workers. Config objects add retries, fallbacks, load balancing, conditional routing, timeouts and 40+ guardrails; a console at /public shows local logs. Semantic caching, prompt management and RBAC are hosted or enterprise features.

Strengths

  • Runs with npx in Node.js; 122 KB footprint, sub-millisecond overhead claimed
  • Fallbacks, retries, load balancing, conditional routing and timeouts via config
  • 40+ built-in guardrails plus bring-your-own
  • Works with OpenAI, LangChain, LlamaIndex, CrewAI and Autogen SDKs

Weaknesses

  • Semantic caching, prompt management and provider optimization are hosted or enterprise only
  • Last commit 2026-05-25; Gateway 2.0 enterprise merge still pre-release
  • Docs links are portkey.wiki short links
  • RBAC, PII redaction and compliance features are enterprise
  • no GPU
  • Docker + Compose
  • Models: OpenAI, Azure OpenAI, Anthropic, Gemini, Cohere, Mistral, Together, Perplexity, Ollama, Bedrock, Groq and 45+ providers
  • port 8787
#1237 out of 100

MetaMCP

Aggregates MCP servers into namespaced endpoints with auth and middleware

2.7k stars, MIT, last commit Jun 2026

MetaMCP groups MCP servers into namespaces and publishes each as one MCP endpoint over SSE, Streamable HTTP or OpenAPI, with API-key or OAuth auth, per-tool toggles, name overrides and middleware. It runs with Docker Compose beside PostgreSQL on port 12008, adds OIDC SSO, multi-tenancy and rate limits, and includes an inspector with saved configs. The author reports maintenance delays.

Strengths

  • Namespaces group servers, toggle tools and override names and annotations
  • Endpoints over SSE, Streamable HTTP and OpenAPI with API key or MCP OAuth
  • OIDC SSO, multi-tenancy and registration controls for organizations
  • Built-in inspector with saved server configs

Weaknesses

  • Author notes maintenance delays; a community fork exists
  • Endpoints are remote-only; stdio clients like Claude Desktop need mcp-proxy
  • Rate-limit counters are in-memory per instance, not cluster-wide
  • MCP servers needing more than uvx or npx require a custom Dockerfile
  • no GPU
  • Docker + Compose
  • Needs PostgreSQL
  • port 12008
#1334 out of 100

CoAI

Multi-user chat site plus OpenAI-compatible proxy with billing

9.3k stars, Apache-2.0, last commit Mar 2026

CoAI pairs a multi-user chat frontend with an OpenAI-compatible API proxy and billing for operators of commercial AI sites. A Go backend on MySQL and Redis routes across channels with priority, weight, retries and model redirection for OpenAI, Anthropic, Gemini, Midjourney, Ollama and more; the React UI adds file parsing, SearXNG search and image generation. Docker Compose serves it on port 8000.

Strengths

  • Chat UI and OpenAI-compatible proxy in one deployment
  • Channel priorities, weights, retries and model redirection for routing
  • Subscription and per-token billing with gift and redemption codes
  • Midjourney, DALL-E and Stable Diffusion image generation in chat

Weaknesses

  • Default admin login root / chatnio123456 must be changed after deploy
  • RAG, TTS/STT, OAuth login and rate limiting are in the paid Pro version
  • Needs MySQL and Redis
  • Last commit 2026-03-12
  • no GPU
  • Docker + Compose
  • Needs MySQL, Redis, SearXNG (optional web search), CoAI blob-service (optional file parsing)
  • Models: OpenAI, Azure OpenAI, Anthropic, Gemini, Midjourney, SparkDesk, Zhipu, Qwen, Hunyuan, Baichuan, Moonshot, DeepSeek, Skylark, Groq, OpenRouter, 360, LocalAI, Ollama
  • port 8000
#1429 out of 100

mcpo

Exposes any MCP server as an OpenAPI HTTP endpoint

4.4k stars, MIT, last commit Feb 2026

mcpo wraps an MCP server command, SSE or Streamable HTTP endpoint and exposes its tools as an OpenAPI REST server on port 8000 with generated docs, so HTTP clients such as Open WebUI can call MCP tools. A Claude Desktop-style config serves several servers under separate routes with hot reload; OAuth 2.1 dynamic client registration handles protected upstreams. Runs via uvx, pip or Docker.

Strengths

  • One command turns any MCP server into an OpenAPI server with /docs
  • stdio, SSE and Streamable HTTP upstreams; OAuth 2.1 with dynamic registration
  • Config file in Claude Desktop format with hot reload
  • Docker image and --root-path for reverse proxies

Weaknesses

  • Last commit 2026-02-27
  • Single shared API key; no users or RBAC
  • Converts to OpenAPI only; does not aggregate servers into one MCP endpoint
  • Python 3.8+ process per deployment; no clustering
  • no GPU
  • Docker
  • Needs MCP servers to proxy
  • port 8000

Written from each project's README and checked facts. Spot something wrong? Report it on GitHub (opens in a new tab).