Open-source alternatives to OpenRouter

The highest-scoring open-source projects in Gateways. They are picks, not exact replacements, so read the weaknesses before you switch.

Open-source alternatives to OpenRouter
RankProjectScore
1OmniRouteOpenAI-compatible gateway that routes requests across hundreds of AI providersGateways, 74.9k stars, MIT86 out of 100
2LiteLLMOpenAI-format gateway and Python SDK for calling 100+ LLM providersGateways, 60.9k stars, custom license84 out of 100
3freellmapiOpenAI-compatible router that fails over across free LLM provider tiersGateways, 32.9k stars, MIT76 out of 100
#4ContextForge MCP GatewayRegistry and proxy federating MCP, A2A, REST and gRPC behind one endpointGateways, 4.6k stars, Apache-2.068 out of 100
#5HigressEnvoy-based API gateway with LLM proxy plugins and MCP server hosting Live demo ↗ (opens in a new tab)Gateways, 9.5k stars, Apache-2.061 out of 100
#6BifrostGo AI gateway with web UI, fallbacks, budgets and semantic cachingGateways, 8.7k stars, Apache-2.059 out of 100
#7PlanoEnvoy-based data plane that routes, traces and guards agent trafficGateways, 7.1k stars, Apache-2.058 out of 100
#8GoModelGo AI gateway with OpenAI and Anthropic APIs, caching and budgets Live demo ↗ (opens in a new tab)Gateways, 1.2k stars, MIT57 out of 100
#9agentgatewayOne proxy for LLM, MCP and A2A traffic with auth and RBACGateways, 5.3k stars, Apache-2.054 out of 100
#10optillmOpenAI-compatible proxy applying inference-time reasoning techniques Live demo ↗ (opens in a new tab)Gateways, 4.3k stars, Apache-2.051 out of 100

Reviews

186 out of 100

OmniRoute

OpenAI-compatible gateway that routes requests across hundreds of AI providers

74.9k stars, MIT, last commit Oct 2026

OmniRoute exposes one OpenAI-compatible endpoint at localhost:20128/v1 and routes requests to a catalog of 357+ providers, including many free tiers, with automatic fallback between them. It also accepts Claude, Gemini and Responses API formats, supports MCP and A2A, and has a dashboard for keys, quotas and free-tier usage. Providers are connected with your own accounts or API keys.

Strengths

  • Single endpoint with automatic fallback across many providers and model IDs
  • Dashboard page tracks free-tier pools and remaining quota
  • Install via npm, Docker or Electron; MIT license
  • Documents the free-tier token math and flags providers with risky terms

Weaknesses

  • Provider count in the README varies (290, 357, 370) across sections
  • The 'auto' model needs at least one eligible connected provider to route
  • Providers marked tos:avoid, such as Kiro, are excluded from auto routing by default
  • Free-tier token estimate depends on third-party limits that change
  • no GPU
  • Docker + Compose
  • Compose runs Redis, Qdrant
  • Models: OpenAI API, Claude API, Gemini API, Responses API
  • port 20128
284 out of 100

LiteLLM

OpenAI-format gateway and Python SDK for calling 100+ LLM providers

60.9k stars, custom license, last commit Oct 2026

LiteLLM translates calls to 100+ providers (OpenAI, Anthropic, Gemini, Bedrock, Azure and others) into the OpenAI format, either as a Python SDK or as a proxy server. The proxy adds virtual keys, spend tracking, guardrails, load balancing and an admin dashboard, and it also exposes A2A agent and MCP server gateways. The README reports 8ms P95 latency at 1k RPS.

Strengths

  • One OpenAI-style API across 100+ providers and many endpoint types
  • Proxy includes virtual keys, spend tracking, load balancing and admin dashboard
  • Also gateways A2A agents and MCP servers
  • Use as a Python library or as a standalone proxy

Weaknesses

  • License reported as NOASSERTION; an enterprise tier exists, feature split unclear from README
  • Proxy listens on port 4000 and its database requirements are not stated in the README excerpt
  • Provider coverage varies by endpoint; many providers support only chat-style endpoints
  • Python-based, so latency figures depend on the benchmark setup
  • no GPU
  • Docker + Compose
  • Compose runs PostgreSQL
  • Models: OpenAI, Anthropic, Gemini, AWS Bedrock, Azure
  • port 4000
376 out of 100

freellmapi

OpenAI-compatible router that fails over across free LLM provider tiers

32.9k stars, MIT, last commit Oct 2026

FreeLLMAPI exposes one /v1 endpoint (chat, responses, completions, embeddings, images, video, audio) and routes requests across free tiers from 34 providers, plus custom OpenAI-compatible endpoints. Provider keys are AES-256-GCM encrypted in SQLite, per-key RPM/RPD/TPM/TPD counters keep requests under quotas, and a 429 or 5xx triggers fallover to the next model. It also serves Anthropic Messages, Gemini and opt-in Ollama surfaces, and ships a React dashboard and desktop apps.

Strengths

  • Also speaks Anthropic, Gemini and Ollama formats, so Claude Code and Codex CLI connect
  • Per-key rate counters and automatic fallover on 429/5xx across providers
  • Keys AES-256-GCM encrypted in SQLite; apps only see one unified token
  • Runs on Node 20+ at about 40 MB idle RSS, or via Docker

Weaknesses

  • Free installs get new models 30 days after premium; same-day catalog costs $19/yr
  • Single-user by design; no multi-user setup described
  • Depends on free tiers that providers can change or retire without notice
  • Catalog sync pulls a signed feed from freellmapi.co
  • no GPU
  • Docker + Compose
  • Needs SQLite, Node 20+, provider API keys
  • Models: OpenAI-compatible API, Anthropic Messages API, Gemini API, Ollama API
  • port 3001
#468 out of 100

ContextForge MCP Gateway

Registry and proxy federating MCP, A2A, REST and gRPC behind one endpoint

4.6k stars, Apache-2.0, last commit Oct 2026

ContextForge is IBM's Python registry and proxy that federates MCP servers, A2A agents and REST or gRPC APIs into one MCP-compliant endpoint with auth, rate limiting, retries, an Admin UI and OpenTelemetry tracing. It installs from PyPI (mcpgateway on port 4444), as a GHCR container, via Docker Compose with PostgreSQL, Redis and nginx, or with a Helm chart, and virtualizes legacy REST services as MCP tools.

Strengths

  • Federates MCP, A2A, REST and gRPC (via reflection) behind one MCP endpoint
  • Transports: HTTP, JSON-RPC, WebSocket, SSE, Streamable HTTP, stdio
  • Admin UI with live log viewer; OpenTelemetry to Phoenix, Jaeger, Zipkin
  • Helm chart with HPA, Redis clustering and Grafana dashboards

Weaknesses

  • arm64 containers unsupported in production; Apple Silicon needs Rosetta or PyPI
  • Local Docker builds fail without the CI-only wheel closure; pull the GHCR image
  • Will not start without generated JWT_SECRET_KEY and AUTH_ENCRYPTION_SECRET
  • Large surface: 55+ tables, 40+ plugins, nginx and pgAdmin in the Compose stack
  • no GPU
  • Docker + Compose
  • Needs PostgreSQL (production; SQLite for dev), Redis (caching and federation)
  • Models: A2A agents: OpenAI, Anthropic, custom
  • port 4444
#561 out of 100

Higress

Envoy-based API gateway with LLM proxy plugins and MCP server hosting

Live demo ↗ (opens in a new tab)9.5k stars, Apache-2.0, last commit Oct 2026

Higress is a CNCF sandbox API gateway on Istio and Envoy, extended with Wasm plugins in Go, Rust or JS. Its AI plugins proxy mainstream LLM providers with token rate limiting, load balancing, caching and observability, and host remote MCP servers, including ones generated from OpenAPI specs. A Docker all-in-one image exposes the console on 8001 and the gateway on 8080/8443; Helm covers Kubernetes.

Strengths

  • Envoy-based with millisecond config reloads and no connection drops
  • Hosts MCP servers with auth, rate limits and audit; OpenAPI-to-MCP converter
  • Token rate limiting, multi-model load balancing and caching for LLM routes
  • Also a Kubernetes ingress controller and Gateway API implementation

Weaknesses

  • Images only on Alibaba Cloud registries; pulls can time out outside the mirrors
  • Istio and Envoy underneath; heavier than single-binary LLM proxies
  • AI features are Wasm plugins on a general API gateway
  • Docs split across higress.ai and higress.cn
  • no GPU
  • Docker
  • Models: mainstream LLM providers, domestic and international, via the ai-proxy plugin
  • port 8001
#659 out of 100

Bifrost

Go AI gateway with web UI, fallbacks, budgets and semantic caching

8.7k stars, Apache-2.0, last commit Oct 2026

Bifrost is a Go AI gateway that fronts 23+ providers (OpenAI, Anthropic, Bedrock, Vertex and more) with one OpenAI-compatible API and drop-in paths for the OpenAI, Anthropic and GenAI SDKs. It starts with npx or Docker on port 8080 with a web UI, and adds fallbacks, load balancing, semantic caching, MCP tool access, virtual keys, budgets and Prometheus metrics. Clustering, guardrails and the MCP gateway are enterprise features.

Strengths

  • Single Go binary via npx or Docker with zero-config web UI on 8080
  • Drop-in base URLs for OpenAI, Anthropic and Google GenAI SDKs
  • Virtual keys, team budgets, OIDC provisioning and Prometheus metrics
  • 11 microsecond added latency at 5k RPS in its own benchmark

Weaknesses

  • Guardrails, clustering, adaptive load balancing and MCP gateway are enterprise-only
  • Benchmarks are self-reported on t3 instances
  • 23+ providers, fewer than LiteLLM or Portkey
  • Semantic caching needs a vector store backend
  • no GPU
  • Models: OpenAI, Anthropic, AWS Bedrock, Google Vertex, Azure, Cerebras, Cohere, Mistral, Ollama, Groq and more
  • port 8080
#758 out of 100

Plano

Envoy-based data plane that routes, traces and guards agent traffic

7.1k stars, Apache-2.0, last commit Oct 2026

Plano is an Envoy-based proxy for agentic apps: a YAML file declares agents (HTTP servers with an OpenAI chat endpoint), model providers and listeners, and Plano routes each turn to the right agent with its 4B orchestrator model (hosted free or run locally). It also routes LLM calls by model name, alias or preference, captures OpenTelemetry traces with no instrumentation and applies guardrails via filter chains.

Strengths

  • Declarative multi-agent orchestration; agents are plain OpenAI-compatible HTTP servers
  • Zero-code OpenTelemetry traces and agentic signals for every request
  • Filter chains add moderation, jailbreak checks and memory out of process
  • Model routing by name, alias or preference across providers

Weaknesses

  • Agent routing depends on Plano's own orchestrator model; hosted by default
  • Install prerequisites live in external docs; README shows only YAML and curl
  • Envoy underneath; heavier than a single-binary proxy
  • No port or resource guidance beyond example listeners
  • no GPU
  • Docker
  • Needs Plano-Orchestrator routing model (hosted or local)
  • Models: OpenAI, Anthropic and other providers configured as model_providers
#857 out of 100

GoModel

Go AI gateway with OpenAI and Anthropic APIs, caching and budgets

Live demo ↗ (opens in a new tab)1.2k stars, MIT, last commit Oct 2026

GoModel is a Go AI gateway (install script or container on port 8080) exposing OpenAI-compatible /v1 and Anthropic /v1/messages endpoints in front of OpenAI, Anthropic, Gemini, Bedrock, Azure, Ollama, vLLM, SGLang and more. It adds exact and semantic caching, cost tracking, budgets, rate limits, failover, an MCP gateway, guardrails and a dashboard with playground. Compose adds Redis, PostgreSQL, MongoDB and Prometheus.

Strengths

  • Single Go binary or container; official OpenAI and Anthropic SDKs work unchanged
  • Budgets, rate limits, cost tracking and a usage API per user, team or key
  • Exact and semantic caching, failover with circuit breakers, provider key rotation
  • Dashboard with playground, live request stream, Prometheus and OpenTelemetry

Weaknesses

  • Prompt compression, intelligent routing and OIDC SSO are in the paid Pro build
  • Pre-1.0; roadmap points to an upcoming 0.2.0 release
  • Full Compose stack pulls in Redis, PostgreSQL, MongoDB and Prometheus
  • Benchmarks against LiteLLM and Portkey are self-run
  • no GPU
  • Docker + Compose
  • Needs Redis, PostgreSQL, MongoDB (Compose infrastructure)
  • Models: OpenAI, Anthropic, xAI, Gemini, Vertex AI, Cohere, DeepSeek, Groq, Fireworks, OpenRouter, Azure OpenAI, Bedrock, Ollama, SGLang, vLLM, llm-d, ElevenLabs and any OpenAI-compatible provider
  • port 8080
#954 out of 100

agentgateway

One proxy for LLM, MCP and A2A traffic with auth and RBAC

5.3k stars, Apache-2.0, last commit Oct 2026

Agentgateway is a Linux Foundation proxy for agent traffic: an LLM gateway (OpenAI-compatible API, budgets, failover), an MCP gateway federating tools over stdio, HTTP, SSE and Streamable HTTP, and an A2A gateway. It adds JWT, API key and OAuth auth, CEL RBAC, rate limits, guardrails and OpenTelemetry, and runs standalone from YAML or as a Kubernetes controller with Gateway API.

Strengths

  • One proxy for LLM, MCP and A2A traffic with an OpenAI-compatible API
  • MCP federation over stdio, HTTP, SSE and Streamable HTTP plus OpenAPI tools
  • JWT, API key and OAuth auth with CEL-based RBAC and rate limits
  • Standalone YAML mode or Kubernetes controller with Gateway API

Weaknesses

  • README has no install command, ports or resource needs; quickstart is external
  • Inference routing assumes Kubernetes Inference Gateway extensions
  • Marked in active development; roadmap is the issue tracker
  • Guardrail backends beyond regex are cloud services (OpenAI, Bedrock, Model Armor)
  • no GPU
  • Docker
  • Models: OpenAI, Anthropic, Gemini, Bedrock and other providers; self-hosted models via inference routing
#1051 out of 100

optillm

OpenAI-compatible proxy applying inference-time reasoning techniques

Live demo ↗ (opens in a new tab)4.3k stars, Apache-2.0, last commit Oct 2026

OptiLLM is an OpenAI-compatible proxy (pip or Docker, port 8000) that applies inference-time techniques such as mixture of agents, N-sample selection, self-consistency, MCTS, CePO and MARS to any upstream model, selected by a model-name prefix like moa-gpt-4o-mini. Plugins add an MCP client, memory, PII anonymization, code execution, JSON outputs and provider failover; upstreams are OpenAI, Cerebras, Azure or anything LiteLLM supports.

Strengths

  • 20+ techniques selected by model-name prefix, e.g. moa-gpt-4o-mini
  • Per-technique benchmarks listed (MARS +30 points on AIME 2025 with Gemini 2.5 Flash Lite)
  • Plugins for MCP client, memory, PII anonymization, code execution, JSON output
  • Works with any OpenAI-compatible endpoint; LiteLLM covers other providers

Weaknesses

  • Techniques multiply upstream calls (bon, MoA, MCTS), raising cost and latency
  • Runs on Flask's development server by default
  • Decoding techniques (cot_decoding, AutoThink) need the local inference path
  • Web search plugin drives Chrome through Selenium
  • GPU optional
  • Docker + Compose
  • Models: OpenAI, Cerebras, Azure OpenAI, any OpenAI-compatible endpoint, LiteLLM providers, local models via the built-in inference server
  • port 8000