Open-source alternatives to LM Studio

The highest-scoring open-source projects in Model serving. They are picks, not exact replacements, so read the weaknesses before you switch.

Open-source alternatives to LM Studio
RankProjectScore
1LocalAIOne OpenAI-compatible server for text, speech, image and video modelsModel serving, 49.5k stars, MIT82 out of 100
2llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible APIModel serving, 130.7k stars, MIT79 out of 100
3vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIsModel serving, 93.5k stars, Apache-2.075 out of 100
#4OllamaRuns open-weight models locally behind a CLI and REST APIModel serving, 182.6k stars, MIT74 out of 100
#5colibriC inference engine that runs huge MoE models by streaming experts from diskModel serving, 41k stars, Apache-2.070 out of 100
#6SGLangInference server for LLM, vision-language and diffusion modelsModel serving, 37k stars, Apache-2.068 out of 100
#7LemonadeLocal AI server that targets GPUs and AMD NPUs with OpenAI-style APIsModel serving, 5.9k stars, Apache-2.066 out of 100
#8exoDistributed LLM inference across Macs and other devices, MLX-basedModel serving, 47.8k stars, Apache-2.064 out of 100
#9llama-swapGo proxy that hot-swaps local model servers per requestModel serving, 5.9k stars, MIT63 out of 100
#10XinferenceServes LLM, embedding, speech and image models behind one OpenAI-compatible APIModel serving, 9.6k stars, Apache-2.062 out of 100

Reviews

182 out of 100

LocalAI

One OpenAI-compatible server for text, speech, image and video models

49.5k stars, MIT, last commit Oct 2026

LocalAI is a Go server on port 8080 with OpenAI, Anthropic, ElevenLabs and Ollama-compatible APIs for text, vision, speech, image and video. Backends (llama.cpp, vLLM, SGLang, whisper.cpp, diffusers, MLX, 60+ total) ship as separate OCI images pulled on demand; containers exist for CPU, CUDA, ROCm, Intel and Vulkan. It adds API keys, quotas and OIDC, agents with MCP, and a PostgreSQL/NATS distributed mode.

Strengths

  • Small core; 60+ backends installed on demand as OCI images
  • OpenAI, Anthropic, ElevenLabs and Ollama API compatibility in one server
  • Multi-user: API keys, per-user quotas, role-based access, OIDC
  • Container images for CPU, CUDA 12/13, ROCm, Intel oneAPI, Vulkan, Jetson

Weaknesses

  • First model load pulls backend images; needs network and disk space
  • macOS DMG is unsigned and needs quarantine removal
  • Distributed mode requires PostgreSQL and NATS
  • Very wide scope (agents, biometrics, video) increases configuration surface
  • GPU optional
  • Docker + Compose
  • Needs PostgreSQL and NATS (distributed mode only)
  • Models: GGUF via llama.cpp, vLLM, SGLang, transformers, MLX, diffusers, whisper.cpp backends, models from gallery, Hugging Face, Ollama registry, OCI images, YAML
  • port 8080
279 out of 100

llama.cpp

C/C++ inference engine serving GGUF models over an OpenAI-compatible API

130.7k stars, MIT, last commit Oct 2026

llama.cpp is a C/C++ inference engine for LLMs and VLMs with no dependencies, built on ggml. llama serve starts an OpenAI-compatible API server with a built-in web UI, pulling GGUF models straight from Hugging Face, with 1.5 to 8-bit quantization and CPU+GPU hybrid offload for models larger than VRAM. Backends cover CUDA, HIP, Metal, Vulkan, SYCL, OpenCL, CANN, MUSA and WebGPU.

Strengths

  • Plain C/C++ with no runtime dependencies; prebuilt binaries and Docker
  • Backends for NVIDIA, AMD, Apple Metal, Intel SYCL, Vulkan, Ascend, Moore Threads
  • Hybrid CPU+GPU offload runs models larger than available VRAM
  • Built-in web UI and OpenAI-compatible server via llama serve

Weaknesses

  • GGUF model format only
  • README gives no port, auth or sizing guidance; see tools/server docs
  • Install script is curl piped to sh; otherwise build from source
  • OpenVINO backend still in progress
  • GPU optional
  • Docker
  • Models: GGUF models from Hugging Face (e.g. Qwen3.5-0.8B-GGUF)
375 out of 100

vLLM

High-throughput LLM serving engine with OpenAI and Anthropic APIs

93.5k stars, Apache-2.0, last commit Oct 2026

vLLM is a Python serving engine for Hugging Face models that batches requests continuously with PagedAttention, prefix caching and speculative decoding, exposing an OpenAI-compatible API plus Anthropic Messages API and gRPC. It covers 200+ architectures (dense, MoE, multimodal, embedding) with FP8, INT8, GPTQ, AWQ and GGUF quantization and tensor, pipeline and expert parallelism.

Strengths

  • Continuous batching with PagedAttention for high multi-user throughput
  • 200+ Hugging Face architectures including MoE, multimodal and embedding models
  • OpenAI, Anthropic Messages and gRPC endpoints with tool calling and structured output
  • Runs on NVIDIA, AMD, Intel GPUs, CPUs, TPUs, Gaudi, Ascend via plugins

Weaknesses

  • No web UI; API server only
  • README gives no VRAM, port or model-size guidance
  • Heavy Python, PyTorch and CUDA dependency chain; no single binary
  • Most optimized kernels target NVIDIA and AMD GPUs; CPU path is secondary
  • GPU optional
  • Docker
  • Models: 200+ Hugging Face architectures: Llama, Qwen, Gemma, Mixtral, DeepSeek-V3, GPT-OSS, LLaVA, Qwen-VL, E5-Mistral
#474 out of 100

Ollama

Runs open-weight models locally behind a CLI and REST API

182.6k stars, MIT, last commit Oct 2026

Ollama runs open-weight models locally with a CLI and a REST API on port 11434, pulling models from its own library (for example gemma4) and using llama.cpp as the inference backend. Install scripts cover macOS, Windows and Linux, and an official Docker image exists. The ollama launch command wires it into coding agents such as Claude Code, Codex, Copilot CLI and OpenCode, or into OpenClaw as a chat assistant.

Strengths

  • One command pulls and runs a model; REST API on 11434
  • Official Docker image plus Python and JavaScript libraries
  • ollama launch integrates with Claude Code, Codex, Copilot CLI, OpenCode
  • Broad ecosystem: dozens of web, desktop and IDE clients listed

Weaknesses

  • Single inference backend: llama.cpp
  • Install is a curl piped to sh script
  • README gives no RAM or VRAM guidance per model size
  • Models come from Ollama's own registry; others need import steps
  • GPU optional
  • Docker
  • Models: Ollama library models (e.g. gemma4), GGUF via llama.cpp
  • port 11434
#570 out of 100

colibri

C inference engine that runs huge MoE models by streaming experts from disk

41k stars, Apache-2.0, last commit Oct 2026

colibri runs large mixture-of-experts models such as GLM-5.2 (744B) and Kimi K3 (2.8T) on ordinary hardware. It keeps the dense weights in RAM and reads routed experts from disk through a cache, with optional Vulkan or CUDA offload. It serves a browser dashboard plus OpenAI- and Anthropic-compatible HTTP endpoints, and ships guided setup scripts for Windows, Linux and macOS.

Strengths

  • Pure C engines, no GPU required; 8 GB RAM minimum for the smallest model
  • Serves OpenAI and Anthropic-style APIs on port 8000, plus a web dashboard
  • Vulkan works on AMD, Intel and NVIDIA; CUDA path for NVIDIA on Linux and Windows
  • Setup script detects hardware, recommends a model, and resumes interrupted downloads

Weaknesses

  • Large models stream from disk: 0.05-3 tok/s on typical machines, per README tables
  • Discrete-GPU performance mostly unmeasured by the authors; relies on user reports
  • Supports a fixed list of model families, one engine each; other architectures need new code
  • Some models need manual conversion or preparation steps after download
  • RAM ≥ 8 GB
  • GPU optional
  • Docker + Compose
  • Models: GLM-5.2, GLM-5.3, GLM-5.3-Flash, Inkling, Kimi K3
  • port 8000
#668 out of 100

SGLang

Inference server for LLM, vision-language and diffusion models

37k stars, Apache-2.0, last commit Oct 2026

SGLang is a Python inference framework for serving large language, vision-language and diffusion models, aimed at agentic workloads, large-scale serving and RL rollouts. It runs on NVIDIA, AMD, Google TPU, Intel, Apple Silicon, Huawei Ascend and Moore Threads hardware, and includes a built-in engine for image and video generation. Install via the lmsysorg/sglang Docker image or uv pip.

Strengths

  • Supports NVIDIA, AMD, TPU, Intel, Apple Silicon, Ascend and Moore Threads hardware
  • Image and video diffusion engine ships in the same package
  • Integrated by RL training frameworks such as verl and slime for rollouts
  • Apache-2.0 license, with a Docker image and a cookbook of launch commands

Weaknesses

  • Audio TTS/ASR serving lives in a separate project, SGLang Omni
  • Install requires --prerelease=allow with uv, suggesting prerelease dependencies
  • Several hardware backends (Trainium, Cambricon, MetaX) are still in progress
  • README gives no RAM or VRAM requirements; sizing depends on model and hardware
  • Docker + Compose
  • Models: LLMs, vision-language models, diffusion models
#766 out of 100

Lemonade

Local AI server that targets GPUs and AMD NPUs with OpenAI-style APIs

5.9k stars, Apache-2.0, last commit Oct 2026

Lemonade is a local AI server with OpenAI, Anthropic and Ollama-compatible APIs on port 13305 that runs GGUF, FLM and ONNX models, Whisper transcription, Kokoro speech and Stable Diffusion images. It picks the backend for the hardware: llama.cpp on CPU, CUDA, Vulkan, ROCm or Metal, plus AMD XDNA2 NPU paths for Ryzen AI. Packages exist for Windows, macOS, Debian, Fedora, Ubuntu, Arch, Snap and Docker.

Strengths

  • NPU backends for AMD XDNA2 (Ryzen AI) alongside CUDA, ROCm, Vulkan, Metal
  • Chat, speech-to-text, text-to-speech, image and audio generation in one server
  • Native packages: msi, pkg, deb, rpm, Arch, Snap, PPA, Docker
  • Model aliases enable active-standby failover between models

Weaknesses

  • Many engines (vllm, ds4, openmoss, trellis) are marked experimental
  • NPU support covers AMD XDNA2 only
  • macOS gets Metal only; several backends are Windows or Linux only
  • Cloud offload to OpenAI-compatible providers is experimental
  • GPU optional
  • Docker
  • Models: GGUF, FLM and ONNX LLMs (e.g. Gemma 4, Qwen3), Whisper, Kokoro, SDXL-Turbo
  • port 13305
#864 out of 100

exo

Distributed LLM inference across Macs and other devices, MLX-based

47.8k stars, Apache-2.0, last commit Aug 2026

exo joins devices on a network into one inference cluster, splitting a model across them with pipeline or tensor parallelism based on a live view of topology. It uses MLX for inference and exposes OpenAI Chat Completions, Claude Messages, OpenAI Responses and Ollama-compatible APIs plus a built-in dashboard on port 52415. On Macs with Thunderbolt 5 and macOS 26.2 or later, it can use RDMA between nodes.

Strengths

  • Runs models too large for one machine by sharding across devices
  • Devices discover each other automatically, no manual cluster config
  • Serves OpenAI, Claude Messages, Responses and Ollama-style APIs
  • Tensor parallelism reported up to 1.8x on 2 devices, 3.2x on 4

Weaknesses

  • On Linux, inference currently runs on CPU only; GPU support is in development
  • RDMA needs macOS 26.2+, Thunderbolt 5, a Recovery-mode setting and identical OS versions
  • Source install needs Rust nightly, Node and uv; no Docker image detected
  • Models are MLX-format; GGUF support is not mentioned in the README
  • GPU optional
  • Needs uv, Node.js 18+, Rust nightly, MLX, macmon (macOS)
  • Models: MLX, HuggingFace custom models
  • port 52415
#963 out of 100

llama-swap

Go proxy that hot-swaps local model servers per request

5.9k stars, MIT, last commit Oct 2026

llama-swap is one Go binary that proxies OpenAI and Anthropic API calls to local servers (llama-server, vLLM, stable-diffusion.cpp, whisper.cpp) and starts, stops or swaps the right one per model ID from a YAML file. It adds a web UI, log streaming, Prometheus metrics, API keys, TTL unload and a matrix DSL for concurrent models. Unified Docker images bundle the servers for CUDA and Vulkan.

Strengths

  • One binary, one YAML file, zero dependencies
  • Hot-swaps any OpenAI or Anthropic-compatible upstream per model ID, with ttl unload
  • Unified images bundle llama-server, stable-diffusion.cpp, whisper.cpp, audio.cpp
  • Web UI with playground, token metrics, request inspection and live logs

Weaknesses

  • Basic mode runs one model at a time; concurrency needs the matrix DSL
  • Python servers like vLLM or tabbyAPI should run in containers for clean SIGTERM
  • nginx needs proxy_buffering off or SSE streaming breaks
  • Container listen address must stay 0.0.0.0 when publishing ports
  • GPU optional
  • Docker
  • Needs an upstream inference server (llama-server, vLLM, etc.)
  • Models: any model served by the configured upstream (GGUF via llama-server, etc.)
  • port 8080
#1062 out of 100

Xinference

Serves LLM, embedding, speech and image models behind one OpenAI-compatible API

9.6k stars, Apache-2.0, last commit Oct 2026

Xinference is a Python library and server that runs language, embedding, speech recognition, image and multimodal models from a single command, with built-in model definitions and support for custom ones. It exposes an OpenAI-compatible REST API with function calling, plus RPC, a CLI and a web UI, and can spread models across multiple workers. It runs on GPUs and CPUs through backends including vLLM and its own llama.cpp binding.

Strengths

  • One server covers LLM, embedding, audio, image and multimodal models
  • OpenAI-compatible REST API with function calling
  • Distributed inference across workers; Helm chart for Kubernetes
  • Install via pip, one-line script, Docker or Helm

Weaknesses

  • Enterprise, Cloud and managed Model API offerings are commercial; edition differences not listed
  • 3.0.0 release notes mention breaking changes and migration steps
  • RAM and VRAM requirements are not stated; they depend on the model
  • README feature list is dense; hardware and backend support per model is unclear
  • GPU optional
  • Models: vLLM, ggml, llama.cpp (xllamacpp), TensorRT, embedding models
  • port 9997