#88th of 20 in Model serving
exo
Distributed LLM inference across Macs and other devices, MLX-based
- Stars
- 47.8k
- License
- Apache-2.0
- Last commit
- Aug 2026
- Last release
- Apr 2026
- Language
- Python
Overview
exo joins devices on a network into one inference cluster, splitting a model across them with pipeline or tensor parallelism based on a live view of topology. It uses MLX for inference and exposes OpenAI Chat Completions, Claude Messages, OpenAI Responses and Ollama-compatible APIs plus a built-in dashboard on port 52415. On Macs with Thunderbolt 5 and macOS 26.2 or later, it can use RDMA between nodes.
Who it is for: Self-hosters pooling several Macs to run large models
Strengths
- Runs models too large for one machine by sharding across devices
- Devices discover each other automatically, no manual cluster config
- Serves OpenAI, Claude Messages, Responses and Ollama-style APIs
- Tensor parallelism reported up to 1.8x on 2 devices, 3.2x on 4
Weaknesses
- On Linux, inference currently runs on CPU only; GPU support is in development
- RDMA needs macOS 26.2+, Thunderbolt 5, a Recovery-mode setting and identical OS versions
- Source install needs Rust nightly, Node and uv; no Docker image detected
- Models are MLX-format; GGUF support is not mentioned in the README
What it needs
- GPU optional
- Needs uv, Node.js 18+, Rust nightly, MLX, macmon (macOS)
- Models: MLX, HuggingFace custom models
- port 52415
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |