#1111th of 20 in Model serving
mistral.rs
Rust inference server with OpenAI and Anthropic APIs and agent tools
- Stars
- 7.7k
- License
- MIT
- Last commit
- Oct 2026
- Last release
- Sep 2026
Overview
mistral.rs is a Rust engine whose single binary runs and serves Hugging Face, GGUF and UQFF models (text, vision, video, audio, speech, image generation; 45+ architectures) with auto-detected architecture and chat template. The serve command exposes OpenAI /v1 and Anthropic Messages endpoints, a web UI at /ui and Prometheus metrics on port 1234, with paged attention, ISQ, LoRA and a built-in agent loop.
Who it is for: developers wanting a fast Rust server with agent features
Strengths
- One binary for chat, server, benchmarks and web UI; prebuilt for Metal, CUDA, CPU
- In-situ quantization of any Hugging Face model plus GGUF 2-8 bit, GPTQ, AWQ, FP8
- Server-side agent loop with Python, shell, web search, skills and MCP client
- mistralrs tune recommends quantization and device mapping for your hardware
Weaknesses
- BF16 prefill trails vLLM by 5-10x on the 26B MoE in its own benchmarks
- Install script is curl piped to sh, falling back to a source build
- cuTile acceleration needs NVIDIA's separately installed tileiras tool
- Not affiliated with Mistral AI despite the name
What it needs
- GPU optional
- Docker
- Models: Hugging Face safetensors, GGUF, UQFF, Qwen3, Gemma 4, Muse Glimmer, DiffusionGemma and 45+ architectures
- port 1234
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |