#1818th of 20 in Model serving
Text Embeddings Inference
Rust server for embedding, reranker and classification models
- Stars
- 5.1k
- License
- Apache-2.0
- Last commit
- Oct 2026
- Last release
- Sep 2026
Overview
TEI is a Rust server from Hugging Face for embedding, reranker and sequence-classification models (BERT, XLM-RoBERTa, Nomic, Jina, GTE, Qwen3, ModernBERT, Gemma3) with token-based dynamic batching and Flash Attention. The router listens on port 3000 with /embed and OpenAI-compatible routes, gRPC, OpenTelemetry tracing and Prometheus metrics. Images cover CPU x86/arm64 and NVIDIA Turing through Blackwell.
Who it is for: RAG builders needing a fast embedding and rerank endpoint
Strengths
- Token-based dynamic batching with Flash Attention, Candle and cuBLASLt
- Small images and fast boot; no graph compilation step
- Rerankers and classifiers served alongside embeddings
- OpenTelemetry tracing, Prometheus metrics, API key auth, gRPC
Weaknesses
- No Volta support; Turing image is experimental with Flash Attention off
- GPU images need drivers compatible with CUDA 12.2 or higher
- Only CamemBERT and XLM-RoBERTa for sequence classification
- 7B embedders such as Qwen3-Embedding-8B are flagged very expensive
What it needs
- GPU optional
- Docker
- Models: Qwen3-Embedding, gte-Qwen2, multilingual-e5, embeddinggemma, arctic-embed, nomic-embed, ModernBERT, jina-embeddings-v2, bge-reranker, gte rerankers
- port 3000
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |