#1414th of 20 in Model serving
Triton Inference Server
NVIDIA inference server for TensorRT, PyTorch, ONNX and more over HTTP/gRPC
- Stars
- 11.1k
- License
- BSD-3-Clause
- Last commit
- Oct 2026
- Last release
- Sep 2026
Overview
Triton serves TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL models over HTTP/REST and gRPC (KServe v2), with concurrent execution, dynamic and sequence batching, ensembles and Business Logic Scripting. NVIDIA ships it as NGC containers (2.73.0 / 26.09) for NVIDIA GPUs, x86 and ARM CPUs, Jetson and AWS Inferentia, with C and Java in-process APIs and a metrics endpoint.
Who it is for: ML platform teams serving many model types in production
Strengths
- Serves TensorRT, PyTorch, ONNX, OpenVINO, Python and FIL models together
- Dynamic and sequence batching, ensembles and BLS pipelines
- HTTP/REST and gRPC (KServe v2) plus C and Java in-process APIs
- Metrics for GPU utilization, throughput and latency
Weaknesses
- No OpenAI-compatible endpoint in the README; clients speak KServe v2
- Model repository and per-model config files are hand-written
- Containers track NVIDIA's monthly NGC release cycle
- Not every backend is supported on every platform
What it needs
- GPU optional
- Docker
- Models: TensorRT, PyTorch, ONNX, OpenVINO, Python, RAPIDS FIL backends
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |