#1414th of 20 in Model serving

Triton Inference Server

NVIDIA inference server for TensorRT, PyTorch, ONNX and more over HTTP/gRPC

Stars
11.1k
License
BSD-3-Clause
Last commit
Oct 2026
Last release
Sep 2026

Overview

Triton serves TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL models over HTTP/REST and gRPC (KServe v2), with concurrent execution, dynamic and sequence batching, ensembles and Business Logic Scripting. NVIDIA ships it as NGC containers (2.73.0 / 26.09) for NVIDIA GPUs, x86 and ARM CPUs, Jetson and AWS Inferentia, with C and Java in-process APIs and a metrics endpoint.

Who it is for: ML platform teams serving many model types in production

Strengths

  • Serves TensorRT, PyTorch, ONNX, OpenVINO, Python and FIL models together
  • Dynamic and sequence batching, ensembles and BLS pipelines
  • HTTP/REST and gRPC (KServe v2) plus C and Java in-process APIs
  • Metrics for GPU utilization, throughput and latency

Weaknesses

  • No OpenAI-compatible endpoint in the README; clients speak KServe v2
  • Model repository and per-model config files are hand-written
  • Containers track NVIDIA's monthly NGC release cycle
  • Not every backend is supported on every platform

What it needs

  • GPU optional
  • Docker
  • Models: TensorRT, PyTorch, ONNX, OpenVINO, Python, RAPIDS FIL backends

Also in Model serving

See all 20
Also in Model serving
RankProjectScore
1LocalAIOne OpenAI-compatible server for text, speech, image and video models49.5k stars, MIT82 out of 100
2llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API130.8k stars, MIT79 out of 100
3vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs93.5k stars, Apache-2.075 out of 100
#4OllamaRuns open-weight models locally behind a CLI and REST API182.7k stars, MIT74 out of 100
#5colibriC inference engine that runs huge MoE models by streaming experts from disk41k stars, Apache-2.070 out of 100