Inference engines and model servers that expose local models over an API.
LocalAI leads with 82, ahead of llama.cpp (79) and vLLM (75).20 projects ranked by score.
The ranking
Model serving: full ranking
Rank
Project
Adoption
Freshness
Maintenance
Easy to run
Agent-ready
Score
1
LocalAIOne OpenAI-compatible server for text, speech, image and video models49.5k stars, MIT, last commit Oct 2026
83
100
89
67
70
82 out of 100score
2
llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API130.8k stars, MIT, last commit Oct 2026
94
100
89
50
70
79 out of 100score
3
vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs93.6k stars, Apache-2.0, last commit Oct 2026
91
100
81
50
45
75 out of 100score
#4
OllamaRuns open-weight models locally behind a CLI and REST API182.7k stars, MIT, last commit Oct 2026
98
100
82
33
70
74 out of 100score
#5
colibriC inference engine that runs huge MoE models by streaming experts from disk41.1k stars, Apache-2.0, last commit Oct 2026
72
100
98
50
0
70 out of 100score
#6
SGLangInference server for LLM, vision-language and diffusion models37k stars, Apache-2.0, last commit Oct 2026
67
100
85
50
15
68 out of 100score
#7
LemonadeLocal AI server that targets GPUs and AMD NPUs with OpenAI-style APIs5.9k stars, Apache-2.0, last commit Oct 2026
20
100
84
67
70
66 out of 100score
#8
exoDistributed LLM inference across Macs and other devices, MLX-based47.8k stars, Apache-2.0, last commit Aug 2026
79
100
12
50
70
64 out of 100score
#9
llama-swapGo proxy that hot-swaps local model servers per request5.9k stars, MIT, last commit Oct 2026
22
100
89
50
70
63 out of 100score
#10
XinferenceServes LLM, embedding, speech and image models behind one OpenAI-compatible API9.6k stars, Apache-2.0, last commit Oct 2026
36
100
98
33
70
62 out of 100score
#11
mistral.rsRust inference server with OpenAI and Anthropic APIs and agent tools7.7k stars, MIT, last commit Oct 2026
28
100
74
50
70
62 out of 100score
#12
Text Generation Web UILocal LLM chat UI and API with five switchable loader backends47.7k stars, AGPL-3.0, last commit Aug 2026
76
100
12
50
0
58 out of 100score
#13
KTransformersCPU-GPU hybrid inference and fine-tuning for very large MoE models19.6k stars, Apache-2.0, last commit Oct 2026
51
100
78
33
0
57 out of 100score
#14
Triton Inference ServerNVIDIA inference server for TensorRT, PyTorch, ONNX and more over HTTP/gRPC11.1k stars, BSD-3-Clause, last commit Oct 2026
40
100
85
33
0
56 out of 100score
#15
GPUStackGPU cluster manager that deploys models on vLLM, SGLang and TensorRT-LLM5.8k stars, Apache-2.0, last commit Oct 2026
17
100
87
33
70
56 out of 100score
#16
LMDeployLLM and VLM serving toolkit with the TurboMind and PyTorch engines8.1k stars, Apache-2.0, last commit Oct 2026
31
100
82
33
15
54 out of 100score
#17
llamafileSingle-file executables that bundle llama.cpp with model weights26.2k stars, custom license, last commit Oct 2026
57
100
80
0
0
49 out of 100score
#18
Text Embeddings InferenceRust server for embedding, reranker and classification models5.1k stars, Apache-2.0, last commit Oct 2026
13
100
49
50
0
49 out of 100score
#19
TabbyAPIOpenAI-compatible API server for running ExLlamaV3 models1.5k stars, AGPL-3.0, last commit Oct 2026
1
100
57
33
0
42 out of 100score
#20
OpenLLMOne-command OpenAI-compatible endpoints for curated open LLMs12.6k stars, Apache-2.0, last commit May 2026
44
59
25
33
0
38 out of 100score
Momentum, Verified build, Docs and Privacy are not measured yet; their weight goes to the signals shown. A dash means the signal is not scored for that kind of project. Hover a number for its rating in words.
One OpenAI-compatible server for text, speech, image and video models
49.5k stars, MIT, last commit Oct 2026
LocalAI is a Go server on port 8080 with OpenAI, Anthropic, ElevenLabs and Ollama-compatible APIs for text, vision, speech, image and video. Backends (llama.cpp, vLLM, SGLang, whisper.cpp, diffusers, MLX, 60+ total) ship as separate OCI images pulled on demand; containers exist for CPU, CUDA, ROCm, Intel and Vulkan. It adds API keys, quotas and OIDC, agents with MCP, and a PostgreSQL/NATS distributed mode.
Strengths
Small core; 60+ backends installed on demand as OCI images
OpenAI, Anthropic, ElevenLabs and Ollama API compatibility in one server
Multi-user: API keys, per-user quotas, role-based access, OIDC
Container images for CPU, CUDA 12/13, ROCm, Intel oneAPI, Vulkan, Jetson
Weaknesses
First model load pulls backend images; needs network and disk space
macOS DMG is unsigned and needs quarantine removal
Distributed mode requires PostgreSQL and NATS
Very wide scope (agents, biometrics, video) increases configuration surface
C/C++ inference engine serving GGUF models over an OpenAI-compatible API
130.8k stars, MIT, last commit Oct 2026
llama.cpp is a C/C++ inference engine for LLMs and VLMs with no dependencies, built on ggml. llama serve starts an OpenAI-compatible API server with a built-in web UI, pulling GGUF models straight from Hugging Face, with 1.5 to 8-bit quantization and CPU+GPU hybrid offload for models larger than VRAM. Backends cover CUDA, HIP, Metal, Vulkan, SYCL, OpenCL, CANN, MUSA and WebGPU.
Strengths
Plain C/C++ with no runtime dependencies; prebuilt binaries and Docker
Backends for NVIDIA, AMD, Apple Metal, Intel SYCL, Vulkan, Ascend, Moore Threads
Hybrid CPU+GPU offload runs models larger than available VRAM
Built-in web UI and OpenAI-compatible server via llama serve
Weaknesses
GGUF model format only
README gives no port, auth or sizing guidance; see tools/server docs
Install script is curl piped to sh; otherwise build from source
OpenVINO backend still in progress
GPU optional
Docker
Models: GGUF models from Hugging Face (e.g. Qwen3.5-0.8B-GGUF)
High-throughput LLM serving engine with OpenAI and Anthropic APIs
93.6k stars, Apache-2.0, last commit Oct 2026
vLLM is a Python serving engine for Hugging Face models that batches requests continuously with PagedAttention, prefix caching and speculative decoding, exposing an OpenAI-compatible API plus Anthropic Messages API and gRPC. It covers 200+ architectures (dense, MoE, multimodal, embedding) with FP8, INT8, GPTQ, AWQ and GGUF quantization and tensor, pipeline and expert parallelism.
Strengths
Continuous batching with PagedAttention for high multi-user throughput
200+ Hugging Face architectures including MoE, multimodal and embedding models
OpenAI, Anthropic Messages and gRPC endpoints with tool calling and structured output
Runs on NVIDIA, AMD, Intel GPUs, CPUs, TPUs, Gaudi, Ascend via plugins
Weaknesses
No web UI; API server only
README gives no VRAM, port or model-size guidance
Heavy Python, PyTorch and CUDA dependency chain; no single binary
Most optimized kernels target NVIDIA and AMD GPUs; CPU path is secondary
Runs open-weight models locally behind a CLI and REST API
182.7k stars, MIT, last commit Oct 2026
Ollama runs open-weight models locally with a CLI and a REST API on port 11434, pulling models from its own library (for example gemma4) and using llama.cpp as the inference backend. Install scripts cover macOS, Windows and Linux, and an official Docker image exists. The ollama launch command wires it into coding agents such as Claude Code, Codex, Copilot CLI and OpenCode, or into OpenClaw as a chat assistant.
Strengths
One command pulls and runs a model; REST API on 11434
Official Docker image plus Python and JavaScript libraries
ollama launch integrates with Claude Code, Codex, Copilot CLI, OpenCode
Broad ecosystem: dozens of web, desktop and IDE clients listed
Weaknesses
Single inference backend: llama.cpp
Install is a curl piped to sh script
README gives no RAM or VRAM guidance per model size
Models come from Ollama's own registry; others need import steps
GPU optional
Docker
Models: Ollama library models (e.g. gemma4), GGUF via llama.cpp
C inference engine that runs huge MoE models by streaming experts from disk
41.1k stars, Apache-2.0, last commit Oct 2026
colibri runs large mixture-of-experts models such as GLM-5.2 (744B) and Kimi K3 (2.8T) on ordinary hardware. It keeps the dense weights in RAM and reads routed experts from disk through a cache, with optional Vulkan or CUDA offload. It serves a browser dashboard plus OpenAI- and Anthropic-compatible HTTP endpoints, and ships guided setup scripts for Windows, Linux and macOS.
Strengths
Pure C engines, no GPU required; 8 GB RAM minimum for the smallest model
Serves OpenAI and Anthropic-style APIs on port 8000, plus a web dashboard
Vulkan works on AMD, Intel and NVIDIA; CUDA path for NVIDIA on Linux and Windows
Setup script detects hardware, recommends a model, and resumes interrupted downloads
Weaknesses
Large models stream from disk: 0.05-3 tok/s on typical machines, per README tables
Discrete-GPU performance mostly unmeasured by the authors; relies on user reports
Supports a fixed list of model families, one engine each; other architectures need new code
Some models need manual conversion or preparation steps after download
RAM ≥ 8 GB
GPU optional
Docker + Compose
Models: GLM-5.2, GLM-5.3, GLM-5.3-Flash, Inkling, Kimi K3
Inference server for LLM, vision-language and diffusion models
37k stars, Apache-2.0, last commit Oct 2026
SGLang is a Python inference framework for serving large language, vision-language and diffusion models, aimed at agentic workloads, large-scale serving and RL rollouts. It runs on NVIDIA, AMD, Google TPU, Intel, Apple Silicon, Huawei Ascend and Moore Threads hardware, and includes a built-in engine for image and video generation. Install via the lmsysorg/sglang Docker image or uv pip.
Strengths
Supports NVIDIA, AMD, TPU, Intel, Apple Silicon, Ascend and Moore Threads hardware
Image and video diffusion engine ships in the same package
Integrated by RL training frameworks such as verl and slime for rollouts
Apache-2.0 license, with a Docker image and a cookbook of launch commands
Weaknesses
Audio TTS/ASR serving lives in a separate project, SGLang Omni
Install requires --prerelease=allow with uv, suggesting prerelease dependencies
Several hardware backends (Trainium, Cambricon, MetaX) are still in progress
README gives no RAM or VRAM requirements; sizing depends on model and hardware
Local AI server that targets GPUs and AMD NPUs with OpenAI-style APIs
5.9k stars, Apache-2.0, last commit Oct 2026
Lemonade is a local AI server with OpenAI, Anthropic and Ollama-compatible APIs on port 13305 that runs GGUF, FLM and ONNX models, Whisper transcription, Kokoro speech and Stable Diffusion images. It picks the backend for the hardware: llama.cpp on CPU, CUDA, Vulkan, ROCm or Metal, plus AMD XDNA2 NPU paths for Ryzen AI. Packages exist for Windows, macOS, Debian, Fedora, Ubuntu, Arch, Snap and Docker.
Strengths
NPU backends for AMD XDNA2 (Ryzen AI) alongside CUDA, ROCm, Vulkan, Metal
Chat, speech-to-text, text-to-speech, image and audio generation in one server
Native packages: msi, pkg, deb, rpm, Arch, Snap, PPA, Docker
Model aliases enable active-standby failover between models
Weaknesses
Many engines (vllm, ds4, openmoss, trellis) are marked experimental
NPU support covers AMD XDNA2 only
macOS gets Metal only; several backends are Windows or Linux only
Cloud offload to OpenAI-compatible providers is experimental
Distributed LLM inference across Macs and other devices, MLX-based
47.8k stars, Apache-2.0, last commit Aug 2026
exo joins devices on a network into one inference cluster, splitting a model across them with pipeline or tensor parallelism based on a live view of topology. It uses MLX for inference and exposes OpenAI Chat Completions, Claude Messages, OpenAI Responses and Ollama-compatible APIs plus a built-in dashboard on port 52415. On Macs with Thunderbolt 5 and macOS 26.2 or later, it can use RDMA between nodes.
Strengths
Runs models too large for one machine by sharding across devices
Devices discover each other automatically, no manual cluster config
Serves OpenAI, Claude Messages, Responses and Ollama-style APIs
Tensor parallelism reported up to 1.8x on 2 devices, 3.2x on 4
Weaknesses
On Linux, inference currently runs on CPU only; GPU support is in development
RDMA needs macOS 26.2+, Thunderbolt 5, a Recovery-mode setting and identical OS versions
Source install needs Rust nightly, Node and uv; no Docker image detected
Models are MLX-format; GGUF support is not mentioned in the README
Go proxy that hot-swaps local model servers per request
5.9k stars, MIT, last commit Oct 2026
llama-swap is one Go binary that proxies OpenAI and Anthropic API calls to local servers (llama-server, vLLM, stable-diffusion.cpp, whisper.cpp) and starts, stops or swaps the right one per model ID from a YAML file. It adds a web UI, log streaming, Prometheus metrics, API keys, TTL unload and a matrix DSL for concurrent models. Unified Docker images bundle the servers for CUDA and Vulkan.
Strengths
One binary, one YAML file, zero dependencies
Hot-swaps any OpenAI or Anthropic-compatible upstream per model ID, with ttl unload
Serves LLM, embedding, speech and image models behind one OpenAI-compatible API
9.6k stars, Apache-2.0, last commit Oct 2026
Xinference is a Python library and server that runs language, embedding, speech recognition, image and multimodal models from a single command, with built-in model definitions and support for custom ones. It exposes an OpenAI-compatible REST API with function calling, plus RPC, a CLI and a web UI, and can spread models across multiple workers. It runs on GPUs and CPUs through backends including vLLM and its own llama.cpp binding.
Strengths
One server covers LLM, embedding, audio, image and multimodal models
OpenAI-compatible REST API with function calling
Distributed inference across workers; Helm chart for Kubernetes
Install via pip, one-line script, Docker or Helm
Weaknesses
Enterprise, Cloud and managed Model API offerings are commercial; edition differences not listed
3.0.0 release notes mention breaking changes and migration steps
RAM and VRAM requirements are not stated; they depend on the model
README feature list is dense; hardware and backend support per model is unclear
Rust inference server with OpenAI and Anthropic APIs and agent tools
7.7k stars, MIT, last commit Oct 2026
mistral.rs is a Rust engine whose single binary runs and serves Hugging Face, GGUF and UQFF models (text, vision, video, audio, speech, image generation; 45+ architectures) with auto-detected architecture and chat template. The serve command exposes OpenAI /v1 and Anthropic Messages endpoints, a web UI at /ui and Prometheus metrics on port 1234, with paged attention, ISQ, LoRA and a built-in agent loop.
Strengths
One binary for chat, server, benchmarks and web UI; prebuilt for Metal, CUDA, CPU
In-situ quantization of any Hugging Face model plus GGUF 2-8 bit, GPTQ, AWQ, FP8
Server-side agent loop with Python, shell, web search, skills and MCP client
mistralrs tune recommends quantization and device mapping for your hardware
Weaknesses
BF16 prefill trails vLLM by 5-10x on the 26B MoE in its own benchmarks
Install script is curl piped to sh, falling back to a source build
Local LLM chat UI and API with five switchable loader backends
47.7k stars, AGPL-3.0, last commit Aug 2026
TextGen runs local LLMs behind a chat UI and an OpenAI/Anthropic-compatible API with tool calling and MCP, with llama.cpp, ik_llama.cpp, Transformers, ExLlamaV3 or TensorRT-LLM loaders switchable without restart. Portable builds for Linux, Windows and macOS bundle CUDA, Vulkan, ROCm or CPU dependencies for GGUF; the full install adds LoRA training, image generation and extensions. Web UI on port 7860.
Strengths
Portable builds with all dependencies for CUDA, Vulkan, ROCm and CPU
Five loaders switchable without restarting
OpenAI and Anthropic-compatible API with tool calling and MCP servers
LoRA training and diffusers image generation in the same app
Weaknesses
Full install needs ~10 GB disk and PyTorch; portable build is GGUF only
Multi-user mode does not save chat histories; meant for small trusted teams
Docker needs per-GPU Dockerfile symlinks and manual .env edits
AGPL-3.0 license
GPU optional
Docker + Compose
Models: GGUF via llama.cpp and ik_llama.cpp, Transformers safetensors, EXL3 via ExLlamaV3, TensorRT-LLM
CPU-GPU hybrid inference and fine-tuning for very large MoE models
19.6k stars, Apache-2.0, last commit Oct 2026
KTransformers is a research framework for CPU-GPU heterogeneous inference and fine-tuning of large MoE models. Its kt-kernel package provides Intel AMX and AVX512/AVX2 INT4/INT8 kernels with NUMA-aware expert placement, so DeepSeek-V3/R1, Kimi K2.x and GLM-5.x run with hot experts on GPU and cold ones on CPU, served through SGLang. A LlamaFactory integration fine-tunes the same models with LoRA or full parameters.
Strengths
DeepSeek-R1 class models on one 24 GB GPU plus large host RAM
Day-0 support for DeepSeek-V4, Kimi K2.x, GLM-5.x and MiniMax-M3
LoRA and full fine-tuning of MoE models on 4x RTX 4090 via LlamaFactory
Ascend NPU, AMD ROCm and Intel Arc paths beyond NVIDIA
Weaknesses
Serving goes through SGLang (sglang-kt); the standalone framework is archived
DeepSeek-R1 example needs 382 GB DRAM alongside 24 GB VRAM
Fastest kernels need Intel AMX or AVX512; AVX2 support is newer
Research project; some docs and support channels are Chinese-only
NVIDIA inference server for TensorRT, PyTorch, ONNX and more over HTTP/gRPC
11.1k stars, BSD-3-Clause, last commit Oct 2026
Triton serves TensorRT, PyTorch, ONNX, OpenVINO, Python and RAPIDS FIL models over HTTP/REST and gRPC (KServe v2), with concurrent execution, dynamic and sequence batching, ensembles and Business Logic Scripting. NVIDIA ships it as NGC containers (2.73.0 / 26.09) for NVIDIA GPUs, x86 and ARM CPUs, Jetson and AWS Inferentia, with C and Java in-process APIs and a metrics endpoint.
Strengths
Serves TensorRT, PyTorch, ONNX, OpenVINO, Python and FIL models together
Dynamic and sequence batching, ensembles and BLS pipelines
HTTP/REST and gRPC (KServe v2) plus C and Java in-process APIs
Metrics for GPU utilization, throughput and latency
Weaknesses
No OpenAI-compatible endpoint in the README; clients speak KServe v2
Model repository and per-model config files are hand-written
Containers track NVIDIA's monthly NGC release cycle
GPU cluster manager that deploys models on vLLM, SGLang and TensorRT-LLM
5.8k stars, Apache-2.0, last commit Oct 2026
GPUStack is a GPU cluster manager that deploys models across on-prem, Kubernetes and cloud workers, configuring vLLM, SGLang, TensorRT-LLM or custom engines behind OpenAI-compatible APIs with auth, API keys and token metering. The server is one Docker container on port 80 and can run CPU-only; Linux workers join with a privileged Docker command. It supports NVIDIA, AMD, Ascend and six Chinese accelerator families.
Strengths
Multi-cluster: on-prem, Kubernetes and cloud GPUs under one server
Auto-selects and tunes vLLM, SGLang or TensorRT-LLM per model
Built-in auth, API keys, token metering, Grafana and Prometheus dashboards
SSH-accessible GPU instances on demand for fine-tuning
Weaknesses
Workers are Linux-only; macOS cannot be a worker, Windows needs WSL2
Worker container runs privileged with the Docker socket mounted
Cluster topology view is in the paid GPUStack Enterprise
Quick start assumes an NVIDIA GPU; other vendors need extra steps
GPU required
Compose
Models: catalog models (e.g. Qwen3.5-0.8B) via vLLM, SGLang, TensorRT-LLM, LLM, voice, image and video models
LLM and VLM serving toolkit with the TurboMind and PyTorch engines
8.1k stars, Apache-2.0, last commit Oct 2026
LMDeploy compresses and serves LLMs and VLMs with two engines: TurboMind (CUDA, persistent batching, blocked KV cache, AWQ W4A16, MXFP4) and a pure-Python PyTorch engine that also runs on Huawei Ascend. pip install lmdeploy adds an API server plus a proxy for multi-model, multi-machine serving. Models span Llama, Qwen3, DeepSeek-V3/V4, GLM-5, InternVL and Qwen3-VL.
Strengths
TurboMind engine with persistent batching, blocked KV cache and 4-bit AWQ inference
Online INT8/INT4 KV cache quantization and prefix caching usable together
Single-file executables that bundle llama.cpp with model weights
26.2k stars, custom license, last commit Oct 2026
llamafile packages llama.cpp and model weights into one executable using Cosmopolitan Libc, so a downloaded .llamafile runs on Linux, macOS, Windows and BSD across CPU architectures with no install and serves a local web UI and API. Since 0.10 it tracks upstream llama.cpp closely for newer models, and whisperfile applies the same packaging to speech-to-text. Maintained by Mozilla.ai.
Strengths
Single file, no installation, runs across OSes and CPU architectures
0.10 build system tracks upstream llama.cpp for recent model support
Can run external GGUF weights with the bare llamafile binary
whisperfile gives single-file transcription and translation
Weaknesses
Windows cannot run executables above 4 GB; larger models need external weights
0.10.x dropped some classic features; older releases remain for those
Pre-built llamafiles limited to Mozilla.ai's Hugging Face uploads
One model per file; not a multi-model server
GPU optional
Models: GGUF (bundled or external), e.g. Qwen3.5-0.8B
Rust server for embedding, reranker and classification models
5.1k stars, Apache-2.0, last commit Oct 2026
TEI is a Rust server from Hugging Face for embedding, reranker and sequence-classification models (BERT, XLM-RoBERTa, Nomic, Jina, GTE, Qwen3, ModernBERT, Gemma3) with token-based dynamic batching and Flash Attention. The router listens on port 3000 with /embed and OpenAI-compatible routes, gRPC, OpenTelemetry tracing and Prometheus metrics. Images cover CPU x86/arm64 and NVIDIA Turing through Blackwell.
Strengths
Token-based dynamic batching with Flash Attention, Candle and cuBLASLt
Small images and fast boot; no graph compilation step
Rerankers and classifiers served alongside embeddings
OpenTelemetry tracing, Prometheus metrics, API key auth, gRPC
Weaknesses
No Volta support; Turing image is experimental with Flash Attention off
GPU images need drivers compatible with CUDA 12.2 or higher
Only CamemBERT and XLM-RoBERTa for sequence classification
7B embedders such as Qwen3-Embedding-8B are flagged very expensive
OpenAI-compatible API server for running ExLlamaV3 models
1.5k stars, AGPL-3.0, last commit Oct 2026
TabbyAPI is a FastAPI server that loads and serves LLMs through the ExLlamaV3 backend, exposing an OpenAI-compatible HTTP API. It supports runtime model loading and unloading, HuggingFace downloads, embedding models, constrained output (JSON schema, regex, EBNF), and tool calling. The README describes it as a hobby project for a small user base, not for production servers.
Strengths
Continuous batching with paged attention on Nvidia Ampere and newer GPUs
Speculative decoding with draft models; JSON schema, regex and EBNF constraints
Load, unload and download models at runtime without restarting the server
Published Docker images for CUDA 12.8, CUDA 13 and ROCm
Weaknesses
README says it is not meant for production servers
Only EXL3 and FP16/BF16 models; no GGUF support listed
Rolling release with no tagged releases; dependencies may need reinstalling
AGPL-3.0 license may restrict use in some network services
One-command OpenAI-compatible endpoints for curated open LLMs
12.6k stars, Apache-2.0, last commit May 2026
OpenLLM serves open LLMs as OpenAI-compatible APIs with one command: pip install openllm, then openllm serve llama3.2:1b starts a vLLM-backed server on port 3000 with a chat UI at /chat. Models come as prebuilt Bentos from a curated repository (Llama 3.x and 4, Qwen2.5, Mistral, Phi-4, Gemma, DeepSeek R1) and the catalog states the GPU each needs. openllm deploy pushes the same Bento to BentoCloud.
Strengths
One command gives an OpenAI API plus /chat UI on port 3000
Catalog lists the required GPU per model tag (12 GB to 16x80 GB)
vLLM backend for serving
Same Bento deploys to Docker, Kubernetes or BentoCloud
Weaknesses
Every catalog model requires a GPU; no CPU-only entries
Custom model repositories must be public
Catalog tops out around Llama 3.3 and Qwen2.5; last commit 2026-05-29