#1313th of 20 in Model serving
KTransformers
CPU-GPU hybrid inference and fine-tuning for very large MoE models
- Stars
- 19.6k
- License
- Apache-2.0
- Last commit
- Oct 2026
- Last release
- Sep 2026
Overview
KTransformers is a research framework for CPU-GPU heterogeneous inference and fine-tuning of large MoE models. Its kt-kernel package provides Intel AMX and AVX512/AVX2 INT4/INT8 kernels with NUMA-aware expert placement, so DeepSeek-V3/R1, Kimi K2.x and GLM-5.x run with hot experts on GPU and cold ones on CPU, served through SGLang. A LlamaFactory integration fine-tunes the same models with LoRA or full parameters.
Who it is for: researchers running huge MoE models on limited GPUs
Strengths
- DeepSeek-R1 class models on one 24 GB GPU plus large host RAM
- Day-0 support for DeepSeek-V4, Kimi K2.x, GLM-5.x and MiniMax-M3
- LoRA and full fine-tuning of MoE models on 4x RTX 4090 via LlamaFactory
- Ascend NPU, AMD ROCm and Intel Arc paths beyond NVIDIA
Weaknesses
- Serving goes through SGLang (sglang-kt); the standalone framework is archived
- DeepSeek-R1 example needs 382 GB DRAM alongside 24 GB VRAM
- Fastest kernels need Intel AMX or AVX512; AVX2 support is newer
- Research project; some docs and support channels are Chinese-only
What it needs
- GPU required
- Docker
- Needs SGLang (serving), LLaMA-Factory (fine-tuning)
- Models: DeepSeek-V3/R1/V4-Flash, Kimi K2 to K2.6, GLM-5 to 5.3, MiniMax-M2.x/M3, Qwen3-MoE, Qwen3-Next
Also in Model serving
See all 20| Rank | Project | Score |
|---|---|---|
| 1 | LocalAIOne OpenAI-compatible server for text, speech, image and video models | 82 out of 100 |
| 2 | llama.cppC/C++ inference engine serving GGUF models over an OpenAI-compatible API | 79 out of 100 |
| 3 | vLLMHigh-throughput LLM serving engine with OpenAI and Anthropic APIs | 75 out of 100 |
| #4 | OllamaRuns open-weight models locally behind a CLI and REST API | 74 out of 100 |
| #5 | colibriC inference engine that runs huge MoE models by streaming experts from disk | 70 out of 100 |