#44th of 8 in Image and video
AI Toolkit
Fine-tuning suite for image, video and audio diffusion models, with GUI and CLI
- Stars
- 12.3k
- License
- MIT
- Last commit
- Oct 2026
- Language
- Python
Overview
AI Toolkit trains LoRA and LoKr adapters for diffusion models, covering FLUX.1, FLUX.2, Qwen-Image, Wan 2.1/2.2, LTX-2 and SDXL, plus audio models like ACE-Step 1.5. Jobs are defined in YAML and run with `python run.py`, or started and monitored from a web UI on port 8675. The UI can be protected with an AI_TOOLKIT_AUTH token, and training can also run on Modal or RunPod.
Who it is for: Self-hosters fine-tuning LoRAs for image and video diffusion models
Strengths
- Supports a wide range of image, edit, video and audio models in one tool
- Same YAML configs work from the CLI or the web UI
- Training resumes from the last checkpoint after interruption
- Layer targeting via only_if_contains and ignore_if_contains, plus LoKr support
Weaknesses
- Manual install requires an Nvidia GPU; Mac support is experimental
- Install manager is labeled experimental by the author
- Dataset images must be jpg, jpeg or png; webp has known issues
- Ctrl+C during a checkpoint save can corrupt that checkpoint
What it needs
- GPU required
- Docker + Compose
- Needs PyTorch, Node.js, Hugging Face
- Models: FLUX.1, FLUX.2, Qwen-Image, Wan 2.1/2.2, LTX-2
- port 8675
Also in Image and video
See all 8| Rank | Project | Score |
|---|---|---|
| 1 | ComfyUINode-graph engine for diffusion image, video, audio and 3D models | 75 out of 100 |
| 2 | MoneyPrinterTurboGenerates short videos from a topic with script, footage, voice and subtitles | 69 out of 100 |
| 3 | InvokeAICanvas-first web UI for Stable Diffusion and Flux image generation | 62 out of 100 |
| #5 | Kohya's GUIGradio GUI and CLI for Kohya diffusion training scripts | 52 out of 100 |
| #6 | Pixelle-VideoTopic-to-short-video pipeline built on ComfyUI workflows and TTS | 51 out of 100 |