#44th of 8 in Image and video

AI Toolkit

Fine-tuning suite for image, video and audio diffusion models, with GUI and CLI

Stars
12.3k
License
MIT
Last commit
Oct 2026
Language
Python

Overview

AI Toolkit trains LoRA and LoKr adapters for diffusion models, covering FLUX.1, FLUX.2, Qwen-Image, Wan 2.1/2.2, LTX-2 and SDXL, plus audio models like ACE-Step 1.5. Jobs are defined in YAML and run with `python run.py`, or started and monitored from a web UI on port 8675. The UI can be protected with an AI_TOOLKIT_AUTH token, and training can also run on Modal or RunPod.

Who it is for: Self-hosters fine-tuning LoRAs for image and video diffusion models

Strengths

  • Supports a wide range of image, edit, video and audio models in one tool
  • Same YAML configs work from the CLI or the web UI
  • Training resumes from the last checkpoint after interruption
  • Layer targeting via only_if_contains and ignore_if_contains, plus LoKr support

Weaknesses

  • Manual install requires an Nvidia GPU; Mac support is experimental
  • Install manager is labeled experimental by the author
  • Dataset images must be jpg, jpeg or png; webp has known issues
  • Ctrl+C during a checkpoint save can corrupt that checkpoint

What it needs

  • GPU required
  • Docker + Compose
  • Needs PyTorch, Node.js, Hugging Face
  • Models: FLUX.1, FLUX.2, Qwen-Image, Wan 2.1/2.2, LTX-2
  • port 8675

Also in Image and video

See all 8
Also in Image and video
RankProjectScore
1ComfyUINode-graph engine for diffusion image, video, audio and 3D models136.8k stars, GPL-3.075 out of 100
2MoneyPrinterTurboGenerates short videos from a topic with script, footage, voice and subtitles129.5k stars, MIT69 out of 100
3InvokeAICanvas-first web UI for Stable Diffusion and Flux image generation28.5k stars, Apache-2.062 out of 100
#5Kohya's GUIGradio GUI and CLI for Kohya diffusion training scripts12.6k stars, Apache-2.052 out of 100
#6Pixelle-VideoTopic-to-short-video pipeline built on ComfyUI workflows and TTS28.8k stars, Apache-2.051 out of 100