Developer tools#Model training
LlamaFactory: one config tunes 100+ LLMs
hiyouga's Apache-2.0 framework fine-tunes 100+ LLMs and VLMs from one config, from LoRA to full-parameter, with 4-bit QLoRA on 7B at about 6 GB.
Project facts
GitHub Ecosystem- License
- Apache-2.0
- Language
- Python
- Stars
- 75,235
- Data checked
- 2026-09-30
Snapshot figures reflect the check date and may change over time.
When a local fine-tune stalls, it is rarely the GPU — it is the environment: data formats, chat templates, and out-of-memory errors that force a restart from scratch. LlamaFactory, an open-source project by hiyouga, collapses all of that into one config file: 100+ LLMs and VLMs share a single training entry point, and switching between LoRA, QLoRA, full-parameter and preference training means changing fields, not rewriting scripts. It ships under Apache-2.0, the paper appeared at ACL 2024, it had about 75.2k stars as of September 30, 2026, and Amazon, NVIDIA and Alibaba Cloud all appear on its users list; the WeChat account 开源软件社 recently walked through it in detail.

Core features
- Model coverage: the supported list runs from Llama, Qwen3, DeepSeek, Gemma to GLM and Phi, plus multimodal LLaVA, InternVL and Qwen3-VL and audio/video models; new releases land fast — Qwen3, Gemma 3, GLM-4.1V, InternLM 3 and MiniCPM-o-2.6 are all in the Day 0 support table.
- Training methods: pre-training, SFT, reward modeling, PPO, DPO, KTO, ORPO and SimPO are implemented, combinable with full-parameter, frozen, LoRA, QLoRA and OFT; memory-saving optimizers like GaLore, BAdam, APOLLO, Adam-mini and Muon are one line in the config.
- VRAM floor: per the official estimates, 16-bit LoRA on 7B needs about 16 GB, 4-bit QLoRA about 6 GB, and 2-bit fits in 4 GB (table below).
- Acceleration switches: FlashAttention-2, Unsloth, Liger Kernel and KTransformers plug in as flags —
use_unsloth: true,enable_liger_kernel: true— and training curves go to LlamaBoard, TensorBoard, Wandb, MLflow or SwanLab. - Zero-code entry:
llamafactory-cli webuiopens a web UI where train, evaluate and chat are all forms; no training code required. - After training: merge and export LoRA weights (an Ollama modelfile comes along), or serve an OpenAI-style API for other programs.
Typical use cases
- One 16–24 GB GPU and a private dataset: the individual or small team fine-tuning a 7B–14B model on their own domain.
- Experimenters comparing methods — LoRA vs DPO vs full-parameter is a config change, not a rewrite.
- Teaching: run the whole fine-tune loop point-and-click, with curves on TensorBoard or SwanLab.
Quick start
git clone --depth 1 https://github.com/hiyouga/LlamaFactory.git
cd LlamaFactory
pip install -e .
pip install -r requirements/metrics.txt
You need Python 3.11+, PyTorch 2.0+ (2.6.0 recommended) and transformers 4.49+. The official quick start then fine-tunes, chats and merges a Qwen3-4B LoRA:
llamafactory-cli train examples/train_lora/qwen3_lora_sft.yaml
llamafactory-cli chat examples/inference/qwen3_lora_sft.yaml
llamafactory-cli export examples/merge_lora/qwen3_lora_sft.yaml
Prefer no CLI at all? llamafactory-cli webui puts everything in forms, uv run llamafactory-cli webui builds an isolated environment for you, and the official Docker image ships a fixed Ubuntu 22.04 + CUDA 12.4 + PyTorch 2.6.0 stack.
Benchmarks
VRAM estimates from the official README (marked estimated; 7B/14B/70B shown):
| Method | Bits | 7B | 14B | 70B |
|---|---|---|---|---|
| Full (bf16/fp16) | 32 | 120 GB | 240 GB | 1200 GB |
| Freeze / LoRA / GaLore / APOLLO / BAdam / OFT | 16 | 16 GB | 32 GB | 160 GB |
| QLoRA / QOFT | 4 | 6 GB | 12 GB | 48 GB |
| QLoRA / QOFT | 2 | 4 GB | 8 GB | 24 GB |
The same table draws the line: full-parameter bf16 tuning of a 70B model needs 1200 GB — a single machine won’t carry it, so check this table before planning a full fine-tune.
Summary
LlamaFactory is the pragmatic starting point for individuals and small teams with a GPU and a dataset who would rather not spend a day on environment setup; teams with a mature training platform get more from it as a fast baseline-comparison tool. The honest caveats: it does nothing about data quality — cleaning and labeling will take longer than training; on Windows you install CUDA-enabled PyTorch manually and compile FlashAttention-2 yourself; models outside the support list need a hand-written chat template; and when HuggingFace is unreachable, switch to ModelScope with USE_MODELSCOPE_HUB=1. Apache-2.0, written in Python, and one rule that explains most “inconsistent results”: training and inference must use the same template (_nothink suffix for reasoning models).