Developer tools#Model training

LlamaFactory: one config tunes 100+ LLMs

hiyouga's Apache-2.0 framework fine-tunes 100+ LLMs and VLMs from one config, from LoRA to full-parameter, with 4-bit QLoRA on 7B at about 6 GB.

Project facts

GitHub Ecosystem
Repositorygithub.com/hiyouga/LlamaFactory
License
Apache-2.0
Language
Python
Stars
75,235
Data checked
2026-09-30

Snapshot figures reflect the check date and may change over time.

When a local fine-tune stalls, it is rarely the GPU — it is the environment: data formats, chat templates, and out-of-memory errors that force a restart from scratch. LlamaFactory, an open-source project by hiyouga, collapses all of that into one config file: 100+ LLMs and VLMs share a single training entry point, and switching between LoRA, QLoRA, full-parameter and preference training means changing fields, not rewriting scripts. It ships under Apache-2.0, the paper appeared at ACL 2024, it had about 75.2k stars as of September 30, 2026, and Amazon, NVIDIA and Alibaba Cloud all appear on its users list; the WeChat account 开源软件社 recently walked through it in detail.

LlamaFactory’s web UI: the Train tab with dataset selection, learning rate and epochs, plus LoRA, GaLore and APOLLO configuration sections below

Core features

  • Model coverage: the supported list runs from Llama, Qwen3, DeepSeek, Gemma to GLM and Phi, plus multimodal LLaVA, InternVL and Qwen3-VL and audio/video models; new releases land fast — Qwen3, Gemma 3, GLM-4.1V, InternLM 3 and MiniCPM-o-2.6 are all in the Day 0 support table.
  • Training methods: pre-training, SFT, reward modeling, PPO, DPO, KTO, ORPO and SimPO are implemented, combinable with full-parameter, frozen, LoRA, QLoRA and OFT; memory-saving optimizers like GaLore, BAdam, APOLLO, Adam-mini and Muon are one line in the config.
  • VRAM floor: per the official estimates, 16-bit LoRA on 7B needs about 16 GB, 4-bit QLoRA about 6 GB, and 2-bit fits in 4 GB (table below).
  • Acceleration switches: FlashAttention-2, Unsloth, Liger Kernel and KTransformers plug in as flags — use_unsloth: true, enable_liger_kernel: true — and training curves go to LlamaBoard, TensorBoard, Wandb, MLflow or SwanLab.
  • Zero-code entry: llamafactory-cli webui opens a web UI where train, evaluate and chat are all forms; no training code required.
  • After training: merge and export LoRA weights (an Ollama modelfile comes along), or serve an OpenAI-style API for other programs.

Typical use cases

  • One 16–24 GB GPU and a private dataset: the individual or small team fine-tuning a 7B–14B model on their own domain.
  • Experimenters comparing methods — LoRA vs DPO vs full-parameter is a config change, not a rewrite.
  • Teaching: run the whole fine-tune loop point-and-click, with curves on TensorBoard or SwanLab.

Quick start

git clone --depth 1 https://github.com/hiyouga/LlamaFactory.git
cd LlamaFactory
pip install -e .
pip install -r requirements/metrics.txt

You need Python 3.11+, PyTorch 2.0+ (2.6.0 recommended) and transformers 4.49+. The official quick start then fine-tunes, chats and merges a Qwen3-4B LoRA:

llamafactory-cli train examples/train_lora/qwen3_lora_sft.yaml
llamafactory-cli chat examples/inference/qwen3_lora_sft.yaml
llamafactory-cli export examples/merge_lora/qwen3_lora_sft.yaml

Prefer no CLI at all? llamafactory-cli webui puts everything in forms, uv run llamafactory-cli webui builds an isolated environment for you, and the official Docker image ships a fixed Ubuntu 22.04 + CUDA 12.4 + PyTorch 2.6.0 stack.

Benchmarks

VRAM estimates from the official README (marked estimated; 7B/14B/70B shown):

Method Bits 7B 14B 70B
Full (bf16/fp16) 32 120 GB 240 GB 1200 GB
Freeze / LoRA / GaLore / APOLLO / BAdam / OFT 16 16 GB 32 GB 160 GB
QLoRA / QOFT 4 6 GB 12 GB 48 GB
QLoRA / QOFT 2 4 GB 8 GB 24 GB

The same table draws the line: full-parameter bf16 tuning of a 70B model needs 1200 GB — a single machine won’t carry it, so check this table before planning a full fine-tune.

Summary

LlamaFactory is the pragmatic starting point for individuals and small teams with a GPU and a dataset who would rather not spend a day on environment setup; teams with a mature training platform get more from it as a fast baseline-comparison tool. The honest caveats: it does nothing about data quality — cleaning and labeling will take longer than training; on Windows you install CUDA-enabled PyTorch manually and compile FlashAttention-2 yourself; models outside the support list need a hand-written chat template; and when HuggingFace is unreachable, switch to ModelScope with USE_MODELSCOPE_HUB=1. Apache-2.0, written in Python, and one rule that explains most “inconsistent results”: training and inference must use the same template (_nothink suffix for reasoning models).