Four open source projects currently dominate LLM fine-tuning: Unsloth, Axolotl, TRL, and LLaMA-Factory. They all wrap the same underlying PyTorch and Hugging Face stack. They differ on where they spend engineering effort.
In this article
Unsloth rewrites kernels. Axolotl composes parallelism strategies. TRL defines the trainer APIs the others build on. LLaMA-Factory optimizes for breadth of model coverage and zero-code operation.
This comparison covers three axes engineers actually hit: training throughput, peak VRAM, and multi-GPU scaling.
Understand each framework
- TRL is the reference implementation layer. It ships
SFTTrainer,DPOTrainer,GRPOTrainer,KTOTrainer,RewardTrainer, andRLOOTrainer. Axolotl and LLaMA-Factory both call into it. The current stable release line is v1.8.0. - Unsloth replaces parts of the modeling code with hand-written Triton kernels. Backpropagation steps are manually derived rather than autograd-generated. Hugging Face’s own writeup notes accuracy degradation is 0% versus standard QLoRA, because no approximations are introduced.
- Axolotl is a YAML-driven wrapper over Transformers, PEFT, TRL, Accelerate, and DeepSpeed. Its differentiator is composability of parallelism strategies, not kernel work.
- LLaMA-Factory is an ACL 2024 system demonstration paper with a Gradio web UI called LlamaBoard. The repository covers 100+ LLMs and VLMs.
Speed
Unsloth: kernel-level gains on a single GPU
Unsloth’s published benchmarks show 2x training speed for Llama 3.1 8B and Llama 3.3 70B. The setup used the Alpaca dataset, batch size 2, and gradient accumulation 4. QLoRA ran at rank 32 on all linear layers.
The MoE results are larger. Unsloth fine-tuned unsloth/gpt-oss-20b-BF16 on an NVIDIA B200. It reports 712.33 ms per step at 8K context, versus 5,226.86 ms for Transformers v5. That is a 7.3x gap. At 4K the gap is 4.82x, and at 1K it is only 1.37x.
The trend direction is model-dependent, and Unsloth’s docs scope this claim to gpt-oss. There, the speedup grows with sequence length, credited to Flex Attention and the MoE kernels.
Qwen3-30B-A3B on B200 runs the other way. Its reported speedup falls from 1.7x at 1K to 1.1x at 16K. Memory savings move the opposite direction, rising from about 2% to 15%.
Qwen3-30B-A3B on H100 reaches up to 1.77x. GLM-4.7-Flash on RTX PRO 6000 reaches 2.1x. A collaboration with AMD measured Llama-3.1-8B LoRA SFT at 2.07 s/step. TRL plus FlashAttention-2 took 2.87 s/step, a 1.39x gap with matching loss curves.
Axolotl: kernels borrowed, parallelism native
Axolotl added custom Triton kernels and autograd functions for LoRA in February 2025, explicitly citing Unsloth as inspiration. They are opt-in through lora_mlp_kernel, lora_qkv_kernel, and lora_o_kernel.
Recent release notes add SonicMoE LoRA support. It delivers up to 1.45x speedup and 30% memory reduction over a grouped_mm baseline. That figure is for Qwen3.5-35B-A3B 8-bit LoRA on a single H100 SXM.
Axolotl also ships FlashAttention 2/3/4, xFormers, Flex Attention, SageAttention, Liger Kernel, Cut Cross Entropy, and ScatterMoE.
TRL: the baseline everyone measures against
TRL is usually the reference point rather than the winner on raw single-GPU throughput. It compensates with breadth of memory and speed levers documented in Reducing Memory Usage and Speeding Up Training.
Those levers include packing, padding-free batching, truncation, Liger Kernel, and vLLM sleep mode for GRPO. TRL also has a first-party Unsloth integration, so the two are not mutually exclusive.
LLaMA-Factory: speed by delegation
LLaMA-Factory does not write its own kernels. It exposes other people’s work through config flags.
Setting use_unsloth: true activates the Unsloth patch. The project’s changelog reports 170% relative speed from that path. Unsloth’s long-sequence training is listed at 117% speed and 50% memory. It also supports enable_liger_kernel: true and FlashAttention-2 via flash_attn: fa2.
VRAM
Reported memory floors
Unsloth publishes a VRAM requirements table sorted by parameter count. It lists 6 GB for an 8B model in 4-bit QLoRA and 41 GB for 70B. LoRA at 16-bit costs 22 GB and 164 GB for the same models.
LLaMA-Factory’s README hardware table covers the same regime for 4-bit QLoRA. It lists 6 GB at 7B, 24 GB at 30B, and 48 GB at 70B. Full bf16 fine-tuning of 70B is listed at 600 GB.
Both tables describe minimums. Batch size, sequence length, and optimizer choice move the real number.
Context length is the sharper differentiator
Peak VRAM at a fixed context is less interesting than the maximum context a given VRAM budget allows. Unsloth’s context length benchmarks for Llama 3.1 8B QLoRA at rank 32 and batch size 1 are stark.
| GPU VRAM | Unsloth context | Transformers + FA2 context |
|---|---|---|
| 8 GB | 2,972 | OOM |
| 16 GB | 40,724 | 2,551 |
| 24 GB | 78,475 | 5,789 |
| 48 GB | 191,728 | 15,502 |
| 80 GB | 342,733 | 28,454 |
Unsloth attributes this to its gradient checkpointing algorithm combined with Apple’s Cut Cross Entropy. For Llama 3.3 70B on an 80 GB A100, it reports 89,389 tokens. The FA2 baseline reaches 6,916.
The MoE memory story
MoE training is where memory behavior has shifted most in 2026. Unsloth reports gpt-oss-20b fine-tuning inside 12.8 GB, while Qwen3-30B-A3B at 16-bit LoRA needs 63 GB.
Its B200 gpt-oss run used 47.43 GB at 8K context where Transformers v5 used 73.80 GB. At 16K, Transformers v5 went out of memory and Unsloth used 55.13 GB.
The mechanism is a split-LoRA formulation. PEFT materializes the LoRA delta across all experts before the MoE matmul. Unsloth reorders the operations instead, which is mathematically identical but avoids the materialization.
Axolotl attacks the same problem differently. Its MoE expert




