LoRA & QLoRA
Knowledge
Full fine-tuning of an LLM means: All billions of parameters are retrained with new data. For Llama 4 Maverick with 400 billion parameters, you'd need multiple high-end GPUs with hundreds of gigabytes of VRAM and days of training time. That's unrealistic for most teams.
LoRA (Low-Rank Adaptation) solves this problem radically: Instead of modifying all parameters, small "adapter matrices" are placed alongside the existing weights. Only these adapters are trained -- typically 0.01% to 0.1% of total parameters.
How Does LoRA Work?
In a Transformer model, most parameters are stored in large weight matrices -- for example, in the attention layers. These matrices typically have dimensions like 4096 x 4096.
LoRA is based on a mathematical observation: The changes made to these matrices during fine-tuning have a low rank. This means the change to a 4096 x 4096 matrix can be approximated by two much smaller matrices:
Original matrix W: 4096 x 4096 = 16.7 million parameters
LoRA decomposition (rank r=16):
Matrix A: 4096 x 16 = 65,536 parameters
Matrix B: 16 x 4096 = 65,536 parameters
Total: 131,072 parameters (0.8% of original)
During training, the original weights remain frozen. Only the small adapter matrices A and B are trained. During inference, the result is added: Output = W*x + B*A*x.
Why Does This Work So Well?
The intuition: When you fine-tune a model, you don't fundamentally change "how" it thinks. You only slightly adjust "what" it pays attention to. This adjustment can be expressed with very few parameters.
Analogy: Imagine an orchestra that has rehearsed a piece of music. For a new piece in the same genre, you don't need to retrain every musician. You give the conductor (LoRA adapter) new instructions, and the orchestra (base model) plays differently.
QLoRA: Even More Efficient
QLoRA (Quantized LoRA) goes one step further: The base model is quantized to 4-bit precision -- compressed -- before the LoRA adapters are applied.
| Method | VRAM for 7B model | VRAM for 70B model | Trainable parameters |
|---|---|---|---|
| Full Fine-Tuning (fp16) | ~28 GB | ~280 GB | 100% |
| LoRA (fp16) | ~16 GB | ~160 GB | 0.01-0.1% |
| QLoRA (4-bit) | ~6 GB | ~48 GB | 0.01-0.1% |
QLoRA makes it possible to fine-tune a 70-billion-parameter model on a single consumer GPU (48 GB). This has democratized fine-tuning.
Understanding
Where in the Model LoRA Actually Attaches
Before looking at the adapters themselves: LoRA does not hook into the model just anywhere, it attaches at clearly named places. The stations carrying the LoRA badge are the usual targets -- the attention projections are the default, the MLP layers are added when more capacity is needed.
The LoRA Architecture Visualized
LoRA Architecture: Adapter-Based Fine-Tuning
Click a layer to see details. Green blocks indicate trainable LoRA adapters.
Base Model (e.g. Llama 4, 70B)
Adapter
Adapter
Click a layer for details
Parameter Comparison: Training Effort
LoRA trains only ~0.1% of parameters. QLoRA combines this with 4-bit quantization for even lower memory requirements.
Hyperparameters You Need to Know
Rank (r): Determines the size of the adapter matrices. Typical values: 8, 16, 32, 64. Higher rank = more trainable parameters = more capacity, but also more memory and overfitting risk.
Alpha: A scaling factor that controls how strongly the LoRA adapters influence the output. Typical: alpha = 2 * r. Too high = unstable training, too low = barely any effect.
Target Modules: Which layers of the model get LoRA adapters? Typically the query and value matrices of the attention layers. Some setups also include the feed-forward layers.
# Typical LoRA configuration with PEFT
from peft import LoraConfig
config = LoraConfig(
r=16, # Rank
lora_alpha=32, # Scaling (2 * r)
target_modules=[
"q_proj", "v_proj", # Attention Query + Value
"k_proj", "o_proj", # Attention Key + Output
],
lora_dropout=0.05, # Regularization
bias="none",
task_type="CAUSAL_LM"
)
Complete the LoRA configuration for a model where you want maximum adapter capacity with moderate memory usage:
Application
LoRA vs. QLoRA: When to Use Which?
Choose LoRA when:
- You have access to professional GPUs (A100, H100)
- Maximum training quality is important
- You want to deploy the model at full precision later
Choose QLoRA when:
- You're training on consumer hardware (RTX 4090, Apple Silicon)
- Budget is limited
- A minimal quality loss from quantization is acceptable
In practice: The quality difference between LoRA and QLoRA is negligible for most tasks (under 1% on common benchmarks), while the efficiency gain is enormous.
A startup wants to fine-tune a 13B model for medical text summarization. The team has an RTX 4090 (24 GB VRAM). Which method is realistic?
Reflect
LoRA and QLoRA have transformed fine-tuning from a privilege of large AI labs into an accessible tool. The core idea -- "change only what's necessary" -- is elegant and effective. In the next section, we'll look at which open-source models are the best candidates for fine-tuning.