Zum Inhalt springen

LoRA & QLoRA

Knowledge

Full fine-tuning of an LLM means: All billions of parameters are retrained with new data. For Llama 4 Maverick with 400 billion parameters, you'd need multiple high-end GPUs with hundreds of gigabytes of VRAM and days of training time. That's unrealistic for most teams.

LoRA (Low-Rank Adaptation) solves this problem radically: Instead of modifying all parameters, small "adapter matrices" are placed alongside the existing weights. Only these adapters are trained -- typically 0.01% to 0.1% of total parameters.

How Does LoRA Work?

In a Transformer model, most parameters are stored in large weight matrices -- for example, in the attention layers. These matrices typically have dimensions like 4096 x 4096.

LoRA is based on a mathematical observation: The changes made to these matrices during fine-tuning have a low rank. This means the change to a 4096 x 4096 matrix can be approximated by two much smaller matrices:

Original matrix W: 4096 x 4096 = 16.7 million parameters

LoRA decomposition (rank r=16):
  Matrix A: 4096 x 16 = 65,536 parameters
  Matrix B: 16 x 4096 = 65,536 parameters
  Total: 131,072 parameters (0.8% of original)

During training, the original weights remain frozen. Only the small adapter matrices A and B are trained. During inference, the result is added: Output = W*x + B*A*x.

Why Does This Work So Well?

The intuition: When you fine-tune a model, you don't fundamentally change "how" it thinks. You only slightly adjust "what" it pays attention to. This adjustment can be expressed with very few parameters.

Analogy: Imagine an orchestra that has rehearsed a piece of music. For a new piece in the same genre, you don't need to retrain every musician. You give the conductor (LoRA adapter) new instructions, and the orchestra (base model) plays differently.

QLoRA: Even More Efficient

QLoRA (Quantized LoRA) goes one step further: The base model is quantized to 4-bit precision -- compressed -- before the LoRA adapters are applied.

MethodVRAM for 7B modelVRAM for 70B modelTrainable parameters
Full Fine-Tuning (fp16)~28 GB~280 GB100%
LoRA (fp16)~16 GB~160 GB0.01-0.1%
QLoRA (4-bit)~6 GB~48 GB0.01-0.1%

QLoRA makes it possible to fine-tune a 70-billion-parameter model on a single consumer GPU (48 GB). This has democratized fine-tuning.

Understanding

Where in the Model LoRA Actually Attaches

Before looking at the adapters themselves: LoRA does not hook into the model just anywhere, it attaches at clearly named places. The stations carrying the LoRA badge are the usual targets -- the attention projections are the default, the MLP layers are added when more capacity is needed.

Click a station -- you will see what happens to your text there.
Transformer block · 12×
AttentionShape of the data: 9 × 768
In plain words
Now every token may look at all the previous ones. The model derives three roles from each token: a question (query), a description (key) and a content (value). Where question and description match, a lot of that content flows back.
Technically
Scaled dot-product attention with a causal mask, split across 12 heads of 64 dimensions each. A softmax over the scores produces the weights -- exactly the ones in the attention visualisation.
LoRA freezes these weights and instead trains two small extra matrices alongside them -- usually on the attention projections (Q, K, V, O), often on the MLP layers as well. Everything else in the model stays untouched.
The numbers come from GPT-2 small: 124M parameters, 12 blocks, 12 attention heads, 768 dimensions, a vocabulary of 50,257 tokens. Larger models have more blocks, more heads and wider vectors -- the path through the model stays the same.

The LoRA Architecture Visualized

LoRA Architecture: Adapter-Based Fine-Tuning

Click a layer to see details. Green blocks indicate trainable LoRA adapters.

Base Model (e.g. Llama 4, 70B)

LoRA
Adapter
LoRA
Adapter

Parameter Comparison: Training Effort

Full Fine-Tuning70B parameters
LoRA70M parameters
QLoRA35M + 4-bit quantization

LoRA trains only ~0.1% of parameters. QLoRA combines this with 4-bit quantization for even lower memory requirements.

Hyperparameters You Need to Know

Rank (r): Determines the size of the adapter matrices. Typical values: 8, 16, 32, 64. Higher rank = more trainable parameters = more capacity, but also more memory and overfitting risk.

Alpha: A scaling factor that controls how strongly the LoRA adapters influence the output. Typical: alpha = 2 * r. Too high = unstable training, too low = barely any effect.

Target Modules: Which layers of the model get LoRA adapters? Typically the query and value matrices of the attention layers. Some setups also include the feed-forward layers.

# Typical LoRA configuration with PEFT
from peft import LoraConfig

config = LoraConfig(
    r=16,                          # Rank
    lora_alpha=32,                 # Scaling (2 * r)
    target_modules=[
        "q_proj", "v_proj",        # Attention Query + Value
        "k_proj", "o_proj",        # Attention Key + Output
    ],
    lora_dropout=0.05,             # Regularization
    bias="none",
    task_type="CAUSAL_LM"
)

Complete the LoRA configuration for a model where you want maximum adapter capacity with moderate memory usage:

config = LoraConfig( r=,
lora_alpha=,

Application

LoRA vs. QLoRA: When to Use Which?

Choose LoRA when:

  • You have access to professional GPUs (A100, H100)
  • Maximum training quality is important
  • You want to deploy the model at full precision later

Choose QLoRA when:

  • You're training on consumer hardware (RTX 4090, Apple Silicon)
  • Budget is limited
  • A minimal quality loss from quantization is acceptable

In practice: The quality difference between LoRA and QLoRA is negligible for most tasks (under 1% on common benchmarks), while the efficiency gain is enormous.

A startup wants to fine-tune a 13B model for medical text summarization. The team has an RTX 4090 (24 GB VRAM). Which method is realistic?

Reflect

LoRA and QLoRA have transformed fine-tuning from a privilege of large AI labs into an accessible tool. The core idea -- "change only what's necessary" -- is elegant and effective. In the next section, we'll look at which open-source models are the best candidates for fine-tuning.