Fine-Tuning & Custom Models
In the Advanced path, you learned how to get the most out of pre-trained models using Prompt Engineering and RAG. But what do you do when that's not enough? When your model needs to master a specific domain language that no prompt can cover? When latency and costs of a RAG system are too high?
This is where Fine-Tuning comes in -- the third pillar of LLM customization.
The Three Pillars of LLM Customization
Prompt Engineering adapts a model's behavior at runtime. You provide instructions, examples, and context. The model itself remains unchanged.
Retrieval-Augmented Generation (RAG) extends the model's knowledge through external data sources. The model remains unchanged but receives relevant documents as context at runtime.
Fine-Tuning changes the model itself. You train it with your own data so that it internalizes new behavior, knowledge, or style. The result is a customized model that delivers the desired responses without additional context.
When to Use Which Approach?
| Criterion | Prompt Engineering | RAG | Fine-Tuning |
|---|---|---|---|
| Model change | None | None | Weights modified |
| Knowledge updatable | Instantly (change prompt) | Quickly (update documents) | Slowly (retraining needed) |
| Latency | Low | Medium (retrieval step) | Low (no retrieval needed) |
| Cost (setup) | Minimal | Medium (vector database) | High (GPU training) |
| Cost (inference) | Higher (long prompts) | Higher (context tokens) | Lower (shorter prompts) |
| Data needed | None | Documents | 100-10,000+ training examples |
| Domain expertise | Limited | Good (current documents) | Very good (internalized) |
*Learning Objective
After this module, you'll be able to evaluate when fine-tuning is the right strategy, understand LoRA/QLoRA as efficient training methods, know the most important open-source models as fine-tuning bases, and understand what an MLOps workflow for custom models looks like. Bloom's level: Evaluate and Create.
What to Expect
- When to Fine-Tune? -- Decision guide with interactive Decision Tree
- LoRA & QLoRA -- How to adapt a model using just 0.1% of its parameters
- Open-Source Models -- Llama 4, Mistral 3, DeepSeek V4, and Qwen 3.8 as base models
- MLOps Basics -- From data preparation to monitoring in production
Let's start -- with the crucial question: Do you even need fine-tuning?
Reflect
Fine-tuning is a powerful tool, but not always the right answer. In this module, you will learn when it pays off, how LoRA and QLoRA dramatically reduce costs, and what MLOps means in practice. Let us begin with the crucial question: Do you even need fine-tuning?