When to Fine-Tune?
Knowledge
The most common question in practice isn't "How do I fine-tune?" but "Do I even need fine-tuning?" The answer depends on your specific use case -- and the wrong decision either costs too much money or delivers poor results.
The Decision Criteria
Prompt Engineering is sufficient when:
- Your task works well with a few examples (Few-Shot)
- You want to flexibly switch between different tasks
- You don't have your own training data
- The task is generic (summarization, translation, classification)
RAG is better when:
- Your knowledge changes frequently (product catalogs, documentation, news)
- You need traceable source references
- The data volume is large (thousands of documents)
- Timeliness is critical
Fine-Tuning is necessary when:
- The model must consistently maintain a specific style or tone
- Domain-specific terminology or reasoning is required
- Latency and per-request costs need to be minimized
- You need a small, specialized model instead of a large general-purpose model
- Prompt Engineering can't reliably achieve the required quality
Data Requirements
Fine-Tuning is data-driven. Without high-quality training data, every project fails.
| Data Volume | Suitable For | Quality Requirement |
|---|---|---|
| 50-200 examples | Style adaptation, formatting | Very high -- every example counts |
| 200-1,000 examples | Domain adaptation, terminology | High -- consistent quality |
| 1,000-10,000 examples | New capabilities, complex reasoning | Medium -- diversity matters more than perfection |
| 10,000+ examples | Full fine-tuning, specialist models | Automated QA processes needed |
Cost Comparison
A realistic cost comparison for a customer service chatbot with 10,000 requests per day:
Prompt Engineering (GPT-5): ~500 tokens system context per request = ~$150/month for inference.
RAG (GPT-5-mini + vector database): ~200 tokens context + retrieval = ~$80/month for inference + $50/month infrastructure.
Fine-Tuning (Llama 4 Maverick, LoRA): One-time training ~$50 + ~$30/month for self-hosting (or pay-per-token via providers).
Inference costs drop dramatically with fine-tuning because shorter prompts are sufficient and smaller models can be deployed.
Understanding
The Decision Tree in Practice
Use the interactive decision tree to find the right strategy for your use case:
Decision Tree: Prompt Engineering vs RAG vs Fine-Tuning
How much of your own training data do you have?
A fintech startup wants to train its chatbot so that it correctly uses regulatory terminology (MiFID II, KYC, AML) and always responds in formal banking language. Which approach is best?
Application
The Hybrid Strategy
In practice, the answer is often not "either/or" but a combination:
- Prompt Engineering as a baseline -- quickly tested, immediately usable
- RAG for current knowledge -- integrate documents, reference sources
- Fine-Tuning for specialization -- when the baseline isn't enough
You'll find this triad in almost every successful enterprise AI project. A fine-tuned model that additionally accesses current data via RAG and is configured for the specific task via system prompt.
Your team has fine-tuned a model on medical terminology. Now it needs to consider current treatment guidelines that change quarterly. What do you recommend?
Reflect
The decision between Prompt Engineering, RAG, and Fine-Tuning isn't just a technical question -- it's a business decision. Costs, data quality, latency requirements, and long-term maintainability determine the right approach. And most of the time, the best solution is a combination of all three pillars.