MLOps for Fine-Tuning
Knowledge
Fine-tuning doesn't end when training is complete. In practice, most projects fail not because of model training, but because of everything around it: poor data quality, missing evaluation, no monitoring in production. MLOps (Machine Learning Operations) is the discipline that systematizes this "everything around it."
The Five Phases of a Fine-Tuning Project
Phase 1: Data Preparation
The foundation. Your data determines the quality of your model.
- Collect data: Internal documents, customer service logs, domain text corpora
- Clean data: Remove duplicates, standardize formatting, remove PII (personally identifiable information)
- Format data: Instruction-tuning format (prompt-completion pairs), chat format, or preference data (for DPO/RLHF)
- Train/Eval/Test split: Typically 80/10/10 or 90/5/5
{"messages": [
{"role": "system", "content": "You are an insurance advisor."},
{"role": "user", "content": "What does my homeowner's insurance cover for water damage?"},
{"role": "assistant", "content": "Your homeowner's insurance covers pipe water damage..."}
]}
Phase 2: Training
With prepared data and chosen base model, the actual training begins.
- Set hyperparameters: Learning rate (typical: 1e-4 to 2e-5), epochs (1-5), batch size, LoRA rank
- Start training: Hugging Face Trainer, Unsloth, or Axolotl
- Monitoring: Watch the loss curve -- is it steadily decreasing or oscillating?
- Checkpoints: Save intermediate states regularly
Phase 3: Evaluation
The finished model must prove it's better than the baseline.
- Automatic metrics: Perplexity, ROUGE (for summaries), BLEU (for translations), Pass@k (for code)
- Human evaluation: Experts rate output quality on a scale
- A/B comparison: Fine-tuned model vs. base model + Prompt Engineering
- Regression tests: Can the model still handle general tasks after fine-tuning? (Catastrophic Forgetting)
Phase 4: Deployment
The evaluated model goes into production.
- Model Registry: Model versioning (MLflow Model Registry, Hugging Face Hub)
- Serving infrastructure: vLLM, TGI (Text Generation Inference), NVIDIA Triton
- Quantization for inference: GPTQ, AWQ, or GGUF for efficient serving
- API layer: FastAPI, LitServe, or cloud provider endpoints
Phase 5: Monitoring
In production, the model must be continuously monitored.
- Performance metrics: Latency, throughput, error rate
- Quality metrics: User feedback, automated quality checks
- Drift detection: Are the input data changing? Is the output quality changing?
- Cost tracking: Token consumption, GPU utilization, cost per request
Understanding
Tooling Landscape
Weights & Biases (W&B): The de facto standard for experiment tracking. Logs hyperparameters, loss curves, metrics, and model artifacts. Free for individuals.
import wandb
wandb.init(project="fine-tuning-customer-service")
wandb.config.update({
"base_model": "meta-llama/Llama-4-Maverick",
"lora_r": 16,
"learning_rate": 2e-5,
"epochs": 3
})
# Training loop -- W&B logs automatically
trainer.train()
wandb.finish()
MLflow: Open-source alternative to W&B. Focus on Model Registry and deployment. Good for teams that don't want to send data to the cloud.
Hugging Face Hub: Not just model downloads, but also Model Registry and deployment. Integrates seamlessly with the Transformers ecosystem.
The Most Common Mistakes
- Too little evaluation effort: Teams train for days but evaluate for only 5 minutes. Invest at least 20% of project time in evaluation.
- No baseline comparison: Without a baseline, you don't know if fine-tuning even helped. Always test Prompt Engineering + RAG first.
- Ignoring Catastrophic Forgetting: Your model can lose general capabilities through fine-tuning. Always test general tasks too.
- No monitoring in production: Models degrade over time as usage patterns change. Without monitoring, you notice too late.
Your team fine-tuned a model and the eval loss is significantly lower than the base model's training loss. But in A/B testing with real users, the model performs worse than expected. What's the most likely cause?
Application
Minimal MLOps Stack for Getting Started
You don't need enterprise tooling right away. This stack is enough for your first projects:
- Data preparation: Python script + Pandas for cleaning, Hugging Face Datasets for formatting
- Training: Unsloth (fast) or Hugging Face Trainer (flexible) with LoRA/QLoRA
- Experiment tracking: Weights & Biases (free account)
- Evaluation: lm-eval-harness for automatic benchmarks + manual samples
- Deployment: vLLM on a cloud GPU or Hugging Face Inference Endpoints
- Monitoring: W&B Monitoring or simple logging + dashboards
The Workflow as a Checklist
- Training data collected and cleaned
- Train/Eval/Test split created
- Baseline performance measured with Prompt Engineering
- LoRA configuration set (rank, alpha, target modules)
- Training started and loss curve monitored
- Evaluation on eval set and manual samples
- A/B comparison: Fine-Tuned vs. Baseline
- Regression tests for general capabilities
- Model quantized and deployed
- Monitoring set up
Reflect
MLOps isn't an optional add-on, but the foundation for productive AI systems. Most teams invest 80% in training and 20% in everything else -- successful teams almost reverse this ratio. Data quality, systematic evaluation, and continuous monitoring distinguish a successful fine-tuning project from an expensive experiment.