Systematically Testing Hallucinations
Wissen
LLMs hallucinate. They generate text that is grammatically correct, logically consistent, and convincingly worded -- but factually wrong. This isn't a bug in the traditional sense. It's an unavoidable property of probabilistic language models: they optimize for "probable-sounding next tokens," not "factually correct statements."
At the expert level, the relevant question isn't "Do LLMs hallucinate?" (yes, always) but: How do you systematically measure the hallucination rate? And how do you reduce it to an acceptable level?
Types of Hallucinations
1. Factual Hallucinations (Factual Confabulations) -- The model invents facts: cites never-published studies, provides fabricated statistics, describes nonexistent software features.
2. Logical Hallucinations (Reasoning Errors) -- Correct premises with wrong conclusion, inconsistency within a response, circular reasoning disguised as argumentation.
3. Contextual Hallucinations (Context Drift) -- Answers a different question than asked, adds information not in the given context, contradicts explicitly provided context.
4. Identity Hallucinations -- Claims capabilities it doesn't have, provides false information about its training, "remembers" conversations that never happened.
Benchmark Approaches
TruthfulQA: 817 questions designed to provoke common misconceptions. Answers checked against ground truth. Limitation: Static benchmark subject to data contamination.
HaluEval: Benchmark specifically for hallucination detection in various task types (QA, summarization, dialog).
RAGAS (RAG Assessment): Framework with metrics Faithfulness, Answer Relevancy, and Context Precision for evaluating RAG systems.
Custom Benchmarks: For production systems you need domain-specific test sets:
1. Collect 200-500 domain-specific questions with verified answers
2. Categorize by difficulty and hallucination risk
3. Automate evaluation (LLM-as-judge or rule-based)
4. Run the benchmark with every model change and prompt update
5. Track hallucination rate over time (monitoring)
Verstehen
Click a card to see the exampleTap a card to see the example
Automated Testing
LLM-as-Judge -- A second LLM evaluates the first's output for hallucinations. Sounds circular, but works surprisingly well when the judge model has access to ground truth, receives explicit evaluation criteria, and is stronger than the tested model.
judge_prompt = """
Given:
- Question: {question}
- Context (ground truth): {context}
- System's answer: {answer}
Rate the answer on a scale of 1-5:
1 = Completely hallucinated, no facts are correct
2 = Mostly hallucinated, individual facts correct
3 = Mixed, some hallucinations
4 = Mostly correct, minimal inaccuracies
5 = Fully correct and supported by context
Justify your rating point by point.
"""
Assertion-Based Testing -- Explicit rules for every answer: Must-contain (key facts), Must-not-contain (known misinformation), Consistency check (numbers/dates), Format check (JSON schema).
CI/CD Integration -- Hallucination tests belong in your pipeline: regression tests with prompt updates, full benchmark on model change, weekly sampling from production logs.
Grounding Strategies
1. RAG (Retrieval-Augmented Generation) -- Reduces hallucinations by 40-70%. But: the model can also misinterpret RAG context.
2. Citation Forcing -- The model must provide a source for every statement. Statements without sources are automatically filtered.
System prompt: "Support every factual claim with a source reference
from the provided context. Use the format [Source: Document X,
Section Y]. If you cannot find a source, explicitly write:
'No source available -- this statement is not verified.'"
3. Constrained Generation -- Structured output (JSON schema), enum fields, temperature 0 (reduces variability but does not eliminate hallucinations).
4. Multi-Model Verification -- Two or more models answer the same question independently. Only matching statements are accepted.
5. Human-in-the-Loop Verification -- For the highest reliability tier: a human reviews every output. Doesn't scale, but for critical applications often the only responsible option.
*Defense in Depth
No single strategy eliminates hallucinations. In practice, you combine multiple strategies: RAG + citation forcing + assertion tests + monitoring. Each layer reduces the error rate further.
Anwenden
A company deploys RAG and finds that the hallucination rate dropped by 50%. Yet the system still hallucinates on 15% of queries. What is the most sensible next step?
Reflect
Hallucinations are not a bug but an inherent property of probabilistic language models. Defense in depth -- combining citation forcing, assertion tests, and monitoring -- is the most effective approach. In the next section, we will look at the energy consumption and environmental impact of LLMs.