AI Hallucinations: How to Measure and Reduce Error Rates

11 min read

356
AI Hallucinations: How to Measure and Reduce Error Rates

Measuring Hallucination Error

AI hallucinations show up when a model generates content that does not match the evidence available to it. In health contexts, that mismatch can be subtle: a medication name may be correct while the dosing interval, contraindication, or monitoring parameter is wrong. A second failure mode is fabricated citations, where the text sounds plausible but the referenced guideline cannot be verified. A third mode is “reasoning drift,” where the model follows a chain of logic that ignores key constraints from the prompt or the provided documents.

To measure hallucinations, you need a test set and a scoring rule. The test set should include prompts that resemble real use, such as “summarize the risks of drug X for kidney disease” or “compare two treatment options for mild hypertension.” The scoring rule should separate “unsupported” from “incorrect,” because both matter but require different fixes. In my experience reviewing quality reports, teams often track only overall accuracy and miss the category of unsupported claims, which is where many health harms begin.

Start by defining what counts as an error. For example, an output sentence is an error if it contradicts a trusted source, if it states a medical fact without a source when sources are required, or if it invents a guideline, trial, or lab value. Then decide whether you score at the sentence level, claim level, or document level. Sentence-level scoring is slower but catches partial failures; document-level scoring is faster but hides which part went wrong.

Once you have a scoring rule, you can compute an error rate. A common choice is the fraction of scored claims that fail your criteria. If you score 200 claims and 18 fail, the hallucination error rate is 9%. Track this rate across model versions and across prompt styles, because small changes in wording can shift the balance between unsupported and incorrect content.

Common Failure Points

People often treat hallucinations as a single problem, but the causes differ. A model may hallucinate because it lacks the needed information, because the prompt asks for a detail that the model cannot infer reliably, or because the model is not constrained to use provided sources. In health reading, the most frequent dependency is retrieval: if the system retrieves outdated or irrelevant material, the model can confidently restate it.

Another dependency is the evaluation setup. If you test only “easy” prompts that match common training patterns, the measured error rate will look better than it is for your actual questions. If you test only short answers, you may miss failure modes that appear when the model produces long structured outputs like “risk table” or “step-by-step plan.” A third dependency is the scoring rubric. Rubrics that allow vague language to pass, such as “generally safe,” tend to undercount errors.

Many users also get tripped up by the difference between fluency and correctness. A model can generate coherent text that still violates clinical constraints, such as contraindications, dosing adjustments, or monitoring requirements. When the output includes numbers, the risk rises because numbers are easy to fabricate and hard to verify without a source. I once saw an internal audit where fabricated lab thresholds accounted for a minority of claims but a majority of high-severity errors.

Finally, “helpful formatting” can hide mistakes. If the model outputs a table, the reader may assume the table is grounded. If the model outputs bullet points that look like guideline language, readers may treat them as citations. Your measurement should therefore include a check for source grounding when the workflow expects it.

Reducing Errors With Controls

Define Claims And Rubrics

Write a rubric before you measure. Split outputs into atomic claims, then label each claim as correct, incorrect, unsupported, or ambiguous. Ambiguous claims are those that could be correct but lack enough context to verify, such as “monitor closely” without specifying what to monitor. A practical rubric for health-related text often treats “unsupported” as an error when the user expects evidence-based statements.

Use a small pilot set first, then refine the rubric until two reviewers agree on labels. If you see low agreement, the rubric is too subjective. In one team workflow I reviewed (tooling: a simple spreadsheet plus a two-person label pass), tightening the definition of “unsupported” reduced label disagreements after the first 30 samples.

Measure Error Rates Repeatedly

Run the same test prompts across model versions and prompt templates. Track at least three metrics: unsupported-claim rate, contradiction rate (claims that conflict with trusted sources), and high-severity rate (errors that could cause harm if acted on). High-severity examples include wrong contraindications, wrong dosing frequency, or missing “seek urgent care” guidance when symptoms suggest emergency risk.

For realistic numbers, many teams start with 50–200 prompts and accept wide confidence intervals. If you want tighter estimates, increase sample size and keep the prompt set stable. A simple approach is to compute error rate per 50 prompts and watch whether it moves by more than a few percentage points after changes. If it swings wildly, the test set is too small or too heterogeneous.

Ground With Retrieval And Citations

When your workflow can retrieve documents, require the model to use retrieved text. A typical pattern is: retrieve relevant guideline sections, then ask the model to answer using only those sections, citing the retrieved snippets. If the system cannot retrieve enough context, the model should respond with “insufficient evidence in provided sources” rather than filling gaps.

Retrieval quality matters more than prompt wording. If the retriever returns a paragraph about general safety but misses kidney dosing guidance, the model can still hallucinate the missing dosing adjustment. A practical check is to score “evidence coverage,” meaning the retrieved snippets contain the key facts needed for each claim. If coverage is low, your error rate will stay high even with better prompting.

Add Human Review For High-Risk Outputs

Human review should target the highest-risk categories, not every sentence. A common triage rule is: if the output includes medication changes, dosing, contraindications, or emergency guidance, route it to a reviewer. If the output is purely educational and explicitly framed as general information, you can accept a lower review threshold.

Use a review checklist that mirrors your rubric: verify each numeric value, verify each contraindication, and verify that any “based on guideline” statement can be traced to a source. This review step is where many teams catch fabricated citations and reasoning drift. It also gives you labeled data to improve your evaluation loop.

Case Examples For Audits

Scenario A: Medication risk summary with missing sources. A user asks an AI to summarize “metformin risks for people with reduced kidney function.” The system returns a structured list of risks and monitoring steps but does not cite sources. During evaluation, reviewers split the output into 60 claims. They label 12 as unsupported and 3 as incorrect due to wrong monitoring frequency. After switching the workflow to require retrieval of kidney dosing guidance and to cite retrieved snippets, the unsupported-claim rate drops from 20% to 8%, while the incorrect-claim rate drops from 5% to 2%. The remaining errors cluster around edge cases like acute illness wording, where retrieval coverage is weaker.

Scenario B: Comparing two treatment options with table output. A user requests a comparison table for “two antihypertensive options for mild hypertension,” expecting a risk/benefit breakdown. The model produces a table with side effects and a “who should avoid” column. In scoring, reviewers find that 9 out of 40 table entries are unsupported, even though the narrative paragraph around the table sounds careful. The fix is not only prompting; the team adds a rule that every table cell must be traceable to a retrieved snippet or marked “not covered in provided sources.” After the change, unsupported table entries fall to 2 out of 40, and reviewers spend less time debating whether the model “meant” something.

Checklist For Decision Support

Goal What To Measure How To Reduce When To Escalate
Grounding Unsupported-claim rate Require citations to retrieved text; block “fill in gaps” High unsupported rate or missing evidence coverage
Correctness Contradiction rate vs trusted sources Tighten rubric; add retrieval for key constraints Any wrong contraindication or wrong numeric value
Safety High-severity error rate Route high-risk outputs to human review Emergency guidance, dosing changes, or treatment escalation
Stability Error rate across prompt variants Lock prompt template; test paraphrases Large swings after minor wording changes

Step-by-step checklist for a small audit run:

  1. Collect 50–200 prompts that match your real questions, including ones that require numbers or contraindications.
  2. Split outputs into claims and label each claim using a rubric that distinguishes unsupported from incorrect.
  3. Compute three rates: unsupported-claim rate, contradiction rate, and high-severity rate.
  4. Repeat the same test after each change to retrieval, prompt template, or output format.
  5. Review the top 10% worst outputs to identify which dependency failed (retrieval coverage, rubric ambiguity, or missing context).
  6. Lock the workflow for production use only after the error rates stop moving in the direction you do not want.

Common Mistakes That Skew Results

One frequent mistake is scoring only “final answers” while ignoring intermediate text like rationales. If the model includes a rationale that invents a guideline, the final answer may still look correct, yet the rationale can mislead users who read it. Your scoring should cover the parts users actually rely on.

A second mistake is using a trusted-source set that does not match the question scope. For example, evaluating a pediatric dosing claim against adult-only guidance can label correct statements as incorrect. Keep the evaluation sources aligned with the population and context in the prompt.

A third mistake is letting the model’s own citations decide correctness. If the model says “according to guideline X,” but guideline X is not present in your evidence set, the claim should be labeled unsupported. I have seen audits where reviewers treated the model’s citation text as proof, then later discovered the cited document did not contain the stated recommendation.

A fourth mistake is measuring with too few samples. With 20 prompts, a single bad output can swing the error rate by 5–10 percentage points. That makes it easy to overreact to changes that are just sampling noise. Increase sample size or run multiple batches across weeks.

A fifth mistake is confusing “less hallucination” with “better safety.” A workflow can reduce unsupported claims while still producing wrong numeric values. Track contradiction and high-severity rates, not only unsupported claims.

FAQ

What counts as a hallucination in health text?

A hallucination is a claim that is unsupported by the provided evidence, contradicts trusted sources, or invents details like guideline names, trial results, or numeric thresholds that cannot be verified.

How do I measure hallucination error rates without special tools?

Use a spreadsheet to split outputs into claims, label each claim with a rubric, and compute rates such as unsupported-claim rate and contradiction rate across a fixed prompt set.

How many test prompts do I need for a useful estimate?

For early tuning, 50–200 prompts can show directionally useful changes. For stable comparisons between versions, larger batches reduce sampling noise, especially for rare high-severity errors.

Do citations from the model reduce hallucinations automatically?

Citations reduce hallucinations only when the cited content is actually present in your evidence set and the system enforces grounding. Model-generated citation text alone does not prove correctness.

How can retrieval fail and still produce fluent answers?

Retrieval can return irrelevant or incomplete sections, so the model fills missing constraints with plausible guesses. This often shows up as unsupported claims or wrong numeric values even when the narrative sounds careful.

Author's Insight

Hallucination measurement works best when you treat outputs as collections of verifiable claims rather than as a single “answer.” A rubric that separates unsupported from incorrect claims gives you actionable failure categories. Retrieval-based workflows tend to reduce unsupported claims when the system enforces grounding to retrieved text, but they can still produce contradictions if retrieval misses key constraints. I recommend tracking unsupported-claim rate, contradiction rate, and high-severity rate together, then reviewing the worst outputs to identify which dependency failed.

For a practical starting point, run a small audit on a fixed prompt set, label 100–200 claims, and compare error rates before and after each change. If you see improvements only in one metric, the remaining errors may still be the ones that matter most for safety. I also suggest versioning your prompt template and test set; I have seen teams change wording and then lose the ability to explain why the error rate moved.

Key Takeaways

  • Define hallucination types as unsupported, incorrect, and ambiguous claims, then score at the claim level.
  • Measure multiple rates: unsupported-claim rate, contradiction rate, and high-severity error rate.
  • Reduce errors with evidence grounding, retrieval coverage checks, and output rules that block “fill in gaps.”
  • Use human review for high-risk outputs, and route based on the same rubric you use for measurement.
  • Avoid misleading results by aligning evaluation sources to the prompt scope and using enough samples to reduce noise.

Was this article helpful?

Your feedback helps us improve our editorial quality

Latest Articles

AI Tools 18.09.2026

Structured Outputs: Getting Reliable JSON From AI

This guide explains how to get reliable JSON from AI systems when you need machine-readable outputs for health-related workflows. It covers why models sometimes return malformed JSON, how structured output features work, and what to test in prompts and validators. You’ll learn practical patterns for schema design, error handling, and verification, plus examples of anonymized extraction tasks and a checklist to reduce failures.

Read » 417
AI Tools 12.08.2026

The Best AI Tool Stack for Solo Founders

This article explains how solo founders can build a practical AI tool stack for writing, research, customer support, and internal operations without creating security or compliance gaps. It covers common failure points, the supporting tools behind each workflow, and realistic outcomes you can measure. You’ll get example setups, a decision checklist, and a FAQ focused on privacy, data handling, and cost control.

Read » 306
AI Tools 30.09.2026

Prompt Caching: When It Cuts AI Costs

Prompt caching reduces repeated inference work by reusing previously computed prompt states. This guide is for people evaluating AI tools for writing, support, or analysis who want lower costs without losing accuracy. You’ll learn how caching works, what breaks it, how to measure savings, and how to design prompts and workflows that reuse context safely. Includes realistic examples, a decision checklist, and common mistakes to avoid.

Read » 300
AI Tools 25.08.2026

RAG vs Long Context: Which Works Better for Documents?

This article explains how Retrieval-Augmented Generation (RAG) and long-context prompting handle document-heavy tasks. It’s for readers evaluating document Q&A, policy search, and report drafting systems in health and other regulated settings. You’ll learn where each approach fails, how to test them with measurable checks, and how to choose chunking, retrieval, and context windows without guessing. Includes examples, a decision checklist, and common mistakes to avoid.

Read » 182
AI Tools 31.08.2026

AI Agents vs Workflows: When Should You Use Each?

Explore how AI agents and workflow automation differ in real-world tasks, with examples from customer support, research, and operations. It’s for readers who want reliable, testable automation rather than vague promises. You’ll learn how agents decide and act, how workflows route and transform data, what can fail, and how to choose based on risk, data access, and audit needs. Includes a decision checklist, common mistakes, and practical evaluation steps.

Read » 508
AI Tools 24.09.2026

AI Hallucinations: How to Measure and Reduce Error Rates

AI hallucinations are confident-sounding mistakes where a model produces text that is wrong, fabricated, or unsupported. This article is for readers who use AI for health-related reading, triage, or documentation and want measurable quality. You’ll learn how to define hallucination types, measure error rates with repeatable tests, and reduce errors using retrieval, prompting controls, and human review. Practical examples show how to audit outputs and avoid common traps.

Read » 356