Monology
Back to Blog
AI & Machine LearningFeatured

We Measured Our Chatbot's Hallucination Rate. Here's the Real Number, and the One Answer It Got Wrong

We ran a hallucination evaluation on our own knowledge base: 65 questions, 1 unsupported claim in 49 answered (2.0%), and 15 of 15 should-refuse questions correctly refused. Here's the method, the follow-up question it got wrong, and the caveats that come with a sample this size.

AI Systems Architect

7 min read
#Hallucination#AI Evaluation#Grounding#Insurance#Compliance
Featured image for We Measured Our Chatbot's Hallucination Rate. Here's the Real Number, and the One Answer It Got Wrong

In insurance, the AI failure that should worry you isn’t “I don’t know.” A customer who hears that gets passed to a person, and nothing bad happens. The failure that should worry you is a confident, fluent, wrong answer, like “Yes, that’s covered,” from a chatbot carrying your brand, about a policy that says otherwise.

That kind of answer is what people mean by a hallucination. Monology is designed to avoid them. So we measured how often our own assistant actually produces one, and we’re publishing the result, including the answer it got wrong.

How we built against hallucination

Two design choices do most of the work. (We cover the mechanics in more depth in why AI chatbots hallucinate.)

Grounding. The assistant doesn’t answer from general knowledge. For every question, it first retrieves the relevant passages from the knowledge base it has been given, such as policy wordings, product guides and FAQs. It then answers only from those passages and cites where each answer came from. If the knowledge base doesn’t say it, the assistant shouldn’t say it either.

Abstain. When the retrieved passages don’t contain the answer, the assistant should say it can’t answer, rather than fill the gap with a plausible guess. In a regulated setting, a refusal is an acceptable outcome. A guess isn’t.

Why we measured instead of claiming

It would be easy to stop there and say the system has “near-zero hallucinations by design.” But “by design” describes what we intended, not what happens. Grounding reduces hallucinations; it doesn’t guarantee their absence. Retrieval can pull the wrong passage. The model can blend two passages together. A follow-up question can lose its context.

If you’re evaluating a vendor for coverage questions, you can’t take an intention to your compliance committee. You can take a measured number with its method and limits attached. So we built an evaluation and ran it.

How we tested

  • Knowledge base: Monology’s own product knowledge base, covering our plans, features and integrations. This is not an insurer’s policy documents. More on why that matters below.
  • Questions: 65 in total. 50 were answerable from the knowledge base. 15 were adversarial, meaning the correct response was to refuse because the knowledge base doesn’t support an answer.
  • Judging: a separate language model acted as the judge and scored each answer. It ran at temperature 0 to keep its scoring as consistent as possible. We also used embedding similarity, an automated measure of how closely an answer’s meaning matches the expected answer.
  • Run: a single clean run with 0 errors.

The results

Every figure below comes from that one setup: our knowledge base, 65 questions, one run.

MetricWhat it measuresResult (our KB, 65 questions)
Hallucination rateAnswered questions containing a claim not supported by the retrieved content2.0% (1 of 49 answered)
Answer accuracyAnswers judged correct98%
Correct refusalAdversarial questions correctly refused100% (15 of 15)
Over-refusalAnswerable questions wrongly refused2% (1 case)
Citation accuracyCitations pointing to the right source100%
Retrieval hit@5Expected passage among the top five retrieved86%
Latency (median / 95th percentile)Time to answer2.5s / 3.7s

In plain terms: on our knowledge base, the assistant answered 49 questions and made one unsupported claim. It refused all 15 questions it should have refused, and it wrongly refused one question it should have answered.

One figure needs a note. Retrieval hit@5 is 86%, so for some questions the passage we expected wasn’t among the top five retrieved. Accuracy is higher than that. That can happen when the same fact appears in more than one place in a knowledge base, but we haven’t verified that this is the explanation here. Retrieval is the metric with the most room to improve.

The one answer it got wrong

The hallucination came from a follow-up question. A plan had already been discussed in the conversation, and the next question was: “How many chatbots does it include?”

The assistant replied that “the Professional plan includes 3 chatbots.” The retrieved context didn’t support that statement. The assistant had lost track of which plan “it” referred to, and it answered confidently anyway.

Notice what kind of error this is. It isn’t a fabricated fact in response to a standalone question. It’s the assistant losing the thread across turns. Standalone questions are comparatively easy to ground: retrieve, answer, cite. Follow-ups are harder, because the system has to work out what “it” refers to from the conversation before retrieval can find the right passage. Get that step wrong and the answer can be about the wrong thing entirely.

Now translate that to insurance. A customer asks about a family floater, then follows up with “Does it cover maternity?” An assistant that loses track of “it” could answer for a different policy. An insurer can’t accept that failure. It’s why we now treat multi-turn context tracking, not single-question accuracy, as the hard part of this problem.

The answer it shouldn’t have refused

The second failure went the other way. Asked “Do you integrate with Zendesk?”, the assistant declined to answer, even though the knowledge base covers it.

That’s the safer direction to fail in, because nobody was told anything false. But it’s still a failure. In production, a customer with a legitimate question would have gone without an answer. The lesson is that the abstain path can be too cautious. An over-refusal like this can come from retrieval missing the passage or from the threshold for abstaining being too strict, and the two need different fixes. The goal isn’t to refuse as often as possible. It’s to refuse exactly when the documents don’t support an answer.

Honest caveats

These caveats are as important as the numbers.

  • The judge isn’t ground truth. An automated judge is itself a language model. Running it at temperature 0 makes it consistent, not infallible, and it can be wrong in either direction. Its scores are a strong signal, not a verdict.
  • Faithful isn’t the same as correct. Our hallucination metric asks whether an answer is supported by the retrieved content. If the knowledge base itself is wrong or out of date, the assistant can be perfectly faithful and still wrong. For an insurer, grounding is only as good as the documents behind it.
  • The sample is small. 49 answered questions is enough to catch obvious problems, but not enough to pin down a precise rate. The error bars on one failure in 49 are wide. Read the 2.0% figure as “low on our knowledge base, on this test,” not as a guarantee.
  • One knowledge base, one run. We tested on our own product content, not on policy wordings. Insurance documents are longer and denser, and they’re full of exclusions, sub-limits and definitions that cross-reference each other. Your documents will produce different numbers, possibly better, possibly worse. A repeated run could also vary.

What this means for an insurer

No AI system is right every time, and you should be wary of any vendor who says theirs is. What you can ask for is a setup that makes errors rare, makes them visible, and keeps a person accountable:

  • Grounding, so answers come only from your approved documents, with citations a reviewer can check.
  • Abstain, so that when the documents don’t support an answer, the customer hears “let me get someone” instead of a guess.
  • Measurement on your own documents, before launch and again after significant changes to them.
  • A human in the loop for anything that touches a coverage decision, a claim or advice.

That last point matters most. A well-measured assistant can answer “What’s the waiting period on this plan?” from the policy wording. Whether a specific claim is covered is a decision for your people.

To see how the assistant handles grounding and refusals, watch the demo. And if you’d like this evaluation run on your own policy documents, with the same candour about what it finds, get in touch at contact@monology.io.

Alex Rodriguez profile picture

Alex Rodriguez

AI Systems Architect

AI systems architect focused on retrieval-augmented generation, knowledge-base design, and building reliable, grounded chatbots that answer from verified content rather than guesswork.