Blog

Author Image
Dyllon Johnson
September 17, 2026

What a Preference Score Can Miss In a Health Response

We asked our crowd for their preferences on language model responses to health queries. We found that people rejected dismissive health answers more often than answers with factual errors, and factually flawed answers still won nearly three in ten choices.

We wanted to understand how a human's choice changes when a language model response falls short in a specific way. So we took 188 pairs of AI responses to consumer health queries and rewrote one answer in each pair. Each rewrite targeted one of seven flaws, such as a dismissive tone, factual errors, or downplaying a warning sign.

We kept the model responses similar in length and balanced which response was presented first. Participants weren't told that an answer had been changed, and each person could judge a given pair only once. We asked which response they'd rather receive and collected 2,909 judgments that passed our quality checks. Because the unchanged answer is word-for-word identical in both versions, any shift in which one people pick is caused by the flaw we introduced.

At first, the result looked straightforward. Before the rewrites, participants chose the unchanged answer about half the time. Afterwards, they chose it nearly four times out of five. Every flaw shifted preference toward the unchanged answer, although the strength of that shift varied by flaw type.

Dismissiveness produced the strongest effect. Participants chose the unchanged answer 85.9% of the time when the alternative was dismissive, compared with 70.3% when it contained factual errors. For an eval or AI engineering team, this suggests that preference scores may penalize an obviously cold or dismissive answer more reliably than they penalize a medically incorrect one. The difference held up after accounting for all 21 possible comparisons between flaw types.

The pattern holds beyond those two flaws. We grouped the seven into flaws of manner (dismissive, overconfident, vague) and flaws of substance (factual errors, a downplayed warning sign, missing the point, generic filler), with the grouping fixed by a separate model that saw only the flaw definitions and never the results. Manner flaws cost preference about 6 percentage points more.

Even so, an answer we had made factually worse still won nearly three in ten choices. These were deliberately noticeable flaws, written to be catchable, so subtler errors would likely fare better rather than worse. We can't tell whether participants missed the errors or noticed them and preferred the answer anyway. Either way, choosing an answer wasn't enough to establish that its medical information was correct.

Reading time offered another clue. In pairs with factual errors, participants chose the unchanged answer 61.7% of the time in the fastest third of judgments, compared with 76.5% in the slowest third. One possible explanation is that checking a medical claim requires more attention than judging whether an answer brushes off a concern. We didn't test that explanation directly.

For teams using human preference to guide model training, this is a reason to look closely at what a winning answer gets right. A response can be pleasant to read and still contain an error. We haven't tested how training on these choices would change a model, but the results show why medical accuracy needs its own evaluation.

Pairing preference judgments with clinical review gives us a way to find those gaps: answers people want to receive that still fall short on correctness or safety. It also makes the response more specific. A factual error calls for a different correction than a confusing explanation, so separating those issues can help teams choose training examples that address the underlying problem.

If you're building or evaluating a model for health questions, get in touch. We'd love to help you measure what people prefer alongside medical accuracy and safety, then use those findings to guide better training data.


Want to see what your rubric is missing?

Book a demo and we'll show you what Centaur can do for your team.

Book a demo

Related posts

April 1, 2026

Why Centaur.ai Is Going to HumanX 2026

Centaur.ai is heading to HumanX 2026 to address the biggest challenge in healthcare AI: data quality. Most teams struggle not with labeling, but with measuring accuracy. Using collective intelligence and expert consensus, Centaur delivers faster, higher-quality datasets that improve model performance and support FDA-ready, defensible AI deployment.

Continue reading →
November 10, 2025

Visit Centaur AI at RSNA 2025 | Radiology Conference

Radiology AI models are only as strong as their annotations. Centaur.ai engineers quality through collective intelligence, combining expert crowds, benchmarking, and performance-based incentives to produce validated data for model training and evaluation. Visit our RSNA booth to see how we make radiology AI accuracy inevitable at scale.

Continue reading →
August 11, 2025

AI Data Labeling for Manufacturing Robots | Centaur AI

Autonomous robots in manufacturing rely on high-quality labeled data to function effectively. Precise annotation enables defect detection, precision assembly, and safe collaboration. Continuous labeling prevents performance drift as factories evolve. Centaur.ai delivers expert labeling services that power smarter factories where human insight and machine intelligence work seamlessly together.

Continue reading →