Blog

Author Image
Dyllon Johnson
September 22, 2026

Multi-Turn Healthcare Evaluation Research Note

‍

‍

‍

At Centaur, we're developing an evaluation of how healthcare AI assistants respond as patients add or correct information over several turns in a conversation.

‍

The question behind it is a simple one. If two patients tell a model the same things in a different order, does it give them the same advice?

‍

‍

‍

The full sequence of patient messages, model responses, and changes in context is the conversational trajectory. Two conversations can share an endpoint, meaning the model has been told exactly the same information over the course of the conversation, and still differ completely in trajectory. Our evaluation holds the information constant and varies only the trajectory.

‍

Stability means giving clinically consistent advice once the model has the same relevant patient information. The recommendation should not depend on the order in which the facts arrive, whether one of them was corrected along the way, or how they are split across turns. Degradation caused by a presentation-only change is a defect.

‍

Stability is the property we're testing.

‍

Three behaviors determine whether a model achieves it, and each can fail on its own:

‍

Updating means revising a recommendation when new patient information changes what advice is appropriate. The model's next response should account for the clinical implications of that information.
‍

Retraction means withdrawing or correcting earlier advice that no longer applies, so that a recommendation which has since become unsafe is not left standing.
‍

Persistence means continuing to use relevant patient information in later responses. The model should still account for an established fact when the patient returns to the question, even after the conversation has moved to another topic.
‍

Each of these is graded across the whole conversation, not on the final reply alone.
‍

‍

‍

We’re developing this evaluation across many clinical contexts and patient scenarios. Clinical validation is underway, and we’ll publish the results once that work is complete.

‍

If you want to see how your model performs on Centaur’s private multi-turn healthcare evaluation, get in touch. We’ll work with your team to identify where it loses patient context or fails to correct earlier advice as new information arrives, and develop datasets and expert review protocol to supercharge your next model release. 

‍

‍

‍

Want to see how your model stacks up in this evaluation?

Book a demo and we'll show you where Centaur can help.

Book a demo

Related posts

November 3, 2025

Why Radiology AI Can't Afford Poor Annotation | Centaur AI

Radiology AI requires engineered annotation quality for training and evaluation to avoid dangerous clinical error. Centaur uses collective intelligence to outperform individual annotators and create reliable labels for imaging tasks like stroke detection and tumor classification, producing scientifically trustworthy datasets for LLM evaluation and high stakes medical AI applications.

Continue reading →
May 4, 2026

Superhuman Data: How Data Quality Drives AI Reliability

AI is fast, but accuracy remains the real barrier to safe deployment. This post explains how poor data quality, collapsed expert disagreement, and weak evaluation create false confidence in production AI. It shows how collective intelligence, gold-standard labeling, and human-in-the-loop workflows at Centaur.ai build auditable, high-accuracy datasets for high-stakes applications.

Continue reading →