Blog

.png)
On all three datasets, Centaur’s human consensus was more accurate than every frontier model we tested, including Google’s Gemini family, OpenAI’s GPT family, and Anthropic’s Claude family. The margin between the most performing model and the Centaur human consensus ranged from 7.8 to 35.1 percentage points.
We track each human annotator's accuracy using hidden gold-standard cases with known answers. Those results continually update a quality score, or qscore, for that task, which determines how much the annotator's answers count toward our human consensus.
For the lymph node tissue dataset, we beat the best Frontier model by 7.8 percentage points. 1,733 human annotators reviewed 5,856 microscope images from PatchCamelyon, checking whether the marked area contained cancer cells and contributing 209,828 answers. Our human consensus reached 87.0% accuracy on this task. Gemini 3.1 Pro led the models at 79.2%, while Gemini 3 Flash scored 77.9% and GPT-5.5 reached 71.7%.

The pediatric chest X-ray task showed Centaur beating the top Frontier model by nearly 16 percentage points. We collected 235,498 answers from 738 human annotators across 3,000 images, with each image labeled as normal, bacterial pneumonia, or viral pneumonia. Here, our human consensus reached 74.8% accuracy, compared with 58.9% for Gemini 3.1 Pro, 50.4% for Claude Sonnet 5, and 49.0% for Claude Opus 4.8.

The EEG task revealed the biggest gap between model and human. Our human consensus reached 69.4% accuracy, compared with 34.3% for GPT-5.5 and 19.0% for Claude Opus 4.8. In this dataset, 737 human annotators gave 388,414 answers across 5,000 short recordings, using the tracing and four supporting charts to choose among seven brain activity patterns.

If you're building or evaluating medical AI, explore the datasets and get in touch about testing your model against our human consensus. We'd love to work with you to understand where your model struggles and how better training data could help. In our Arena, you can:
Learn the how to mitigate the impact of medical error in your data labeling pipeline by intelligently aggregating multiple expert opinions together
Learn all about NLP in healthcare - and the medical text datasets that power it - in our new 4-part blog series.
The new AI-powered scientific search engine, Consensus, partners with Centaur.ai to generate high-quality, scalable scientific data labels for research.