Blog

Author Image
Dyllon Johnson
September 15, 2026

3 new datasets added to Centaur’s Arena

    🔬 Lymph node tissue slices
    🩻 Pediatric chest X-rays 
    🧠 EEG recordings

On all three datasets, Centaur’s human consensus was more accurate than every frontier model we tested, including Google’s Gemini family, OpenAI’s GPT family, and Anthropic’s Claude family. The margin between the most performing model and the Centaur human consensus ranged from 7.8 to 35.1 percentage points.

We track each human annotator's accuracy using hidden gold-standard cases with known answers. Those results continually update a quality score, or qscore, for that task, which determines how much the annotator's answers count toward our human consensus.

For the lymph node tissue dataset, we beat the best Frontier model by 7.8 percentage points. 1,733 human annotators reviewed 5,856 microscope images from PatchCamelyon, checking whether the marked area contained cancer cells and contributing 209,828 answers. Our human consensus reached 87.0% accuracy on this task. Gemini 3.1 Pro led the models at 79.2%, while Gemini 3 Flash scored 77.9% and GPT-5.5 reached 71.7%.

The pediatric chest X-ray task showed Centaur beating the top Frontier model by nearly 16 percentage points. We collected 235,498 answers from 738 human annotators across 3,000 images, with each image labeled as normal, bacterial pneumonia, or viral pneumonia. Here, our human consensus reached 74.8% accuracy, compared with 58.9% for Gemini 3.1 Pro, 50.4% for Claude Sonnet 5, and 49.0% for Claude Opus 4.8.

The EEG task revealed the biggest gap between model and human. Our human consensus reached 69.4% accuracy, compared with 34.3% for GPT-5.5 and 19.0% for Claude Opus 4.8. In this dataset, 737 human annotators gave 388,414 answers across 5,000 short recordings, using the tracing and four supporting charts to choose among seven brain activity patterns. 

If you're building or evaluating medical AI, explore the datasets and get in touch about testing your model against our human consensus. We'd love to work with you to understand where your model struggles and how better training data could help. In our Arena, you can:

  1. Buy datasets, including Centaur’s labels
  2. Benchmark your own model
  3. Run your own private Arena with our experts
  4. Source datasets not yet on Arena

Want to see how your model stacks up?

Book a demo and we'll show you what Centaur's benchmarking can do for your team.

Book a demo

Related posts

December 8, 2020

When Medical Experts Disagree: AI Training Data | Centaur AI

Learn the how to mitigate the impact of medical error in your data labeling pipeline by intelligently aggregating multiple expert opinions together

Continue reading →
August 19, 2022

NLP in Healthcare Blog Series | Centaur AI

Learn all about NLP in healthcare - and the medical text datasets that power it - in our new 4-part blog series.

Continue reading →
September 2, 2022

Consensus Partnership for Scientific Data Labels | Centaur AI

The new AI-powered scientific search engine, Consensus, partners with Centaur.ai to generate high-quality, scalable scientific data labels for research.

Continue reading →