HealthBench Professional is an open benchmark for realistic clinical conversations across care consultation, writing and documentation, and medical research. Its physician-authored conversations, detailed rubrics, and multi-stage physician adjudication target real clinical work rather than medical recall alone.
Building a non-saturated eval
The examples were selected for quality, representativeness, and difficulty from a candidate pool of 15,079 conversations. The final benchmark overrepresents difficult cases by roughly 3.5× relative to the candidate pool, and approximately one third of the benchmark involves physicians deliberately testing models adversarially.
Comparing models and clinicians
Specialist-matched physicians answered the same tasks without time limits and with web access, providing a strong human baseline for tracking model progress.