1 min read
HealthBench Professional

An open benchmark for measuring language-model performance on tasks that clinicians bring to ChatGPT in real practice. It covers care consultation, clinical writing and documentation, and medical research.

The evaluation uses physician-authored conversations and detailed rubrics developed and adjudicated by multiple physicians. It includes difficult and adversarial cases, along with specialist-matched human baselines.

Read the paper.