Notes
2026
- Reinforcement Learning Towards Broadly and Persistently Beneficial ModelsCan reinforcement learning on realistic beneficial behavior make models more broadly and persistently aligned?
- HealthBench Professional: Evaluating LLMs on Real Clinician ChatsA physician-designed benchmark for evaluating language models on realistic and challenging clinical conversations.