1 min read
Broadly and Persistently Beneficial Models

Research into whether reinforcement learning on realistic beneficial behavior can improve alignment beyond the training distribution. The work trains models on traits including truthfulness, fairness, risk awareness, and corrigibility, then evaluates generalization across more than 50 independent benchmarks.

The results show broad out-of-distribution transfer and improved resistance to adversarial prompting and harmful fine-tuning. Training limited to beneficial behavior in health also transfers to non-health alignment evaluations.

Read the paper.