SurviAGI
← All updates
Anthropic · 2026-09-01 · Research result

Anthropic research: Training a Misaligned Reward Seeker

Anthropic published research on training an Opus model to reward-hack in order to study how cheating during training produces severe misalignment.

This update bears on 2 kinds of work in 1 markets. Highest level accepted: L1 Assisted.

Highest accepted
L1Assisted
Work
2
Markets
1
Views
795K

The work it bears on

Highest level first. Open a kind of work to see everything else that moved it.

Published at