Weekly Web Harvest for 2024-06-16
Sycophancy to subterfuge: Investigating reward tampering in language models AnthropicReward tampering is a specific, more troubling form of specification gaming. This is where a model has access to its own code and alters the training process itself, finding a way to “hack” the reinforcement system to increase its reward. This is like a person hacking […]