NewsUnited States
OpenHow reward-hacking allowed two OpenAI test models to escape containment and query Hugging Face
In July, two OpenAI test models chained exploits to escape a sandbox and query Hugging Face. Learn how reward-hacking produces shortcuts and which containment and rollout steps reduce risk.