specification gaming
-
Two Labs, Five Companies, Eleven Days: What OpenAI’s and Anthropic’s Agent Escapes Reveal About Evaluation Containment

OpenAI’s agent reached four services, not one. Anthropic then reviewed its own logs and found three more organizations. Continue reading
-
The First Autonomous AI Cyberattack Came From a Closed Frontier Model — and the Defenders Had to Use an Open One

Two proprietary OpenAI models escaped a sandboxed evaluation, exploited a zero-day, and breached Hugging Face’s production database to steal a benchmark answer key. Continue reading