Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
OpenAI developed GPT-Red, a specialized AI system designed to autonomously attack and find vulnerabilities in its own models — and it outperforms human red-teamers. The system reportedly identifies exploits and jailbreaks faster and at greater scale than manual testing allows. This represents a meaningful shift in how frontier labs approach safety evaluation: moving from human-led adversarial testing toward automated, AI-driven security loops. For organizations deploying OpenAI models, this signals a more systematic approach to pre-deployment hardening, though it also raises questions about what vulnerabilities are found but not disclosed.