Red-teaming is essential to discovering vulnerabilities and improving the robustness of our models. However, current approaches are not scalable, creating a bottleneck.
Commonly used robustness evaluations have already been saturated by our latest models.
We need to develop methods that allow safety and alignment to scale alongside model capabilities.
We trained GPT‑Red, an automated red-teaming model that scales our ability to find vulnerabilities so we can fix them before wider deployment.
GPT‑Red is a strong red-teamer, and our previous models are highly vulnerable to its prompt injection attacks.
We use GPT‑Red to adversarially train GPT‑5.6, making it much more robust to prompt injections.
We will continue to scale this approach alongside human and third-party red-teaming, layered safeguards, and real-time monitoring.
AI systems commonly encounter third-party data through browsers, connected apps, local files, and other tools. These affordances are necessary for performing real-world tasks, but they also create more opportunities for malicious actors to influence model behavior. For example, a third party might embed a carefully crafted instruction—designed to trick the model into uploading sensitive data to an external server—in an email, webpage, tool response, or code repository.
Human red-teaming is a critical part of our safety work, helping us uncover these vulnerabilities before deployment and put the right safeguards in place. But human red-teaming alone is difficult to scale. Designing and running these exercises is time-intensive, limiting how quickly we can identify new failure modes and incorporate them into stronger safeguards. Further, while these exercises produce valuable examples of successful attacks, they cannot generate the volume and diversity of adversarial data needed to improve model robustness through training.
Source link







