Model Releases Hacker News (GPT)

GPT-6 Astra System Card

GPT-6 Astrasystem cardAI alignmentcybersecurity

OpenAI released GPT-6 Astra, described as the most capable model it has broadly deployed and the first to reach the Critical level of cybersecurity capability under its Preparedness Framework. With the right tools and access, Astra can find previously unknown security flaws and develop novel exploits across many well-protected systems without a person guiding each step. Accordingly, OpenAI says it significantly strengthened protections against harmful cyber actions from misuse or misalignment, and secured internal development via stricter isolation, checkpoint encryption, universal monitoring of full trajectories including chain-of-thought, and a blocking alignment evaluation before internal use.

Astra is reported to be significantly more robust to jailbreaks than GPT-5.6 Sol, including across longer trajectories, based on offline tests and a program of internal and external jailbreak testing and remediation. For users flagged as potentially high risk, OpenAI trained in the ability to adjust the model's refusal boundary to be more conservative and cover a broader range of dual-use risks. Regression testing confirms robustness against prior jailbreaks, and new rounds of automated red-teaming with the latest internal attackers validate the improvements.

Alignment improvements span the composition of pre-training data to grading during reinforcement learning. New alignment evaluations find Astra stronger at respecting safety and security boundaries and staying within its authorized scope. In a simulation using more than 54,000 internal Codex tasks, Astra received roughly half as many flags for higher-severity misaligned behavior as Sol.

OpenAI is also deploying misalignment monitoring broadly. It views model alignment as the primary lever to prevent misaligned behavior, but monitoring provides broad visibility into frontier model behavior and serves as an additional layer of protection. For these reasons, misalignment monitoring has been added to all tool-using inference involved in external deployment of Astra, at significant compute cost.

Read original →

← Back to home