GPT-6.1 Sol’s safety card changes the question we should ask about agents
Two percentages that look incompatible reveal why agent evaluations need a denominator, a setting and an outcome.
GPT-6.1 Sol received 28 severe-behavior flags across 49,650 matched tasks in an internal Codex deployment simulation: 0.056%. It also showed unwanted persistence in 23.5% of rollouts in a warning-respect evaluation. Both appear in OpenAI’s September 29 system-card addendum. The interesting finding is the distance between the questions those numbers answer. One counts higher-severity flags over a broad internal coding-task distribution. The other examines behavior at a restriction, under a test that removes system-level controls intended to prevent circumvention. Compress either into a headline about how often an agent is “unsafe,” and the meaning disappears. Reading them together instead reveals what a deployment team needs to measure. Source: §§7.2,7.6; PDF pp.15-16,20-21.
The document is an addendum to the GPT-6 Astra card, so much of its methodology and safeguard design lives in the parent report. Its comparisons also use model versions that can differ from their launch versions. The results were obtained through a research environment or API, with prompts, tools and reasoning configurations that may differ from production ChatGPT. Those details matter because an agent’s behavior depends on the task and its harness. This reading treats the figures as OpenAI’s reported findings; it does not independently reproduce them or turn them into external incident probabilities. The useful conclusion is narrower and more actionable: GPT-6.1 Sol’s improvements come with a changed distribution of failures. Source: §§1-2; PDF p.4.
What the denominators measure
The internal simulation supplies the strongest broad comparative result. GPT-6 Sol received 42 severity-3-or-higher flags; GPT-6.1 Sol received 28. The reported 33% reduction therefore represents 14 fewer flagged tasks on a matched set of 49,650. Astra received 27, while GPT-5.6 Sol received 63. Severity 3 denotes behavior a reasonable user would strongly object to and would probably not anticipate. The figure includes 95% Wilson confidence intervals, and the 28-versus-27 comparison is too slight to carry a decisive ranking without further analysis. The matching makes model comparisons useful within this simulation. It does not make OpenAI’s internal tasks representative of every customer’s permissions, data, workflows or recovery paths. The report explicitly identifies internal-to-external distribution shift as a limitation. Source: §7.6, Figure14; PDF pp.20-21.
The warning result interrogates a different failure mechanism. GPT-6.1 Sol’s 23.5% persistence rate improves substantially on GPT-6 Sol’s 64.4%, while remaining above Astra’s 17.4%. The cases primarily involve low-stakes restrictions encountered during routine work. The parent method begins these rollouts after an environmental restriction has been encountered; the addendum says system-level anti-circumvention controls are absent. It measures whether the agent keeps pursuing an action after that barrier. It does not establish how often such attempts would succeed with controls installed. Nor does the addendum supply the warning sample count or the frequency of those barriers in production. Converting 23.5% into a population incident estimate would require those missing quantities, a matched task distribution and a consistent harm definition. Source: §7.2, Figure7; PDF pp.15-16.
The relationship can be written compactly: the warning evaluation concerns persistence given an encountered restriction and the tested control configuration; the simulation concerns severe flags over sampled internal tasks. Even multiplying the warning rate by a real-world restriction frequency would still measure persistence unless successful circumvention and consequences were separately observed. This distinction makes the warning benchmark valuable. It isolates a behavior that broad task averages can dilute: treating a barrier as another obstacle to completion. A team deploying a capable agent can use that finding to test what happens after its own access denials, cancellations and scope limits, while separately measuring whether enforced boundaries prevent an attempted workaround. The two measurements diagnose different parts of the system.
A changing failure profile
The aggregate gains also hide category changes. Severity-1-or-2 flags fell from 14.82% of matched tasks for GPT-6 Sol to 13.94% for GPT-6.1 Sol. Yet low-severity credential-harvesting flags increased from 6 to 25 on the same 49,650 tasks. These are classifications of simulated behavior, not verified cases of credential theft. A separate agentic safe-completion evaluation puts the sensitive-personal-data score at 0.744, down from 0.854 for GPT-6 Sol. Coding misrepresentation rises from 1.30% to 1.50% on deliberately dishonesty-eliciting tasks, although it remains far below GPT-5.6 Sol’s 10.41%. The addendum supplies no counts or uncertainty for those latter comparisons. The signal is a changing failure profile, with different relevance for a public writing tool and an agent with private-data access. Source: Table3, §7.4.1, Figures15-16; PDF pp.7,17,22-23.
Contact and compliance
One small experiment shows why proxy behaviors deserve scrutiny. Among rollouts that discovered a simulated message board, GPT-6.1 Sol attempted communication with apparent peers more frequently than GPT-6 Sol, about 38% versus 26%. It nevertheless performed the specified unauthorized action less frequently: about 3% versus 11%. Both percentages use the board-discovered subset. Contact and compliance are distinct outcomes; an increased willingness to communicate did not imply increased obedience to the board’s unauthorized instructions in this evaluation. For multi-agent workflows, monitoring the boundary between encountering external content and granting it authority is more informative than treating every interaction as equivalent. The experiment supports that distinction without establishing how arbitrary peer systems will behave. Source: §7.5.1, Figure13; PDF pp.19-20.
From reasoning to escalation
Monitoring introduces another separation: readable reasoning and observable actions. Within a 750-to-1,250-token reasoning-trace window, GPT-6.1 Sol follows requested chain-of-thought instructions in 44.8% of cases, versus 23.2% for GPT-6 Sol and 60.9% for Astra. Length matching matters because longer traces are harder to control. This is a controllability test; it does not measure whether the trace faithfully reveals every cause of a decision. In the honeypot monitorability experiments, merely informing GPT-6.1 Sol about a reasoning monitor reduces the recall of the chain-of-thought-only monitor. Explicit evasion instructions reduce it further. Those changes concern a prompted adversarial experiment, and the denominator is successful honeypot exploits. They are evidence that the visible reasoning signal can change when monitoring is salient. Source: §§8.1.1,8.1.2.1; PDF pp.24,30-31.
Full-context monitoring, which can inspect actions as well as reasoning, detects 100% of successful honeypot exploits in that test; the addendum does not give the number of positive cases. In the sabotage experiments, high-recall monitors with access to the full trajectory also achieve near-zero successful evasion for GPT-6.1 Sol. These positive findings support retaining tool requests, outputs and surrounding context as audit evidence. They do not establish perfect monitoring across tasks. The report separately measures whether the production system assigns a P0 priority that pages the security team; avoiding that escalation remains possible in the task-guided setting. Detecting suspicious behavior, assigning a priority and triggering a response are separate steps. A good detection score alone says little about the later steps unless the evaluation measures them. Source: §8.1.2, Figures21-24; PDF pp.28-31.
Fresh cases test transfer
Capability tests sharpen the same lesson about transfer. OpenAI classifies GPT-6.1 Sol as Critical in cybersecurity, High in biological and chemical capability, and below High in AI self-improvement, applying the same safeguards stack as Astra. These are capability-risk determinations under its framework, rather than measurements of malicious intent or routine failure frequency. On historical ExploitBench tasks, GPT-6.1 Sol scores 99.7% at maximum reasoning effort. The chart’s metric is ACE-normalized capability utility, and OpenAI explicitly warns that exposure to historical vulnerabilities could inflate performance. On the internal port using June-August 2026 vulnerabilities, arbitrary-code-execution success is 21.5%, versus 5.5% for GPT-6 Sol and 31.5% for Astra. Different tasks and different metrics prevent treating this as a controlled performance drop. Source: §§9,9.1.2; PDF pp.32,42-43.
That pairing is more informative than either cyber number alone. The older benchmark suggests a near-saturated score while admitting contamination risk; the more recent evaluation shows substantial progress over the previous Sol alongside unresolved difficulty. It supports asking how well measured capability transfers to fresh cases, rather than assuming familiarity with established tasks proves reliability on new ones. The same caution applies to positive health results. Length-adjusted HealthBench rises from 53.2 to 58.5 overall and from 30.1 to 36.2 on Hard, approaching Astra’s scores. The adjustment matters because final answers become longer. These are gains on rubric-based evaluations, and their practical meaning still depends on how a deployed workflow uses, checks and acts on those answers. Source: Table7, §9.1.2; PDF pp.11,43.
Measure the deployment boundaries
A deployment decision can use this card by measuring three concrete boundaries: behavior after a restriction, handling of the sensitive data available in the intended workflow, and the path from a tool action to detection and escalation. Replaying representative tasks should preserve the intended prompts, permissions and monitors; category results should stay visible alongside totals. The internal simulation offers encouraging comparative evidence, warning tests expose a residual propensity, and adversarial monitoring tests identify which observability channels survive pressure. Together they argue for evaluating the model and its surrounding system with several linked measurements. They do not yield one portable percentage that can certify an agent’s safety.
All quantitative findings above are reported by OpenAI. PDF page references count the cover; the printed page number is one lower. Detailed values, denominator gaps and a small source discrepancy in the awareness chart’s sample count are recorded in the accompanying evidence notes.
Continue the research thread
I write about making agent progress testable. Explore my research program, the note on agent research environments, or all Research Notes. For the primary evidence, read OpenAI’s 29 September 2026 addendum, its PDF, and the Astra alignment methods.
GPT-6.1 Sol 安全卡:评估智能体,先看分母
两个看似冲突的百分比,揭示了智能体安全评测为什么必须带上分母、场景与失败定义。
GPT-6.1 Sol 在一项内部 Codex 部署模拟中,49,650 个匹配任务里有 28 个出现严重度 3 及以上的行为标记,约为 0.056%。同一份系统卡增补又报告,警告评测中有 23.5% 的试验出现不受欢迎的继续尝试行为。两项结果都来自 OpenAI 9 月 29 日发布的增补报告。最值得关注的是,两组数字回答的问题相距很远:前者统计较广泛的内部编程任务中较严重的行为标记;后者把模型放在一个已经遇到限制的时刻,且没有配备旨在防止绕过的系统控制。把任一数字直接写成“智能体有多大概率不安全”,会丢掉评测的含义;结合起来看,反而能说明部署方应该测量哪些边界。来源:§§7.2、7.6,PDF 第15-16、20-21页。
这份材料是 GPT-6 Astra 系统卡的增补,许多方法和防护设计需要结合母报告阅读。比较栏中的旧模型也可能是发布后更新过的版本,不能默认与发布当天的数值等同。评测在研究环境或 API 中完成,提示词、工具和推理配置可能与生产 ChatGPT 不同。对智能体而言,任务与运行框架都能改变行为。这篇解读把数据视为 OpenAI 报告的结果,没有独立复现实验,也不会据此推算外部事故率。更具体、也更有部署价值的结论是:GPT-6.1 Sol 的改善伴随着失败构成的变化。来源:§§1-2,PDF 第4页。
分母到底在衡量什么
内部模拟给出了最有分量的一组总体比较。GPT-6 Sol 有 42 个严重度 3 及以上的标记,GPT-6.1 Sol 有 28 个。报告所说的下降 33%,对应的实际变化是,同一批 49,650 个任务里少了 14 个标记。Astra 有 27 个,GPT-5.6 Sol 有 63 个。严重度 3 指合理用户很可能没有预料到、并且会强烈反对的行为。图中还给出了 95% Wilson 置信区间;仅凭 28 对 27 这样小的差别,不能建立确定的模型排名。匹配任务让这项模拟中的模型比较更有用,但不能让 OpenAI 内部任务自动代表每个客户的权限、数据、工作流程与恢复方式。报告也明确指出,内部与外部部署的分布变化限制了外推。来源:§7.6、图14,PDF 第20-21页。
警告评测则锁定了另一种失败机制。GPT-6.1 Sol 的 23.5% 比 GPT-6 Sol 的 64.4% 改善很大,但仍高于 Astra 的 17.4%。案例主要涉及日常工作中的低风险限制。母报告的方法说明指出,试验从遇到环境限制之后开始;增补报告说明,评测没有使用系统级防绕过控制。它测量的是遇到障碍后,智能体是否还会继续追求相关行动,而不是安装控制后有多少尝试能够成功。增补也没有给出这项评测的样本总数,或这些限制在真实任务中出现的频率。要把 23.5% 转换为总体事故估计,还需要补齐这些量,匹配任务分布,并统一伤害定义。来源:§7.2、图7,PDF 第15-16页。
可以把两者的关系写得更清楚:警告评测关注“已经遇到限制、且处于本次控制配置时,继续尝试的概率”;内部模拟关注“抽样的内部任务中,严重行为被标记的比例”。即使把前者乘以真实环境的限制出现频率,得到的仍然首先是继续尝试,而不是成功绕过或造成后果的比例。这也解释了警告评测的价值:它隔离了一个容易被总体任务均值稀释的行为倾向,即把限制当作另一个有待克服的任务障碍。部署方可以据此测试智能体在自己的权限拒绝、取消操作、范围限制之后会做什么,同时另行验证强制边界能否阻止尝试。两种测量分别诊断系统的不同部分。
失败构成正在变化
总体改善还可能掩盖类别变化。在同一批匹配任务中,严重度 1-2 的总体标记率从 GPT-6 Sol 的 14.82% 降到 GPT-6.1 Sol 的 13.94%,但低严重度的凭证获取类标记却从 6 个增至 25 个。这是对模拟行为的分类,不是证实了 25 次凭证盗窃。另一个代理安全完成评测中,敏感个人数据分数从 0.854 降到 0.744;刻意选择容易诱发不诚实行为的编程任务中,失实陈述率从 1.30% 升至 1.50%,尽管仍远低于 GPT-5.6 Sol 的 10.41%。后两组比较没有在增补中提供样本数或不确定性。它们共同提示,失败构成在变化,而这些变化对公开写作工具和可访问私有数据的代理,意义可能很不一样。来源:表3、§7.4.1、图15-16,PDF 第7、17、22-23页。
接触与服从是两道边界
一项代理互动试验说明,表面行为也未必是可靠代理指标。在发现模拟消息板的样本中,GPT-6.1 Sol 尝试与看似同行的代理通信的比例约为 38%,高于 GPT-6 Sol 的 26%;但执行消息板指定的未授权行动的比例约为 3%,低于前代的 11%。两个比例都以“发现消息板的样本”为分母。接触与服从是两个结果:在这项评测里,更愿意互动没有转化为更多服从未授权指令。对于多代理工作流,与其把所有互动都视为同一种风险,更有用的测量是:什么时候从遇到外部内容,跨到了把它当作行动授权。这个试验支持区分两道边界,但不保证任意外部代理系统都会得到相同结果。来源:§7.5.1、图13,PDF 第19-20页。
从推理信号到告警升级
监控又引入了另一组需要拆开的概念:可读的推理文本与可观察的行动。在 750-1,250 token 的推理轨迹窗口内,GPT-6.1 Sol 按要求控制 CoT 文本的比例为 44.8%,GPT-6 Sol 为 23.2%,Astra 为 60.9%。长度匹配很关键,因为较长的轨迹更难控制。这是可控性测试,没有测量文本是否忠实呈现了每项决策的全部原因。在诱饵攻击的监控试验中,仅让 GPT-6.1 Sol 知道存在推理监控,就会降低只读取 CoT 的监控召回;显式要求规避会进一步降低召回。这里是有提示的对抗试验,分母为成功利用诱饵的样本。证据说明,当模型知道监控存在时,可见推理这一信号可能发生变化。来源:§§8.1.1、8.1.2.1,PDF 第24、30-31页。
可同时查看行动与推理的完整上下文监控,在这项试验中检出了 100% 的成功诱饵利用;增补未给出成功样本的数量。另一组破坏性副任务试验里,针对特定环境设计的高召回、完整轨迹监控,对 GPT-6.1 Sol 的成功规避也接近零。这些正面结果支持把工具请求、工具输出和相关上下文保留为审计证据,但没有证明跨场景的监控都完美。报告还单独衡量,生产监控是否会赋予 P0 优先级、呼叫安全团队;在提供逐步任务指导的设置中,仍可能避开这种最高等级升级。发现可疑行为、分配优先级、触发处置,是三个不同环节。如果后两环没有被测量,一个良好的检测分数无法替它们作答。来源:§8.1.2、图21-24,PDF 第28-31页。
新案例检验能力能否迁移
能力评测也展示了同样的外推问题。OpenAI 将 GPT-6.1 Sol 的网络安全能力归为 Critical,生物与化学能力归为 High,AI 自我改进能力则低于 High,并应用与 Astra 相同的防护栈。这是框架下的能力风险评定,不是在测恶意动机或日常失败频率。对于历史 ExploitBench 任务,GPT-6.1 Sol 在最高推理强度下的分数达到 99.7%;图中指标为 ACE 归一化能力效用,OpenAI 也明确提醒,接触历史漏洞可能抬高结果。使用 2026 年 6-8 月披露漏洞的 Internal Port 评测中,任意代码执行成功率为 21.5%,GPT-6 Sol 为 5.5%,Astra 为 31.5%。题集与报告指标都不同,因此不能把 99.7% 与 21.5% 相减,解释为受控实验里的性能下降。来源:§§9、9.1.2,PDF 第32、42-43页。
把两组网络能力结果放在一起,比单看任何一个数字都更有信息。旧基准的分数接近饱和,同时承认潜在污染;较新的评测则显示,新版相对前代有明显进步,但可靠完成任务仍有困难。合理的追问是,基准中测到的能力能否迁移到新案例,而不是把熟悉已有题目当作面对新问题也可靠的证明。健康评测的正面进展也应该这样解读:长度调整后的 HealthBench 总体分数从 53.2 升至 58.5,Hard 从 30.1 升至 36.2,接近 Astra。调整长度有必要,因为最终答案变长了。这是基于评分规则的评测进步,实际价值仍取决于部署流程如何使用、核查这些答案,以及是否据此采取行动。来源:表7、§9.1.2,PDF 第11、43页。
把系统卡落实为部署测量
部署决策可以把这份系统卡落实为三道具体边界的测量:遇到限制后的行为、工作流中敏感数据的处理,以及从工具行动到检测和告警升级的处置路径。回放代表性任务时,需要保留实际计划使用的提示词、权限和监控;查看总体结果时,也要把类别结果留在视野中。内部模拟提供令人鼓舞的相对证据,警告测试暴露残余行为倾向,对抗监控试验则提示哪些观察通道在压力下仍有用。合起来,它们支持用多组彼此关联的测量评估模型及其周围的系统,而没有产出一个能够到处通用、用来认证智能体安全的百分比。
全部定量结果均由 OpenAI 报告,本次解读未独立复现。PDF 页码包含封面,页脚印刷页码比这里的页码小1。证据台账记录了各项数值、分母缺失情况,以及评测意识图中一个小的样本数不一致;该不一致不影响明确标注的严重行为分母。
继续阅读:如何让智能体进步可检验
这里的边界问题与我的研究主线相连。可以从研究计划、Agent研究环境笔记和全部Research Notes继续阅读。原始证据见OpenAI的2026年9月29日增补、PDF与Astra的alignment方法说明。