OpenAI与Ironclad训练合同智能体
OpenAI表示已与Ironclad合作,围绕合同配置、审批和可复用法律条款等复杂工作流训练并评估计算机使用智能体。在OpenAI针对11项任务进行的研究评估中,GPT-6 Astra平均得分为55.0%,GPT-5.6 Sol为41.6%;每次尝试的估计时间则从37.0分钟降至19.2分钟,这些结果来自OpenAI内部评估。
为什么重要:这项合作展示了如何将专业软件工作流转化为企业智能体的训练和评估任务。
When we introducedGPT‑6 Astra, we demonstrated how far our models have come in using computers for professional work, from preparing documents to testing websites. Our next goal is to make agents more capable and efficient at using specialized software to solve complex business problems. We’re exploring how to train models to understand a company’s business rules, execute multi-step workflows, and verify that their work meets the original requirements.
To accelerate this research, we’re partnering directly with a small number of software companies that understand these workflows best. Together, we’re identifying challenging, high-value tasks and turning them into research problems for training and evaluating our models, ultimately making our models more capable and useful in real-world business applications.
Our first partner is Ironclad, a leader in AI contracting. Working closely with Ironclad’s team, we’ve developed tasks that require agents to configure agreements, approvals, and reusable legal terms, and demonstrated progress on these complex workflows. Ironclad’s expertise has been instrumental in defining what success looks like and bringing real customer needs directly into frontier model development. We’re grateful for their partnership and excited to share what we’ve accomplished together.
GPT‑6 Astra is our first frontier model trained on Ironclad tasks. On our research evaluation, its average score was 32% higher than GPT‑5.6 Sol’s, while estimated time per attempt was 48% lower.1
Turning contracting workflows into training tasks
Consider a legal operations team setting up a process for buying software. Finance may need to approve purchases above a certain amount, Security may need to review certain requests, and Legal may need to review nonstandard terms. The person setting up the process has to turn that short list into an intake form, document templates, approval rules, and a record of the final agreement.
An AI agent doing the same work has to keep those requirements in view as it moves through the software. For example, it must configure Finance approval above the spending threshold and check that requests above and below it follow the right paths. Getting individual steps right is not enough: the finished process must work across the situations it was designed to handle.
Ironclad employees and people who use Ironclad at OpenAI helped our researchers identify 11 tasks across legal, commercial, and procurement work. These included tasks like setting up nondisclosure agreements, creating procurement approval processes, and updating a reusable legal clause so that it reflects the jurisdiction a requester selects. We estimate that this work would take an experienced user about 30 to 40 minutes per task, on average.
We evaluated each task against 8 to 50 criteria, depending on its complexity. This let us see which parts a model got right and where it fell short. Ironclad also provided hosted software environments of their product where the models could practice these tasks. Our researchers developed synthetic training tasks2 around representative workflows and used reinforcement learning to help the models improve through practice and feedback. This combination—tasks selected with people who know the work, detailed criteria, and a place for models to practice—ensures we are improving on real-world tasks that are most important to our customers.
How Astra performed
We compared Astra and GPT‑5.6 Sol using Max reasoning for Astra and High reasoning for Sol, the settings where each model scored highest. Across the 11 research tasks, Astra’s average score was 55.0%, compared with 41.6% for GPT‑5.6 Sol, while estimated average time per attempt fell from 37.0 minutes for Sol to 19.2 minutes for Astra.3 An internal model used in the development of Astra achieved an even stronger 63.7% on these tasks, and we aim to bring these further gains to future models.
Astra met more of the task requirements with less simulated time per attempt. The two clips below show how the models handled the same research task. Astra (first) met about 94% of the task criteria in an estimated 20 minutes; GPT‑5.6 Sol (second) met about 85% in an estimated 32 minutes.
GPT‑6 Astra met about 94% of the task criteria in an estimated 20 minutes.
GPT‑5.6 Sol met about 85% of the task criteria in an estimated 32 minutes.
What this collaboration means for Ironclad
Ironclad’s role in this work reflects a challenge its customers know well: a contracting process must handle exceptions while preserving the rules a business depends on. If an agent loses track of one of those rules halfway through a task, that limits what a software company can confidently ask it to do. This underscores why human oversight still matters as agents get better at complex contracting tasks, and why a full contracting platform remains essential. Working with our researchers gives Ironclad a way to bring these problems into the development of the underlying model, where we can study them together.
As models become more capable, Ironclad has an opportunity to use them in more of its own products. For legal and business teams, that could mean AI taking on more of the effort involved in complex contracting while preserving the controls those teams rely on.
Work with us to advance AI for complex professional work
We’re inviting a small number of software companies to work directly with our research and engineering teams on important professional tasks that today’s agents still can’t reliably complete. If your team has one, we’d like to see a concrete example: what you’re asking the agent to do, evidence of where it fails, and how you would judge a successful result. Partners should also be able to bring people who know the work deeply, a secure environment for testing, and data that can be safely used for research. This gives us a starting point for investigating the failure and measuring whether the model improves.
前因后果
- AI智能体沙盒逃逸:必然性、案例与防御清单钛媒体 · OpenAI
- Kevin Roose谈生成式AI史WIRED AI · OpenAI
- 陶哲轩警告AI数学“证明消化不良”量子位(原生 RSS) · OpenAI
- OpenAI“疯狂28天”首日:GPT-6提速引争议量子位(原生 RSS) · OpenAI
- Instinct搅动AI助手赛道,信任成关键虎嗅 AI · OpenAI
- Christian Szegedy:AI将重塑数学研究新智元 · OpenAI