🔍 Read the full analysis: OpenAI Is Training Agents Inside Your Software. Read The Fine Print On Ironclad. on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI says it trained GPT-6 Astra using hosted copies of Ironclad’s contract-management software and 11 legal, commercial and procurement tasks. Astra met an average 55% of task criteria; OpenAI’s time figures are simulations, not measured customer savings. The project also invites selected software companies to supply challenging workflows for agent research.
OpenAI has described training its GPT-6 Astra model inside hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. The results, published October 6, show an average of 55% of evaluation criteria met—not 55% of tasks completed—and OpenAI says its estimated task times are simulations rather than measured customer savings.
Ironclad employees and OpenAI staff who use the product selected 11 workflows, including setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause based on a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. Each task was scored against a rubric of 8 to 50 criteria, depending on its complexity.
OpenAI says Ironclad provided hosted product copies for model practice. It also says it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, after filtering out personal information. According to OpenAI, the work did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.
In OpenAI’s comparison, GPT-6 Astra met an average 55.0% of rubric criteria, against 41.6% for GPT-5.6 Sol in a high-compute configuration. Astra’s estimated time per attempt was 19.2 minutes, compared with 37 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%. On one featured task, Astra met about 94% of the criteria. These are reported evaluation results; they do not establish that the model can reliably complete contract work for customers.
OpenAI is training agents inside your software. Read the fine print on Ironclad.
Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.
legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses
per task, experienced user (OpenAI estimate)
criteria per task — a rubric, not pass/fail
public SEC filings; no customer or non-public Ironclad data
The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.
37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.
Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.
Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.
Averages hide missed approvals.
Narrowest access; no self-escalation.
METR found agents spoofing tool-call records.
Measure the whole loop.
Public filings, not your contracts.
Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.
Why Contract Work Needs More Than Partial Credit
The reported score is difficult to translate into business value because a workflow can fail if it misses even one required control. OpenAI’s example describes a procurement process that must route spending above a threshold to Finance, certain requests to Security and nonstandard terms to Legal. Getting two of those three rules right would still leave a required approval unaddressed; a high average score across criteria does not show whether a particular workflow is safe to use.
That makes the evaluation useful as a measure of progress, but not proof of deployment readiness. OpenAI’s own account says that losing track of a business rule limits what a software company can confidently ask an agent to do, and that human oversight remains necessary. For businesses evaluating agents, the missed criteria and the severity of each failure matter at least as much as an average percentage.
The project may also matter to software vendors beyond Ironclad. OpenAI says it is seeking a small number of software-company partners with difficult agent tasks, knowledgeable staff, secure test environments and data suitable for research. Such work could help models operate specialized products, while increasing the importance of the business rules, records and controls behind those products—not just their screens.
contract management software for legal teams
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How OpenAI Built the Ironclad Test
The collaboration is presented as research into whether models can follow company rules through multi-step work in specialized software, then check their output against the original requirements. Rather than assessing general computer use alone, the test put models in a hosted version of a real product and used tasks chosen by people familiar with its workflows.
OpenAI’s post also frames the work as a way to study where agents fail. It argues that contracting platforms remain important because they provide business rules and controls that an agent must respect. The published results cover 11 selected tasks, however, and should not be read as a broad benchmark of all Ironclad features, all contract work or all customer environments.
The time comparison needs similar care. OpenAI explicitly describes the 19.2- and 37-minute figures as simulated estimates based on assumed processing and generation speeds. They are not observations of employees using the software, and the source says they apply to the research tasks rather than Ironclad workflows generally.
What the Evaluation Does Not Establish
The published summary does not provide enough detail to determine how the 11 tasks were selected beyond the roles involved, how often each was attempted, or which specific criteria Astra missed across the full set. It also does not establish how the model would perform on new workflows, unusual contract language or live customer data. Those details would be needed to judge reliability for particular uses.
The 55% average is a share of rubric criteria met, not a percentage of tasks completed successfully. OpenAI’s simulated times likewise do not show actual productivity gains. The source material does not give a publication year for the October 6 post, nor does it describe a customer deployment, a release timeline or a measured business outcome. It remains unclear which other software companies, if any, will take part in the proposed research partnerships.
What Vendors and Buyers Should Ask
OpenAI says it is inviting a small number of software companies to bring concrete examples of work current agents cannot reliably complete, along with domain experts, secure test environments and research-appropriate data. The source does not identify future partners or give dates for another evaluation, so the next public milestone is not yet known.
For software vendors and buyers, the useful next step is to ask for more than an average score: which requirements failed, how serious those failures were, and whether results hold up on workflows outside the test set. Organizations considering agents for contracts or procurement should also clarify what human review is required and how approval rules, records and controls are maintained. Until such evidence is available, the reported results support continued testing—not an assumption that agents can independently handle these workflows.
Key Questions
What did OpenAI and Ironclad test?
They tested a model on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management product, including NDA setup and procurement approvals.
Does Astra’s 55% score mean it completed 55% of the tasks?
No. OpenAI reported that Astra met an average of 55% of the rubric criteria across the evaluation. That is not a task-completion rate, and the score alone does not show whether a required approval or other control was missed.
Did the test prove that agents save customers time?
No. OpenAI described the time figures as simulated estimates based on assumed processing and generation speeds. They were not measured customer time savings and applied to the 11 research tasks.
What data does OpenAI say it used?
OpenAI says it used synthetic training tasks based on publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It says it did not use OpenAI customer data, internal contracts or non-public Ironclad customer data.
Can companies use these agents for contract work now?
The published results do not establish readiness for independent customer use. OpenAI’s account says human oversight remains necessary, and the reported average leaves open which specific workflow requirements were missed.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
