OpenAI Is Training Agents Inside Your Software. Read The Fine Print On Ironclad.
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Is Training Agents Inside Your Software. Read The Fine Print On Ironclad. on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI says it trained GPT-6 Astra using hosted copies of Ironclad’s contract-management software and 11 legal, commercial and procurement tasks. Astra met an average 55% of task criteria; OpenAI’s time figures are simulations, not measured customer savings. The project also invites selected software companies to supply challenging workflows for agent research.

OpenAI has described training its GPT-6 Astra model inside hosted copies of Ironclad’s contract-management software, using 11 legal, commercial and procurement tasks. The results, published October 6, show an average of 55% of evaluation criteria met—not 55% of tasks completed—and OpenAI says its estimated task times are simulations rather than measured customer savings.

Ironclad employees and OpenAI staff who use the product selected 11 workflows, including setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause based on a requester’s selected jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task. Each task was scored against a rubric of 8 to 50 criteria, depending on its complexity.

OpenAI says Ironclad provided hosted product copies for model practice. It also says it generated synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, after filtering out personal information. According to OpenAI, the work did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.

In OpenAI’s comparison, GPT-6 Astra met an average 55.0% of rubric criteria, against 41.6% for GPT-5.6 Sol in a high-compute configuration. Astra’s estimated time per attempt was 19.2 minutes, compared with 37 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%. On one featured task, Astra met about 94% of the criteria. These are reported evaluation results; they do not establish that the model can reliably complete contract work for customers.

At a glance
reportWhen: Published October 6; the source does no…
The developmentOpenAI published details of a collaboration with Ironclad to train and evaluate an AI model on contract-software workflows, and invited other software vendors to explore similar research partnerships.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Work Needs More Than Partial Credit

The reported score is difficult to translate into business value because a workflow can fail if it misses even one required control. OpenAI’s example describes a procurement process that must route spending above a threshold to Finance, certain requests to Security and nonstandard terms to Legal. Getting two of those three rules right would still leave a required approval unaddressed; a high average score across criteria does not show whether a particular workflow is safe to use.

That makes the evaluation useful as a measure of progress, but not proof of deployment readiness. OpenAI’s own account says that losing track of a business rule limits what a software company can confidently ask an agent to do, and that human oversight remains necessary. For businesses evaluating agents, the missed criteria and the severity of each failure matter at least as much as an average percentage.

The project may also matter to software vendors beyond Ironclad. OpenAI says it is seeking a small number of software-company partners with difficult agent tasks, knowledgeable staff, secure test environments and data suitable for research. Such work could help models operate specialized products, while increasing the importance of the business rules, records and controls behind those products—not just their screens.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How OpenAI Built the Ironclad Test

The collaboration is presented as research into whether models can follow company rules through multi-step work in specialized software, then check their output against the original requirements. Rather than assessing general computer use alone, the test put models in a hosted version of a real product and used tasks chosen by people familiar with its workflows.

OpenAI’s post also frames the work as a way to study where agents fail. It argues that contracting platforms remain important because they provide business rules and controls that an agent must respect. The published results cover 11 selected tasks, however, and should not be read as a broad benchmark of all Ironclad features, all contract work or all customer environments.

The time comparison needs similar care. OpenAI explicitly describes the 19.2- and 37-minute figures as simulated estimates based on assumed processing and generation speeds. They are not observations of employees using the software, and the source says they apply to the research tasks rather than Ironclad workflows generally.

What the Evaluation Does Not Establish

The published summary does not provide enough detail to determine how the 11 tasks were selected beyond the roles involved, how often each was attempted, or which specific criteria Astra missed across the full set. It also does not establish how the model would perform on new workflows, unusual contract language or live customer data. Those details would be needed to judge reliability for particular uses.

The 55% average is a share of rubric criteria met, not a percentage of tasks completed successfully. OpenAI’s simulated times likewise do not show actual productivity gains. The source material does not give a publication year for the October 6 post, nor does it describe a customer deployment, a release timeline or a measured business outcome. It remains unclear which other software companies, if any, will take part in the proposed research partnerships.

What Vendors and Buyers Should Ask

OpenAI says it is inviting a small number of software companies to bring concrete examples of work current agents cannot reliably complete, along with domain experts, secure test environments and research-appropriate data. The source does not identify future partners or give dates for another evaluation, so the next public milestone is not yet known.

For software vendors and buyers, the useful next step is to ask for more than an average score: which requirements failed, how serious those failures were, and whether results hold up on workflows outside the test set. Organizations considering agents for contracts or procurement should also clarify what human review is required and how approval rules, records and controls are maintained. Until such evidence is available, the reported results support continued testing—not an assumption that agents can independently handle these workflows.

Key Questions

What did OpenAI and Ironclad test?

They tested a model on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management product, including NDA setup and procurement approvals.

Does Astra’s 55% score mean it completed 55% of the tasks?

No. OpenAI reported that Astra met an average of 55% of the rubric criteria across the evaluation. That is not a task-completion rate, and the score alone does not show whether a required approval or other control was missed.

Did the test prove that agents save customers time?

No. OpenAI described the time figures as simulated estimates based on assumed processing and generation speeds. They were not measured customer time savings and applied to the 11 research tasks.

What data does OpenAI say it used?

OpenAI says it used synthetic training tasks based on publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It says it did not use OpenAI customer data, internal contracts or non-public Ironclad customer data.

Can companies use these agents for contract work now?

The published results do not establish readiness for independent customer use. OpenAI’s account says human oversight remains necessary, and the reported average leaves open which specific workflow requirements were missed.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Incident postmortem builder for managed service providers

A new incident postmortem builder aimed at small managed service providers is being tested to streamline post-incident analysis and communication.

ChatGPT Rank Monitor

A new ChatGPT rank monitor is being tested to help brands track their AI search presence, offering insights into share-of-voice and citations in AI answers.

Lessons From Other Tech Giants

An analysis of how historical tech giants’ failures from platform shifts offer lessons for current AI industry leaders.

Forge Oder Self-Hosting? Die Wahren Kosten Souveräner KI

Analyse der wahren Kosten von Self-Hosting versus Cloud-Lösungen für souveräne KI, basierend auf aktuellen Markt- und Technologieentwicklungen.