What did OpenAI and Ironclad evaluate?

OpenAI announced Ironclad as its first software-company partner in a research program focused on complex professional computer use. The collaboration turns real contracting work into tasks that can be used to train and evaluate an AI agent. OpenAI says Ironclad employees and people who use Ironclad at OpenAI helped identify 11 tasks across legal, commercial, and procurement work.

The examples go beyond summarizing a contract. They include setting up nondisclosure agreements, configuring procurement approval processes, and updating a reusable legal clause so it changes with the jurisdiction selected by a requester. OpenAI estimates that an experienced Ironclad user would need about 30 to 40 minutes for each task on average.

Why are these tasks harder than clicking the right buttons?

A contracting workflow contains connected rules. A software purchase above a threshold may need Finance approval, a security-sensitive request may need Security review, and unusual terms may need Legal review. An agent must translate those requirements into forms, document templates, approval logic, and final records while keeping the relationships intact. This is the difference between isolated interface actions and a computer-use agent that remains oriented through a multi-step job.

A single correct action is not enough. If an agent creates an approval rule but applies it to the wrong amount, forgets an exception, or fails to test both sides of a condition, the finished workflow can still be wrong. That makes end-state verification especially important in legal and procurement software, where a plausible-looking configuration can hide a broken business rule.

How was the research setup built?

OpenAI says each task was scored against 8 to 50 criteria depending on its complexity. Ironclad provided hosted product environments where the models could practice. OpenAI researchers also created synthetic training tasks based on contracts available through the US Securities and Exchange Commission's EDGAR database, after applying filters intended to remove personal information, and used reinforcement learning to improve performance through practice and feedback.

The design is more informative than a one-number benchmark because it combines domain-selected tasks, detailed rubrics, and an executable environment. Even so, it remains an internal collaboration rather than an independent audit. A useful agent evaluation suite should also include repeat runs, hidden cases, recovery from tool errors, permission boundaries, and tests for whether the system notices incomplete work.

What do the reported scores actually show?

Using the reasoning setting where each model performed best, GPT-6 Astra received a 55.0% mean rubric score across the 11 tasks. GPT-5.6 Sol received 41.6%. OpenAI describes that as a 32% relative improvement. The estimated average time per attempt fell from 37.0 minutes for Sol to 19.2 minutes for Astra, a reduction of about 48%.

Those figures are evidence of progress on this test, not proof that the model can reliably run every contracting process. OpenAI explicitly says the time figures are simulated estimates based on assumed processing and generation speeds, not measured customer time savings. As our benchmark-reading guide explains, the task set, model settings, rubric, environment, and evaluator all shape the result.

OpenAI also reports that one internal model used during Astra's development reached 63.7%. That is a research direction, not a product promise. The published article does not establish when or whether that internal version will reach customers.

Can teams use this Ironclad agent today?

The announcement describes research and a partner program, not a new standalone product with public pricing or a general release date. GPT-6 Astra is available through OpenAI products and the API under OpenAI's existing rollout, but access to Astra does not automatically reproduce the Ironclad training environment, task harness, permissions, or workflow integrations used in this study.

OpenAI is inviting a small number of software companies to propose difficult professional tasks, provide specialists who understand the work, offer a secure testing environment, and supply data that can safely be used for research. For ordinary Ironclad customers, the practical takeaway is to check their existing product plan and administrator settings rather than assume this research agent has appeared in their account.

What are the implications and limitations?

The useful idea is that professional agents may improve faster when software vendors help define what success means. General browser benchmarks can test navigation, but domain experts know the exceptions, dependencies, and final checks that make a workflow trustworthy. That approach could transfer to accounting, customer support, engineering administration, and other structured business systems.

The limitation is reliability. A 55.0% rubric score leaves substantial room for missed requirements, even if it is better than the comparison model. Contracting also involves confidential data, authority limits, and consequences that are not captured by task completion alone. Organizations still need permissions, audit trails, representative testing, and human approval gates before allowing an agent to publish or change a live workflow.

OpenAI says the synthetic training material came from public SEC contracts and that it did not use OpenAI customer data, internal OpenAI contracts, or nonpublic Ironclad customer contracts for this training and evaluation. That narrows one data concern, but each future deployment will still need its own privacy, retention, access-control, and legal review.

Common questions

Did GPT-6 Astra complete the Ironclad tasks perfectly?

No. OpenAI reports a 55.0% average rubric score across the 11 tasks. That was higher than GPT-5.6 Sol's 41.6%, but it still indicates missed criteria.

Does the 48% time reduction mean customers will save 48% of their time?

No. OpenAI says the times are simulated estimates based on assumed model speeds, not observed customer productivity or a controlled workplace study.

Was private customer contract data used for training?

OpenAI says the synthetic tasks used filtered contracts from the public SEC EDGAR database and did not use nonpublic Ironclad customer contracts, OpenAI customer data, or OpenAI's internal contracts for this work.

THE TAKEAWAY

What to remember

OpenAI's Ironclad collaboration is a concrete attempt to train and measure agents on connected business rules rather than isolated clicks. The reported improvement is meaningful within the 11-task study, while the modest absolute score, simulated timing, and lack of a general product launch make human review and deployment-specific testing essential.

Sources & further reading

  1. Advancing computer use with Ironclad ↗
  2. GPT-6 Astra: A new generation of intelligence ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories