Announced 6 Oct 2026 · Sources checked
What did OpenAI and Ironclad evaluate?
OpenAI announced Ironclad as its first software-company partner in a research program focused on complex professional computer use. The collaboration turns real contracting work into tasks that can be used to train and evaluate an AI agent. OpenAI says Ironclad employees and people who use Ironclad at OpenAI helped identify 11 tasks across legal, commercial, and procurement work.
The examples go beyond summarizing a contract. They include setting up nondisclosure agreements, configuring procurement approval processes, and updating a reusable legal clause so it changes with the jurisdiction selected by a requester. OpenAI estimates that an experienced Ironclad user would need about 30 to 40 minutes for each task on average.
How was the research setup built?
OpenAI says each task was scored against 8 to 50 criteria depending on its complexity. Ironclad provided hosted product environments where the models could practice. OpenAI researchers also created synthetic training tasks based on contracts available through the US Securities and Exchange Commission's EDGAR database, after applying filters intended to remove personal information, and used reinforcement learning to improve performance through practice and feedback.
The design is more informative than a one-number benchmark because it combines domain-selected tasks, detailed rubrics, and an executable environment. Even so, it remains an internal collaboration rather than an independent audit. A useful agent evaluation suite should also include repeat runs, hidden cases, recovery from tool errors, permission boundaries, and tests for whether the system notices incomplete work.
What do the reported scores actually show?
Using the reasoning setting where each model performed best, GPT-6 Astra received a 55.0% mean rubric score across the 11 tasks. GPT-5.6 Sol received 41.6%. OpenAI describes that as a 32% relative improvement. The estimated average time per attempt fell from 37.0 minutes for Sol to 19.2 minutes for Astra, a reduction of about 48%.
Those figures are evidence of progress on this test, not proof that the model can reliably run every contracting process. OpenAI explicitly says the time figures are simulated estimates based on assumed processing and generation speeds, not measured customer time savings. As our benchmark-reading guide explains, the task set, model settings, rubric, environment, and evaluator all shape the result.
OpenAI also reports that one internal model used during Astra's development reached 63.7%. That is a research direction, not a product promise. The published article does not establish when or whether that internal version will reach customers.
Can teams use this Ironclad agent today?
The announcement describes research and a partner program, not a new standalone product with public pricing or a general release date. GPT-6 Astra is available through OpenAI products and the API under OpenAI's existing rollout, but access to Astra does not automatically reproduce the Ironclad training environment, task harness, permissions, or workflow integrations used in this study.
OpenAI is inviting a small number of software companies to propose difficult professional tasks, provide specialists who understand the work, offer a secure testing environment, and supply data that can safely be used for research. For ordinary Ironclad customers, the practical takeaway is to check their existing product plan and administrator settings rather than assume this research agent has appeared in their account.
What are the implications and limitations?
The useful idea is that professional agents may improve faster when software vendors help define what success means. General browser benchmarks can test navigation, but domain experts know the exceptions, dependencies, and final checks that make a workflow trustworthy. That approach could transfer to accounting, customer support, engineering administration, and other structured business systems.
The limitation is reliability. A 55.0% rubric score leaves substantial room for missed requirements, even if it is better than the comparison model. Contracting also involves confidential data, authority limits, and consequences that are not captured by task completion alone. Organizations still need permissions, audit trails, representative testing, and human approval gates before allowing an agent to publish or change a live workflow.
OpenAI says the synthetic training material came from public SEC contracts and that it did not use OpenAI customer data, internal OpenAI contracts, or nonpublic Ironclad customer contracts for this training and evaluation. That narrows one data concern, but each future deployment will still need its own privacy, retention, access-control, and legal review.
Common questions
Did GPT-6 Astra complete the Ironclad tasks perfectly?
No. OpenAI reports a 55.0% average rubric score across the 11 tasks. That was higher than GPT-5.6 Sol's 41.6%, but it still indicates missed criteria.
Does the 48% time reduction mean customers will save 48% of their time?
No. OpenAI says the times are simulated estimates based on assumed model speeds, not observed customer productivity or a controlled workplace study.
Was private customer contract data used for training?
OpenAI says the synthetic tasks used filtered contracts from the public SEC EDGAR database and did not use nonpublic Ironclad customer contracts, OpenAI customer data, or OpenAI's internal contracts for this work.
What to remember
OpenAI's Ironclad collaboration is a concrete attempt to train and measure agents on connected business rules rather than isolated clicks. The reported improvement is meaningful within the 11-task study, while the modest absolute score, simulated timing, and lack of a general product launch make human review and deployment-specific testing essential.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





