JOURNAL
GPT-6 Astra Scored 55% on Real Contract Work — What That Number Actually Means
OpenAI and Ironclad published a computer-use benchmark on real contracting tasks. The headline number is a 14-point jump over the prior model. The footnotes tell a more careful story.
On this page
On October 6, 2026, OpenAI and Ironclad — a contract lifecycle management company — published results from running GPT-6 Astra’s computer-use mode against 11 real contracting tasks: drafting an NDA, building a procurement approval chain, updating a reusable legal clause to reflect a new jurisdiction, and similar work a contracts team actually does. Astra averaged a score of 55.0% across those tasks. The prior model, GPT-5.6 Sol, averaged 41.6%. A 14.4-point jump, and it’s a measured one — the same 11 tasks, the same grading rubric, two models.
That’s the part of the report worth taking seriously. The parts worth reading twice are in the footnotes.
55% doesn’t mean what it sounds like
Each of the 11 tasks was graded against somewhere between 8 and 50 individual criteria, depending on how complex the task was. A score of 55% is an average of partial credit across all of those criteria, not a pass rate on whole tasks. It is not “Astra drafted roughly half the contracts correctly.” It’s closer to “across every sub-requirement checked — the right clause present, the right party named, the right approval step inserted — Astra satisfied a bit more than half of them, on average.” That’s a meaningfully different claim, and the distinction matters if you’re the one deciding whether this is ready to touch a real contract queue.
The time-savings number is simulated, and the report says so
The figure that traveled furthest from this report — average task time dropping from 37.0 minutes to 19.2 minutes — carries a footnote in the original publication describing it as “simulated estimates based on assumed model processing and generation speeds, not measured customer time savings.” Nobody timed a contracts specialist doing these 11 tasks by hand with a stopwatch. Nobody timed Astra running them against a live Ironclad account either. The number is a model of what the time difference should be, built from assumptions about processing speed, not a measurement of either side actually doing the work. That’s a legitimate thing to publish as a projection. It is not evidence, and the report itself doesn’t claim it is — the care is in the footnote, not the headline.
The demo clips and the data table tell different stories, and both are true
Demo clips that circulated around the release showed individual task scores of 94% and 85% — numbers that look nothing like the 55.0% average across all 11 tasks. Both figures are accurate. The clips show specific tasks, picked for the demo, where Astra did unusually well; the average sits across all 11, including the tasks it handled worse. This is a standard move in vendor benchmark marketing and also a standard trap for anyone skimming a launch page instead of the data table: a best-case clip and an average score are both real numbers, describing different things, and only one of them tells you what to expect on the next contract that lands in the queue.
OpenAI’s own best model didn’t ship
Buried in the same report: an unreleased internal model scored 63.7% on the identical 11 tasks — nearly nine points above the Astra build that actually shipped to customers. OpenAI doesn’t explain the gap in the published material. Whatever the reason — cost, latency, safety review, something else — the number on the market today is not the best number OpenAI has produced internally against this exact test. Worth remembering the next time a vendor publishes a benchmark result: ask whether it’s the released product’s score, or a lab result that the released product hasn’t caught up to yet.
The honest line is in the report, not the pitch
OpenAI’s own writeup states plainly that the real constraint on what a software company can confidently hand to an agent today is a model “losing track of a business rule halfway through a task.” That’s the actual engineering problem, and it’s a more useful sentence than any percentage on the page. A procurement approval chain has business rules at every branch — who approves what dollar amount, which jurisdiction needs an extra signature, which clause variant applies to which contract type. An agent that’s right 55% of the time on average and wrong in a way that quietly drops a rule halfway through isn’t a contracts team member yet. It’s a drafting assistant that needs every output checked against the rules it might have lost track of.
A checklist for the next agent benchmark a vendor hands you
- Is the top-line percentage an average across graded criteria, or a pass rate on whole tasks? They can look identical in a headline and mean very different things in practice.
- Are the time or cost savings measured on real usage, or simulated/estimated? Check the footnote, not the slide.
- Do the demo clips show the average case or the best case out of many runs?
- Is the model being benchmarked the one you can actually buy, or is there a better internal version the vendor hasn’t shipped?
A 14-point jump in grading-criteria coverage across 11 real contracting tasks is a legitimate result. Publishing task-level detail and an honest footnote about simulated time savings, instead of a single marketing percentage, deserves some credit too — plenty of launch pages don’t bother. But “legitimate result” and “ready to run your legal team’s approval chain unsupervised” are different claims, and the distance between them is sitting in the footnotes this entire report. Read those before you read the time-savings number.
Source: OpenAI and Ironclad’s joint benchmark report, published October 6, 2026.
Discussion
Comments are reviewed before publication. Your email is kept private.