OpenAI starts training agents on partner software, beginning with Ironclad
GPT-6 Astra scores 55% vs 41.6% for GPT-5.6 Sol on contracting workflows
Confirmed
The takeaway
On Oct 6, 2026 OpenAI said it is partnering with a small number of software companies to turn their products’ hard workflows into training and evaluation tasks for computer-use agents.
Status
- Claim
Confirmed
- Our reporting
- High confidence
- Importance
- 2 of 5
- Last verified
- 8 October 2026
Your AI and this story
- GPT-6 Astra159 days after its cutoff
- Claude Opus 5.598 days after its cutoff
- Gemini 3.8 Flash189 days after its cutoff
- Grok 4.7128 days after its cutoff
None of these four assistants can know about it. The closest, Claude Opus 5.5, stops 98 days before it.
Key facts
- 11 research tasks across legal, commercial and procurement work (NDAs, procurement approval flows, jurisdiction-dependent clauses), each ~30–40 minutes for an experienced user, graded on 8–50 criteria
- Ironclad supplied hosted instances of its product for practice; OpenAI added synthetic training tasks and used reinforcement learning
- Results: GPT-6 Astra (Max reasoning) 55.0% vs GPT-5.6 Sol (High) 41.6%; estimated time per attempt 19.2 vs 37.0 minutes; OpenAI’s headline: 32% higher score, 48% less time
- An internal model used in Astra’s development scored 63.7%; OpenAI aims to bring these gains to future models
- OpenAI invites other software companies to partner on ‘complex professional work’
- Same day: OpenAI and Atlassian expanded their partnership (GPT-6-family models power Rovo agents; 3,000+ Atlassian developers use Codex); a Jump Trading customer post described GPT-6 Astra in long-horizon quant research agents
What happened
OpenAI described a program in which software vendors bring their real customer workflows into model training: a vendor defines the tasks and success criteria, provides a practice environment, and OpenAI trains on it with reinforcement learning. Ironclad, an AI contracting platform, is the first such partner.
Why it matters
It shows how frontier labs are getting the “RL environments” for professional software that agents need, by asking vendors directly. It also reports Astra’s computer-use gains on realistic multi-step enterprise work, and gives a score for an internal successor model (63.7%). The scores come from OpenAI’s own evaluation, not an independent one.
Sources
3 sources from 1 site. Numbers match the chips in the text.
3 sources: 3 primary
Primary
- OpenAI: Advancing computer use with Ironcladopenai.com, official
- OpenAI: Atlassian and OpenAI expand partnership to turn enterprise knowledge into actionopenai.com, official
- OpenAI: How Jump Trading is scaling quant research with ChatGPTopenai.com, official
Changes
- Merged a duplicate entry written in parallel the same day (
2026-10-06-openai-ironclad-computer-use-training: Jump Trading link) - Filed (pages read through a text mirror because openai.com returns 403 to fetchers)