GPT-6 Astra launch visual from OpenAI

GPT-6 Astra: what the charts mean for hardware teams

OpenAI launched GPT-6 Astra in September 2026. The useful story for a hardware team is not that it writes nicer chat replies. It operates CAD, KiCad, terminals, and browsers long enough to finish a job. That is closer to a junior engineer at a desk than an autocomplete box.

This note is for procurement, firmware, and founding teams who have to decide what Astra changes on a real OEM program. Scores come from OpenAI’s launch post, ARC Prize, and third-party indexes. We did not run these benches in our lab.

What the charts actually say

Astra leads on the jobs that look like factory and lab work: terminal agents, CAD reconstruction, and computer use. It does not sweep every general-intelligence index. Claude Fable 5.1 still leads Artificial Analysis Intelligence Index v4.1.1 (65.7 vs Astra 61.2). Claude Opus 5 still leads the Coding Agent Index (68.1 vs Astra 67.0). Read the scoreboard as a map of strengths, not a trophy.

Terminal-Bench 4.0 (OpenAI launch scores)
GPT-6 Astra57.9%
Claude Fable 5.155.8%
GPT-5.6 Sol 237.3%
Gemini 3.8 Flash19.1%

Source: OpenAI, GPT-6 Astra launch. Estimated API cost per task is about 9% lower than Sol 2 and 63% lower than Fable 5.1 in the configurations they published.

Jobs that look like engineering desks
Benchmark GPT-6 Astra GPT-5.6 Sol Claude Fable 5.1
Terminal-Bench Science 0.1 64.6% 22.4% 52.6%
BenchCAD (with tools) 95.9% 83.3% 84.3%
OSWorld 2.0 (offline, partial) 72.6% 65.7% n/a
AutomationBench 41.4% 18.1% 31.4%
FrontierMath Tier 4 (v2) 97.6% 83.0% 87.8%

Source: OpenAI launch tables. OSWorld Fable cell was not in the same comparison row. On OSWorld latency simulations, Astra scored 72.6% in about 40 minutes per task vs Sol 65.7% in about 75 minutes.

ARC-AGI-3 leaderboard showing GPT-6 Astra Standard and Provider Adapter results
ARC Prize leaderboard for GPT-6 Astra on ARC-AGI-3. Chart: ARC Prize, 3 Sep 2026.

ARC-AGI-3 is the chart people will screenshot. Treat the 99.9% number with the harness label attached. On ARC Prize’s standard harness (the model keeps only the notes it writes down), Astra max scored 62.7% on the semi-private set for about $26k of API spend. On the Provider Adapter (OpenAI keeps opaque reasoning state between calls), Astra high scored 99.9% for about $19k. Both are state of the art. They are not the same test.

Scatter plot of Astra actions versus the human baseline on completed ARC-AGI-3 levels
Action efficiency vs the human median. Points below the line used fewer actions than people. Chart: ARC Prize.

On the Provider Adapter runs, Astra max used fewer actions than the human baseline on 96% of levels, and about 52% fewer actions per level on average. Greg Kamradt at ARC Prize called that human parity on action efficiency, not a claim of AGI. The environments are closed, deterministic games. A SMT line is not.

Technical breakthroughs: loops, memory, and computer use

OpenAI’s public writeup stresses three product bets: pre-training, reinforcement learning, and alignment. The architecture rumor is more specific.

The Information reported that Astra uses recurrent depth, also called looped transformers. The idea is old. A Universal Transformer (2018) and later papers such as Geiping et al. on latent recurrent depth reuse the same block weights across several passes. You buy extra thinking depth without storing a second copy of the weights. Training still pays for those extra forward and backward passes. At a fixed compute budget, the literature shows a modest quality gain and a useful property at inference: you can spend more loops on hard tokens and fewer on easy ones.

Sebastian Raschka’s technical note is the sober read. Looped stacks are a plausible ingredient. They are not a confirmed dump of Astra’s source. OpenAI chief scientist Jakub Pachocki said the computation-graph depth of current frontier models, Astra included, is still within a factor of two of GPT-4. That leaves room for a loop of two, not a 20-times-deeper secret brain. Treat “Astra is a looped transformer” as a research hypothesis, not a datasheet line.

What OpenAI did confirm in product terms:

  • Computer use as a first-class skill. OSWorld, ScreenSpot-Pro, and Agents’ Last Exam all move. Codex’s harness is faster. OpenAI cites about 1.9x faster completion vs Sol on Mind2Web when the new harness and Astra run together.
  • Notes that survive compaction. In Codex, Astra can keep searchable notes across context windows instead of crushing the whole session into one summary. That matters when a DFM thread lasts three days.
  • Shorter written traces. The system card says written reasoning is harder to monitor than Sol’s, because Astra solves more with fewer visible steps. If you run agents on a factory network, that is a harness problem, not a marketing footnote. You log tools and file diffs, not only the chat.
  • Tighter task boundaries. OpenAI’s impossible-task eval: Sol stepped outside the authorized target 48% of the time without production safeguards. Astra did so in 0% of cases on that eval. Alignment is part of the product, not a sticker.

The science numbers are loud (FrontierMath Tier 4 about 98%, GPQA Diamond 96%). For an OEM, the CAD and PCB demos are the ones that change a weekly.

Use cases people already ran

KiCad, on OpenAI’s own tape. The launch video is a 15-second cut of Astra placing parts and routing copper from a schematic. OpenAI’s claim is a layout pass in under three minutes. That is not a finished Gerber package. It is a first pass a layout engineer would still DFM. For a Hong Kong program that already owns SMT, a faster first board still cuts days off EVT if someone reviews stackup, clearances, and test points.

FreeCAD to Blender. OpenAI showed a written brief becoming a five-speed transmission concept in FreeCAD, then the same gears animated in Blender. BenchCAD (reconstruct 3D from multi-view renders into CAD code) is the numbered version of that demo: Astra 95.9% geometric overlap vs Sol 83.3% and Fable 84.3%.

SOLIDWORKS, outside OpenAI’s studio. Alexandre Senet at MecAgent published an Astra run in their harness against SOLIDWORKS 2026: one prompt, 41 parts, two subassemblies, one top assembly, 57 minutes. Parametric feature tree, still editable. Oleg Shilovitsky covered it on Beyond PLM (20 Sep 2026). The scarce resource after a run like that is not “can the model click Extrude.” It is whether the assembly remembers why a wall is 2.0 mm, and whether that number survives a cost-down.

FRC intake in Inventor. A Chief Delphi write-up connected Codex to Autodesk Inventor through a local MCP server and asked Astra to CAD an offseason slapdown intake on a 27-inch chassis. The model mixed API scripts, computer use, and rendered views. The author called it a first pass, not a week-six robot. The interesting part for us is the loop: brief, CAD, screenshot, revise. That is how DFM reviews already work, with a human owning steel and certs.

Software shops on day one. Cognition put Astra into Devin’s harness at launch. Jane Street reported cleaner agentic coding traces. Lovable saw more verification through browser testing when effort is turned up. Higgsfield said complex creative workflows used up to 20% fewer tokens. Harvey said legal drafts distinguished documents from established records. Those are not factory stories. They are evidence the same computer-use stack is already in paid products.

Unity city from a prompt. OpenAI also showed a 3D scene build. Treat it as a visualization path next to CAD, not as a substitute for a STEP file you can machine.

What this does not replace on an OEM line

Astra can place footprints. It cannot sign an NDA, hold an AVL, or sit an FCC sample. A looped block that thinks twice in latent space still does not know your mold draft or your BSCI audit date unless you put those in the harness.

If you put Astra on firmware or layout:

  1. Keep a harness and eval set that matches your repo, not a public SWE bench.
  2. Log tool calls and CAD diffs. Do not trust the short written chain of thought as the full record.
  3. Run DFM before steel. A pretty KiCad board that fails creepage is still scrap.
  4. Tie agents to devices you can build. That is the Physical AI job: sensors, edge compute, and a plant that can repeat the unit.

Techwall’s part is unchanged. Hong Kong owns the commercial reply. Southern China runs SMT and assembly under ISO 9001 and BSCI. An engineer still answers within one business day. Astra is a new intern who can drive KiCad at 2 a.m. It is not the factory.

Sources: OpenAI, “GPT-6 Astra: A new generation of intelligence”; ARC Prize, “OpenAI’s GPT-6 Astra on ARC-AGI-3”; Sebastian Raschka on looped transformers; Beyond PLM on the MecAgent SOLIDWORKS run; Chief Delphi on the FRC Inventor experiment. Artificial Analysis index numbers as reported in launch roundups. We are not affiliated with OpenAI.

Back to blog