Lime line: successive best scores among tested models. Gray dots: other tested models. 60 photos, three furniture builds; photos + manual + tools. A retrospective comparison, not a physical assembly test.
A practical test of visual reasoning
A photo, an assembly manual and tools. Epoch’s test asks models to judge the assembly and, if something is wrong, identify the relevant steps and explain the mistake. The published benchmark score rose from 28.3% for Claude Opus 4.5 to 80% for GPT-6 Astra.
Not a promise to fix your furniture
The sample covers only 60 photos from three builds and is partly graded by another AI. It does not establish reliability on other furniture, real-time repairs or physical assembly. The score is not simply the proportion of flawed builds flagged, and the tested models are not a complete historical leaderboard.
Sources, Method & Data
The 21 source rows are ordered by model release date as reported by Epoch, not test date. The line retains each new record among those models; it is not a complete history of all available AI. GPT-5 and Gemini 3 are absent. Published scores are plotted unchanged ×100. The endpoints are 28.3333% and 80%, a gain of 51.6667 percentage points. Source standard errors are 5.8665 and 4.9289 percentage points; all model standard errors are in the table below. They are not 95% confidence intervals.
The dataset has 42 intentionally flawed and 18 correct photos from STÄLL, TONSTAD and GULLABERG builds. Models receive instructions and image/Python tools. Correct step identification and an adequate error description matter; a lenient LLM grades descriptions. The score is not simply the fraction of flawed builds flagged. Some exported scores imply half cases when multiplied by 60; the public method does not explain that granularity, so no exact “x out of 60” count is inferred. No human-performance comparison. Accessed October 6, 2026.
Benchmark methodology · Original chart dataset
| Model | Release date (Epoch) | Score (%) | Standard error (pp) |
|---|---|---|---|
| Claude Opus 4.5 | 2025-11-24 | 28.3333 | 5.8665 |
| GPT-5.2 | 2025-12-11 | 38.3333 | 6.3298 |
| Claude Opus 4.6 | 2026-02-05 | 28.3333 | 5.8665 |
| Gemini 3.1 Pro | 2026-02-19 | 26.6667 | 5.6332 |
| GPT-5.4 | 2026-03-05 | 37.5000 | 6.2465 |
| Claude Opus 4.7 | 2026-04-16 | 33.3333 | 6.1372 |
| Kimi K2.6 | 2026-04-20 | 21.6667 | 5.2301 |
| GPT-5.5 | 2026-04-23 | 44.1667 | 6.4101 |
| Claude Opus 4.8 | 2026-05-28 | 42.5000 | 6.3807 |
| Claude Fable 5 | 2026-06-09 | 35.8333 | 6.1859 |
| GPT-5.6 Sol | 2026-07-09 | 56.6667 | 6.3409 |
| GPT-5.6 Terra | 2026-07-09 | 54.1667 | 6.4321 |
| GPT-5.6 Luna | 2026-07-09 | 42.5000 | 6.3807 |
| Kimi K3 | 2026-07-16 | 34.1667 | 6.1170 |
| Gemini 3.6 Flash | 2026-07-21 | 23.3333 | 5.3766 |
| Claude Opus 5 | 2026-07-24 | 60.8333 | 6.1859 |
| Gemini 3.7 Flash | 2026-08-13 | 26.6667 | 5.6332 |
| Claude Fable 5.1 | 2026-09-01 | 70.0000 | 5.7244 |
| Qwen3.8 Max | 2026-09-01 | 20.0000 | 5.2076 |
| Gemini 3.8 Flash | 2026-09-02 | 31.6667 | 5.9383 |
| GPT-6 Astra | 2026-09-03 | 80.0000 | 4.9289 |