Skip to content
← All Charts
In ChartsPublished

Can AI Spot Your IKEA Mistake?

A photo, an assembly manual and tools. The best score among tested models rose from 28% to 80% in Epoch’s furniture test.

Can AI judge an IKEA assembly photo?BEST TESTED SCORE: 28% → 80%

Lime line: successive best scores among tested models. Gray dots: other tested models. 60 photos, three furniture builds; photos + manual + tools. A retrospective comparison, not a physical assembly test.

Data: Aiden Ament & Greg Burnham · Epoch AI · CC BY 4.0 · Chart: VastkindReport: 23 Sep 2026 · Model releases: Nov 2025–Sep 2026 · Accessed: 6 Oct 2026

A practical test of visual reasoning

A photo, an assembly manual and tools. Epoch’s test asks models to judge the assembly and, if something is wrong, identify the relevant steps and explain the mistake. The published benchmark score rose from 28.3% for Claude Opus 4.5 to 80% for GPT-6 Astra.

Not a promise to fix your furniture

The sample covers only 60 photos from three builds and is partly graded by another AI. It does not establish reliability on other furniture, real-time repairs or physical assembly. The score is not simply the proportion of flawed builds flagged, and the tested models are not a complete historical leaderboard.

Sources, Method & Data

The 21 source rows are ordered by model release date as reported by Epoch, not test date. The line retains each new record among those models; it is not a complete history of all available AI. GPT-5 and Gemini 3 are absent. Published scores are plotted unchanged ×100. The endpoints are 28.3333% and 80%, a gain of 51.6667 percentage points. Source standard errors are 5.8665 and 4.9289 percentage points; all model standard errors are in the table below. They are not 95% confidence intervals.

The dataset has 42 intentionally flawed and 18 correct photos from STÄLL, TONSTAD and GULLABERG builds. Models receive instructions and image/Python tools. Correct step identification and an adequate error description matter; a lenient LLM grades descriptions. The score is not simply the fraction of flawed builds flagged. Some exported scores imply half cases when multiplied by 60; the public method does not explain that granularity, so no exact “x out of 60” count is inferred. No human-performance comparison. Accessed October 6, 2026.

Benchmark methodology · Original chart dataset

All 21 tested models; original Epoch scores
ModelRelease date (Epoch)Score (%)Standard error (pp)
Claude Opus 4.52025-11-2428.33335.8665
GPT-5.22025-12-1138.33336.3298
Claude Opus 4.62026-02-0528.33335.8665
Gemini 3.1 Pro2026-02-1926.66675.6332
GPT-5.42026-03-0537.50006.2465
Claude Opus 4.72026-04-1633.33336.1372
Kimi K2.62026-04-2021.66675.2301
GPT-5.52026-04-2344.16676.4101
Claude Opus 4.82026-05-2842.50006.3807
Claude Fable 52026-06-0935.83336.1859
GPT-5.6 Sol2026-07-0956.66676.3409
GPT-5.6 Terra2026-07-0954.16676.4321
GPT-5.6 Luna2026-07-0942.50006.3807
Kimi K32026-07-1634.16676.1170
Gemini 3.6 Flash2026-07-2123.33335.3766
Claude Opus 52026-07-2460.83336.1859
Gemini 3.7 Flash2026-08-1326.66675.6332
Claude Fable 5.12026-09-0170.00005.7244
Qwen3.8 Max2026-09-0120.00005.2076
Gemini 3.8 Flash2026-09-0231.66675.9383
GPT-6 Astra2026-09-0380.00004.9289

Explore Vastkind

Topics
Subscribe

Stay curious. Stay informed.

For another moment

Your reading list

Saved on this device. No account needed.

Something catch your eye? Use the bookmark beside an article to keep it here.