Skip to content

The future is arriving badly explained.

Let’s make sense of it.

The ideas shaping the future, explained for you.

A minute of perspective

In brief.

Read all editions

Updated on · Morning edition. Four signals worth your attention.

  1. A quantum link connects an ion and a diamond a thousand times a second

    Researchers report entangling a trapped barium ion with a silicon-vacancy center in diamond at an average 1.03 kilohertz, including photon-conversion losses, with 87.9% fidelity. A detected photon signals each successful link. Connecting unlike qubits this quickly could help join processor modules with different strengths. This is a laboratory entanglement result, not a working distributed computer: the fidelity and the rest of the networking system still matter alongside speed.

    arXiv preprint ·

  2. Giving an AI agent more time does not guarantee better work

    In tests of Qwen-model agents on selected machine-learning competitions and the text game Zork, adding clock feedback, deadline enforcement or budget-aware training helped agents finish on time. But punctuality did not reliably turn extra minutes into better results; trained game agents sometimes filled the gap with repeated actions. The study separates two problems often bundled together: meeting a deadline and using the time well. These limited tasks and models do not establish how every production agent behaves.

    arXiv preprint ·

  3. The same robot motion can leave cloth in different places

    Researchers recorded 269 robot rollouts across four fast cloth-handling tasks using three fabrics. Repeating the same trajectory still produced substantial variation, influenced by the fabric and execution speed. Four calibrated cloth simulators failed to reproduce both the observed variability's magnitude and its ordering. The result exposes a problem for training manipulation in simulation: a repeatable robot command does not imply a repeatable material response. This was a controlled experiment, not a test of complete household laundry handling.

    arXiv preprint ·

  4. A network switch helps rebuild data before a storage node fails

    IAPRepair coordinates data migration, reconstruction and computation inside a programmable network switch to speed repair in erasure-coded storage. Its authors tested a Tofino switch with 16 storage nodes and ran larger simulations, reporting faster repair than the compared methods. The practical idea is to avoid overwhelming healthy machines with repair traffic. It assumes an at-risk node has already been identified and can still send data; this is not a new failure predictor or evidence of deployment in a commercial cloud.

    arXiv preprint ·

These preprints were first announced in arXiv's October 9 listings after submission October 7 or 8. Their experiments predate publication; the findings have not been independently replicated here.

A World in Perspective

In Charts.

All Charts
Repeat the same case. The result can change.97 SIMULATED CASES PER MODEL · FOUR RUNS EACH

Numbers inside the bars count cases, not individual runs. “Every time” means success in all four recorded runs, not four chances to succeed once. Two examples, not a model ranking.

Data: Sierra Research · τ-knowledge · Chart: VastkindRecorded: 3–4 Aug 2026 · Submission: 4 Aug 2026 · Retrieved: 7 Oct 2026

Getting it right once is the easier bar

Imagine asking customer service to replace a lost card. A successful interaction is encouraging. The harder question is whether the same request still works when the conversation takes a different path.

In these published simulations, Claude Opus 5 succeeded at least once on 64 of 97 cases; Qwen 3.8 Max did so on 71. Require success in all four runs, and those counts fall to 31 and 34. The middle segment shows cases that sometimes worked and sometimes did not.

A benchmark, not a bank

These are recorded evaluations in a fictional banking environment, with an AI playing the customer. Vastkind grouped the original scores; we did not run the agents again or independently regrade their conversations. The chart measures neither real customer satisfaction nor financial safety.

Both agents received the same embedded case definitions and tool access, but used different reasoning settings. These two examples do not establish a general model ranking or equal-cost comparison. Four runs per case are a small sample, not a guarantee about the next run.

A further reproducibility limit: 15 of the 97 embedded task definitions differ from the code revision named in the run files. They match between the two files analyzed here. The reason for the repository discrepancy is unresolved.

Sources, Method & Data

What we counted. The τ-knowledge banking_knowledge AllTools run files contain 97 case definitions and 388 runs per model, four distinct trial IDs per case. We treated a logged reward within 0.000001 of 1 as success, matching the project metric, then grouped cases by zero, one to three, or four successes. Both selected files contain no infrastructure-error terminations. Bar widths are group counts divided by 97; labels give the counts.

Settings and timing. Qwen 3.8 Max used xhigh reasoning; Claude Opus 5 used max. Both used GPT-5.2 with low reasoning as the customer simulator and a 200-step limit. The metadata evaluation date is August 3, 2026; raw timestamps span August 3–4 without a stated time zone. The submission date is August 4; the hosted files also carry August 4 modification timestamps. This is a new explanation of that snapshot, not a new October test or a comprehensive current leaderboard.

Denominators matter. Individual-run success was 189/388 for Claude and 214/388 for Qwen. That average is not the first attempt. “At least once” counts cases with one or more successes; “every time” counts cases with four. We make no independence assumption or statistical-significance claim.

Original records. Repository snapshot; MIT code license. Own aggregation and graphic, not a copied source figure. The hosted conversation files carry no separate license notice; full dialogs are not republished here.

Task-version boundary. The run files name code commit fc0055dc4e0a316c3f83133267fbd6faaa770992. Fifteen embedded task definitions differ from that revision and from the repository snapshot above. Our counts describe the scores in the published run files. They do not establish an exact reproduction from the named code revision. Other model results were not included because their comparability was not fully verified.

Recorded outcomes: cases, not runs
Model0 of 41 of 42 of 43 of 44 of 4Total
Claude Opus 5331112103197
Qwen 3.8 Max26145183497

Related: Can AI spot an IKEA assembly mistake? A different capability test, with its own limits.

Two AI agents, 97 simulated cases each, four runs per case. Most cases succeed at least once. Far fewer succeed every time.

Explore the Chart
In Charts

Can AI Spot Your IKEA Mistake?

Can AI judge an IKEA assembly photo?BEST TESTED SCORE: 28% → 80%

Lime line: successive best scores among tested models. Gray dots: other tested models. 60 photos, three furniture builds; photos + manual + tools. A retrospective comparison, not a physical assembly test.

Data: Aiden Ament & Greg Burnham · Epoch AI · CC BY 4.0 · Chart: VastkindReport: 23 Sep 2026 · Model releases: Nov 2025–Sep 2026 · Accessed: 6 Oct 2026

A practical test of visual reasoning

A photo, an assembly manual and tools. Epoch’s test asks models to judge the assembly and, if something is wrong, identify the relevant steps and explain the mistake. The published benchmark score rose from 28.3% for Claude Opus 4.5 to 80% for GPT-6 Astra.

Not a promise to fix your furniture

The sample covers only 60 photos from three builds and is partly graded by another AI. It does not establish reliability on other furniture, real-time repairs or physical assembly. The score is not simply the proportion of flawed builds flagged, and the tested models are not a complete historical leaderboard.

Sources, Method & Data

The 21 source rows are ordered by model release date as reported by Epoch, not test date. The line retains each new record among those models; it is not a complete history of all available AI. GPT-5 and Gemini 3 are absent. Published scores are plotted unchanged ×100. The endpoints are 28.3333% and 80%, a gain of 51.6667 percentage points. Source standard errors are 5.8665 and 4.9289 percentage points; all model standard errors are in the table below. They are not 95% confidence intervals.

The dataset has 42 intentionally flawed and 18 correct photos from STÄLL, TONSTAD and GULLABERG builds. Models receive instructions and image/Python tools. Correct step identification and an adequate error description matter; a lenient LLM grades descriptions. The score is not simply the fraction of flawed builds flagged. Some exported scores imply half cases when multiplied by 60; the public method does not explain that granularity, so no exact “x out of 60” count is inferred. No human-performance comparison. Accessed October 6, 2026.

Benchmark methodology · Original chart dataset

All 21 tested models; original Epoch scores
ModelRelease date (Epoch)Score (%)Standard error (pp)
Claude Opus 4.52025-11-2428.33335.8665
GPT-5.22025-12-1138.33336.3298
Claude Opus 4.62026-02-0528.33335.8665
Gemini 3.1 Pro2026-02-1926.66675.6332
GPT-5.42026-03-0537.50006.2465
Claude Opus 4.72026-04-1633.33336.1372
Kimi K2.62026-04-2021.66675.2301
GPT-5.52026-04-2344.16676.4101
Claude Opus 4.82026-05-2842.50006.3807
Claude Fable 52026-06-0935.83336.1859
GPT-5.6 Sol2026-07-0956.66676.3409
GPT-5.6 Terra2026-07-0954.16676.4321
GPT-5.6 Luna2026-07-0942.50006.3807
Kimi K32026-07-1634.16676.1170
Gemini 3.6 Flash2026-07-2123.33335.3766
Claude Opus 52026-07-2460.83336.1859
Gemini 3.7 Flash2026-08-1326.66675.6332
Claude Fable 5.12026-09-0170.00005.7244
Qwen3.8 Max2026-09-0120.00005.2076
Gemini 3.8 Flash2026-09-0231.66675.9383
GPT-6 Astra2026-09-0380.00004.9289
Explore the Chart
In Charts

Who Is Gaining Ground in Energy?

1. Energy supply, beyond electricity2025 · EXAJOULES (EJ)

Power, transport, heat and industrial uses. Fossil = oil + gas + coal. EI coverage excludes some traditional biomass. More supply is not automatically greater efficiency or independence.

Data: Energy Institute · Statistical Review 2026 · Chart: VastkindData: 2025 · Comparison: 2020–2025 · Released: 30 Jun 2026
2. What is driving power growth?CHANGE IN ANNUAL GENERATION · 2020 → 2025 · TWh

+ more / − less annual generation. China adds 1,583 TWh of wind and solar output; the other three together add 874 TWh. Other includes bioenergy, other renewables and other fossil.

Data: Ember · Global Electricity Review 2026 · CC BY 4.0 · Chart: Vastkind2025 generation estimates may be revised · Report: 21 Apr 2026 · Accessed: 6 Oct 2026

Different sources of growth

The U.S. adds solar and gas. India adds coal. In the EU, wind and solar rise as other output falls. China’s wind and solar gain is larger than the other three combined, but its coal generation grows too.

The comparison concerns changes in annual output, not the cumulative electricity generated over five years.

Energy is bigger than electricity

The first panel includes fuel use beyond power stations. The second examines electricity generation only. The two panels use different measures and cannot be added. Higher energy supply alone is not a ranking of technology, efficiency or energy independence.

2020 was a pandemic year. These are 2020–2025 comparisons, not forecasts or claims about 2026 growth.

Sources, Method & Data

Panel 1 uses EI Statistical Review 2026 total energy supply under its physical-energy-content method, in EJ. Fossil = oil + gas + coal. Non-fossil = nuclear + hydro + solar + wind + other renewables, as covered by EI; commercially recorded energy including modern renewables, not complete coverage of traditional/noncommercial biomass. TES is not useful final energy, domestic production or energy independence. Wind/solar enter as electricity output; thermal sources include conversion losses, so source shares are not equivalent to electricity shares. China means mainland China; U.S. territories are excluded. The official Total EU aggregate covers the current 27 members. 2025 annual values; 2020–2025 change is a ratio of annual totals, not EI’s leap-year-adjusted growth series. Numbers rounded for display. Direct EI downloads were blocked; unchanged EI CSV/XLSX snapshots archived by Our World in Data were used, with matching source checksums.

Panel 2 uses Ember’s separate gross-electricity-generation dataset. 2025 values are estimates based on monthly generation data and may be revised. Each cell is annual 2025 output minus annual 2020 output, in TWh, not cumulative generation over five years. Other = bioenergy + other renewables + other fossil. The net row is the total-generation difference; displayed rounded cells may not add exactly. An output increase is not the same as installed capacity: weather, plant use and outages matter. Renewables and thermal energy have different conversion losses. The four regions are selected comparisons, not a world total or a country ranking of technological quality.

China’s wind/solar change: 913.83 + 668.95 = 1,582.78 TWh. Other three: 384.55 + 303.08 + 186.63 = 874.26 TWh. EU net output grows 43.80 TWh; wind/solar grows 303.08 while fossil output falls 200.09. Sources published June 30 and April 21, 2026; accessed October 6, 2026.

EI data and methodology · Ember source and methods

2025 energy supply and change since 2020
Region2025 EJFossil share %2020 EJChange %
China162.19387.55135.448+19.75
USA93.82983.1786.185+8.87
EU2752.19373.2053.458-2.37
India39.09992.9930.055+30.09
Energy supply by source, 2025 (EJ)
RegionOilGasCoalNuclearHydroSolarWindOther renewables
China33.84515.91092.2395.2935.0124.2244.0611.611
USA36.42432.8848.7309.0120.8791.4141.6892.797
EU2721.76611.8524.5897.1031.1551.3401.7282.660
India10.8922.39623.0710.5870.6420.6070.3740.530
Electricity change: annual 2025 minus annual 2020 (TWh)
SourceChinaU.S.EU-27India
Solar+913.83+258.10+222.44+143.08
Wind+668.95+126.45+80.64+43.55
Coal+835.09-36.24-95.54+330.37
Gas+81.92+183.17-93.24-26.25
Nuclear+121.98-5.10-31.10+9.22
Hydro+77.57-38.25-20.72+14.03
Other+113.60-11.52-18.68+4.17
Explore the Chart
All ChartsEvery chart. Every source.
Explore by topic

Pick a frontier.

Seven fields. A clear place to begin.

The Vastkind Briefing

Stay ahead
of the story.

A personal take on the week. By Dorian.

Check your inbox.

Click the confirmation link in your email to complete your subscription. If it hasn’t arrived, check your spam folder.

Find something to read

Occasional editions. Free to subscribe.
Unsubscribe anytime.

Explore Vastkind

Topics
Subscribe

Stay curious. Stay informed.

For another moment

Your reading list

Saved on this device. No account needed.

Something catch your eye? Use the bookmark beside an article to keep it here.