Skip to content

The future is arriving badly explained.

Let’s make sense of it.

The ideas shaping the future, explained for you.

Updated on · Morning edition. Four signals worth your attention.

  1. A hydrogen catalyst pairs tiny interfaces with long endurance

    A nickel-based catalyst ran a laboratory water electrolyzer at 2 amps per square centimeter and 1.62 volts at 80°C. A separate endurance test exceeded 2,200 hours at 25°C. The results suggest promising materials performance, but efficiency and durability were measured under different conditions—not proof of cheaper industrial hydrogen.

    Nature Communications ·

  2. Small language models suggest better columns for a dataset

    Small language models proposed extra inputs for tabular machine-learning datasets, with an expert checking sources. One dataset showed an improved random-forest classification score, but results varied across tasks. The approach could help researchers design better inputs; it does not automatically discover verified facts or remove the need for human review.

    Scientific Reports ·

  3. A chemical test makes fungal pigment in soil easier to measure

    Researchers used liquid chromatography to measure melanin markers in soil, adding a chemical measure to existing genetic clues. Results were repeatable in tested samples, and melanin levels correlated with soil carbon. The assay could improve soil research; it does not show that adding fungi or pigment would store more carbon.

    Nature Communications ·

  4. The same lymphoma cells meet different immune surroundings

    In mice, the same lymphoma cells produced different immune profiles in brain tumors and tumors under the skin. The comparison helps isolate location’s influence from the cancer cells themselves. It measured immune markers, not human treatment effects, and did not directly test improved T-cell killing or antigen presentation.

    Scientific Reports ·

These four journal papers were published on October 11. Their experiments and underlying datasets predate publication. Laboratory performance, analytical measurements and mouse immune profiles are not commercial or clinical validation.

A World in Perspective

In Charts.

All Charts
Repeat the same case. The result can change.97 SIMULATED CASES PER MODEL · FOUR RUNS EACH

Numbers inside the bars count cases, not individual runs. “Every time” means success in all four recorded runs, not four chances to succeed once. Two examples, not a model ranking.

Data: Sierra Research · τ-knowledge · Chart: VastkindRecorded: 3–4 Aug 2026 · Submission: 4 Aug 2026 · Retrieved: 7 Oct 2026

Getting it right once is the easier bar

Imagine asking customer service to replace a lost card. A successful interaction is encouraging. The harder question is whether the same request still works when the conversation takes a different path.

In these published simulations, Claude Opus 5 succeeded at least once on 64 of 97 cases; Qwen 3.8 Max did so on 71. Require success in all four runs, and those counts fall to 31 and 34. The middle segment shows cases that sometimes worked and sometimes did not.

A benchmark, not a bank

These are recorded evaluations in a fictional banking environment, with an AI playing the customer. Vastkind grouped the original scores; we did not run the agents again or independently regrade their conversations. The chart measures neither real customer satisfaction nor financial safety.

Both agents received the same embedded case definitions and tool access, but used different reasoning settings. These two examples do not establish a general model ranking or equal-cost comparison. Four runs per case are a small sample, not a guarantee about the next run.

A further reproducibility limit: 15 of the 97 embedded task definitions differ from the code revision named in the run files. They match between the two files analyzed here. The reason for the repository discrepancy is unresolved.

Sources, Method & Data

What we counted. The τ-knowledge banking_knowledge AllTools run files contain 97 case definitions and 388 runs per model, four distinct trial IDs per case. We treated a logged reward within 0.000001 of 1 as success, matching the project metric, then grouped cases by zero, one to three, or four successes. Both selected files contain no infrastructure-error terminations. Bar widths are group counts divided by 97; labels give the counts.

Settings and timing. Qwen 3.8 Max used xhigh reasoning; Claude Opus 5 used max. Both used GPT-5.2 with low reasoning as the customer simulator and a 200-step limit. The metadata evaluation date is August 3, 2026; raw timestamps span August 3–4 without a stated time zone. The submission date is August 4; the hosted files also carry August 4 modification timestamps. This is a new explanation of that snapshot, not a new October test or a comprehensive current leaderboard.

Denominators matter. Individual-run success was 189/388 for Claude and 214/388 for Qwen. That average is not the first attempt. “At least once” counts cases with one or more successes; “every time” counts cases with four. We make no independence assumption or statistical-significance claim.

Original records. Repository snapshot; MIT code license. Own aggregation and graphic, not a copied source figure. The hosted conversation files carry no separate license notice; full dialogs are not republished here.

Task-version boundary. The run files name code commit fc0055dc4e0a316c3f83133267fbd6faaa770992. Fifteen embedded task definitions differ from that revision and from the repository snapshot above. Our counts describe the scores in the published run files. They do not establish an exact reproduction from the named code revision. Other model results were not included because their comparability was not fully verified.

Recorded outcomes: cases, not runs
Model0 of 41 of 42 of 43 of 44 of 4Total
Claude Opus 5331112103197
Qwen 3.8 Max26145183497

Related: Can AI spot an IKEA assembly mistake? A different capability test, with its own limits.

Two AI agents, 97 simulated cases each, four runs per case. Most cases succeed at least once. Far fewer succeed every time.

Explore the Chart
In Charts

Can AI Spot Your IKEA Mistake?

Can AI judge an IKEA assembly photo?BEST TESTED SCORE: 28% → 80%

Lime line: successive best scores among tested models. Gray dots: other tested models. 60 photos, three furniture builds; photos + manual + tools. A retrospective comparison, not a physical assembly test.

Data: Aiden Ament & Greg Burnham · Epoch AI · CC BY 4.0 · Chart: VastkindReport: 23 Sep 2026 · Model releases: Nov 2025–Sep 2026 · Accessed: 6 Oct 2026

A practical test of visual reasoning

A photo, an assembly manual and tools. Epoch’s test asks models to judge the assembly and, if something is wrong, identify the relevant steps and explain the mistake. The published benchmark score rose from 28.3% for Claude Opus 4.5 to 80% for GPT-6 Astra.

Not a promise to fix your furniture

The sample covers only 60 photos from three builds and is partly graded by another AI. It does not establish reliability on other furniture, real-time repairs or physical assembly. The score is not simply the proportion of flawed builds flagged, and the tested models are not a complete historical leaderboard.

Sources, Method & Data

The 21 source rows are ordered by model release date as reported by Epoch, not test date. The line retains each new record among those models; it is not a complete history of all available AI. GPT-5 and Gemini 3 are absent. Published scores are plotted unchanged ×100. The endpoints are 28.3333% and 80%, a gain of 51.6667 percentage points. Source standard errors are 5.8665 and 4.9289 percentage points; all model standard errors are in the table below. They are not 95% confidence intervals.

The dataset has 42 intentionally flawed and 18 correct photos from STÄLL, TONSTAD and GULLABERG builds. Models receive instructions and image/Python tools. Correct step identification and an adequate error description matter; a lenient LLM grades descriptions. The score is not simply the fraction of flawed builds flagged. Some exported scores imply half cases when multiplied by 60; the public method does not explain that granularity, so no exact “x out of 60” count is inferred. No human-performance comparison. Accessed October 6, 2026.

Benchmark methodology · Original chart dataset

All 21 tested models; original Epoch scores
ModelRelease date (Epoch)Score (%)Standard error (pp)
Claude Opus 4.52025-11-2428.33335.8665
GPT-5.22025-12-1138.33336.3298
Claude Opus 4.62026-02-0528.33335.8665
Gemini 3.1 Pro2026-02-1926.66675.6332
GPT-5.42026-03-0537.50006.2465
Claude Opus 4.72026-04-1633.33336.1372
Kimi K2.62026-04-2021.66675.2301
GPT-5.52026-04-2344.16676.4101
Claude Opus 4.82026-05-2842.50006.3807
Claude Fable 52026-06-0935.83336.1859
GPT-5.6 Sol2026-07-0956.66676.3409
GPT-5.6 Terra2026-07-0954.16676.4321
GPT-5.6 Luna2026-07-0942.50006.3807
Kimi K32026-07-1634.16676.1170
Gemini 3.6 Flash2026-07-2123.33335.3766
Claude Opus 52026-07-2460.83336.1859
Gemini 3.7 Flash2026-08-1326.66675.6332
Claude Fable 5.12026-09-0170.00005.7244
Qwen3.8 Max2026-09-0120.00005.2076
Gemini 3.8 Flash2026-09-0231.66675.9383
GPT-6 Astra2026-09-0380.00004.9289
Explore the Chart
In Charts

Who Is Gaining Ground in Energy?

1. Energy supply, beyond electricity2025 · EXAJOULES (EJ)

Power, transport, heat and industrial uses. Fossil = oil + gas + coal. EI coverage excludes some traditional biomass. More supply is not automatically greater efficiency or independence.

Data: Energy Institute · Statistical Review 2026 · Chart: VastkindData: 2025 · Comparison: 2020–2025 · Released: 30 Jun 2026
2. What is driving power growth?CHANGE IN ANNUAL GENERATION · 2020 → 2025 · TWh

+ more / − less annual generation. China adds 1,583 TWh of wind and solar output; the other three together add 874 TWh. Other includes bioenergy, other renewables and other fossil.

Data: Ember · Global Electricity Review 2026 · CC BY 4.0 · Chart: Vastkind2025 generation estimates may be revised · Report: 21 Apr 2026 · Accessed: 6 Oct 2026

Different sources of growth

The U.S. adds solar and gas. India adds coal. In the EU, wind and solar rise as other output falls. China’s wind and solar gain is larger than the other three combined, but its coal generation grows too.

The comparison concerns changes in annual output, not the cumulative electricity generated over five years.

Energy is bigger than electricity

The first panel includes fuel use beyond power stations. The second examines electricity generation only. The two panels use different measures and cannot be added. Higher energy supply alone is not a ranking of technology, efficiency or energy independence.

2020 was a pandemic year. These are 2020–2025 comparisons, not forecasts or claims about 2026 growth.

Sources, Method & Data

Panel 1 uses EI Statistical Review 2026 total energy supply under its physical-energy-content method, in EJ. Fossil = oil + gas + coal. Non-fossil = nuclear + hydro + solar + wind + other renewables, as covered by EI; commercially recorded energy including modern renewables, not complete coverage of traditional/noncommercial biomass. TES is not useful final energy, domestic production or energy independence. Wind/solar enter as electricity output; thermal sources include conversion losses, so source shares are not equivalent to electricity shares. China means mainland China; U.S. territories are excluded. The official Total EU aggregate covers the current 27 members. 2025 annual values; 2020–2025 change is a ratio of annual totals, not EI’s leap-year-adjusted growth series. Numbers rounded for display. Direct EI downloads were blocked; unchanged EI CSV/XLSX snapshots archived by Our World in Data were used, with matching source checksums.

Panel 2 uses Ember’s separate gross-electricity-generation dataset. 2025 values are estimates based on monthly generation data and may be revised. Each cell is annual 2025 output minus annual 2020 output, in TWh, not cumulative generation over five years. Other = bioenergy + other renewables + other fossil. The net row is the total-generation difference; displayed rounded cells may not add exactly. An output increase is not the same as installed capacity: weather, plant use and outages matter. Renewables and thermal energy have different conversion losses. The four regions are selected comparisons, not a world total or a country ranking of technological quality.

China’s wind/solar change: 913.83 + 668.95 = 1,582.78 TWh. Other three: 384.55 + 303.08 + 186.63 = 874.26 TWh. EU net output grows 43.80 TWh; wind/solar grows 303.08 while fossil output falls 200.09. Sources published June 30 and April 21, 2026; accessed October 6, 2026.

EI data and methodology · Ember source and methods

2025 energy supply and change since 2020
Region2025 EJFossil share %2020 EJChange %
China162.19387.55135.448+19.75
USA93.82983.1786.185+8.87
EU2752.19373.2053.458-2.37
India39.09992.9930.055+30.09
Energy supply by source, 2025 (EJ)
RegionOilGasCoalNuclearHydroSolarWindOther renewables
China33.84515.91092.2395.2935.0124.2244.0611.611
USA36.42432.8848.7309.0120.8791.4141.6892.797
EU2721.76611.8524.5897.1031.1551.3401.7282.660
India10.8922.39623.0710.5870.6420.6070.3740.530
Electricity change: annual 2025 minus annual 2020 (TWh)
SourceChinaU.S.EU-27India
Solar+913.83+258.10+222.44+143.08
Wind+668.95+126.45+80.64+43.55
Coal+835.09-36.24-95.54+330.37
Gas+81.92+183.17-93.24-26.25
Nuclear+121.98-5.10-31.10+9.22
Hydro+77.57-38.25-20.72+14.03
Other+113.60-11.52-18.68+4.17
Explore the Chart
All Charts
Explore by topic

Pick a frontier.

Seven fields. A clear place to begin.

The Vastkind Briefing

Stay ahead
of the story.

A personal take on the week. By Dorian.

Check your inbox.

Click the confirmation link in your email to complete your subscription. If it hasn’t arrived, check your spam folder.

Find something to read

Occasional editions. Free to subscribe.
Unsubscribe anytime.

Explore Vastkind

Topics
Subscribe

Stay curious. Stay informed.

For another moment

Your reading list

Saved on this device. No account needed.

Something catch your eye? Use the bookmark beside an article to keep it here.