Another attempt can change the problem itself.
Genentech describes a revealing episode in Anthropic’s August 27 research-preview announcement.
- Physical failure. Bubbles during mixing triggered errors.
- Unhelpful retry. Claude changed parameters in the same well, agitating the liquid and creating more bubbles.
- Human explanation. Scientists identified the cause and directed it to a clean well with fewer mixing cycles.
- Retained lesson. Claude kept that context for the run; the team later encoded it in reusable skills.
This was a BCA protein-assay proof of concept. The team also reported successful transfer optimization against expert-performed comparisons. Vastkind has not reproduced it.
The episode interests us because it makes an abstract question tangible: when an automated system encounters an error, what does it know about the thing it is trying to change?
A number needs a chain of reasons
A BCA assay estimates protein concentration through a color-producing chemical reaction. Thermo Fisher’s assay guide describes comparing absorbance with a calibration curve made from known concentrations. Preparation, mixing, incubation and measurement all matter; protein composition and incubation conditions can affect the response. This is background on the method, not identification of the exact commercial kit used by Genentech.
Think of that chain as a set of dependencies. To interpret a reading, you need to know what reached the well. To compare samples, you need to know how they were prepared. A tidy final chart can conceal a problem several steps upstream.
For automation, the interesting capability is therefore maintaining that chain of reasons as the experiment changes. A useful run record would connect the intended operation, the observed device state, the response to an error and the decision to keep or discard a measurement. That is what we would ask to inspect.
What the new connection layer contributes
The Model Hardware Standard began with Anthropic and HHMI’s Janelia Research Campus. Its official project page presents a common approach for AI agents operating physical equipment and, as of September 6, an application-only research preview. Partners are helping develop evaluations and practices ahead of a planned open-source release.
Instrument communication already has a substantial history. SiLA’s documentation dates its first standards to 2009. SiLA 2 exposes discoverable commands and properties and addresses error handling, authentication and security. It describes coordinating a plate handler and reader through defined operations and completion messages.
That context gives us a better question for MHS: how much effort does it remove when a laboratory connects equipment, changes a procedure or diagnoses an interruption? Those are comparisons someone could measure. A shared interface is useful infrastructure; the quality of the resulting experiment still needs its own evidence.
This is also why a robot’s response to a missed pick belongs in the same conversation. Both stories invite us to look at what happens after an intended action fails, and how the next action is chosen.
The next experiment should travel
The demonstration we would most like to see next is a procedure transferred to another laboratory. Give a second team the instructions, the software and a clear measurement target. Record what must be adapted and how much expert attention that requires. Then compare the results using criteria fixed in advance.
Three records would make that comparison especially revealing:
- The complete attempt log. Include setup, rejected measurements and unsuccessful runs, so a successful ending has a visible history.
- The intervention log. Record what a person changed and why, allowing readers to see where specialist judgment entered.
- The transfer log. Show which instructions worked elsewhere, which assumptions broke and what the second team had to rebuild.
These are proposed evaluation criteria, not results established by the preview. They would help answer the practical question for a laboratory: how much dependable experimental work can this system deliver with the people and equipment actually available?
For a software version of that question, our Astra delegation brief starts with the task, its boundaries and the evidence required at handoff. Across both settings, those details determine what “finished” can mean.
Reporting note: Source-based analysis, rechecked September 6, 2026. The experiment is the participating team’s account, published by Anthropic. Vastkind did not observe or reproduce it. The sequence is an explanatory reconstruction; the proposed transfer test has not been conducted.
Produced with AI-assisted research, drafting and editorial checks; publication and this update authorized by Vastkind’s publisher. No separate human fact-check was performed.



