A library robot understood enough to recommend a book, but not enough to say where the book was. Its conversation stalled, it misread commands and it never completed the farewell. The researchers gave that interaction a 43.18% dialogue-success score—better than the lab test, yet a vivid example of why “worked” is too blunt a verdict for a robot in public.
A new preprint accepted by Frontiers in Robotics and AI describes seven benchmarking events from 2023 to 2025 across a library, a park and a pedestrian underpass. The three robots were not compared on one universal leaderboard. Experts instead evaluated task completion, interaction quality, safety and economic viability. The evidence is exploratory and sometimes tiny, but the framework exposes the questions that laboratory accuracy scores leave out.
The same robot can pass and fail at once
In Duisburg’s public library, visitors could approach a Pepper robot for book recommendations. The one systematically scored public interaction achieved its practical aim only partially: the visitor received information about the requested title, but not its location. Repeated misunderstandings and long processing loops made the exchange cumbersome. A single conversation cannot estimate a population-level success rate, and the authors say informal interactions informed only qualitative observations.
The cleaning robots produced cleaner numbers—and new ambiguity. In a restricted 67.6-square-metre area of Ulm’s main-station underpass, a robot collected all ten coffee spills and nine of ten paper pieces in five minutes, reported as 95% completeness and quality. Broken glass was excluded. Two people and a piece of luggage tested obstacle avoidance, but the experiment was still a short, managed run rather than an open-ended shift.
A real-world robot is a bundle of capabilities. It may clean efficiently but struggle around people; answer a question but confuse the user; avoid a collision yet cost too much to operate. A useful benchmark therefore needs several scores and a description of the setting—not one percentage marketed as “real-world performance.”
At Munich’s Glyptothek museum, the outdoor cleaner missed litter when it detoured around people and failed to lift bottle caps embedded in grass. On a gravel sub-area without active disturbances, it collected every distributed item. The task did not suddenly become easier because the robot learned; the surface and social environment changed.
Public space adds people, rules and money
The project’s panel included robotics, human–robot interaction, safety and economics expertise. That mix matters because technical performance is only one reason a public deployment survives. A library needs understandable dialogue and a graceful recovery when speech recognition fails. A park needs a machine that behaves around children, dogs and unpredictable debris. A cleaning contractor needs throughput, maintenance cost and staffing requirements that beat the alternative.
Interaction quality included comprehensibility, predictability, trustworthiness, perceived presence and acceptance. Those concepts are hard to standardize across a social robot and a floor scrubber. The researchers concluded that their broad interaction guideline was more useful as a modular toolkit than as a universal instrument.
Safety proved harder still. Germany has no single safety standard for everyday public interaction with robots, the paper notes. Existing test objects may represent an adult leg without representing a child, a walking aid or a thin grill leg that sensors struggle to detect. The panel ended safety testing after the second phase, explaining that the robots had not changed for the final public phase. Those tests therefore do not establish safety across every public operating condition.
Economic viability was treated as a process rather than a comparable outcome. Teams used adapted business-model tools, stakeholder maps and target costing. That cannot tell a city whether one robot pays back in three years. It does force the deployment team to name who buys, who operates, who benefits and which costs appear only after a pilot becomes a service.
A benchmark can be valuable by exposing its own limits
This study does not prove that any of the three robots is ready for unsupervised public operation. The tests were short. Park cleaning used two runs in each phase. The library percentages came from one systematically scored interaction per phase. Safety was not carried into the final public phase, and economic work produced planning artifacts rather than measured return on investment. The authors explicitly identify low sample rates and the absence of long-term trials as limitations.
Those weaknesses would be fatal if the paper claimed a product ranking. They are more informative in a methodology study. A narrow trial can reveal which metrics break when a robot leaves the lab: a projected hourly cleaning rate ignores interruptions; dialogue completion ignores confusion; obstacle avoidance does not establish safety for every body or object; a successful pilot does not establish an affordable service.
The next benchmark should run long enough for weather, crowd patterns, maintenance and operator workarounds to appear. It should pre-register scenario definitions, record failure recovery and report both machine time and human support. Most importantly, it should preserve the differences among settings instead of collapsing them into a single score. The library robot’s half-finished answer is not a reason to abandon public robots. It is a reason to measure the part after the demo.
Keep exploring
- This robot learned from demonstrations humans kept failing — why the assistance around a robot remains part of its capability.
- Large behavior models matter because they could change robotics’ real bottleneck — the case for broader training data and harder evaluation.
- Digit 5 makes robot safety a cooperative problem — how machines and workplaces share the burden of safe operation.
AI-assisted. Sources checked.




