Skip to content
Completed

Language as an Uncertain Sensor for Robot Navigation

Research question. If a navigation instruction is treated as an uncertain observation rather than a command, can a robot recover from instructions that are wrong, ambiguous, or contradicted by what it sees?

Dates
August 2026–September 2026
Research areas
Embodied AI, Multimodal Intelligence, Autonomous Systems, Uncertainty Calibration

The question

Most instruction-following treats the instruction as ground truth: the robot is told to turn left at the fire extinguisher, so it looks for a fire extinguisher and turns left. The failure mode is obvious once stated — if there is no fire extinguisher, or there are two, or the speaker misremembered, the instruction is wrong and the system has no representation for that.

This study takes the other position. An instruction is an observation, not a command. It carries information about where to go, it is generated by a fallible process, and it should therefore enter the system with a confidence attached and be weighed against what the robot actually sees.

That reframing is the thread back to the rest of this work: Research 1 asked whether calibrated perceptual uncertainty should change navigation, Research 2 asked whether failure can be anticipated, and this asks where else uncertainty enters the system. The answer it tests is: the instruction itself.

But weighing an instruction against what the robot sees only works if what the robot sees carries a trustworthy confidence. That is the part this study reached, and the part where it stopped.

Why the study narrowed

The preregistered plan called for a calibrated four-class grounding stage — chair, laboratory_entrance, office_entrance, doorway — admitted to runtime only after passing a fixed coverage gate. Neither of the two things that gate required turned out to be obtainable. Both reasons are reported below, because both are results.

1. The required low-confidence bin cannot exist

The admission criterion asked for at least 5 scoreable emissions in the [0, 0.5) confidence bin for each class. The pinned OCR component enforces text_score >= 0.5 internally and discards weaker text proposals before it emits anything. The entrance localizer copies that retained score through unchanged.

Under that pass-through, entrance scores are mathematically confined to [0.5, 1.0], so the maximum attainable number of emissions in [0, 0.5) is identically zero against a requirement of five. This was checked both by auditing the filter's source and by synthetic boundary testing against the real filter, not inferred from the observed data.

2. Clean-view grounding saturated, leaving nothing to fit

A fixed serial 400-attempt protocol ran across ten development environments, yielding 398 successful captures (2 Nav2 lifecycle failures), of which 287 were scoreable — alongside 72 detector or depth abstentions and 39 intact-frame nondetections. Every one of the 287 was reviewed by hand for category, physical instance identity, and planar reference localization to 0.35 m: 281 correct, 6 incorrect, 0 unreviewable.

ClassCorrectIncorrectViews
chair99099
laboratory_entrance84084
office_entrance78078
doorway20626

All six errors fell in a single class, and all six were reference-point errors — the category and the physical instance were right, the located point was not. The planned per-class models required at least 5 incorrect outcomes each to fit. Three classes returned zero. The models could not be fitted, for the same structural reason as above rather than a modelling failure.

A proposed explanation for the doorway errors is that depth rays pass through the open portal rather than striking the jamb; the implemented estimator uses side-jamb support. The labels establish which dimension failed. They do not establish the physical cause, and it is recorded as a proposed mechanism rather than a demonstrated one.

What was retained

A frozen hybrid scoring candidate is kept as an exploratory artifact only. Its doorway component is a five-feature regularised logistic regression fitted on the 26 valid doorway emissions, converging after 578 updates. Its chair and signage components are identity pass-throughs of the raw matching scores and are explicitly labelled uncalibrated — they are not models, and calling them calibrated would be the exact error this study exists to avoid.

Of seven leave-one-map-out folds, three failed the unchanged negative-outcome floor. They are reported as unfit rather than dropped, because silently excluding the folds that would not fit is how a weak result is made to look clean.

The earlier graph-world campaign

An authorised create-once campaign produced 240 records across six routes, three test maps, eight conditions, five systems and one seed, every record SHA-256 verified, with 48 complete five-system blocks.

Between two of the variants the completion difference was +25 percentage points (hierarchical bootstrap 95% interval +12.5 to +37.5; 5,000 resamples, analysis seed 303) and the critical-failure difference −25 points (−37.5 to −12.5). Two qualifications travel with that number and are not separable from it: a third variant also completed every case, so these results do not show the proposed system beating the strongest baseline; and the protected partition has since been accessed, so it is no longer sealed and must not be reused to support a further confirmatory claim.

Verification

The closure bundle carries a per-file SHA-256 manifest, the recorded authority for the scope change, the full human review return, and the audit that establishes the threshold limitation. The regression suite passes 1,339 tests with no failures and one skip.

The bundle deliberately references raw captures and model weights by hash rather than containing them; it is an evidence archive, not a self-contained simulator distribution.

The repository is public, so the claims above can be checked directly rather than taken on trust: the evidence report states the findings and their limits, and the closure bundle contains the manifest, the recorded authority for the scope change, the human review return and the threshold audit.

What this study does not claim

It does not claim a validated four-class runtime calibration. It does not claim the retained candidate is fit for navigation. It does not claim the observed clean-view precision is a property of the classes rather than of these scenes and this emission policy. It does not claim the doorway error mechanism has been demonstrated. And it does not claim the language layer was shown to improve navigation — the study closed before reaching that question.

What it does establish is narrower and, for a grounding stage anyone else might build on the same components, more immediately useful: a frozen upstream filter silently determines what calibration is possible downstream, and a benchmark whose clean views are too easy will hide that by handing you no errors to learn from.