Skip to content
Completed

Early Failure Prediction and Recovery for Mobile Robot Navigation

Research question. How early can mobile robot navigation failure be predicted from onboard signals, and can a guarded recovery act on that prediction without degrading outcomes?

Dates
August 2026–September 2026
Research areas
Embodied AI, Autonomous Systems, Failure Prediction, Uncertainty Calibration

The question

Research 1 asked whether calibrated perception uncertainty should change how a planner moves, and found that it did not improve safety. This study asks the adjacent question: rather than acting on uncertainty continuously, can the system anticipate a specific failure early enough to do something about it — and does acting on that anticipation help?

Two halves, deliberately separated:

  1. Prediction. How early can failure be predicted from onboard signals, and at what cost in false alarms?
  2. Recovery. Given a prediction, does a guarded recovery improve the outcome, or does intervening on an uncertain warning make things worse?

The second half is the one that matters, and it is the one most work in this area skips.

Design

The study is an overlay on the platform frozen at the end of Research 1, running isolated on its own ROS domain so the two cannot interfere. Reusing a pinned, already-validated platform means the failure behaviour under test is not confounded by changes to the navigation stack itself.

Fault families. Seven fault families are frozen, admitted after a 21/21 family-and-severity treatment-integrity gate.

Development campaign. Complete at 648 of 648 artefact-usable episodes across a frozen six-map design — 125 terminal events and 523 non-events. Two pre-goal infrastructure invalids are retained rather than discarded, each replaced exactly once by a same-cell, same-seed episode under a linked replacement campaign.

Validation campaign. Complete at 324 of 324 episodes, hash-addressed, and barred from selecting models, calibration or thresholds until the human admission gate passes.

Baseline. A transparent threshold baseline, so any learned predictor has to beat something interpretable rather than nothing.

  1. Onboard signals (2 Hz decisions)
  2. Causal feature extraction
  3. Predictor P1–P6 or threshold baseline
  4. Calibration and alarm policy
  5. Event-level evaluation
  6. Guarded recovery (recommendation-only)
Figure 1 — The evaluated pipeline. The recovery manager is recommendation-only in the current stage: it proposes, it does not act.

How the gates work

The part worth describing is not the model but the sequence of locks, because they are what will make any eventual result credible.

Feature extraction for model fitting stayed fail-closed until two conditions were met independently: a researcher threshold sign-off, and a review of all twenty audit-sample episodes with exact agreement on the causal fields. That gate passed under a recorded protocol amendment. No inter-rater reliability is claimed, because only one reviewer has completed the review — a second reviewer must be a different real person, and the interface records attestations append-safe rather than allowing edits.

A feature definition was corrected before any model was fitted — goal distance now uses the map-frame pose rather than odometry-frame coordinates — and the superseded derived artefacts are kept rather than deleted, under a dated directory, with the correction recorded in the research log.

Reuse of Research 1 data is staged through an outcome-blind structural catalogue. Of the structurally eligible development bags, an audit admitted 647; 646 of those end in a terminal event, because Research 1 retained failure bags, so they are excluded from the fitting pool and kept as a separately reported natural-failure and domain-shift set. Zero protected-test rows are exported.

Results

Six hypotheses, all evaluated against a threshold of 0.235 and a model checkpoint frozen before the protected split was touched. Two were confirmatory; four were prespecified supporting hypotheses.

HypothesisEstimate95% CIOutcome
H1Learned predictor beats the transparent baseline at the fixed false-alert budget0.197[−0.000, 0.447]not supported
H2Median useful lead time of detected failures ≥ 3 s3.818 s[1.745, 6.120]supported
H3Calibration reduces Brier score and ECE on validation−0.019—supported
H4Removing planner and localisation health features degrades early warning0.102—reported; no numeric threshold was prespecified
H5Leave-one-family-out recall beats baseline in ≥ 5 of 7 fault families4 of 7—not supported
H6Prediction-triggered recovery improves mission completion without more collisions——not supported

Both confirmatory hypotheses failed. H1's interval touches zero at the lower bound, which is the whole question: at a false-alert budget fixed on validation, the learned predictor could not be shown to beat a transparent threshold baseline. H6, the one that mattered operationally, also failed — acting on the prediction did not improve mission completion.

What did hold is worth stating precisely, because it is the part that survives. Failures in this system are visible in advance: the median useful lead time on detected events is roughly 3.8 seconds, with the interval comfortably above the 3-second bar set in advance. Calibration behaved as expected. The signal is real; what could not be established is that a learned model exploits it better than a simple rule, or that acting on it helps.

Generalisation is the sharper negative. Held out one fault family at a time, the predictor beat the baseline in four of seven families against a preregistered bar of five. A predictor that works on the faults it has seen and not reliably on the ones it has not is not yet an early-warning system.

Demonstrations

Two clips, and the difference between them matters more than either one.

Figure 2 — Faithful replay of a held-out episode, rendered from its recorded bag as a top-down view, with the risk trace beneath. The predictor raises a useful warning 7.9 seconds before the failure. This is an episode from the protected evaluation — the run whose numbers are in the hypothesis table above.
Figure 3 — An engineering re-run, not evidence. Validation episode val_02, route 0, planner oscillation, same seed and fault, with the GUI camera following the robot and the frozen predictor running live in recommendation mode. It shows two things honestly: an early alert at 62 s, which counts as a false alert under the preregistered 10-second window, and a useful warning at 72.9 s — 6.3 seconds before the planner failure at 79.1 s. The risk trace is that run's own, so both halves are synchronised to within a frame.

Why both. The first clip is drawn from the protected evaluation and is the one the reported numbers come from. The second is a fresh re-run made to be watchable: a live camera view rather than a top-down replay, with the predictor running in recommendation mode. It is a demonstration and nothing in the hypothesis table depends on it.

It is included precisely because it contains a false alert. A single cherry-picked run showing a clean save would misrepresent a system whose confirmatory result was that it could not be shown to beat a transparent baseline. What this run shows — one spurious early alert, then a genuine warning six seconds out — is a fairer picture of the behaviour than the successful replay alone.

Reproducibility

The release record checksums 67 manifests, 37 configurations, 2 model checkpoints, 773 reports and 11 documents. An independent rerun was executed from a clean worktree against the frozen checkpoint, with the full test suite passing. Protected-split use is recorded explicitly in both the release record and the rerun record rather than left implicit.

Three earlier rerun attempts failed and are kept alongside the successful one — the first because tables were not shared, the second and third on tests. They are retained rather than deleted, because a reproduction record that shows only the attempt that worked is not a reproduction record.

What this says next to Research 1

Two studies, two interventions that did not deliver. Research 1 found that calibrated perception uncertainty did not make navigation safer; Research 2 finds that predicting failure — which is genuinely possible several seconds ahead — did not translate into a demonstrated operational benefit either.

That is not a discouraging pattern. It is the same pattern twice: a component metric improves, and the system-level outcome does not follow. Establishing that carefully, twice, with the analysis frozen in advance both times, is a more useful contribution than a positive result obtained without those constraints.

Review status

The study was released without an independent second review. A reviewer packet was prepared and the evidence-review application made available; no second reviewer completed one, and rather than wait indefinitely or substitute something that is not an independent review, the decision was recorded and the work released as it stands.

Two things follow, and both were already in the manuscript rather than added afterwards. No inter-rater reliability is claimed or estimable — the twenty-episode label audit carries a completed review by one human reviewer, and AI comparison was retained as supporting evidence only, never counted as a human rating. And no external technical review informs the conclusions: the hypothesis outcomes stand on the frozen protocol, the completion audit and the byte-identical independent rerun, all of which are machine-checkable, rather than on anyone's judgement.

The disposition records this, and states that any review received later will be logged with its date rather than presented as having informed the release.

The manuscript remains a draft and has not been submitted anywhere.