Early Failure Prediction and Recovery for Mobile Robot Navigation
Research question. How early can mobile robot navigation failure be predicted from onboard signals, and can a guarded recovery act on that prediction without degrading outcomes?
- Dates
- August 2026–September 2026
- Research areas
- Embodied AI, Autonomous Systems, Failure Prediction, Uncertainty Calibration
The question
Research 1 asked whether calibrated perception uncertainty should change how a planner moves, and found that it did not improve safety. This study asks the adjacent question: rather than acting on uncertainty continuously, can the system anticipate a specific failure early enough to do something about it — and does acting on that anticipation help?
Two halves, deliberately separated:
- Prediction. How early can failure be predicted from onboard signals, and at what cost in false alarms?
- Recovery. Given a prediction, does a guarded recovery improve the outcome, or does intervening on an uncertain warning make things worse?
The second half is the one that matters, and it is the one most work in this area skips.
Design
The study is an overlay on the platform frozen at the end of Research 1, running isolated on its own ROS domain so the two cannot interfere. Reusing a pinned, already-validated platform means the failure behaviour under test is not confounded by changes to the navigation stack itself.
Fault families. Seven fault families are frozen, admitted after a 21/21 family-and-severity treatment-integrity gate.
Development campaign. Complete at 648 of 648 artefact-usable episodes across a frozen six-map design — 125 terminal events and 523 non-events. Two pre-goal infrastructure invalids are retained rather than discarded, each replaced exactly once by a same-cell, same-seed episode under a linked replacement campaign.
Validation campaign. Complete at 324 of 324 episodes, hash-addressed, and barred from selecting models, calibration or thresholds until the human admission gate passes.
Baseline. A transparent threshold baseline, so any learned predictor has to beat something interpretable rather than nothing.
- Onboard signals (2 Hz decisions)
- Causal feature extraction
- Predictor P1–P6 or threshold baseline
- Calibration and alarm policy
- Event-level evaluation
- Guarded recovery (recommendation-only)
How the gates work
The part worth describing is not the model but the sequence of locks, because they are what will make any eventual result credible.
Feature extraction for model fitting stayed fail-closed until two conditions were met independently: a researcher threshold sign-off, and a review of all twenty audit-sample episodes with exact agreement on the causal fields. That gate passed under a recorded protocol amendment. No inter-rater reliability is claimed, because only one reviewer has completed the review — a second reviewer must be a different real person, and the interface records attestations append-safe rather than allowing edits.
A feature definition was corrected before any model was fitted — goal distance now uses the map-frame pose rather than odometry-frame coordinates — and the superseded derived artefacts are kept rather than deleted, under a dated directory, with the correction recorded in the research log.
Reuse of Research 1 data is staged through an outcome-blind structural catalogue. Of the structurally eligible development bags, an audit admitted 647; 646 of those end in a terminal event, because Research 1 retained failure bags, so they are excluded from the fitting pool and kept as a separately reported natural-failure and domain-shift set. Zero protected-test rows are exported.
Results
Six hypotheses, all evaluated against a threshold of 0.235 and a model checkpoint frozen before the protected split was touched. Two were confirmatory; four were prespecified supporting hypotheses.
| Hypothesis | Estimate | 95% CI | Outcome | |
|---|---|---|---|---|
| H1 | Learned predictor beats the transparent baseline at the fixed false-alert budget | 0.197 | [−0.000, 0.447] | not supported |
| H2 | Median useful lead time of detected failures ≥ 3 s | 3.818 s | [1.745, 6.120] | supported |
| H3 | Calibration reduces Brier score and ECE on validation | −0.019 | — | supported |
| H4 | Removing planner and localisation health features degrades early warning | 0.102 | — | reported; no numeric threshold was prespecified |
| H5 | Leave-one-family-out recall beats baseline in ≥ 5 of 7 fault families | 4 of 7 | — | not supported |
| H6 | Prediction-triggered recovery improves mission completion without more collisions | — | — | not supported |
Both confirmatory hypotheses failed. H1's interval touches zero at the lower bound, which is the whole question: at a false-alert budget fixed on validation, the learned predictor could not be shown to beat a transparent threshold baseline. H6, the one that mattered operationally, also failed — acting on the prediction did not improve mission completion.
What did hold is worth stating precisely, because it is the part that survives. Failures in this system are visible in advance: the median useful lead time on detected events is roughly 3.8 seconds, with the interval comfortably above the 3-second bar set in advance. Calibration behaved as expected. The signal is real; what could not be established is that a learned model exploits it better than a simple rule, or that acting on it helps.
Generalisation is the sharper negative. Held out one fault family at a time, the predictor beat the baseline in four of seven families against a preregistered bar of five. A predictor that works on the faults it has seen and not reliably on the ones it has not is not yet an early-warning system.
Demonstrations
Two clips, and the difference between them matters more than either one.
Why both. The first clip is drawn from the protected evaluation and is the one the reported numbers come from. The second is a fresh re-run made to be watchable: a live camera view rather than a top-down replay, with the predictor running in recommendation mode. It is a demonstration and nothing in the hypothesis table depends on it.
It is included precisely because it contains a false alert. A single cherry-picked run showing a clean save would misrepresent a system whose confirmatory result was that it could not be shown to beat a transparent baseline. What this run shows — one spurious early alert, then a genuine warning six seconds out — is a fairer picture of the behaviour than the successful replay alone.
Reproducibility
The release record checksums 67 manifests, 37 configurations, 2 model checkpoints, 773 reports and 11 documents. An independent rerun was executed from a clean worktree against the frozen checkpoint, with the full test suite passing. Protected-split use is recorded explicitly in both the release record and the rerun record rather than left implicit.
Three earlier rerun attempts failed and are kept alongside the successful one — the first because tables were not shared, the second and third on tests. They are retained rather than deleted, because a reproduction record that shows only the attempt that worked is not a reproduction record.
What this says next to Research 1
Two studies, two interventions that did not deliver. Research 1 found that calibrated perception uncertainty did not make navigation safer; Research 2 finds that predicting failure — which is genuinely possible several seconds ahead — did not translate into a demonstrated operational benefit either.
That is not a discouraging pattern. It is the same pattern twice: a component metric improves, and the system-level outcome does not follow. Establishing that carefully, twice, with the analysis frozen in advance both times, is a more useful contribution than a positive result obtained without those constraints.
Code, data and links
- Code — github.com/alabiemmanuel177/failure-prediction. MIT for code; CC BY 4.0 for data, results, figures, checkpoints and documentation.
- Core evidence dataset — 10.5281/zenodo.22723258. A 39.6 MB archive of 2,225 files: every report and manifest, the frozen configurations, and all 34 trained checkpoints with their normalisation and training records.
- Release — v1.0.0
- Reviewer packet —
docs/review/reviewer-packet.md - Hypothesis table —
reports/confirmatory/final/hypotheses.md, with the source file named for every row
Review status
The study was released without an independent second review. A reviewer packet was prepared and the evidence-review application made available; no second reviewer completed one, and rather than wait indefinitely or substitute something that is not an independent review, the decision was recorded and the work released as it stands.
Two things follow, and both were already in the manuscript rather than added afterwards. No inter-rater reliability is claimed or estimable — the twenty-episode label audit carries a completed review by one human reviewer, and AI comparison was retained as supporting evidence only, never counted as a human rating. And no external technical review informs the conclusions: the hypothesis outcomes stand on the frozen protocol, the completion audit and the byte-identical independent rerun, all of which are machine-checkable, rather than on anyone's judgement.
The disposition records this, and states that any review received later will be logged with its date rather than presented as having informed the release.
The manuscript remains a draft and has not been submitted anywhere.