IROS 2026 Workshop · Interoceptive Perception for Resilient Robotics
Panel Discussions
Two debates, one thread: what role should the body’s own signals play in robot intelligence?
11:10 AM · Before a Robot Can Model the World, Must It Model Itself?
4:30 PM · Explicit or Implicit? The Future of IMU Learning in Robot Perception
Use ← → arrow keys, space, or click to navigate · Ctrl/Cmd+P to export PDF
Morning Panel · 11:10 AM – 12:00 PM
Before a Robot Can Model the World, Must It Model Itself?
World models and vision-language-action policies condition on a body state they cannot produce themselves. Is a robot’s self-model, learned from inertial, proprioceptive, and tactile signals, a prerequisite for modeling the world — or does it emerge on its own from end-to-end training at scale?
Morning Panel · The Two Positions
Prerequisite, or emergent?
The self-model is a prerequisite
- Every observation comes from a moving body; the world only appears once the self is subtracted.
- A world model predicts the consequences of action — and the actor is the body.
- Biology agrees: the brain predicts the sensory consequences of its own motion and perceives the world as deviations.
- Safety demands it: the forces a robot exerts cannot be seen, only felt.
It emerges from scale
- The bitter lesson: end-to-end learning at scale beats hand-designed structure.
- Body state can live as an implicit latent inside the policy — no dedicated model needed.
- Stateless navigation already shows competence without explicit state estimation.
- A separate self-model is one more module to calibrate, maintain, and get wrong.
Morning Panel · Questions for Discussion
Where the debate gets decided
- What failure would only a self-model prevent — and has anyone observed it?
- If self-knowledge emerges at scale, where does the training data for the body’s signals come from?
- Can a self-model transfer across embodiments, or must every robot learn its body from scratch?
- What does each answer imply for the foundation-model roadmap: one giant policy, or a perception substrate beneath it?
From Morning to Afternoon
The morning panel asks whether a robot needs a model of itself.
The closing panel asks how to build one:
as a dedicated, interpretable module — or dissolved inside an end-to-end policy.
Closing Panel · 4:30 – 5:00 PM
Explicit or Implicit? The Future of IMU Learning in Robot Perception
Should robots model inertial sensing explicitly, through dedicated and interpretable estimation modules, or implicitly, inside end-to-end learned policies? What does each path mean for accuracy, generalization, and resilience when exteroceptive sensing degrades or fails?
Closing Panel · The Two Positions
Explicit module, or implicit latent?
Explicit estimation
- Interpretable and certifiable — you can inspect, bound, and debug the state estimate.
- Sensor physics and calibration priors are known; throwing them away is wasteful.
- Modular: one estimator serves many downstream tasks and platforms.
- Degrades diagnosably — you know when and why the estimate is failing.
Implicit, end-to-end
- Optimizes the true task objective instead of a proxy — trajectory error is not the same as staying upright.
- No interface bottleneck: the policy keeps correlations a hand-designed state would discard.
- Less per-platform engineering; the representation adapts with the data.
- Learned features can exploit regularities no filter designer anticipated.
Closing Panel · Questions for Discussion
Where the debate gets decided
- When vision dies mid-task, which fails more gracefully — a filter or a policy?
- Do explicit modules cap performance, or anchor generalization to new environments?
- Is the hybrid middle ground — differentiable filters, learned priors inside estimators — the best or the worst of both worlds?
- What benchmark result would convince you the other side is right?
One Thread Through the Day
Before a robot can model the world,
it must model itself.
Robots that feel in order to move — and stay safe when they cannot see.
superodometry.com/interoception