We invite submissions on methods, evaluations, systems, and perspectives that advance trustworthy embodied foundation models for robotics, spanning both safety by design and safety by practice.
Contributions may include empirical studies, benchmarks, negative results, formal analyses, deployment lessons, and tooling that improves robustness, evaluation, and accountability of FM-based robotic autonomy.
We encourage researchers to submit work in the following areas:
Submissions are now closed. See the Accepted Papers.
| Abstract Submission Deadline | |
| Paper Submission Deadline | |
| Author Notification | |
| Camera-ready Version Due | |
| Workshop | July 17, 2026 |
| Morning Schedule | Evening Schedule | Activity |
| 9:00 - 9:05 | Welcome by organizers | |
| 9:05 - 9:25 | Invited Talk by Andrea Bajscy (15 min + 3 min Q&A) | |
| 9:25 - 9:55 | Invited Talk by Glen Chou (15 min + 3 min Q&A) | |
| 10:00 - 10:30 | Coffee break + poster session (overlap) | |
| 10:30 - 10:50 | Invited Talk by Ian Abraham (15 min + 3 min Q&A) | |
| 10:50 - 11:10 | Invited Talk by Ralf Römer (15 min + 3 min Q&A) | |
| 11:10 - 11:35 | Interactive Breakout Session | |
| 11:35 - 12:00 | Oral presentations | |
| 12:00 - 12:55 | Panel discussion | |
| 12:55 - 13:00 | Closing remarks |
Click a talk title to expand the abstract
There is a growing claim that generalist robot policies are becoming more robust. But what does robustness actually mean in this new era? In this talk, I argue that we are overlooking an important notion of “outcome robustness”: even for the same robot action, outcomes can vary depending on the physical world (e.g., dropping an open bag on a table may or may not spill its contents). But, for generalist robots that operate on high-dimensional perceptual inputs, reasoning about this rich space of outcomes requires predicting future observations, which is extremely challenging to model and optimize. To address this, I will present StressDream, an inference-time optimization algorithm that steers diffusion-based video world models toward high-impact yet plausible outcomes by optimizing over the initial noise. In manipulation and autonomous driving with state-of-the-art world models, I will discuss how StressDream enables robust policy evaluation and improvement.
Reliable deployment of robotic foundation models (FMs) requires mechanisms for recovery when learned policies leave the regimes covered by their data. In this talk, I will argue that massively parallel model-based control can provide a common backbone for both safety by design and safety by practice: generating safety-critical data before deployment, and enabling fast intervention during deployment. I will first discuss our work on GPU-parallelized robust MPC, where physics simulation and parallel reachability analysis are used to synthesize controllers and safety certificates for high-dimensional robotic systems, from humanoids to deformable-object manipulation. These robust controllers can be used to produce safe recovery demonstrations for FM finetuning. I will then discuss our recent work showing that classical linear control can be surprisingly effective inside foundation models themselves: by modeling local activation dynamics, we can steer the internal states of world-action models at deployment time to improve robustness under out-of-distribution perturbations. Finally, I will close with perspectives and future opportunities.
Random sampling methods for robot control has demonstrated immense potential in expanding the capabilities of robotic manipulation and locomotion. In this talk, I will show case studies on how this same random sampling can be used to instill robustness in robot behavioral control for manipulation and effectively counteract unobservable environmental uncertainty. With the use of variational inference methods, I show how this robustness can be harnessed to improve the responsiveness of robots in real-time, enabling highly dynamic manipulation with limited sensor-feedback. Last, I overview recent findings on sampling-based hybrid mode composition that draws from hybrid systems theory to extend the capabilities on whole-body robotic control through fast sequencing of diverse control modes.
Generative policies, such as diffusion policies and VLAs, have demonstrated impressive capabilities in solving complex, long-horizon tasks. Yet deploying them in the real world, where robots must continually adapt to novel situations without compromising safety, remains a major challenge. In this talk, I argue for addressing safety and adaptation through a unifying lens: uncertainty awareness. I first show how safety can be enforced on generative policies at deployment time without sacrificing task performance. I then discuss estimating uncertainty in these policies and how it can be leveraged for runtime failure prediction and collecting the most informative data. Finally, I present recent work toward generalist policies that can actively and continually improve.
The panelists will be tackling three questions during the debate-style panel with pre-decided positions.
Resolution: "FM-powered general-purpose robots should ship to consumers the moment they're commercially viable. Waiting for safety guarantees is paternalistic and just hands the market to less careful players."
Position A) Ship it, iterate in the wild: real-world deployment data is the only path to real safety; lab guarantees are vanity metrics that don't survive contact with actual kitchens. Every year of delay is a year of learning lost.
Position B) Moratorium: no unsupervised FM robot in a home until it clears independent verification. "Move fast and break things" is a fine motto until the thing being broken is a toddler or a countertop. Cars and drugs aren't beta-tested on customers.
Resolution: "Runtime monitoring, shielding, and red-teaming are a fig leaf. If safety isn't baked in at training time, no amount of deployment-time patching makes an embodied FM trustworthy."
Position A) Safety by design is the only real safety: you can't bolt trust onto a finished black box; monitors and shields are theater that let people feel safe while shipping something unsafe.
Position B) Safety by design is a fantasy: you cannot verify a billion-parameter policy, full stop. All actual safety lives at runtime, the monitor is the seatbelt, and pretending the engine is so well-built you don't need one is how people get hurt.
Resolution: "'Trustworthy embodied foundation model' is an oxymoron. End-to-end learned policies will never be trustworthy enough for safety-critical physical settings — the future is classical control and verification, with FMs kept on a short advisory leash."
Position A) FMs are fundamentally untrustworthy in the physical world: hallucination and poor grounding are load-bearing features of how these models work, not bugs to be patched. Keep them as suggesters; never let them close the loop on a motor.
Position B) Classical verification is nostalgia: it never scaled to the open world and never will. Betting the future on formal methods is like insisting on hand-written assembly. Scale plus learning is the only route that reaches "fold my laundry."