BayesContact: Uncertain Pose Estimation via Visuo-Tactile Proposals and Simulation-based Inference

2026-07-17Robotics

Robotics
AI summary

The authors developed BayesContact, a method that combines vision and touch to better estimate the position of objects during peg-in-hole tasks. Instead of just using depth cameras, their approach uses simulations to predict both visual and contact information, which helps keep track of where an object is more accurately. They also use this combined information to decide the best way to explore the object for clearer insights. Their tests show that BayesContact improves accuracy and success in insertion tasks compared to using only vision.

pose estimationvisuo-tactile sensingpeg-in-hole insertionsimulation-based inferenceparticle filterdepth sensingforce/torque sensingactive perceptioninformation gainrobot manipulation
Authors
Aditya Kamireddypalli, Matias Mattamala, Joao Moura, Russell Buchanan, Sethu Vijayakumar, Subramanian Ramamoorthy
Abstract
Contact-rich manipulation requires pose estimates that are often more accurate than what depth-only sensing provides. Existing methods, relying on vision and contact, employ costly offline training procedures that need to be retrained for new environments and geometries. We propose BayesContact, a Simulation-Based Inference framework for visuo-tactile pose estimation in peg-in-hole insertion. BayesContact maintains a particle belief over object pose and fuses depth observations with force/torque-derived contact evidence. We employ simulation based forward models to approximate these observation likelihoods. For each pose hypothesis, a renderer predicts depth measurements and a physics simulator predicts contact outcomes under guarded probing actions; both are scored against real observations to update the belief. The resulting multimodal belief also enables information-gain-based probing for active disambiguation. Across simulated geometries and real-robot experiments, BayesContact improves pose observability and insertion success over vision-only inference by 30%