Vision and Language Guided Robotic Manipulation

TU Darmstadt, Germany
June 2025 - Ongoing
Collaborators: Vignesh Prasad (TU Darmstadt), Alex Chapin (ECL Lyon), Liming Chen (ECL Lyon), Georgia Chalvatzaki (TU Darmstadt), Alap Kshirsagar (IIT Delhi - Abu Dhabi), and Jan Peters (TU Darmstadt)

For robots to work autonomously in real-world environments, they need to understand both what they see and what people tell them to do. This project studies how vision and language can guide robots to plan and perform manipulation tasks in complex, open-ended environments. We explore the full process from understanding visual and language inputs, to reasoning about tasks, to planning and controlling physical actions. The goal is to develop versatile robotic systems that can follow natural instructions and perform a wide range of single-arm, bimanual, and human-robot collaborative tasks.

Publications

Li, Y., Chapin, A., Chen, L., Peters, J., & Kshirsagar, A., More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning, RSS 2026 Workshop From Perception to Action, 2026
Hahne, F., Prasad, V., Chalvatzaki, G., Peters, J., & Kshirsagar, A., Task-Aware Bimanual Affordance Prediction via VLM-Guided Semantic-Geometric Reasoning, ICRA 2026 Workshop on Semantics for Reliable Robot Autonomy, 2026