Multimodal Perception and Language Grounding for Safe Human-Robot Collaboration
Multimodal Perception and Language Grounding for Safe Human-Robot Collaboration
Description
Details
Context and Problem Statement
Natural-language programming could make collaborative robots easier to use in small-batch manufacturing, but ambiguous interpretation must never override safety. Language-model behavior cannot be fully verified, so safety must remain independent from language understanding.
Research Question
How can a system be architected so that language controls task intent without ever bypassing an independent safety layer, and how robust is this separation under ambiguous or contradictory instructions?
Proposed Approach
Strictly separate language-driven planning from deterministic safety logic based on distance, speed, and forbidden zones. Build a taxonomy of risky instructions and measure interception rate, clarification quality, and residual risk.
Expected Contribution
A benchmark for evaluating separation of responsibilities between language interpretation and robotic safety.
Expected Prototype
A controlled collaborative-robot cell with multimodal perception, natural-language instruction, and full logging of safety interventions.
Datasets
Public robotic perception and manipulation datasets, operator instructions, and adversarial scenarios.
Challenges
Language ambiguity, perception latency, lack of formal guarantees for LLM behavior, hardware cost, and machine-safety requirements.
Research Question
Innovation
Expected Deliverable
Technologies
Required Skills
- Robotics and ROS
- Computer Vision and Multi-Sensor Fusion
- Language Models and Semantic Grounding
- Machine Safety and Risk Analysis Fundamentals
Datasets
- Public robotic perception and manipulation datasets
- Operator instruction corpus collected with consent
- Adversarial scenarios constructed for experimentation