Quote of the day

A person who never made a mistake never tried anything new.

- Albert Einstein

Contact

Multimodal Perception and Language Grounding for Safe Human-Robot Collaboration

Open Industry & Manufacturing

Multimodal Perception and Language Grounding for Safe Human-Robot Collaboration

Multimodal Perception and Language Grounding for Safe Human-Robot Collaboration

Description

Enable natural-language instruction of a collaborative robot while ensuring that an independent and verifiable safety layer remains authoritative. This is an exploratory research topic.

Details

Context and Problem Statement

Natural-language programming could make collaborative robots easier to use in small-batch manufacturing, but ambiguous interpretation must never override safety. Language-model behavior cannot be fully verified, so safety must remain independent from language understanding.

Research Question

How can a system be architected so that language controls task intent without ever bypassing an independent safety layer, and how robust is this separation under ambiguous or contradictory instructions?

Proposed Approach

Strictly separate language-driven planning from deterministic safety logic based on distance, speed, and forbidden zones. Build a taxonomy of risky instructions and measure interception rate, clarification quality, and residual risk.

Expected Contribution

A benchmark for evaluating separation of responsibilities between language interpretation and robotic safety.

Expected Prototype

A controlled collaborative-robot cell with multimodal perception, natural-language instruction, and full logging of safety interventions.

Datasets

Public robotic perception and manipulation datasets, operator instructions, and adversarial scenarios.

Challenges

Language ambiguity, perception latency, lack of formal guarantees for LLM behavior, hardware cost, and machine-safety requirements.

Research Question

Is an independent deterministic safety layer sufficient to bound the risk introduced by erroneous language interpretation in a collaborative robot cell?

Innovation

The project reframes natural-language human-robot interaction as a safety-architecture problem and evaluates ambiguous commands adversarially.

Expected Deliverable

A controlled demonstration cell, a taxonomy of risky instructions, and an evaluation report for the safety-interception layer.

Technologies

Vision-Language-Action Models Multimodal Perception and Sensor Fusion Robot Planning and Control Deterministic Safety Architectures ROS 2

Required Skills

  • Robotics and ROS
  • Computer Vision and Multi-Sensor Fusion
  • Language Models and Semantic Grounding
  • Machine Safety and Risk Analysis Fundamentals

Datasets

  • Public robotic perception and manipulation datasets
  • Operator instruction corpus collected with consent
  • Adversarial scenarios constructed for experimentation

Morocco & Africa Relevance

Flexible automation is increasingly important for Moroccan manufacturing, while robotics-programming expertise remains scarce; lowering the programming barrier without compromising safety supports adoption.