arXiv ScienceSearch

arXiv · 2505.01862

ReLI: Cross-Lingual Language-to-Action Grounding for Human-Robot Interaction

Abstract

Adapting autonomous agents for real-world industrial, domestic, and other daily tasks is currently gaining momentum. However, in global or cross-lingual application contexts, the ability to instruct these agents in one's native language remains until today a formidable challenge. Existing language-conditioned human-robot interaction frameworks typically support only a handful of high-resource languages, e.g., English and Chinese, limiting accessibility for billions of potential end users. To address this gap, we propose ReLI, a cross-lingual framework that enables autonomous agents to converse naturally, reason semantically about their environment, and execute downstream tasks, regardless of the tasks' instruction linguistic origin or input modalities. We ground large-scale pre-trained foundation models and transform them into language-to-action models that can directly provide common-sense reasoning and high-level robot control through free-form conversational interactions. We then perform an implicit language-conditioned cross-lingual adaptation of the models to ensure that ReLI generalises effectively across diverse global languages. We conducted extensive empirical evaluation on a diverse set of short- and long-horizon tasks, including zero-shot and few-shot spatial navigation, scene information retrieval, and query-oriented tasks, and then benchmarked the performance across more than $70K+$ multi-turn conversations in over $140$ languages spanning high-resource, low-resource, and vulnerable/creole tiers. Across the benchmarked languages, ReLI achieved consistently high instruction-parsing accuracy, task success rate, and rapid response time. Further, we c..

Explore related subjects

Keep this discovery

BibTeXRIS

Linus Nwankwo, Bjoern Ellensohn, Ozan Özdenizci, Elmar Rueckert. 2026-09-05. ReLI: Cross-Lingual Language-to-Action Grounding for Human-Robot Interaction. https://arxiv.org/abs/2505.01862

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

SMILE: Smooth Motion for Improved Long-Horizon VLA Execution

Vision-Language-Action (VLA) models reduce inference cost by executing multiple actions per call, but longer horizons often degrade accuracy because raw chunks contain jitter and outliers. We introduce SMILE, an architecture-preserving interface that predicts B-spline coefficients and decodes them into smooth action sequences. SMILE changes only the action representation, enabling longer fixed horizons while retaining each baseline's backbone and model scale. We apply SMILE to SmolVLA, Evo1, VPP, and DAWN, improving accuracy and amortized inference efficiency across LIBERO, CALVIN, and real-world experiments. SMILE-Evo1 reaches 98.0% with a 1.1x speedup on LIBERO, while SMILE-VPP reaches an average length of 4.42 with a 1.5x speedup on CALVIN. At a matched execution horizon of 10, SMILE-SmolVLA reduces non-boundary acceleration by 78.6% and velocity sign-change rate by 42.3%. Real-world xArm tests show higher success, fewer drops, and fewer contacts. These results establish smooth coefficient-space generation as a route to accurate, efficient long-horizon VLA execution. Project page: jongwoopark7978.github.io/smilevla

cs.RO

Establishing a Dynamic Multimodal HRI Dataset for Engagement Analysis with a Humanoid Robot

This paper presents an experimental design for constructing a multimodal dataset to analyze user engagement in human-robot interaction (HRI). Prior studies have mainly relied on observable behavioral cues, with limited frameworks integrating physiological signals. We therefore propose a structured data-collection protocol to build a multimodal dataset that includes wearable physiological signals, behavioral data, and self-report measures under different levels of task complexity defined in this experiment.

cs.RO

Connectivity-Aware Graph Extension for Decentralized Multi-Robot Exploration

Exploring unknown environments with multiple UAVs requires coordination under intermittent communication, making decentralized operation a baseline assumption. We propose, within a decentralized framework, a novel exploration graph extension strategy based on frontier connectivity to extend exploration plans and maintain area partitioning among agents stable and robust to disconnections and changes in spatial layout. The proposed extension method is applied to two state-of-the-art area partitioning methods and evaluated in simulation. Experiments show improved performance over existing graph extension approaches with higher exploration efficiency under low communication rate.

cs.RO