日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
手話生成arXiv:2609.32250

RoboSTAR: ヒューマノイドロボットのための次スケール自己回帰手話翻訳

RoboSTAR: Next-Scale Autoregressive Sign Language Translation for Humanoid Robots

シェア:XThreadsFacebookLINEはてブBluesky

手話モーションを粗い時間解像度から細かく生成し、ヒューマノイドロボットで実行できるよう再ターゲティングするテキスト条件付き手話生成フレームワークを提案。

著者: Yujia Zeng, Chensheng Peng, Yuxin Chen, Alex Shao, Nathan Jew, Masayoshi Tomizuka

分類: cs.CV

原文アブストラクト

Sign-language interpretation in public communication relies on qualified professional interpreters and can be difficult to scale, motivating robotic signing as a complementary accessibility interface. We present RoBoSTAR, a text-conditioned sign language production (SLP) framework for generating human-centric sign motion that can be retargeted for robotic execution, with speech supported optionally through an external ASR front end. Conventional autoregressive approaches flatten motion into a single full-resolution token sequence, forcing long-range and local dependencies to be modeled at a uniform temporal granularity. RoBoSTAR instead combines part-wise Finite Scalar Quantization with next-scale autoregression, generating motion over progressively finer temporal resolutions while predicting synchronized body and hand tokens in parallel within each step. This coarse-to-fine formulation provides compact long-range context before progressively refining motion details, while self-conditioning and context corruption improve robustness to cross-scale prediction errors. The generated motion is subsequently retargeted for physical humanoid execution. Extensive qualitative and quantitative evaluations are conducted to demonstrate the effectiveness of RoBoSTAR.

PR本紙発行元 EmplifAI