Xiaomi launched Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. This model views embodied genera...
论文研究HuggingFace Daily Papers(社区热门论文)
Today AI Intelligence Brief
Xiaomi launched Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for
unified embodied synthesis. This model views embodied generation as an extension of basic image and video generation, jointly optimizing text-to-image, image editing, embodied scene generation, embodied migration, and embodied video generation. It is the first model to support high-quality multi-view scene generation across various robot forms, introducing structured, controllable embodiment migration. The model achieves SOTA in single-step and sequential generation tasks, surpasses GPT-Image-2.0 in human evaluation of embodied scene generation and transfer, ranks first in embodied video generation in World Arena, and increases pi_0.5's out-of-distribution success rate on real-world manipulation tasks from 36.9% to 63.2%. The code and checkpoints are open source.
Original Article Excerpt
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs :https://info.arxiv.org/labs/index.html.
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
Xiaomi launched Xiaomi-Robotics-U0, a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. This model views embodied genera...