The team led by Fei-Fei Li at Stanford University proposed Masked Visual Actions, representing robot actions as pixel-space masks of object trajectories in...
论文研究HuggingFace Daily Papers(社区热门论文)
Today AI Intelligence Brief
The team led by Fei-Fei Li at Stanford University proposed Masked Visual Actions, representing robot
actions as pixel-space masks of object trajectories in videos. With only 15 hours of fine-tuning on real and simulated data, a single video model can serve both as a forward dynamics model and an inverse model, used for policy evaluation and model-based planning.
Original Article Excerpt
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs :https://info.arxiv.org/labs/index.html.
The readable text on this page was extracted from the public source and organized with attribution, publication time and the original link. Copyright remains with the original author and publisher.
The team led by Fei-Fei Li at Stanford University proposed Masked Visual Actions, representing robot actions as pixel-space masks of object trajectories in...