[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"public-people-navigation":3,"public-project-v1-ufo":13,"project-content-0-zh-ufo":33},{"data":4,"meta":8},{"visible":5,"photos":6,"people":7},false,[],[],{"source":9,"releaseId":10,"releaseVersion":10,"contentRevision":11,"checksum":12},"static",null,0,"838a5b62d9a4bad8485720b81d5524140cb1e1b230aa9b32177c8299aef29090",{"data":14,"meta":31},{"project":15,"markdown":28},{"slug":16,"titleZh":17,"titleEn":18,"summaryZh":19,"summaryEn":20,"status":21,"coverImage":10,"homepageDesktopDemo":22,"homepageMobileDemo":23,"projectListDemo":23,"articleHeroDemo":24,"demoUrl":10,"githubUrl":25,"paperUrl":26,"homepageUrl":10,"sortOrder":27},"ufo","UFO无监督强化学习框架","UFO","Beyond FB，从训练到真机的开源无监督控制框架。","A fully open-source development framework for humanoid unsupervised RL control — covering training infrastructure, data pipelines, algorithm research, and inference deployment.","published","\u002Fvideos\u002Fprojects\u002Fufo\u002Fufo-home-desktop-preview.mp4","\u002Fvideos\u002Fprojects\u002Fufo\u002Fufo-high-dynamic-card-preview.mp4","\u002Fvideos\u002Fprojects\u002Fufo\u002F20260705-231849.mp4","https:\u002F\u002Fgithub.com\u002FRoboparty\u002FUFO","\u002Fdocuments\u002Fufo\u002Fufo-technical-report.pdf",20,{"zh":29,"en":30},"一套面向研发的、完全开源的人形机器人无监督强化学习控制（Unsupervised RL Control）开发框架，覆盖训练基础设施、数据管线、算法研究和推理部署全流程。\n\n框架致力于降低无监督强化学习控制的研发门槛，使研究者能够快速复现 SOTA 方法、探索新的行为表示（representation）、适配不同机器人平台，并实现从训练到真实机器人遥操作部署的一体化开发。\n\n## 快速训练 Infra\n\nUFO 利用更轻量级的 MJLab 作为 backend，兼容单卡与多卡并行训练。在 8 张 RTX 4090 GPU 上，不到 12 小时即可完成 BFM-Zero 算法训练，摆脱对单张大显存 GPU 的依赖；在 8 张 H200 GPU 上，约 6–8 小时即可完成训练，性能也持续优于 BFM-Zero。\n\n## 通用可扩展框架\n\n统一的 codebase 不再受限于特定机器人，可无缝适配不同机器人形态，大幅降低新平台的迁移成本。框架支持来自不同来源的数据混合训练（data mixture），并提供灵活的数据调度与配比机制。通过合理的数据分布设计，无监督强化学习不仅能够学习稳定的通用运动，还已展现出侧手翻等高动态动作的学习能力。\n\n\u003Cvideo src=\"\u002Fvideos\u002Fprojects\u002Fufo\u002Fhigh-dynamic-cartwheel.mp4\" controls muted playsinline preload=\"none\">\u003C\u002Fvideo>\n\n## 新表征集成\n\n除集成经典 BFM-Zero（FB Representation）外，框架支持多种行为表示（representation）的无监督学习研究。我们已探索 TeCH（Temporal Distance Modeling via Contrastive Representation Learning for Humanoid Whole-Body Control）等新型表示，并取得良好的控制效果，为更通用的无监督控制算法提供统一实验平台。\n\n::article-video-grid\n---\nvideos:\n  - src: \"\u002Fvideos\u002Foutputs\u002Ftech\u002Ftech-under-disturbance.mp4\"\n    label: \"TeCH · 抗扰动控制\"\n  - src: \"\u002Fvideos\u002Fprojects\u002Fufo\u002Ftech-motion-tracking-cut.mp4\"\n    label: \"TeCH · 全身动作跟踪\"\n---\n::\n\n## 真实世界遥操作\n\n首次开源无监督强化学习控制的遥操作（Teleoperation）代码及完整验证方案，支持真实机器人部署。机器人能够自然完成深蹲、半蹲、跪地、打滚、跌倒恢复以及抗外力扰动等复杂全身动作，为无监督强化学习在真实场景中的应用提供完整参考实现。\n\n\u003Cvideo src=\"\u002Fvideos\u002Fprojects\u002Fufo\u002Fteleoperation-real-world.mp4\" controls muted playsinline preload=\"none\">\u003C\u002Fvideo>\n\n\u003Cvideo src=\"\u002Fvideos\u002Fprojects\u002Fufo\u002F20260705-231849.mp4\" controls muted playsinline preload=\"none\">\u003C\u002Fvideo>\n\n\u003Cvideo src=\"\u002Fvideos\u002Fprojects\u002Fufo\u002F20260705-231945.mp4\" controls muted playsinline preload=\"none\">\u003C\u002Fvideo>","A fully open-source development framework for humanoid unsupervised reinforcement-learning control (Unsupervised RL Control), covering training infrastructure, data pipelines, algorithm research, and inference deployment.\n\nThe framework is dedicated to lowering the barrier to unsupervised RL control research — enabling researchers to quickly reproduce SOTA methods, explore novel behavior representations, adapt across robot platforms, and achieve integrated development from training to real-robot teleoperation deployment.\n\n## Fast Training Infrastructure\n\nUFO uses the lightweight MJLab backend and supports both single-GPU and multi-GPU parallel training. It completes BFM-Zero training in under 12 hours on eight RTX 4090 GPUs without relying on a single GPU with very large memory; on eight H200 GPUs, training takes approximately 6–8 hours while consistently outperforming BFM-Zero.\n\n## General and Extensible Framework\n\nThe unified codebase is not tied to a specific robot and adapts seamlessly across morphologies, dramatically reducing migration cost to new platforms. It supports mixed training with data from different sources and provides flexible scheduling and mixture controls. With a well-designed data distribution, unsupervised reinforcement learning can acquire stable general-purpose motion while also learning highly dynamic behaviors such as cartwheels.\n\n\u003Cvideo src=\"\u002Fvideos\u002Fprojects\u002Fufo\u002Fhigh-dynamic-cartwheel.mp4\" controls muted playsinline preload=\"none\">\u003C\u002Fvideo>\n\n## New Representation Integration\n\nBeyond the classic BFM-Zero (FB Representation), the framework supports research into multiple unsupervised behavior representations. We have explored TeCH (Temporal Distance Modeling via Contrastive Representation Learning for Humanoid Whole-Body Control) and other new representations, achieving strong control results and providing a unified experimental platform for more general unsupervised control algorithms.\n\n::article-video-grid\n---\nvideos:\n  - src: \"\u002Fvideos\u002Foutputs\u002Ftech\u002Ftech-under-disturbance.mp4\"\n    label: \"TeCH · Disturbance Robustness\"\n  - src: \"\u002Fvideos\u002Fprojects\u002Fufo\u002Ftech-motion-tracking-cut.mp4\"\n    label: \"TeCH · Whole-Body Motion Tracking\"\n---\n::\n\n## Teleoperation in the Real World\n\nFor the first time, the teleoperation code and complete validation scheme for unsupervised RL control are open-sourced, supporting real-robot deployment. The robot naturally performs squats, half-squats, kneeling, rolling, fall recovery, and disturbance resistance — providing a complete reference implementation for unsupervised RL in real-world scenarios.\n\n\u003Cvideo src=\"\u002Fvideos\u002Fprojects\u002Fufo\u002Fteleoperation-real-world.mp4\" controls muted playsinline preload=\"none\">\u003C\u002Fvideo>\n\n\u003Cvideo src=\"\u002Fvideos\u002Fprojects\u002Fufo\u002F20260705-231849.mp4\" controls muted playsinline preload=\"none\">\u003C\u002Fvideo>\n\n\u003Cvideo src=\"\u002Fvideos\u002Fprojects\u002Fufo\u002F20260705-231945.mp4\" controls muted playsinline preload=\"none\">\u003C\u002Fvideo>",{"source":9,"releaseId":10,"releaseVersion":10,"contentRevision":11,"checksum":32},"6a4cca82ad40bc8b7277e21d35892510bfb8dca9e8891c72caae92c9574c09ff",{"title":34,"description":35,"body":36},"","一套面向研发的、完全开源的人形机器人无监督强化学习控制（Unsupervised RL Control）开发框架，覆盖训练基础设施、数据管线、算法研究和推理部署全流程。",{"type":37,"children":38,"toc":128},"root",[39,46,51,58,63,68,73,83,88,93,98,103,108,115,121],{"type":40,"tag":41,"props":42,"children":43},"element","p",{},[44],{"type":45,"value":35},"text",{"type":40,"tag":41,"props":47,"children":48},{},[49],{"type":45,"value":50},"框架致力于降低无监督强化学习控制的研发门槛，使研究者能够快速复现 SOTA 方法、探索新的行为表示（representation）、适配不同机器人平台，并实现从训练到真实机器人遥操作部署的一体化开发。",{"type":40,"tag":52,"props":53,"children":55},"h2",{"id":54},"快速训练-infra",[56],{"type":45,"value":57},"快速训练 Infra",{"type":40,"tag":41,"props":59,"children":60},{},[61],{"type":45,"value":62},"UFO 利用更轻量级的 MJLab 作为 backend，兼容单卡与多卡并行训练。在 8 张 RTX 4090 GPU 上，不到 12 小时即可完成 BFM-Zero 算法训练，摆脱对单张大显存 GPU 的依赖；在 8 张 H200 GPU 上，约 6–8 小时即可完成训练，性能也持续优于 BFM-Zero。",{"type":40,"tag":52,"props":64,"children":66},{"id":65},"通用可扩展框架",[67],{"type":45,"value":65},{"type":40,"tag":41,"props":69,"children":70},{},[71],{"type":45,"value":72},"统一的 codebase 不再受限于特定机器人，可无缝适配不同机器人形态，大幅降低新平台的迁移成本。框架支持来自不同来源的数据混合训练（data mixture），并提供灵活的数据调度与配比机制。通过合理的数据分布设计，无监督强化学习不仅能够学习稳定的通用运动，还已展现出侧手翻等高动态动作的学习能力。",{"type":40,"tag":41,"props":74,"children":75},{},[76],{"type":40,"tag":77,"props":78,"children":82},"video",{"src":79,"controls":80,"muted":80,"playsInline":80,"preload":81},"\u002Fvideos\u002Fprojects\u002Fufo\u002Fhigh-dynamic-cartwheel.mp4",true,"none",[],{"type":40,"tag":52,"props":84,"children":86},{"id":85},"新表征集成",[87],{"type":45,"value":85},{"type":40,"tag":41,"props":89,"children":90},{},[91],{"type":45,"value":92},"除集成经典 BFM-Zero（FB Representation）外，框架支持多种行为表示（representation）的无监督学习研究。我们已探索 TeCH（Temporal Distance Modeling via Contrastive Representation Learning for Humanoid Whole-Body Control）等新型表示，并取得良好的控制效果，为更通用的无监督控制算法提供统一实验平台。",{"type":40,"tag":94,"props":95,"children":97},"article-video-grid",{":videos":96},"[{\"src\":\"\u002Fvideos\u002Foutputs\u002Ftech\u002Ftech-under-disturbance.mp4\",\"label\":\"TeCH · 抗扰动控制\"},{\"src\":\"\u002Fvideos\u002Fprojects\u002Fufo\u002Ftech-motion-tracking-cut.mp4\",\"label\":\"TeCH · 全身动作跟踪\"}]",[],{"type":40,"tag":52,"props":99,"children":101},{"id":100},"真实世界遥操作",[102],{"type":45,"value":100},{"type":40,"tag":41,"props":104,"children":105},{},[106],{"type":45,"value":107},"首次开源无监督强化学习控制的遥操作（Teleoperation）代码及完整验证方案，支持真实机器人部署。机器人能够自然完成深蹲、半蹲、跪地、打滚、跌倒恢复以及抗外力扰动等复杂全身动作，为无监督强化学习在真实场景中的应用提供完整参考实现。",{"type":40,"tag":41,"props":109,"children":110},{},[111],{"type":40,"tag":77,"props":112,"children":114},{"src":113,"controls":80,"muted":80,"playsInline":80,"preload":81},"\u002Fvideos\u002Fprojects\u002Fufo\u002Fteleoperation-real-world.mp4",[],{"type":40,"tag":41,"props":116,"children":117},{},[118],{"type":40,"tag":77,"props":119,"children":120},{"src":24,"controls":80,"muted":80,"playsInline":80,"preload":81},[],{"type":40,"tag":41,"props":122,"children":123},{},[124],{"type":40,"tag":77,"props":125,"children":127},{"src":126,"controls":80,"muted":80,"playsInline":80,"preload":81},"\u002Fvideos\u002Fprojects\u002Fufo\u002F20260705-231945.mp4",[],{"title":34,"searchDepth":129,"depth":129,"links":130},2,[131,132,133,134],{"id":54,"depth":129,"text":57},{"id":65,"depth":129,"text":65},{"id":85,"depth":129,"text":85},{"id":100,"depth":129,"text":100}]