arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

机器人 / 具身智能

机器人、具身智能、机器人学习、操作、导航和具身世界模型。

共收录 3091 信号源:cs.RO, cs.AI, cs.CV, cs.LG

1. 具身推理 3091 篇

2602.03242 2026-02-04 cs.CV 79%

InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation

InstaDrive: 为真实和一致的视频生成实例感知的驾驶世界模型

Zhuoran Yang, Xi Guo, Chenjing Ding, Chiyu Wang, Wei Wu, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学) SenseAuto(感etime) Tsinghua University(清华大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 InstaDrive通过实例感知机制提升驾驶视频生成质量,增强自动驾驶任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02002 2026-02-03 cs.CV 79%

UniDriveDreamer: A Single-Stage Multimodal World Model for Autonomous Driving

UniDriveDreamer: 一种用于自动驾驶的单阶段多模态世界模型

Guosheng Zhao, Yaozeng Wang, Xiaofeng Wang, Zheng Zhu, Tingdong Yu, Guan Huang, Yongchen Zai, Ji Jiao, Changliang Xue, Xiaole Wang, Zhen Yang, Futang Zhu, Xingang Wang

机构 * GigaAI CASIA BYD

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 UniDriveDreamer是一种用于自动驾驶的单阶段多模态世界模型,通过统一的多模态生成方法提升视频和激光雷达数据合成的性能。

Comments 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01630 2026-02-03 cs.CV 79%

Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks

世界模型的研究并非仅仅是将世界知识注入特定任务

Bohan Zeng, Kaixin Zhu, Daili Hua, Bozhou Li, Chengzhuo Tong, Yuran Wang, Xinyi Huang, Yifan Dai, Zixiang Zhang, Yifan Yang, Zhou Liu, Hao Liang, Xiaochen Ma, Ruichuan An, Tianyi Bai, Hongcheng Gao, Junbo Niu, Yang Shi, Xinlong Chen, Yue Ding, Minglei Shi, Kai Zeng, Yiwen Tang, Yuanxing Zhang, Pengfei Wan, Xintao Wang, Wentao Zhang

机构 * Peking University(北京大学) Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 本文提出统一的世界模型设计规范,强调其应整合交互、感知、符号推理和空间表示,以实现更通用和稳健的世界模型。

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01270 2026-02-03 cs.LG 79%

Mixture-of-World Models: Scaling Multi-Task Reinforcement Learning with Modular Latent Dynamics

混合世界模型:通过模块化潜在动态扩展多任务强化学习

Boxuan Zhang, Weipu Zhang, Zhaohan Feng, Wei Xiao, Jian Sun, Jie Chen, Gang Wang

机构 * School of Automation, Beijing Institute of Technology(自动化学院,北京理工大学) Jiangxing Intelligence Inc.(江行智能有限公司)

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 混合世界模型通过模块化架构和任务聚类策略,在多任务强化学习中实现了高效的参数利用和样本效率,展示了在Atari和Meta-World基准上的卓越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10498 2026-02-03 cs.CV 79%

The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey

世界模型在塑造自动驾驶中的作用:全面综述

Sifan Tu, Xin Zhou, Dingkang Liang, Xingyu Jiang, Yumeng Zhang, Xiaofan Li, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Baidu Inc(百度公司)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 本文综述了驾驶世界模型在自动驾驶中的作用,分析了其生态系统、分类方法及性能表现,探讨了当前研究的局限性与未来发展方向。

Comments For continuous updates, please follow the repository: https://github.com/LMD0311/Awesome-World-Model

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22128 2026-01-30 cs.AI cs.CE q-bio.QM 79%

The Patient is not a Moving Document: A World Model Training Paradigm for Longitudinal EHR

患者并非静态文档:一种用于纵向电子健康记录的world model训练范式

Irsyad Adam, Zekai Chen, David Laprade, Shaun Porwal, David Laub, Erik Reinertsen, Arda Pekis, Kevin Brown

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 本文提出SMB-Structure模型,通过结合联合嵌入预测架构和next-token预测,学习捕捉疾病动态的嵌入,实现对高异质性复杂任务的竞争力表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22086 2026-01-30 physics.flu-dyn cs.CV 79%

Learning Transient Convective Heat Transfer with Geometry Aware World Models

利用几何感知的世界模型学习瞬态对流热传递

Onur T. Doganay, Alexander Klawonn, Martin Eigel, Hanno Gottschalk

机构 * Institute of Mathematics, TU Berlin(柏林技术大学数学研究所) Siemens Energy AG(西门子能源有限公司) Weierstrass Institute for Applied Analysis and Stochastics(魏尔斯特拉斯应用分析与概率研究所)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 本文提出一种几何感知的世界模型架构,用于学习瞬态对流热传递,通过双重条件机制和架构适应性提升物理模拟的可控性和精度。

Comments 36 pages, 18 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19834 2026-01-28 cs.AI 79%

Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models

视觉生成通过多模态世界模型解锁类人推理

Jialong Wu, Xiaoying Zhang, Hongyi Yuan, Xiangcheng Zhang, Tianhao Huang, Changjing He, Chaoyi Deng, Renrui Zhang, Youbin Wu, Mingsheng Long

机构 * Tsinghua University(清华大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 本文提出视觉生成在特定任务中优于纯语言推理,通过构建VisWorld-Eval评估套件验证了多模态世界模型提升类人推理的能力。

Comments Project page: https://thuml.github.io/Reasoning-Visual-World

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18620 2026-01-27 cs.LG 79%

CASSANDRA: Programmatic and Probabilistic Learning and Inference for Stochastic World Modeling

CASSANDRA:面向随机世界建模的程序化与概率学习与推理

Panagiotis Lymperopoulos, Abhiramon Rajasekharan, Ian Berlot-Attwell, Stéphane Aroca-Ouellette, Kaheer Suleman

机构 * Skyfall AI Tufts University(塔夫茨大学) University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Toronto(多伦多大学) Vector Instiute(Vector研究所)

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 CASSANDRA通过结合LLM的知识先验和概率图模型结构学习,提升随机世界建模中的转移预测与规划能力。

Comments 28 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15533 2026-01-23 cs.AI 79%

From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models

从生成引擎到可操作模拟器:世界模型中物理基础的必要性

Zhikang Chen, Tingting Zhu

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 本文提出将世界模型重新定义为可操作模拟器,强调因果结构和约束意识,以提升医疗决策中的反事实推理和长期预见能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10905 2026-01-19 cs.LG stat.ME 79%

Action Shapley: A Training Data Selection Metric for World Model in Reinforcement Learning

动作夏普利:用于强化学习世界模型的训练数据选择度量

Rajat Ghosh, Debojyoti Dutta

机构 * Nutanix Inc.(Nutanix公司)

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 本文提出动作夏普利度量,用于高效选择强化学习世界模型的训练数据,通过随机动态算法提升计算效率并改进数据选择性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09452 2026-01-15 cs.CV 79%

MAD: Motion Appearance Decoupling for efficient Driving World Models

MAD:用于高效驾驶世界模型的运动外观解耦

Ahmad Rahimi, Valentin Gerard, Eloi Zablocki, Matthieu Cord, Alexandre Alahi

机构 * Sorbonne Université(索邦大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 MAD通过解耦运动学习与外观合成,高效地将通用视频扩散模型转化为可控的驾驶世界模型,实现低计算成本和高性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07964 2026-01-14 cs.AI 79%

Executable Ontologies in Game Development: From Algorithmic Control to Semantic World Modeling

可执行本体在游戏开发中的应用:从算法控制到语义世界建模

Alexander Boldachev

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 本文提出利用可执行本体实现游戏开发中的语义世界建模,通过数据流条件实现任务中断,解决传统AI架构中的语义-过程鸿沟。

Comments 25 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04035 2026-01-08 cs.AI 79%

MobileDreamer: Generative Sketch World Model for GUI Agent

MobileDreamer: 用于GUI代理的生成式草图世界模型

Yilin Cao, Yufeng Zhong, Zhixiong Zeng, Liming Zheng, Jing Huang, Haibo Qiu, Peng Shi, Wenji Mao, Wan Guanglu

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Meituan(美团)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 MobileDreamer通过生成式草图世界模型和rollout想象策略,提升GUI代理在长周期任务中的决策能力,任务成功率提升5.25%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03517 2026-01-08 cs.CV 79%

Semantic Belief-State World Model for 3D Human Motion Prediction

语义信念状态世界模型用于3D人体运动预测

Sarim Chaudhry

机构 * Department of Computer Science, Purdue University(计算机科学系,普渡大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 本文提出语义信念状态世界模型,通过将人体视为世界模型状态空间的一部分,实现更稳定的长期运动预测与模拟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00930 2026-01-06 cs.IR cs.AI 79%

AlignUSER: Human-Aligned LLM Agents via World Models for Recommender System Evaluation

AlignUSER: 通过世界模型对齐LLM代理以评估推荐系统

Nicolas Bougie, Gian Maria Marconi, Tony Yip, Narimasa Watanabe

机构 * Woven by Toyota(丰田织物)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 AlignUSER通过世界模型驱动的代理学习,利用人类交互数据提升推荐系统评估的准确性与真实用户行为的对齐程度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00051 2026-01-05 cs.CV 79%

TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model

TeleWorld:面向动态多模态合成的4D世界模型

Yabo Chen, Yuanzhi Liang, Jiepeng Wang, Tingxi Chen, Junfei Cheng, Zixiao Gu, Yuyang Huang, Zicheng Jiang, Wei Li, Tian Li, Weichen Li, Zuoxin Li, Guangce Liu, Jialun Liu, Junqi Liu, Haoyuan Wang, Qizhen Weng, Xuan'er Wu, Xunzhi Xiang, Xiaoyan Yang, Xin Zhang, Shiwen Zhang, Junyu Zhou, Chengcheng Zhou, Haibin Huang, Chi Zhang, Xuelong Li

机构 * TeleWorld Team(TeleWorld团队)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 TeleWorld提出了一种实时多模态4D世界建模框架,通过生成-重建-引导范式实现动态场景重建与长期记忆,提升世界模型的交互性和计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22336 2025-12-30 cs.AI cs.CL 79%

Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback

Agent2World: 通过自适应多智能体反馈学习生成符号世界模型

Mengkang Hu, Bowei Xia, Yuran Wu, Ailing Yu, Yude Zou, Qiguang Chen, Shijian Wang, Jiarui Jin, Kexin Li, Wenxiang Jiao, Yuan Lu, Ping Luo

机构 * The University of Hong Kong(香港大学) Xiaohongshu Inc.(小红书公司) UESTC(电子科技大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.AI

AI总结 Agent2World通过自适应多智能体反馈学习生成符号世界模型,提升推理和微调性能,实现30.95%的相对提升。

Comments 48 pages, 15 tables, 7 figures, Project page: https://agent2world.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19855 2025-12-29 cs.LG cs.HC 79%

Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning

在LLM中诱导因果世界模型以实现零样本物理推理

Aditya Sharma, Ananya Gupta, Chengyu Wang, Chiamaka Adebayo, Jakub Kowalski

机构 * Department of Electrical Engineering(电气工程系) Indian Institute of Technology Bombay(印度理工学院孟买分校) Department of Computer Science and Automation(计算机科学与自动化系) Indian Institute of Science(印度科学研究所) Department of Computer Science(计算机科学系) San Francisco State University(旧金山州立大学) University of Lagos(拉各斯大学) Faculty of Mathematics, Informatics, and Mechanics(数学、信息学与力学学院) University of Warsaw(华沙大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 本文提出CWMI框架,通过引入因果物理模块和因果干预损失,使LLM具备因果推理能力,在零样本物理推理任务中取得显著成效。

Comments 12 pages, 4 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20615 2025-12-24 cs.CV 79%

Active Intelligence in Video Avatars via Closed-loop World Modeling

通过闭环世界建模实现视频虚拟角色的主动智能

Xuanhua He, Tianyu Yang, Ke Cao, Ruiqi Wu, Cheng Meng, Yong Zhang, Zhuoliang Kang, Xiaoming Wei, Qifeng Chen

机构 * The Hong Kong University of Science and Technology(香港科技大学) Meituan(美团) University of Science and Technology of China(中国科学技术大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 本文提出ORCA框架,通过闭环世界建模和双系统架构,实现视频虚拟角色的主动智能与目标导向行为。

Comments Project Page: https://xuanhuahe.github.io/ORCA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08139 2025-12-23 cs.IT cs.LG math.IT 79%

SCA-LLM: Spectral-Attentive LLM-Based Wireless World Modeling for Agentic Communications

SCA-LLM:基于频谱-注意力的LLM无线世界建模用于智能通信

Ke He, Le He, Lisheng Fan, Xianfu Lei, Thang X. Vu, George K. Karagiannidis, Symeon Chatzinotas

机构 * Interdisciplinary Centre for Security, Reliability and Trust (SnT), University of Luxembourg(安全、可靠性与信任跨学科研究中心(SnT),卢森堡大学) School of Computer Science of Guangzhou University(广州大学计算机科学学院) School of Information Science and Technology, Institute of Mobile Communications, Southwest Jiaotong University(信息科学与技术学院,移动通信研究所,西南交通大学) Department of Electrical and Computer Engineering, Aristotle University of Thessaloniki(电气与计算机工程系,塞萨洛尼基阿瑞斯托大学) Cyber Security Systems and Applied AI Research Center, Lebanese American University (LAU)(网络安全与应用人工智能研究中心,黎巴嫩美国大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 SCA-LLM通过频谱-注意力适配器将信道状态信息与LLM结合,实现无线世界建模,提升预测性能和零样本泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17152 2025-12-22 cs.CV 79%

PhysFire-WM: A Physics-Informed World Model for Emulating Fire Spread Dynamics

PhysFire-WM: 一种融合物理信息的世界模型用于模拟火灾扩散动力学

Nan Zhou, Huandong Wang, Jiahao Li, Yang Li, Xiao-Ping Zhang, Yong Li, Xinlei Chen

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 PhysFire-WM通过融合物理信息和跨任务协作训练策略,提升火灾扩散预测的物理真实性和几何准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13604 2025-12-16 cs.CV 79%

LongVie 2: Multimodal Controllable Ultra-Long Video World Model

LongVie 2: 多模态可控超长视频世界模型

Jianxiong Gao, Zhaoxi Chen, Xian Liu, Junhao Zhuang, Chengming Xu, Jianfeng Feng, Yu Qiao, Yanwei Fu, Chenyang Si, Ziwei Liu

机构 * FDU(福建大学) NJU(南京大学) NTU(National University of Taiwan) NVIDIA(英伟达) THU(清华大学) Shanghai AI Laboratory(上海人工智能实验室)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 LongVie 2通过多模态引导、退化感知训练和历史-上下文引导,实现了超长视频的可控生成,支持连续视频生成长达五分钟,推动视频世界建模的统一发展。

Comments Project Page: https://vchitect.github.io/LongVie2-project/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12751 2025-12-16 cs.CV 79%

GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation

GenieDrive: 向具有物理意识的驾驶世界模型迈进:基于4D占用的视频生成

Zhenya Yang, Zhe Liu, Yuxiang Lu, Liping Hou, Chenxuan Miao, Siyi Peng, Bailan Feng, Xiang Bai, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Huawei Noah’s Ark Lab(华为诺亚实验室) Huazhong University of Science and Technology(华中科技大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 GenieDrive通过4D占用引导的视频生成,实现物理意识的驾驶视频生成,提升预测精度和视频质量。

Comments The project page is available at https://huster-yzy.github.io/geniedrive_project_page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11226 2025-12-15 cs.CV 79%

FutureX: Enhance End-to-End Autonomous Driving via Latent Chain-of-Thought World Model

FutureX: 通过潜在链式思维世界模型增强端到端自动驾驶

Hongbin Lin, Yiming Yang, Yifan Zhang, Chaoda Zheng, Jie Feng, Sheng Wang, Zhennan Wang, Shijia Chen, Boyang Wang, Yu Zhang, Xianming Liu, Shuguang Cui, Zhen Li

机构 * FNii-Shenzhen(FNii深圳分部) SSE, CUHK-Shenzhen(SSE,CUHK深圳分校) Xpeng Motors(小鹏汽车) MiroMind AI(MiroMind人工智能) Xidian University(西安电子科技大学)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 FutureX通过链式思维世界模型增强端到端自动驾驶,提升复杂场景下的运动规划质量与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10016 2025-12-12 cs.LG 79%

Latent Action World Models for Control with Unlabeled Trajectories

潜在动作世界模型用于带未标记轨迹的控制

Marvin Alles, Xingyuan Zhang, Patrick van der Smagt, Philip Becker-Ehmck

机构 * Technical University of Munich(慕尼黑技术大学) Machine Learning Research Lab, Volkswagen Group(大众集团机器学习研究实验室) Eötvös Loránd University Budapest(布达佩斯欧多立大学) Foundation Robotics Labs(基础机器人实验室)

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 本文提出了一种潜在动作世界模型,通过结合动作条件和无动作数据,提升离线强化学习在未标记轨迹上的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01821 2025-12-09 cs.CV 79%

Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling

通过想象看见:通过隐式空间世界建模学习场景几何

Meng Cao, Haokun Lin, Haoyuan Li, Haoran Tang, Rongtao Xu, Dong An, Xue Liu, Ian Reid, Xiaodan Liang

机构 * MBZUAI SYSU(南方科技大学) PKU(北京大学) Spatialtemporal AI(时空人工智能)

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 本文提出MILO和RePE,通过隐式空间世界建模提升多模态大语言模型的空间推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06890 2025-12-08 cs.LG stat.ML 79%

SPARTAN: A Sparse Transformer World Model Attending to What Matters

SPARTAN:一种稀疏变换器世界模型,聚焦关键要素

Anson Lei, Bernhard Schölkopf, Ingmar Posner

机构 * Applied AI Lab University of Oxford, UK(应用人工智能实验室牛津大学,英国) MPI for Intelligent Systems Tübingen, Germany(智能系统Max Planck研究所,图宾根,德国)

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 SPARTAN通过稀疏正则化学习稀疏的交互图,提升世界模型对动态变化的适应能力与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05361 2025-12-08 physics.optics cs.LG physics.comp-ph 79%

FieldSeer I: Physics-Guided World Models for Long-Horizon Electromagnetic Dynamics under Partial Observability

FieldSeer I:基于物理的世界模型用于在部分可观测性下的长时间电磁动态

Ziheng Guo, Fang Wu, Maoxiong Zhao, Chaoqun Fang, Yang Bu

专题命中 具身推理 :world model(title,abstract);分类 cs.LG

AI总结 FieldSeer I通过几何条件的世界模型实现长时电磁动态预测,适用于光子设计的交互式数字双。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04040 2025-12-04 cs.CV 79%

RELIC: Interactive Video World Model with Long-Horizon Memory

RELIC:具有长时间记忆的交互式视频世界模型

Yicong Hong, Yiqun Mei, Chongjian Ge, Yiran Xu, Yang Zhou, Sai Bi, Yannick Hold-Geoffroy, Mike Roberts, Matthew Fisher, Eli Shechtman, Kalyan Sunkavalli, Feng Liu, Zhengqi Li, Hao Tan

专题命中 具身推理 :world model(title,abstract);分类 cs.CV

AI总结 RELIC通过统一框架实现长时记忆与实时交互,提升视频世界建模的准确性与稳定性。

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏