arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

世界模型

面向环境建模、时序预测、仿真规划、具身智能和自动驾驶的世界模型方法与应用。

共收录 6450 信号源:cs.AI, cs.LG, cs.CV, cs.RO, cs.MA

1. 通用世界模型 4308 篇

2607.17523 2026-07-21 cs.CV cs.AI cs.CL 新提交 82%

Thinking in Video: Can Video Generators Really Reason About the Real World?

视频中的思考:视频生成器真的能对现实世界进行推理吗?

Yongheng Zhang, Guang Yang, Ruihan Hou, Qiguang Chen, Ziang Liu, Xiaolong Liu, Manman Zhang, Yanchao Hao, Zheng Wei, Hao Wu, Libo Qin, Peishan Dai, Yinghui Li, Di Yin, Xing Sun

机构 * Central South University(中南大学) Tencent(腾讯) Tsinghua University(清华大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 探讨视频生成器能否对现实世界推理,引入因果生成双判断(CGDJ)评估,发现开源模型无明确因果感知却有合理动态,先进闭源系统推理与生成一致性有限,还揭示了视听失调问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13017 2026-07-15 cs.RO cs.CV 新提交 82%

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

FlowWAM:光流作为世界动作模型的统一动作表示

Yixiang Chen, Peiyan Li, Yuan Xu, Qisen Ma, Jiabing Yang, Kai Wang, Jianhua Yang, Dong An, He Guan, Gaoteng Liu, Jianlou Si, Jun Huang, Jing Liu, Nianfeng Liu, Yan Huang, Liang Wang

机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) FiveAges(无) MBZUAI(无) Alibaba Group(阿里巴巴集团)

专题命中 通用世界模型 :world model(abstract);world-model(abstract);world model(abstract);world-model(abstract)

AI总结 研究针对世界动作模型控制中动作表示难题,提出FlowWAM双流扩散框架,以光流为统一动作表示。该框架可实现WAMs两种模式,能利用无动作标签视频预训练,实验表明在操纵和世界建模任务中表现优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16533 2026-07-07 cs.AI cs.CV 新提交 82%

Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI

Kairos: 面向物理AI的原生世界模型栈

Kairos Team, Fei Wang, Shan You, Qiming Zhang, Tao Huang, Zuoyi Fu, Zhisheng Zheng, Yunlong Xi, Feng Lv, Xiaoming Wu, Zeyu Liu, Cong Wan, Pu Li, Ruiqing Yang, Xiaoou Li, Wei Wang, Kangkang Zhu, Yuwei Zhang, Shi Fu, Zheng Zhang, Xiaoning Wu, Xuzeng Fan, Dacheng Tao, Xiaogang Wang

机构 * Kairos Team(Kairos团队)

专题命中 通用世界模型 :world model(abstract);world-model(abstract);world model(abstract);world-model(abstract)

AI总结 提出Kairos原生世界模型栈,通过跨具身数据课程、混合线性时间注意力架构和部署感知系统协同设计,实现世界知识获取、长时程状态保持与高效执行,在具身世界模型等基准上达到顶级性能。

Comments Kairos Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02376 2026-07-03 cs.AI cs.MA 新提交 82%

Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems

面向安全关键实时自主系统的硬件强制语义协调

Uwe M. Borghoff, Paolo Bottoni, Remo Pareschi

机构 * Department of Computer Science University of the Bundeswehr Munich Neubiberg, Germany Department of Computer Science Sapienza University of Rome Rome, Italy STAKE Lab University of Molise Campobasso, Italy

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对异构组件并发运行下的协调问题,提出基于FPGA的硬件强制语义协调架构,将TB-CSPN协调机制映射到硬件原语,实现确定性协调和安全保障。

Comments 1 figure, 6 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26964 2026-06-29 cs.AI cs.CV 新提交 82%

Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds

Look-Before-Move:动态3D故事世界中的叙事驱动世界视觉注意力

Jiaming Bian, Bingliang Li, Yuehao Wu, Pichao Wang, Zhi Wang, Hailan Ma, Huadong Mo, Zhenhong Sun

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出Look-Before-Move框架,通过语义观察合约、蒙特卡洛视点搜索和语义轨迹接地,在动态3D故事世界中实现叙事驱动的视觉注意力规划,提升主体感知、意图一致性和轨迹质量。

Comments 25 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.14622 2026-06-25 cs.LG cs.AI 版本更新 82%

ACT-JEPA: Novel Joint-Embedding Predictive Architecture for Efficient Policy Representation Learning

ACT-JEPA:用于高效策略表征学习的新型联合嵌入预测架构

Aleksandar Vujinovic, Aleksandar Kovacevic

机构 * Faculty of Technical Sciences, University of Novi Sad(技术科学学院,诺维萨德大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出ACT-JEPA架构,结合模仿学习与自监督学习,通过联合预测动作序列和潜在观测序列学习鲁棒世界模型,在多个环境中任务成功率提升达10%。

Comments Published version

Journal ref IEEE Access, vol. 14, pp. 78895-78906, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19476 2026-06-19 cs.LG cs.AI 新提交 82%

Can In-Context Learning Support Intrinsic Curiosity?

上下文学习能否支持内在好奇心?

Eric Elmoznino, Sangnie Bhardwaj, Johannes von Oswald, Rajai Nasser, Blaise Agüera y Arcas, João Sacramento, Rif A. Saurous, Guillaume Lajoie

机构 * Google – Paradigms of Intelligence Team(Google – 智能范式团队) Google DeepMind(谷歌DeepMind)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 研究利用序列模型的上下文学习能力作为即时无更新世界模型,以消除传统内在好奇心方法中梯度下降的计算瓶颈,理论证明在非时间设置下可渐近收敛到真实学习进度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19408 2026-06-19 cs.LG cs.RO 新提交 82%

FlexLAM: Resolving the Bottleneck Trade-off in Latent Action Learning

FlexLAM: 解决潜在动作学习中的瓶颈权衡

Takanori Yoshimoto, Yang Hu, Naruya Kondo, Tatsuya Matsushima

机构 * University of Tsukuba(筑波大学) The University of Tokyo(东京大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对潜在动作模型中固定容量瓶颈导致的权衡问题,提出FlexLAM,通过嵌套dropout实现变长潜在动作,在不增加架构或损失的情况下,在稀缺标签和低回报任务中优于固定容量模型,并支持推理时调整令牌预算。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06212 2026-06-16 cs.CV cs.AI 版本更新 82%

Akasha 2: Hamiltonian State Space Duality and Visual-Language Joint Embedding Predictive Architectur

Akasha 2: 哈密顿状态空间对偶与视觉-语言联合嵌入预测架构

Yani Meziani

机构 * Independent AI Researcher(独立AI研究员) Québec (QC), Canada(魁北克(QC),加拿大)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出 Akasha 2 多模态架构,结合哈密顿状态空间对偶与视觉-语言联合嵌入预测,通过稀疏混合哈密顿专家和哈密顿流匹配实现超低延迟视频预测与合成,在保持能量守恒下取得 SOTA 性能。

Comments No supporting claims were validated in this automated agentic R&D research run

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.03413 2026-06-05 cs.LG cs.AI 82%

Learning to Theorize the World from Observation

从观察中学习理论化世界

Doojin Baek, Gyubin Lee, Junyeob Baek, Hosung Lee, Sungjin Ahn

机构 * University of Washington(华盛顿大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 受认知科学启发,提出Learning-to-Theorize范式,通过神经理论家(NEO)模型从原始非文本观测中推断显式解释性理论,实现基于解释的泛化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19312 2026-06-05 cs.LG cs.AI 82%

LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

LeWorldModel:从像素稳定端到端联合嵌入预测架构

Lucas Maes, Quentin Le Lidec, Damien Scieur, Yann LeCun, Randall Balestriero

机构 * Mila & Université de Montréal(Mila与蒙特利尔大学) New York University(纽约大学) Samsung SAIL(三星SAIL) Brown University(布朗大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出LeWorldModel,一种通过仅使用两个损失项从原始像素稳定端到端训练的联合嵌入预测架构,显著减少了可调损失超参数,并在多种2D和3D控制任务中表现出色,同时在物理结构编码和物理不合理的事件检测方面展示了其能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
1610.05984 2026-06-04 cs.NE cs.AI cs.LG cs.SY eess.SY 82%

Particle Swarm Optimization for Generating Interpretable Fuzzy Reinforcement Learning Policies

粒子群优化用于生成可解释的模糊强化学习策略

Daniel Hein, Alexander Hentschel, Thomas Runkler, Steffen Udluft

机构 * Siemens AG, Corporate Technology(西门子公司企业技术部)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出一种基于模糊粒子群强化学习(FPSRL)的方法,通过训练参数在模拟真实系统动态的世界模型上生成可解释的模糊强化学习策略,适用于无法进行在线学习的领域。

Journal ref Engineering Applications of Artificial Intelligence, Volume 65C, October 2017, Pages 87-98

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13428 2026-06-02 cs.CL cs.AI cs.LG 82%

Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models

Softplus注意力与重加权提升大语言模型的长度外推能力

Bo Gao, Michael W. Spratling, Letizia Gionfrida

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 提出一种两阶段注意力机制,用Softplus和l1归一化替代Softmax,并引入基于不变熵的动态缩放因子和重加权机制,以提升数值稳定性、缓解注意力下沉现象,并显著改善长度外推性能。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31558 2026-06-01 cs.LG cs.AI 82%

Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization

位置注意力头与符号注意力头:学习动态、RoPE几何和长度泛化

Felipe Urrutia, Juan José Alegría, Cinthia Sanchez Macias, Jorge Salas, Cristian B. Calderon, Cristobal Rojas

机构 * CENIA & Faculty of Mathematics UC Santiago(CENIA与圣托里尼大学数学系) IMC UC & CENIA Santiago(UC IMC与圣托里尼CENIA)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 通过控制实验研究Transformer注意力头在位置推理和符号推理任务中的学习动态,发现位置和符号注意力头的不同机制及其对长度泛化的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29357 2026-05-29 cs.AI cs.LG cs.PL 82%

PassNet: Scaling Large Language Models for Graph Compiler Pass Generation

PassNet: 为图编译器通生成扩展大型语言模型

Yiqun Liu, Yingsheng Wu, Ruqi Yang, Enrong Zheng, Honglei Qiu, Sijun He, Tai Liang, Jingjing Wu, Yuhan Zhou, Yiwei Zhang, Dongyan Chen, Weihan Yi, Xinqi Li, Siqi Bao

机构 * Baidu, Inc.(百度公司)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 针对编译器默认优化在长尾子图上性能不佳的问题,提出PassNet生态系统,包含大规模数据集和基准测试,通过微调小模型在少量轨迹上即可接近前沿模型性能。

Comments Code and data available at https://github.com/PaddlePaddle/PassNet

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17923 2026-05-19 cs.DC cs.AI cs.LG 82%

AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training

AdaptiveLoad: 向高效视频扩散变换器训练迈进

Yucheng Guo, Yongjian Guo, Zhong Guan, Haoran Sun, Wen Huang, Wanting Xu, Jing Long, Shuai Di, Junwu Xiong

机构 * Tsinghua University(清华大学) Peking University(北京大学) Tianjin University(天津大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出AdaptiveLoad框架,通过双约束自适应负载平衡系统和融合LayerNorm-Modulate CUDA内核,解决视频生成模型中大规模视频扩散变换器(如DiT和MMDiT)训练中的计算不平衡问题,实验显示其在Wan 2.1世界模型上提升了计算效率和训练吞吐量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16154 2026-05-18 cs.LG cs.RO 82%

Learn Where Outcomes Diverge: Efficient VLA RL via Probabilistic Chunk Masking

学习结果分歧之处:通过概率块掩码实现高效的VLA强化学习

Vaidehi Bagaria, Nikshep Grampurohit, Pulkit Verma

机构 * Indian Institute of Technology Madras(印度理工学院马德拉斯学院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出概率块掩码(PCM),通过选择性分配梯度计算来提升GRPO-based VLA RL的效率,实现更快的训练速度和更低的内存消耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08944 2026-05-14 cs.LG cs.MA 82%

Multi-Agent Decision-Focused Learning via Value-Aware Sequential Communication

多智能体决策导向学习 via 值感知序列通信

Benjamin Amoh, Geoffrey Parker, Wesley Marrero

机构 * Thayer School of Engineering, Dartmouth College(达特茅斯大学泰勒工程学院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出SeqComm-DFL方法,通过值感知序列通信与决策导向学习结合,提升多智能体任务性能。方法引入序列Stackelberg条件生成消息,利用信息论界限证明收敛性,并在协作医疗和StarCraft多智能体挑战中取得显著奖励和胜率提升。

Comments 9 pages, 2 figues, 1 table, neurips 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.28122 2026-05-01 cs.CV cs.LG 82%

Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces

超越高斯瓶颈:基于拓扑对齐的视觉Transformer特征空间编码

Andrew Bond, Ilkin Umut Melanlioglu, Erkut Erdem, Aykut Erdem

机构 * Department of Computer Engineering, Koç University, Istanbul, Turkey(科克大学计算机工程系,伊斯坦布尔,土耳其) Department of Computer Engineering, Hacettepe University, Ankara, Turkey(哈恰塔佩大学计算机工程系,安卡拉,土耳其) KUIS AI Research Center, Istanbul, Turkey(KUIS人工智能研究中心,伊斯坦布尔,土耳其) Department of Electrical and Electronics Engineering, Koç University, Istanbul, Turkey(科克大学电气与电子工程系,伊斯坦布尔,土耳其)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出S²VAE框架,通过压缩和表示场景的3D状态,包括相机运动、深度和点结构,以提升视觉模型的几何一致性。实验显示,几何对齐的超球面隐空间在高压缩条件下优于传统高斯瓶颈。

Comments 16 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20714 2026-04-30 cs.CV cs.AI 82%

Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation

Inferix:基于块扩散的下一代世界模拟推理引擎

Inferix Team, Tianyu Feng, Yizeng Han, Jiahao He, Yuanyu He, Xi Lin, Teng Liu, Hanfeng Lu, Jiasheng Tang, Wei Wang, Zhiyuan Wang, Jichao Wu, Mingyang Yang, Yinghao Yu, Zeyu Zhang, Bohan Zhuang

机构 * Inferix Team(Inferix 团队)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 Inferix 是一种基于块扩散的下一代推理引擎,通过优化半自回归解码过程实现沉浸式世界合成,支持交互式视频流和性能分析,结合 LV-Bench 提供高效评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00911 2026-04-23 cs.CR cs.AI cs.ET cs.HC cs.LG 82%

Device-Native Autonomous Agents for Privacy-Preserving Negotiations

设备本机自主代理用于隐私保护的谈判

Joyjit Roy, Samaresh Kumar Singh

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出基于设备的自主代理系统,通过零知识证明和本地推理实现隐私保护的谈判,实验显示在保险和B2B采购场景中成功率高,隐私保护效果显著。

Comments 9 pages, 6 figures, 9 tables. This version updates metadata after publication in IEEE Xplore

Journal ref 2026 IEEE SoutheastCon, Huntsville, AL, USA, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20627 2026-04-23 cs.LG cs.RO 82%

Occupancy Reward Shaping: Improving Credit Assignment for Offline Goal-Conditioned Reinforcement Learning

占用奖励塑造:改进离线目标导向强化学习中的信用分配

Aravind Venugopal, Jiayu Chen, Xudong Wu, Chongyi Zheng, Benjamin Eysenbach, Jeff Schneider

机构 * Carnegie Mellon University(卡内基梅隆大学) The University of Hong Kong(香港大学) INFIFORCE Intelligent Technology(INFIFORCE智能技术) Princeton University(普林斯顿大学)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出占用奖励塑造方法,通过提取世界模型中的时间信息,改进稀疏奖励下的信用分配问题,在13种长周期运动和操作任务中提升了2.2倍性能。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16320 2026-04-21 cs.SE cs.AI cs.LG 82%

How Robustly do LLMs Understand Execution Semantics?

大语言模型如何理解执行语义的鲁棒性?

Claudio Spiess, Prem Devanbu, Earl T. Barr

机构 * University College London(伦敦大学学院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 研究通过代码输出预测任务评估大语言模型对代码理解的鲁棒性,发现开源模型在代码变换和输入扰动下表现稳定,而前沿模型GPT-5.2表现出显著的脆弱性,且异常处理能力不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02263 2026-04-20 cs.CV cs.AI 82%

Social-JEPA: Emergent Geometric Isomorphism

Social-JEPA:涌现的几何同构

Haoran Zhang, Youjin Wang, Yi Duan, Rong Fu, Dianyu Zhao, Sicheng Fan, Shuaishuai Cao, Wentao Guo, Xiao Zhou

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 Social-JEPA通过让不同视角的独立代理学习环境模型,发现其潜在空间近似线性同构,从而实现跨代理的透明转换与高效迁移学习。

Comments This preprint is withdrawn due to significant errors in the emergent geometric isomorphism results that necessitate full rewriting, coupled with unresolved author disagreement on authorship. A corrected and revised manuscript will be released separately

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06168 2026-04-16 cs.CV cs.RO 82%

Action Images: End-to-End Policy Learning via Multiview Video Generation

动作图像:通过多视图视频生成实现端到端策略学习

Haoyu Zhen, Zixian Gao, Qiao Sun, Yilin Zhao, Yuncong Yang, Yilun Du, Pengsheng Guo, Tsun-Hsuan Wang, Yi-Ling Qiao, Chuang Gan

机构 * UMass Amherst(马萨诸塞大学阿默斯特分校) NVIDIA(英伟达) Harvard University(哈佛大学) Genesis AI

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出Action Images,通过多视图视频生成实现策略学习,利用像素化的动作表示,使视频模型本身成为零样本策略,提升视频-动作联合生成质量。

Comments Project Page: https://actionimages.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02289 2026-04-03 cs.CV cs.AI 82%

Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation

Omni123:通过统一文本到2D和3D生成探索有限3D数据的3D原生基础模型

Chongjie Ye, Cheng Cao, Chuanyu Pan, Yiming Hao, Yihao Zhi, Yuanming Hu, Xiaoguang Han

机构 * FNii-Shenzhen(FNii-深圳) SSE, CUHK(SZ)(香港中文大学(深圳)理工学院) Meshy AI

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 Omni123通过统一文本到2D和3D生成,利用文本-图像-3D的跨模态一致性作为隐式约束,提升3D生成的几何一致性与语义对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29572 2026-04-01 cs.GR cs.AI cs.CV 82%

Turbo4DGen: Ultra-Fast Acceleration for 4D Generation

Turbo4DGen: 4D生成的超快加速

Yuanbin Man, Ying Huang, Zhile Ren, Miao Yin

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出Turbo4DGen框架,通过时空缓存机制和动态语义感知注意力剪枝,显著减少4D生成中的冗余计算,实现9.7倍速度提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14375 2026-03-30 cs.CV cs.AI 82%

The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics

运动的脉搏:从视觉动态测量物理帧率

Xiangbo Gao, Mingyang Wu, Siyuan Yang, Jiongze Yu, Pardis Taghavi, Fangzhou Lin, Zhengzhong Tu

机构 * Texas A\&M University

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 本文提出Visual Chronometer,通过视觉动态直接恢复物理帧率,解决视频生成中因训练数据帧率不一致导致的物理运动速度模糊问题,提升生成视频的自然度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24541 2026-03-26 cs.CV cs.AI 82%

SEGAR: Selective Enhancement for Generative Augmented Reality

SEGAR:生成增强现实的定向增强

Fanjun Bu, Chenyang Yuan, Hiroshi Yasuda

机构 * Cornell University(康奈尔大学) Cornell Tech(康奈尔科技) Toyota Research Institute(电装研究院)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 SEGAR结合扩散世界模型与定向修正阶段,实现生成增强现实中的区域特定编辑与安全区域对齐,为未来帧的生成、缓存和定向修正提供基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19219 2026-03-20 cs.CV cs.LG 82%

DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding

DriveTok: 3D驾驶场景分词用于统一多视角重建与理解

Dong Zhuo, Wenzhao Zheng, Sicheng Zuo, Siming Yan, Lu Hou, Jie Zhou, Jiwen Lu

机构 * Tsinghua University(清华大学) Yinwang Intelligent Technology Co. Ltd.(云网智能科技有限公司)

专题命中 通用世界模型 :world model(abstract);world models(abstract);world model(abstract);world models(abstract)

AI总结 DriveTok通过3D变形交叉注意力实现高效多视角重建与理解,整合语义、几何与纹理信息,提升多视角分词效率与一致性。

Comments Project Page: https://paryi555.github.io/DriveTok/ Code: https://github.com/paryi555/DriveTok

详情

展开后加载摘要…

URL PDF HTML 收藏