GeoWorld-VLM: Geometry from World Models for Vision-Language Models
GeoWorld-VLM:从世界模型中获取几何结构用于视觉-语言模型
Renjie Gu, Kaichen Zhou, Yan Luo, Mengyu Wang
机构
*
Harvard AI and Robotics Lab(哈佛人工智能与机器人实验室)
;
Kempner Institute for the Study of Natural and Artificial Intelligence(凯普纳自然与人工智能研究 institute)
;
Harvard University(哈佛大学)
PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments
PatchWorld:可执行世界模型的免梯度优化
Jiaxin Bai, Yue Guo, Yifei Dong, Jiaxuan Xiong, Tianshi Zheng, Yixia Li, Tianqing Fang, Yufei Li, Yisen Gao, Haoyu Huang, Zhongwei Xie, Hong Ting Tsang, Zihao Wang, Lihui Liu, Jeff Z. Pan, Yangqiu Song
机构
*
Hong Kong Baptist University(香港 Baptist 大学)
;
Independent Researcher(独立研究员)
;
HKUST(香港科技大学)
;
Beijing Institute of Technology(北京理工大学)
;
Southern University of Science and Technology(南方科技大学)
;
Wayne State University(韦恩州立大学)
;
University of Edinburgh(爱丁堡大学)
SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation
SparseWorld: 通过具有稀疏场景表示的世界模型增强端到端自动驾驶
Ruoyu Wang, Jingke Wang, Yukai Ma, Yuehao Huang, Shuangming Lei, Guanglin Xu, Aixue Ye, Yong Liu
机构
*
Institute of Cyber-Systems and Control, Zhejiang University(浙江大学控制系统研究所)
;
Labs, Huawei(华为2012实验室)
;
State Key Laboratory of Industrial Control Technology(国家工业控制技术重点实验室)
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Eastern Institute of Technology(东部技术研究所)
;
PhiGent
;
National University of Singapore(新加坡国立大学)
;
Tsinghua University(清华大学)
;
Wuhan University(武汉大学)
Bootstrap Theory of Representational Emergence: Explanatory Insufficiency as a Driver of Representation Learning and World Models
表征涌现的自举理论:解释不充分性作为表征学习与世界模型的驱动力
Jacques Raynal, Pierre Slangen, Elsa Raynal, Jacques Margerit
机构
*
Laboratory of Bioengineering and Nanosciences (LBN), University of Montpellier(生物工程与纳米科学实验室(LBN),蒙彼利埃大学)
;
EuroMov Digital Health in Motion, University of Montpellier, IMT Mines Alès(EuroMov数字健康运动,蒙彼利埃大学,IMT矿山阿尔勒)
;
Certified Sophrologist, Sensorimotor Practice, Montpellier, France(认证Sophrologist,运动觉实践,蒙彼利埃,法国)
;
Emeritus Professor, University of Montpellier(荣誉教授,蒙彼利埃大学)
Comments27 pages, 25 references, no figures or tables. Conceptual framework on representational emergence and explanatory insufficiency, with implications for representation learning, world models, autonomous AI, and scientific discovery
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
ThinkJEPA:赋予潜在世界模型大型视觉-语言推理能力
Haichao Zhang, Yijiang Li, Shwai He, Tushar Nagarajan, Mingfei Chen, Jianglin Lu, Ang Li, Yun Fu
机构
*
Northeastern University(东北大学)
;
University of California San Diego(加州大学圣地亚哥分校)
;
University of Maryland(马里兰大学)
;
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
;
University of Washington(华盛顿大学)
Abdulaziz Alyahya, Abdallah Al Siyabi, Markus R. Ernst, Luke Yang, Levin Kuhlmann, Gideon Kowadlo
机构
*
Imam Mohammad Ibn Saud Islamic University (IMSIU)(伊玛姆·穆罕默德·本·沙特伊斯兰大学)
;
Monash University(莫纳什大学)
;
University of New South Wales, Sydney(新南威尔士大学,悉尼)
;
Cerenaut
Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models
使用潜在空间生成世界模型减轻自动驾驶模仿学习中的协变量转移
Alexander Popov, Alperen Degirmenci, David Wehr, Shashank Hegde, Ryan Oldja, Alexey Kamenev, Bertrand Douillard, David Nistér, Urs Muller, Ruchi Bhargava, Stan Birchfield, Nikolai Smolyanskiy
Comments8 pages, 6 figures, original September 2024, accepted at ICRA 2025 Workshop "Robots in the Wild", for associated video file, see https://youtu.be/7m3bXzlVQvU
CommentsRevised version: removed an inconclusive development-only event-relative phase analysis; the main rank-4 carrier, fresh-checkpoint replication, B1/B2 temporal reuse, position-edit, and joint-edit conclusions are unchanged