ADAPT: An Autonomous Forklift for Construction Site Operation
ADAPT:一种用于建筑工地作业的自主叉车
Johannes Huemer, Markus Murschitz, Matthias Schörghuber, Lukas Reisinger, Thomas Kadiofsky, Christoph Weidinger, Mario Niedermeyer, Benedikt Widy, Marcel Zeilinger, Csaba Beleznai, Tobias Glück, Andreas Kugi, Patrik Zips
机构
*
Center for Vision, Automation and Control(视觉、自动化与控制中心)
;
AIT Austrian Institute of Technology GmbH(奥地利技术研究所)
;
Automation and Control Institute(自动化与控制研究所)
;
Technische Universität Wien(维也纳技术大学)
Harnessing Embodied Agents: Runtime Governance for Policy-Constrained Execution
利用具身体 agent:运行时治理以实现政策约束执行
Xue Qin, Simin Luan, John See, Zeyd Boukhers, Cong Yang, Zhijun Li
机构
*
School of Software, Harbin Institute of Technology(哈尔滨工业大学软件学院)
;
School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)
;
School of Mathematical and Computer Sciences, Heriot-Watt University, Malaysia Campus(赫瑞-沃德大学马来西亚校区数学与计算机科学学院)
;
School of Future Science and Engineering, Soochow University(苏州大学未来科学与工程学院)
;
Fraunhofer Institute for Applied Information Technology(弗劳恩霍夫应用信息技术研究所)
TraversalBench: Challenging Paths to Follow for Vision Language Models
TraversalBench: 为视觉语言模型设计的复杂路径挑战测试集
Clara Petrova, Zhuo Chen, Marin Soljačić
机构
*
Massachusetts Institute of Technology, Department of Physics(麻省理工学院物理系)
;
Massachusetts Institute of Technology, Institute for Data, Systems, and Society(麻省理工学院数据、系统与社会研究所)
;
NSF AI Institute for Artificial Intelligence and Fundamental Interactions(国家科学基金会人工智能与基本相互作用AI研究所)
MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models
MIRAGE: 具有隐式推理和生成世界模型的移动智能体
Zhichao Yang, Yuanze Hu, Haojie Hao, Longkun Hao, Dongshuo Huang, Hongyu Lin, Gen Li, Lanqing Hong, Yihang Lou, Yan Bai
机构
*
Beihang University(北京航空航天大学)
;
Northwestern Polytechnical University(西北工业大学)
;
Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
;
National University of Singapore(新加坡国立大学)
;
Peking University(北京大学)
PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation
PixDLM:一种用于无人机推理分割的双路径多模态语言模型
Shuyan Ke, Yifan Mei, Changli Wu, Yonghan Zheng, Jiayi Ji, Liujuan Cao, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学)
IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
IndustryNav:探索动态工业导航中具身智能体的空间推理
Yifan Li, Lichi Li, Anh Dao, Xinyu Zhou, Wenjun Huang, Tianyi Ma, Yicheng Qiao, Zheda Mai, Daeun Lee, Zichen Chen, Pan Wang, Lehan Yang, Tianlong Wang, Zhen Tan, Sheng Li, Mohit Bansal, Yang Ni, Yu Kong
机构
*
Michigan State University(密歇根州立大学)
;
Independent Researcher(独立研究者)
;
University of California, Irvine(加州大学尔湾分校)
;
Ohio State University(俄亥俄州立大学)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
;
University of California, Santa Barbara(加州大学圣巴巴拉分校)
;
University of Pittsburgh(匹兹堡大学)
;
University of Virginia(弗吉尼亚大学)
;
Arizona State University(亚利桑那州立大学)
;
Purdue University Northwest(普渡大学西北分校)
CommentsPreprint. Accepted at NeurIPS 2025 Workshops on SPACE in Vision, Language, and Embodied AI (SpaVLE) as Oral, Embodied World Models for Decision Making (EWM), Aligning Reinforcement Learning Experimentalists and Theorists (ARLET), and Scaling Environments for Agents (SEA)
Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
计算机视觉中的推理:分类、模型、任务与方法论
Ayushman Sarkar, Zhenyu Yu, Mohd Yamani Idna Idris
机构
*
Department of Computer Science and Engineering, Birbhum Institute of Engineering and Technology(计算机科学与工程系,比罗尔理工学院)
;
College of Computer Science and Artificial Intelligence, Fudan University(计算机科学与人工智能学院,复旦大学)
;
Faculty of Computer Science and Information Technology, Universiti Malaya(计算机科学与信息技术学院,马来亚大学)