arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 3363 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 3363 篇

2511.06201 2025-11-11 cs.CV cs.HC 50%

Scene-Aware Urban Design: A Human-AI Recommendation Framework Using Co-Occurrence Embeddings and Vision-Language Models

Rodrigo Gallardo, Oz Fishman, Alexander Htet Kyaw

机构 * Department of Architecture(建筑系) Department of EECS(电子工程与计算机科学系) Massachusetts Institute of Technology(麻省理工学院)

专题命中 视觉空间推理 :planning(abstract)

Comments Accepted to NEURIPS 2025 Creative AI Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01322 2025-11-11 cs.CV 50%

FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors

Chenxi Li, Weijie Wang, Qiang Li, Bruno Lepri, Nicu Sebe, Weizhi Nie

机构 * Tianjin University(天津大学) University of Trento(特伦托大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Accepted by ACMMM2025, Our project webpage: https://tjulcx.github.io/FreeInsert/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05491 2025-11-10 cs.CV 50%

Visual Spatial Tuning

Rui Yang, Ziyu Zhu, Yanwei Li, Jingjia Huang, Shen Yan, Siyuan Zhou, Zhe Liu, Xiangtai Li, Shuangye Li, Wenqian Wang, Yi Lin, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Tsinghua University(清华大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.03325 2025-11-07 cs.CV 50%

SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding

Mauro Orazio Drago, Luca Carlini, Pelinsu Celebi Balyemez, Dennis Pierantozzi, Chiara Lena, Cesare Hassan, Danail Stoyanov, Elena De Momi, Sophia Bano, Mobarak I. Hoque

机构 * Dipartimento di Elettronica, Informazione e Bioingegneria (DEIB)(电子、信息与生物工程系) Politecnico di Milano(米兰理工大学) IRCCS Humanitas Research Hospital(IRCCS人类itas研究医院) UCL Hawkes Institute and Department of Computer Science(UCL Hawkes研究所和计算机科学系) University College London(伦敦大学学院) University of Manchester(曼彻斯特大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12795 2025-11-07 cs.CV 50%

EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting

Wei Zhang, Miaoxin Cai, Yaqian Ning, Tong Zhang, Yin Zhuang, Shijian Lu, He Chen, Jun Li, Xuerui Mao

机构 * School of Interdisciplinary Science, Beijing Institute of Technology(交叉科学学院,北京理工大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing, Beijing Institute of Technology(空间智能信息处理国家重点实验室,北京理工大学) School of Optics and Photonics, Beijing Institute of Technology(光学与 photonics 学院,北京理工大学) State Key Laboratory of Explosion Science and Safety Protection, Beijing(爆炸科学与安全防护国家重点实验室,北京)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13492 2025-11-05 cs.CV 50%

GeoSDF: Plane Geometry Diagram Synthesis via Signed Distance Field

Chengrui Zhang, Maizhen Ning, Tianyi Liu, Zihao Zhou, Jie Sun, Qiufeng Wang, Kaizhu Huang

机构 * Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) Duke Kunshan University(杜克大学昆山分校)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23603 2025-11-04 cs.CV 50%

PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity

Yuqian Yuan, Wenqiao Zhang, Xin Li, Shihao Wang, Kehan Li, Wentong Li, Jun Xiao, Lei Zhang, Beng Chin Ooi

机构 * Zhejiang University(浙江大学) DAMO Academy, Alibaba Group(阿里云图研究院) Hupan Lab(鸿篇实验室) The Hong Kong Polytechnic University(香港理工大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments 22 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26536 2025-10-31 cs.RO 50%

RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration

Huajie Tan, Cheng Chi, Xiansheng Chen, Yuheng Ji, Zhongxia Zhao, Xiaoshuai Hao, Yaoxu Lyu, Mingyu Cao, Junkai Zhao, Huaihai Lyu, Enshen Zhou, Ning Chen, Yankai Fu, Cheng Peng, Wei Guo, Dong Liang, Zhuo Chen, Mengsi Lyu, Chenrui He, Yulong Ao, Yonghua Lin, Pengwei Wang, Zhongyuan Wang, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(1 多媒体信息处理国家重点实验室,计算机学院,北京大学) Beijing Academy of Artificial Intelligence(2 北京人工智能研究院) Institute of Automation, Chinese Academy of Sciences(3 中国科学院自动化研究所) Beihang University(4 北航大学)

专题命中 视觉空间推理 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25070 2025-10-30 cs.CV 50%

Vision-Language Integration for Zero-Shot Scene Understanding in Real-World Environments

Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi

专题命中 视觉空间推理 :reasoning(abstract)

Comments Preprint under review at IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04897 2025-10-29 cs.CV 50%

From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes

Tianxu Wang, Zhuofan Zhang, Ziyu Zhu, Yue Fan, Jing Xiong, Pengxiang Li, Xiaojian Ma, Qing Li

机构 * State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) Tsinghua University(清华大学) Peking University(北京大学) Beijing Institute of Technology(北京理工大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Update v3 of the NeurIPS 2025 Datasets and Benchmarks paper (v2), including additional evaluations of state-of-the-art multimodal large language models. Project page: https://anywhere-3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23203 2025-10-28 cs.CV 50%

DecoDINO: 3D Human-Scene Contact Prediction with Semantic Classification

Lukas Bierling, Davide Pasero, Fleur Dolmans, Helia Ghasemi, Angelo Broere

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05795 2025-10-28 physics.ed-ph 50%

Creating a customisable Socratic AI physics tutor

Eugenio Tufino, Bor Gregorcic

专题命中 视觉空间推理 :reasoning(abstract)

Comments 7 pages, 3 figures

Journal ref Phys. Educ. 60 065037 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23449 2025-10-28 cs.MM cs.CV cs.IR 50%

CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection

Fanxiao Li, Jiaying Wu, Canyuan He, Wei Zhou

机构 * School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院) National University of Singapore(新加坡国立大学) Engineering Research Center of Cyberspace, Yunnan University(云南大学网络空间研究院)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21771 2025-10-28 cs.RO 50%

Improving the performance of AI-powered Affordable Robotics for Assistive Tasks

Dharunish Yugeswardeenoo

机构 * Central Bucks High School South(中央巴克斯高中南校区)

专题命中 视觉空间推理 :reasoning(abstract)

Comments 6 pages, 5 figures. Accepted to Conference on Robot Learning (CoRL 2025), Seoul, Korea

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21160 2025-10-27 cs.CV 50%

Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study

Guanlin Wu, Boyan Su, Yang Zhao, Pu Wang, Yichen Lin, Hao Frank Yang

机构 * Johns Hopkins University(约翰霍普金斯大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments NeurIPS 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21069 2025-10-27 cs.CV cs.RO 50%

ZING-3D: Zero-shot Incremental 3D Scene Graphs via Vision-Language Models

Pranav Saxena, Jimmy Chiun

机构 * Birla Institute of Technology and Science Pilani, K.K Birla Goa Campus(比拉理工学院和科学学院,比拉戈阿校园) Department of Mechanical Engineering, College of Design and Engineering, National University of Singapore(设计与工程学院机械工程系,新加坡国立大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20835 2025-10-27 q-bio.NC 50%

Rethinking the Simulation vs. Rendering Dichotomy: No Free Lunch in Spatial World Modelling

Dezhi Luo, Qingying Gao, Hokin Deng

专题命中 视觉空间推理 :reasoning(abstract)

Comments Accepted at NeurIPS 2025 Workshop on Space in Vision, Language, and Embodied AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20385 2025-10-24 cs.CV 50%

Positional Encoding Field

Yunpeng Bai, Haoxiang Li, Qixing Huang

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) Pixocial Technology(Pixocial技术)

专题命中 视觉空间推理 :reasoning(abstract)

Comments 8 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00441 2025-10-22 cs.RO 50%

Seeing through Uncertainty: Robust Task-Oriented Optimization in Visual Navigation

Yiyuan Pan, Yunzhe Xu, Zhe Liu, Hesheng Wang

机构 * Shanghai Jiao Tong University(上海交通大学) School of Automation and Intelligent Sensing(自动化与智能感知学院) Key Laboratory of System Control and Information Processing(系统控制与信息处理重点实验室) Ministry of Education of China(中华人民共和国教育部) National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家级重点实验室) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院)

专题命中 视觉空间推理 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17274 2025-10-21 cs.CV 50%

Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models

Katie Luo, Jingwei Ji, Tong He, Runsheng Xu, Yichen Xie, Dragomir Anguelov, Mingxing Tan

机构 * Computer and Information Sciences Department, Cornell University(康奈尔大学计算机与信息科学系) Waymo LLC(Waymo公司) UC Berkeley(伯克利大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments In proceedings of IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17034 2025-10-21 cs.CV 50%

Where, Not What: Compelling Video LLMs to Learn Geometric Causality for 3D-Grounding

Yutong Zhong

机构 * New York University(纽约大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00682 2025-10-21 cs.RO 50%

Immersive Explainability: Visualizing Robot Navigation Decisions through XAI Semantic Scene Projections in Virtual Reality

Jorge de Heuvel, Sebastian Müller, Marlene Wessels, Aftab Akhtar, Christian Bauckhage, Maren Bennewitz

机构 * University of Bonn(波恩大学) Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔人工智能与机器学习研究所) Center for Robotics(机器人中心) University of Mainz(美因茨大学) Fraunhofer Institute for Intelligent Analysis and Information Systems IAIS(弗劳恩霍夫智能分析与信息系统研究所)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15831 2025-10-20 cs.CV 50%

VISTA: A Test-Time Self-Improving Video Generation Agent

Do Xuan Long, Xingchen Wan, Hootan Nakhost, Chen-Yu Lee, Tomas Pfister, Sercan Ö. Arık

机构 * Google(谷歌)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07299 2025-10-20 cs.RO cs.CV 50%

MLFM: Multi-Layered Feature Maps for Richer Language Understanding in Zero-Shot Semantic Navigation

Sonia Raychaudhuri, Enrico Cancelli, Tommaso Campari, Lamberto Ballan, Manolis Savva, Angel X. Chang

机构 * Simon Fraser University(西蒙弗雷泽大学) University of Padova(帕多瓦大学) Fondazione Bruno Kessler (FBK)(布鲁诺·克塞勒基金会) Alberta Machine Intelligence Institute (Amii)(阿尔伯塔人工智能研究所)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14851 2025-10-17 cs.RO cs.MA 50%

SADCHER: Scheduling using Attention-based Dynamic Coalitions of Heterogeneous Robots in Real-Time

Jakob Bichler, Andreu Matoses Gimenez, Javier Alonso-Mora

机构 * Department for Cognitive Robotics, ME, Delft University of Technology(认知机器人系,代尔夫特理工大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments 7 pages, 5 figures. 2025 IEEE Int. Symposium on Multi-Robot and Multi-Agent Systems (MRS 2025). Website and Code: https://autonomousrobots.nl/paper_websites/sadcher_MRTA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15693 2025-10-17 cs.CV cs.MM 50%

SCENEFORGE: Enhancing 3D-text alignment with Structured Scene Compositions

Cristian Sbrolli, Matteo Matteucci

机构 * Department of Electronics, Information and Bioengineering(电子、信息与生物工程系)

专题命中 视觉空间推理 :reasoning(abstract)

Comments to appear in NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12777 2025-10-15 cs.CV 50%

What If : Understanding Motion Through Sparse Interactions

Stefan Andreas Baumann, Nick Stracke, Timy Phan, Björn Ommer

机构 * Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Project page and code: https://compvis.github.io/flow-poke-transformer

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07984 2025-10-15 cs.CV 50%

OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding

Jingli Lin, Chenming Zhu, Runsen Xu, Xiaohan Mao, Xihui Liu, Tai Wang, Jiangmiao Pang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) The University of Hong Kong(香港大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments 30 pages, a benchmark designed to evaluate Online Spatio-Temporal understanding from the perspective of an agent actively exploring a scene. Project Page: https://rbler1234.github.io/OSTBench.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08307 2025-10-15 cs.CV cs.RO 50%

DSM: Constructing a Diverse Semantic Map for 3D Visual Grounding

Qinghongbing Xie, Zijian Liang, Fuhao Li, Long Zeng

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China(清华大学深圳国际研究生院,清华大学,深圳,中国)

专题命中 视觉空间推理 :reasoning(abstract)

Comments 8 pages, 6 figures, Project Page: https://binicey.github.io/DSM

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04479 2025-10-14 cs.CV 50%

VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery

Nonghai Zhang, Zeyu Zhang, Jiazi Wang, Yang Zhao, Hao Tang

机构 * Peking University(北京大学) Beijing Jiaotong University(北京交通大学) La Trobe University(拉特罗布大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏