arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 3363 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 3363 篇

2510.09483 2025-10-13 cs.RO 50%

FOGMACHINE -- Leveraging Discrete-Event Simulation and Scene Graphs for Modeling Hierarchical, Interconnected Environments under Partial Observations from Mobile Agents

Lars Ohnemus, Nils Hantke, Max Weißer, Kai Furmans

机构 * Institute for Material Handling and Logistics, Karlsruhe Institute of Technology(材料搬运与物流研究所,卡尔斯鲁厄技术大学)

专题命中 视觉空间推理 :planning(abstract)

Comments submitted to the IEEE for possible publication; 8 pages, 3 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15117 2025-10-13 cs.CV 50%

Diffusion-based RGB-D Semantic Segmentation with Deformable Attention Transformer

Minh Bui, Kostas Alexis

机构 * Norwegian University of Science and Technology (NTNU)(挪威科学技术大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08967 2025-10-13 eess.IV cs.CV 50%

SAM2-3dMed: Empowering SAM2 for 3D Medical Image Segmentation

Yeqing Yang, Le Xu, Lixia Tian

机构 * Beijing Jiaotong University(北京交通大学)

专题命中 视觉空间推理 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07979 2025-10-13 cs.CV 50%

Visual Representation Alignment for Multimodal Large Language Models

Heeji Yoon, Jaewoo Jung, Junwan Kim, Hyungyu Choi, Heeseong Shin, Sangbeom Lim, Honggyu An, Chaehyun Kim, Jisang Han, Donghyun Kim, Chanho Eom, Sunghwan Hong, Seungryong Kim

机构 * KAIST AI(韩国科学技术院人工智能研究所) New York University(纽约大学) Chung-Ang University(Chung-Ang 大学) Korea University(韩国大学) ETH Zürich(苏黎世联邦理工学院)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Project Page: https://cvlab-kaist.github.io/VIRAL/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05551 2025-10-08 cs.CV 50%

When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding

Yan Shu, Hangui Lin, Yexin Liu, Yan Zhang, Gangyan Zeng, Yan Li, Yu Zhou, Ser-Nam Lim, Harry Yang, Nicu Sebe

机构 * University of Trento(特伦托大学) Hong Kong University of Science and Technology(香港科学与技术大学) University of International Relations(国际关系大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Nanjing University of Science and Technology(南京理工大学) VCIP & TMCC & DISSec, College of Computer Science, Nankai University(南开大学计算机学院) University of Central Florida(佛罗里达中央大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01862 2025-10-07 cs.RO 50%

ReLI: A Language-Agnostic Approach to Human-Robot Interaction

Linus Nwankwo, Bjoern Ellensohn, Ozan Özdenizci, Elmar Rueckert

机构 * Chair of Cyber-Physical Systems, Technical University of Leoben(技术大学莱博恩首席物理系统 Chair) Institute of Machine Learning and Neural Computation, Graz University of Technology(格拉茨技术大学机器学习与神经计算研究所)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01991 2025-10-03 cs.CV 50%

4DGS-Craft: Consistent and Interactive 4D Gaussian Splatting Editing

Lei Liu, Can Wang, Zhenghao Chen, Dong Xu

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22010 2025-10-02 cs.CV 50%

CoFFT: Chain of Foresight-Focus Thought for Visual Language Models

Xinyu Zhang, Yuxuan Dong, Lingling Zhang, Chengyou Jia, Zhuohang Dang, Basura Fernando, Jun Liu, Mike Zheng Shou

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) Ministry of Education Key Laboratory of Intelligent Networks and Network Security, China(教育部智能网络与网络安全重点实验室) Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, China(陕西省大数据知识工程重点实验室) IHPC, Agency for Science, Technology and Research, Singapore(新加坡科技研究局IHPC) Show Lab, National University of Singapore(新加坡国立大学Show实验室) College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算与数据科学学院)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17349 2025-10-02 cs.CV 50%

Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models

Jianing Qi, Jiawei Liu, Hao Tang, Zhigang Zhu

机构 * CUNY Graduate Center(纽约大学研究生中心) Borough of Manhattan Community College(曼哈顿社区学院) The City College of New York(纽约城市学院)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02723 2025-10-02 cs.RO 50%

ImpedanceGPT: VLM-driven Impedance Control of Swarm of Mini-drones for Intelligent Navigation in Dynamic Environment

Faryal Batool, Yasheerah Yaqoot, Malaika Zafar, Roohan Ahmed Khan, Muhammad Haris Khan, Aleksey Fedoseev, Dzmitry Tsetserukou

机构 * Intelligent Space Robotics Laboratory, Skolkovo Institute of Science and Technology(智能空间机器人实验室,斯克尔科沃科学与技术研究所)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Accepted in IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26039 2025-10-01 cs.CV 50%

SGS: Segmentation-Guided Scoring for Global Scene Inconsistencies

Gagandeep Singh, Samudi Amarsinghe, Urawee Thani, Ki Fung Wong, Priyanka Singh, Xue Li

专题命中 视觉空间推理 :reasoning(abstract)

Comments 6 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20873 2025-10-01 cs.CV 50%

Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models

Chaeyoung Jung, Youngjoon Jang, Jongmin Choi, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21432 2025-10-01 cs.RO cs.CV 50%

UAV-VLN: End-to-End Vision Language guided Navigation for UAVs

Pranav Saxena, Nishant Raghuvanshi, Neena Goveas

机构 * Birla Institute of Technology and Science Pilani, K.K Birla Goa Campus(比拉理工学院和科学学院,K.K比拉果阿校区)

专题命中 视觉空间推理 :reasoning(abstract)

Journal ref Proc. European Conference on Mobile Robots (ECMR), 2025, pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25190 2025-09-30 cs.CV 50%

Visual Jigsaw Post-Training Improves MLLMs

Penghao Wu, Yushan Zhang, Haiwen Diao, Bo Li, Lewei Lu, Ziwei Liu

机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) Linköping University(_linköping大学) SenseTime Research(商汤科技研究院)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23641 2025-09-30 cs.CV cs.RO 50%

From Static to Dynamic: a Survey of Topology-Aware Perception in Autonomous Driving

Yixiao Chen, Ruining Yang, Xin Chen, Jia He, Dongliang Xu, Yue Yao

机构 * Sems Shandong University(山东大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20705 2025-09-26 cs.RO 50%

Building Information Models to Robot-Ready Site Digital Twins (BIM2RDT): An Agentic AI Safety-First Framework

Reza Akhavian, Mani Amani, Johannes Mootz, Robert Ashe, Behrad Beheshti

机构 * Department of Civil, Construction, and Environmental Engineering, San Diego State University, San Diego, CA, United States(土木、建设与环境工程系,圣地亚哥州立大学) Department of Electrical and Computer Engineering, University of California, San Diego, San Diego, CA, United States(电气与计算机工程系,加州大学圣地亚哥分校) Department of Mechanical and Aerospace Engineering, University of California, San Diego, San Diego, CA, United States(机械与航空航天工程系,加州大学圣地亚哥分校) Department of Computer Science, San Diego State University, San Diego, CA, United States(计算机科学系,圣地亚哥州立大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17435 2025-09-23 cs.RO cs.SY eess.SY 50%

GPS Denied IBVS-Based Navigation and Collision Avoidance of UAV Using a Low-Cost RGB Camera

Xiaoyu Wang, Yan Rui Tan, William Leong, Sunan Huang, Rodney Teo, Cheng Xiang

机构 * Electrical and Computer Engineering Dept., National University of Singapore(新加坡国立大学电子与计算机工程系) National University of Singapore(新加坡国立大学)

专题命中 视觉空间推理 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16805 2025-09-23 cs.CV 50%

Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models

Md. Atabuzzaman, Ali Asgarov, Chris Thomas

专题命中 视觉空间推理 :reasoning(abstract)

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15772 2025-09-22 cs.CV 50%

Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation

Weimin Bai, Yubo Li, Weijian Luo, Wenzheng Chen, He Sun

机构 * Peking University(北京大学) Xiaohongshu Inc(小红书公司)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15717 2025-09-22 cs.RO 50%

Imagination at Inference: Synthesizing In-Hand Views for Robust Visuomotor Policy Inference

Haoran Ding, Anqing Duan, Zezhou Sun, Dezhen Song, Yoshihiko Nakamura

机构 * Department of Robotics, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(机器人系,Mohamed bin Zayed人工智能大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Submitted to IEEE for possible publication, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14967 2025-09-22 cs.RO cs.HC 50%

Affordance-Based Disambiguation of Surgical Instructions for Collaborative Robot-Assisted Surgery

Ana Davila, Jacinto Colan, Yasuhisa Hasegawa

机构 * Nagoya University, Japan(名古屋大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments To be presented at the 1st Workshop on Intelligent Cobodied Assistance and Robotic Empowerment (iCARE). 2025 Conference on Robot Learning (CoRL)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15132 2025-09-19 cs.CY cs.CV 50%

From Pixels to Urban Policy-Intelligence: Recovering Legacy Effects of Redlining with a Multimodal LLM

Anthony Howell, Nancy Wu, Sharmistha Bagchi, Yushim Kim, Chayn Sun

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10426 2025-09-18 cs.RO cs.MA 50%

DECAMP: Towards Scene-Consistent Multi-Agent Motion Prediction with Disentangled Context-Aware Pre-Training

Jianxin Shi, Zengqi Peng, Xiaolong Chen, Tianyu Wo, Jun Ma

机构 * School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) Robotics and Autonomous Systems Thrust, The Hong Kong University of Science and Technology (GZ)(香港科技大学(广州)机器人与自主系统方向) Division of Emerging Interdisciplinary Areas, The Hong Kong University of Science and Technology(香港科技大学新兴交叉领域研究所)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.13756 2025-09-18 cs.CV 50%

PlaneRecTR++: Unified Query Learning for Joint 3D Planar Reconstruction and Pose Estimation

Jingjia Shi, Shuaifeng Zhi, Kai Xu

机构 * National University of Defense Technology(国防科技大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments To be published in IEEE T-PAMI 2025. This is the journal extension of our ICCV 2023 paper "PlaneRecTR", which expands from single view reconstruction to simultaneous multi-view reconstruction and camera pose estimation. Note that the ICCV2023 PlaneRecTR paper could be found in the previous arxiv version [v2](arXiv:2307.13756v2)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13317 2025-09-17 cs.CV 50%

3D Aware Region Prompted Vision Language Model

An-Chieh Cheng, Yang Fu, Yukang Chen, Zhijian Liu, Xiaolong Li, Subhashree Radhakrishnan, Song Han, Yao Lu, Jan Kautz, Pavlo Molchanov, Hongxu Yin, Xiaolong Wang, Sifei Liu

机构 * MIT(麻省理工学院) NVIDIA(英伟达)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Project Website: https://www.anjiecheng.me/sr3d

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11292 2025-09-17 cs.CV 50%

Leveraging Geometric Priors for Unaligned Scene Change Detection

Ziling Liu, Ziwei Chen, Mingqi Gao, Jinyu Yang, Feng Zheng

机构 * Southern University of Science and Technology(南方科技大学) University of Sheffield(谢菲尔德大学) Spatialtemporal AI(时空AI)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08679 2025-09-17 cs.CV 50%

ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way

Rajarshi Roy, Devleena Das, Ankesh Banerjee, Arjya Bhattacharjee, Kousik Dasgupta, Subarna Tripathi

机构 * Kalyani Government Engineering College(卡利尼政府工程学院) Intel Labs(英特尔实验室)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11124 2025-09-16 cs.SD eess.AS 50%

STASE: A spatialized text-to-audio synthesis engine for music generation

Tutti Chi, Letian Gao, Yixiao Zhang

专题命中 视觉空间推理 :reasoning(abstract)

Comments Accepted to LLM4Music @ ISMIR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10748 2025-09-16 cs.CV 50%

SCOPE: Speech-guided COllaborative PErception Framework for Surgical Scene Segmentation

Jecia Z. Y. Mao, Francis X Creighton, Russell H Taylor, Manish Sahu

机构 * LCSR, Johns Hopkins University(约翰霍普金斯大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08757 2025-09-11 cs.RO cs.CV 50%

SocialNav-SUB: Benchmarking VLMs for Scene Understanding in Social Robot Navigation

Michael J. Munje, Chen Tang, Shuijing Liu, Zichao Hu, Yifeng Zhu, Jiaxun Cui, Garrett Warnell, Joydeep Biswas, Peter Stone

机构 * Department of Computer Science, The University of Texas at Austin(德克萨斯大学计算机科学系) Army Research Laboratory(陆军研究实验室) Sony AI(索尼人工智能)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Conference on Robot Learning (CoRL) 2025 Project site: https://larg.github.io/socialnav-sub

详情

展开后加载摘要…

URL PDF HTML 收藏