arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 3363 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 3363 篇

2506.19312 2025-06-25 cs.CV cs.AI 57%

Capturing Fine-Grained Alignments Improves 3D Affordance Detection

Junsei Tokumitsu, Yuiga Wada

机构 * Keio AI Research Center(庆应AI研究中心)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments MVA 2025 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17425 2025-06-25 eess.IV cs.AI cs.CV 57%

Trans${^2}$-CBCT: A Dual-Transformer Framework for Sparse-View CBCT Reconstruction

Minmin Yang, Huantao Ren, Senem Velipasalar

机构 * Department of Electrical Engineering and Computer Science at Syracuse University(苏利文大学电气工程与计算机科学系)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18683 2025-06-24 cs.CV cs.AI 57%

SIM-Net: A Multimodal Fusion Network Using Inferred 3D Object Shape Point Clouds from RGB Images for 2D Classification

Youcef Sklab, Hanane Ariouat, Eric Chenin, Edi Prifti, Jean-Daniel Zucker

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 25 pages, 9 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13318 2025-06-24 cs.CV cs.LG eess.IV 57%

VesselGPT: Autoregressive Modeling of Vascular Geometry

Paula Feldman, Martin Sinnona, Claudio Delrieux, Viviana Siless, Emmanuel Iarussi

机构 * Consejo Nacional de Investigaciones Científicas y Técnicas(阿根廷国家科学与技术研究院) Universidad Nacional del Sur(国立南方大学) Universidad Torcuato Di Tella(托克托迪塔拉大学)

专题命中 视觉空间推理 :planning(abstract);分类 cs.LG

Comments Accepted for MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16623 2025-06-23 cs.RO cs.AI 57%

History-Augmented Vision-Language Models for Frontier-Based Zero-Shot Object Navigation

Mobin Habibpour, Fatemeh Afghah

机构 * Holcombe Department of Electrical and Computer Engineering, Clemson University, Clemson, SC, USA(霍尔科姆电气与计算机工程系,克莱姆森大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14782 2025-06-23 cs.LG q-bio.QM 57%

Integrating Dynamical Systems Learning with Foundational Models: A Meta-Evolutionary AI Framework for Clinical Trials

Joseph Geraci, Bessi Qorri, Christian Cumbaa, Mike Tsay, Paul Leonczyk, Luca Pani

机构 * NetraMark Corp.(NetraMark公司) Queen’s University(女王大学) Tandem Centre for Pharmacogenetics(药基因学中心) UC San Diego(圣地亚哥大学) University of Miami(迈阿密大学) University of Modena and Reggio Emilia(摩德纳和雷焦艾米利亚大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

Comments 27 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20499 2025-06-19 cs.LG 57%

Data Distributional Properties As Inductive Bias for Systematic Generalization

Felipe del Rio, Alain Raymond-Saez, Daniel Florea, Rodrigo Toro Icarte, Julio Hurtado, Cristian B. Calderon, Alvaro Soto

机构 * Pontificia Universidad Católica de Chile(天主教智利大学) University of Warwick(沃里克大学) CENIA

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.21015 2025-06-16 cs.CL cs.HC 57%

MapQaTor: An Extensible Framework for Efficient Annotation of Map-Based QA Datasets

Mahir Labib Dihan, Mohammed Eunus Ali, Md Rizwan Parvez

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Bangladesh University of Engineering and Technology (BUET)(孟加拉工程与技术大学) Faculty of Information Technology(信息技术学院) Monash University(墨尔本大学) Qatar Computing Research Institute (QCRI)(卡塔尔计算研究中心)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL

Comments ACL 2025 (Demo)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14254 2025-06-12 cs.RO cs.AI 57%

Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation

Lingfeng Zhang, Yuecheng Liu, Zhanguang Zhang, Matin Aghaei, Yaochen Hu, Hongjian Gu, Mohammad Ali Alomrani, David Gamaliel Arcos Bravo, Raika Karimi, Atia Hamidizadeh, Haoping Xu, Guowei Huang, Zhanpeng Zhang, Tongtong Cao, Weichao Qiu, Xingyue Quan, Jianye Hao, Yuzheng Zhuang, Yingxue Zhang

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08189 2025-06-11 cs.CV cs.CL 57%

Open World Scene Graph Generation using Vision Language Models

Amartya Dutta, Kazi Sajeed Mehrab, Medha Sawhney, Abhilash Neog, Mridul Khurana, Sepideh Fatemi, Aanish Pradhan, M. Maruf, Ismini Lourentzou, Arka Daw, Anuj Karpatne

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL

Comments Accepted in CVPR 2025 Workshop (CVinW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.03969 2025-06-05 cs.AI 57%

Craftium: Bridging Flexibility and Efficiency for Rich 3D Single- and Multi-Agent Environments

Mikel Malagón, Josu Ceberio, Jose A. Lozano

机构 * Department of Computer Science and Artificial Intelligence, University of the Basque Country UPV/EHU(计算机科学与人工智能系,巴斯克国家大学UPV/EHU) Basque Center for Applied Mathematics (BCAM)(巴斯克应用数学中心)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments ICML 2025. Project's website: https://github.com/mikelma/craftium/

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00927 2025-06-03 cs.SD cs.AI eess.AS 57%

In-the-wild Audio Spatialization with Flexible Text-guided Localization

Tianrui Pan, Jie Liu, Zewen Huang, Jie Tang, Gangshan Wu

机构 * State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments Accepted by ACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07089 2025-06-03 cs.CV cs.CL 57%

OmniCaptioner: One Captioner to Rule Them All

Yiting Lu, Jiakang Yuan, Zhen Li, Shitian Zhao, Qi Qin, Xinyue Li, Le Zhuo, Licheng Wen, Dongyang Liu, Yuewen Cao, Xiangchao Yan, Xin Li, Tianshuo Peng, Shufei Zhang, Botian Shi, Tao Chen, Zhibo Chen, Lei Bai, Peng Gao, Bo Zhang

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL

Comments More visualizations on Homepage: https://alpha-innovator.github.io/OmniCaptioner-project-page and Official code: https://github.com/Alpha-Innovator/OmniCaptioner

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12821 2025-05-30 cs.CV cs.AI 57%

From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration

Mingyang Song, Xiaoye Qu, Jiawei Zhou, Yu Cheng

机构 * Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Stony Brook University(石溪大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments Accepted by CVPR 2025. Project Page: https://vlmlt.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20718 2025-05-29 cs.CV cs.AI 57%

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models

Kui Wu, Shuhang Xu, Hao Chen, Churan Wang, Zhoujun Li, Yizhou Wang, Fangwei Zhong

机构 * State Key Laboratory of Complex & Critical Software Environment, Beihang University(北京航空航天大学复杂与关键软件环境国家重点实验室) School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院) City University of Macau(澳门城市大学) Center on Frontiers of Computing Studies, School of Computer Science, Nat’l Eng. Research Center of Visual Technology, Peking University(北京大学计算机学院前沿计算研究中心、国家工程视觉技术研究中心)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09546 2025-05-26 cs.RO cs.AI 57%

How Secure Are Large Language Models (LLMs) for Navigation in Urban Environments?

Congcong Wen, Jiazhao Liang, Shuaihang Yuan, Hao Huang, Geeta Chandra Raju Bethala, Yu-Shen Liu, Mengyu Wang, Anthony Tzes, Yi Fang

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13466 2025-05-21 cs.AI 57%

AgentSGEN: Multi-Agent LLM in the Loop for Semantic Collaboration and GENeration of Synthetic Data

Vu Dinh Xuan, Hao Vo, David Murphy, Hoang D. Nguyen

机构 * University of Information Technology, VNU–HCM(越南胡志明市信息技术大学) University College Cork(科尔克大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03368 2025-05-13 cs.LG 57%

Geospatial Mechanistic Interpretability of Large Language Models

Stef De Sabbata, Stefano Mizzaro, Kevin Roitero

机构 * University of Leicester, UK(利兹大学) University of Udine, Italy(乌迪内大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

Comments Figures 2 and 3: fixed issue with min boundary in colorbar

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05485 2025-05-13 cs.RO cs.AI cs.CV 57%

HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation

Yi Li, Yuquan Deng, Jesse Zhang, Joel Jang, Marius Memmel, Raymond Yu, Caelan Reed Garrett, Fabio Ramos, Dieter Fox, Anqi Li, Abhishek Gupta, Ankit Goyal

机构 * NVIDIA University of Washington(华盛顿大学) University of Southern California(南加州大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments update related work and results on VQA benchmarks

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04718 2025-05-09 cs.CV cs.LG 57%

Lay-Your-Scene: Natural Scene Layout Generation with Diffusion Transformers

Divyansh Srivastava, Xiang Zhang, He Wen, Chenru Wen, Zhuowen Tu

机构 * UC San Diego(加州大学圣地亚哥分校) Tsingua University(清华大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07113 2025-05-07 cs.CV cs.AI 57%

Beyond Bare Queries: Open-Vocabulary Object Grounding with 3D Scene Graph

Sergey Linok, Tatiana Zemskova, Svetlana Ladanova, Roman Titkov, Dmitry Yudin, Maxim Monastyrny, Aleksei Valenkov

机构 * Center for Cognitive Modeling, Moscow Institute of Physics and Technology(认知建模中心,莫斯科物理技术学院) AIRI Sberbank of Russia, Robotics Center(俄罗斯储蓄银行机器人中心)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 6 pages, 6 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.07937 2025-05-07 cs.RO cs.AI 57%

Co-NavGPT: Multi-Robot Cooperative Visual Semantic Navigation Using Vision Language Models

Bangguo Yu, Qihao Yuan, Kailai Li, Hamidreza Kasaei, Ming Cao

机构 * Faculty of Science and Engineering, University of Groningen(工程学院,格罗宁根大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02829 2025-05-06 cs.AI 57%

LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery

Jerome Quenum, Wen-Han Hsieh, Tsung-Han Wu, Ritwik Gupta, Trevor Darrell, David M. Chan

机构 * Department of Electrical Engineering and Computer Sciences, University of California-Berkeley, Berkeley, CA, USA(电气工程与计算机科学系,加州大学伯克利分校)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 28 pages, 10 figures, 19 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11564 2025-05-01 cs.RO cs.CL cs.CV cs.HC 57%

Learning 6-DoF Fine-grained Grasp Detection Based on Part Affordance Grounding

Yaoxian Song, Penglei Sun, Piaopiao Jin, Yi Ren, Yu Zheng, Zhixu Li, Xiaowen Chu, Yue Zhang, Tiefeng Li, Jason Gu

机构 * Center for X-Mechanics, School of Aeronautics and Astronautics, Zhejiang University(浙江大学航空航天学院X力学中心) Information Hub, Data Science and Analytics Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)信息中心、数据科学与分析方向) Advanced Manufacturing Lab, Huawei Technologies(华为技术先进制造实验室) Research Institute, UBTECH Robotics Inc.(UBTECH机器人研究院) School of Information and School of Smart Governance, Renmin University of China(中国人民大学信息学院和智能治理学院) School of Engineering, Westlake University and the Institute of Advanced Technology, Westlake Institute for Advanced Study(西湖大学工程学院和西湖先进科技研究院) Department of Electrical and Computer Engineering, Dalhousie University(达尔豪斯大学电气与计算机工程系)

专题命中 视觉空间推理 :planning(abstract);分类 cs.CL

Comments 15 pages, 8 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11419 2025-04-29 cs.AI cs.NE 57%

Embodied World Models Emerge from Navigational Task in Open-Ended Environments

Li Jin, Liu Jia

机构 * Tsinghua Laboratory of Brain and Intelligence(清华大学脑科学与智能实验室)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments Research on explainable meta-reinforcement learning AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.12533 2025-04-25 cs.RO cs.AI 57%

To Help or Not to Help: LLM-based Attentive Support for Human-Robot Group Interactions

Daniel Tanneberg, Felix Ocker, Stephan Hasler, Joerg Deigmoeller, Anna Belardinelli, Chao Wang, Heiko Wersing, Bernhard Sendhoff, Michael Gienger

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16791 2025-04-22 cs.HC cs.AI 57%

"The Diagram is like Guardrails": Structuring GenAI-assisted Hypotheses Exploration with an Interactive Shared Representation

Zijian Ding, Michelle Brachman, Joel Chan, Werner Geyer

机构 * College of Information, University of Maryland USA(信息学院,马里兰大学美国分校) IBM Research USA(IBM美国研究)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13942 2025-04-22 cs.HC cs.AI cs.ET 57%

Intelligence of Things: A Spatial Context-Aware Control System for Smart Devices

Sukanth Kalivarathan, Muhmmad Abrar Raja Mohamed, Aswathy Ravikumar, S Harini

机构 * School of Computer Science and Engineering, Vellore Institute of Technology(计算机科学与工程学院,维杰里学院技术学院) Sustainable Living Labs(可持续生活实验室)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 16 pages, 8 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12817 2025-04-18 cs.RO cs.AI cs.CV 57%

Explainable Scene Understanding with Qualitative Representations and Graph Neural Networks

Nassim Belmecheri, Arnaud Gotlieb, Nadjib Lazaar, Helge Spieker

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments Workshop "Advancing Automated Driving in Highly Interactive Scenarios through Behavior Prediction, Trustworthy AI, and Remote Operations" @ 36th IEEE Intelligent Vehicles Symposium (IV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17385 2025-04-18 cs.CL cs.CV 57%

Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities

Zheyuan Zhang, Fengyuan Hu, Jayjun Lee, Freda Shi, Parisa Kordjamshidi, Joyce Chai, Ziqiao Ma

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL

Comments Accepted to ICLR 2025 (Oral) | Project page: https://spatial-comfort.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏