arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 3363 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 3363 篇

2510.17157 2025-10-21 cs.CV cs.AI 57%

GACO-CAD: Geometry-Augmented and Conciseness-Optimized CAD Model Generation from Single Image

Yinghui Wang, Xinyu Zhang, Peng Du

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19131 2025-10-21 cs.RO cs.AI cs.CV 57%

ZeST: an LLM-based Zero-Shot Traversability Navigation for Unknown Environments

Shreya Gummadi, Mateus V. Gasparino, Gianluca Capezzuto, Marcelo Becker, Girish Chowdhary

机构 * Field Robotics Engineering and Science Hub (FRESH), Illinois Autonomous Farm, University of Illinois at Urbana-Champaign (UIUC), IL(伊利诺伊大学厄巴纳-香槟分校) Mobile Robotics Group, São Carlos School of Engineering, University of São Paulo (EESC-USP), São Carlos, SP, Brazil(圣保罗大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17092 2025-10-20 cs.CV cs.CL 57%

Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI

Syed Abdul Gaffar Shakhadri, Kruthika KR, Kartik Basavaraj Angadi

机构 * SandLogic Technologies Pvt Ltd(沙德逻辑技术 Pvt Ltd)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08589 2025-10-13 cs.CV cs.AI 57%

Beyond CNNs: Efficient Fine-Tuning of Multi-Modal LLMs for Object Detection on Low-Data Regimes

Nirmal Elamon, Rouzbeh Davoudi

机构 * Artificial Creative intelligence (ACI)(人工创意智能(ACI)) Expedia Group(Expedia集团)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01795 2025-10-13 cs.RO cs.AI 57%

Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving

Haibo Hu, Lianming Huang, Xinyu Wang, Yufei Cui, Shangyu Wu, Nan Guan, Chun Jason Xue

机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) Department of Computer Science, McGill University(麦吉尔大学计算机科学系) Department of Computer Science, Mohamed bin Zayed University of Artificial Intelligence(马尔代夫穆罕默德· bin·扎耶德人工智能大学计算机科学系)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08011 2025-10-10 cs.CV cs.CL 57%

Play to Generalize: Learning to Reason Through Game Play

Yunfei Xie, Yinsong Ma, Shiyi Lan, Alan Yuille, Junfei Xiao, Chen Wei

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL

Comments Project Page: https://yunfeixie233.github.io/ViGaL/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07988 2025-10-10 cs.AI 57%

ReInAgent: A Context-Aware GUI Agent Enabling Human-in-the-Loop Mobile Task Navigation

Haitao Jia, Ming He, Zimo Yin, Likang Wu, Jianping Fan, Jitao Sang

专题命中 视觉空间推理 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07825 2025-10-10 cs.AI 57%

An LLM-Powered Cooperative Framework for Large-Scale Multi-Vehicle Navigation

Yuping Zhou, Siqi Lai, Jindong Han, Hao Liu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03408 2025-10-08 cs.CL cs.CV 57%

Trajectory Prediction Meets Large Language Models: A Survey

Yi Xu, Ruining Yang, Yitian Zhang, Jianglin Lu, Mingyuan Zhang, Yizhou Wang, Lili Su, Yun Fu

机构 * Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL

Comments 16 pages, GitHub: https://github.com/colorfulfuture/Awesome-Trajectory-Motion-Prediction-Papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.05414 2025-10-08 cs.CL 57%

A Lightweight Large Language Model-Based Multi-Agent System for 2D Frame Structural Analysis

Ziheng Geng, Jiachen Liu, Ran Cao, Lu Cheng, Haifeng Wang, Minghui Cheng

机构 * Department of Civil and Architectural Engineering, University of Miami(迈阿密大学土木与建筑工程系) Department of Electrical and Computer Engineering, University of Miami(迈阿密大学电气与计算机工程系) College of Civil Engineering, Hunan University(湖南大学土木学院) Department of Computer Science, University of Illinois Chicago(伊利诺伊大学芝加哥分校计算机科学系) Department of Civil and Environmental Engineering, Washington State University(华盛顿州立大学土木与环境工程系) School of Architecture, University of Miami(迈阿密大学建筑学院)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02634 2025-10-06 cs.SE cs.AI 57%

Automatic Building Code Review: A Case Study

Hanlong Wan, Weili Xu, Michael Rosenberg, Jian Zhang, Aysha Siddika

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01634 2025-10-03 cs.LG 57%

CAT: Curvature-Adaptive Transformers for Geometry-Aware Learning

Ryan Y. Lin, Siddhartha Ojha, Nicholas Bai

机构 * Division of Engineering and Applied Science(工程与应用科学系) California Institute of Technology(加州理工学院)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05288 2025-10-03 cs.CV cs.AI cs.RO 57%

PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes

Ahmed Abdelreheem, Filippo Aleotti, Jamie Watson, Zawar Qureshi, Abdelrahman Eldesokey, Peter Wonka, Gabriel Brostow, Sara Vicente, Guillermo Garcia-Hernando

机构 * Niantic Spatial KAUST(科威特科学与技术研究中心) UCL(伦敦大学学院)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments ICCV 2025. Project page: https://nianticlabs.github.io/placeit3d/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01141 2025-10-02 cs.AI 57%

Apriel-1.5-15b-Thinker

Shruthan Radhakrishna, Aman Tiwari, Aanjaneya Shukla, Masoud Hashemi, Rishabh Maheshwary, Shiva Krishna Reddy Malay, Jash Mehta, Pulkit Pattnaik, Saloni Mittal, Khalil Slimi, Kelechi Ogueji, Akintunde Oladipo, Soham Parikh, Oluwanifemi Bamgbose, Toby Liang, Ahmed Masry, Khyati Mahajan, Sai Rajeswar Mudumba, Vikas Yadav, Sathwik Tejaswi Madhusudhan, Torsten Scholak, Sagar Davasam, Srinivas Sunkara, Nicholas Chapados

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26161 2025-10-01 cs.AI cs.SE 57%

90% Faster, 100% Code-Free: MLLM-Driven Zero-Code 3D Game Development

Runxin Yang, Yuxuan Wan, Shuqing Li, Michael R. Lyu

机构 * The Chinese University of Hong Kong(香港中文大学)

专题命中 视觉空间推理 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25655 2025-10-01 cs.AI 57%

Landmark-Guided Knowledge for Vision-and-Language Navigation

Dongsheng Yang, Meiling Zhu, Yinfeng Yu

机构 * School of Computer Science and Technology, Xinjiang University, Urumqi, China(计算机科学与技术学院,新疆大学,乌鲁木齐,中国) No. 59 Middle School of Urumqi, Urumqi, China(乌鲁木齐第五十九中学,乌鲁木齐,中国)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments Accepted for publication by International Conference on Intelligent Computing 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22698 2025-09-30 cs.RO cs.AI 57%

Advancing Audio-Visual Navigation Through Multi-Agent Collaboration in 3D Environments

Hailong Zhang, Yinfeng Yu, Liejun Wang, Fuchun Sun, Wendong Zheng

机构 * Xinjiang Multimodal Intelligent Processing and Information Security Engineering Technology Research Center(新疆多模态智能处理与信息安全工程技术研究中心) School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) School of Electrical Engineering and Automation, Tianjin University of Technology(天津理工大学电气工程与自动化学院)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments Main paper (15 pages). Accepted for publication by ICONIP( International Conference on Neural Information Processing) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22404 2025-09-29 cs.CV cs.AI 57%

RAU: Reference-based Anatomical Understanding with Vision Language Models

Yiwei Li, Yikang Liu, Jiaqi Guo, Lin Zhao, Zheyuan Zhang, Xiao Chen, Boris Mailhe, Ankush Mukherjee, Terrence Chen, Shanhui Sun

机构 * United Imaging Intelligence(联合影像智能) School of Computing, University of Georgia(佐治亚大学计算机学院) Department of Electrical and Computer Engineering, Northwestern University(西北大学电气与计算机工程系)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21189 2025-09-26 cs.RO cs.AI cs.CV 57%

Human-like Navigation in a World Built for Humans

Bhargav Chandaka, Gloria X. Wang, Haozhe Chen, Henry Che, Albert J. Zhai, Shenlong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments CoRL 2025. Project website: https://reasonnav.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17664 2025-09-23 cs.CV cs.AI 57%

SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models

Pingyi Chen, Yujing Lou, Shen Cao, Jinhui Guo, Lubin Fan, Yue Wu, Lin Yang, Lizhuang Ma, Jieping Ye

机构 * Zhejiang University(浙江大学) Westlake University(西湖大学) Alibaba Cloud Computing(阿里云 computing) Shanghai Jiao Tong University(上海交通大学)

专题命中 视觉空间推理 :chain-of-thought(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17136 2025-09-23 cs.CV cs.AI 57%

SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM

Yuhao Tian, Zheming Yang

机构 * Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Institute of AI for Industries(工业人工智能研究所)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 5 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04653 2025-09-23 cs.CV cs.CL 57%

LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts

Yimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof Czarnecki

机构 * University of Waterloo(滑铁卢大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL

Comments To appear at EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16330 2025-09-23 cs.AI 57%

Generalizability of Large Language Model-Based Agents: A Comprehensive Survey

Minxing Zhang, Yi Yang, Roy Xie, Bhuwan Dhingra, Shuyan Zhou, Jian Pei

机构 * Duke University(杜克大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16087 2025-09-22 cs.CV cs.AI 57%

See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model

Pengteng Li, Pinhao Song, Wuyang Li, Weiyu Guo, Huizai Yao, Yijie Xu, Dugang Liu, Hui Xiong

机构 * HKUST(GZ)(香港科技大学(广州)) KU Leuven(比利时鲁文大学) EPFL(苏黎世联邦理工学院) SZU(深圳大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03642 2025-09-22 cs.CV cs.AI 57%

Spatial Understanding from Videos: Structured Prompts Meet Simulation Data

Haoyu Zhang, Meng Liu, Zaijing Li, Haokun Wen, Weili Guan, Yaowei Wang, Liqiang Nie

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Pengcheng Laboratory(鹏城实验室) Shandong Jianzhu University(山东建筑大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments Accepted by NeurIPS 2025 as a Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14966 2025-09-19 cs.CV cs.AI cs.RO 57%

RoboEye: Enhancing 2D Robotic Object Identification with Selective 3D Geometric Keypoint Matching

Xingwu Zhang, Guanxuan Li, Zhuocheng Zhang, Zijun Long

机构 * Hunan University(湖南大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13372 2025-09-18 eess.IV cs.AI cs.CV cs.ET q-bio.QM 57%

Generative AI Pipeline for Interactive Prompt-driven 2D-to-3D Vascular Reconstruction for Fontan Geometries from Contrast-Enhanced X-Ray Fluoroscopy Imaging

Prahlad G Menon

机构 * University of Pittsburgh(匹兹堡大学) The American Association for Thoracic Surgery (AATS)(美国胸外科医师协会)

专题命中 视觉空间推理 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00419 2025-09-15 physics.geo-ph cs.AI 57%

Geological Everything Model 3D: A Promptable Foundation Model for Unified and Zero-Shot Subsurface Understanding

Yimin Dou, Xinming Wu, Nathan L Bangs, Harpreet Singh Sethi, Jintao Li, Hang Gao, Zhixiang Guo

机构 * University of Science and Technology of China, School of Earth and Space Sciences(中国科学技术大学地球和空间科学学院) University of Texas at Austin, UT Institute for Geophysics(德克萨斯大学奥斯汀分校UT地球物理研究所) NVIDIA(NVIDIA公司)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17813 2025-09-12 cs.RO cs.LG 57%

Safe Multi-Agent Navigation guided by Goal-Conditioned Safe Reinforcement Learning

Meng Feng, Viraj Parimi, Brian Williams

机构 * Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology(计算机科学与人工智能实验室,麻省理工学院)

专题命中 视觉空间推理 :planning(abstract);分类 cs.LG

Comments Due to the limitation "The abstract field cannot be longer than 1,920 characters", the abstract here is shorter than that in the PDF file

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09154 2025-09-12 cs.AI cs.CV 57%

Mind Meets Space: Rethinking Agentic Spatial Intelligence from a Neuroscience-inspired Perspective

Bui Duc Manh, Soumyaratna Debnath, Zetong Zhang, Shriram Damodaran, Arvind Kumar, Yueyi Zhang, Lu Mi, Erik Cambria, Lin Wang

机构 * School of EEE, Nanyang Technological University(南洋理工大学电子工程系) Department of Civil Engineering, Tsinghua University(清华大学土木工程系) Dr. B. R. Ambedkar National Institute of Technology, Jalandhar(B.R.阿姆贝卡尔国家理工学院,贾兰德赫) KTH Royal Institute of Technology(皇家理工学院) College of AI, Tsinghua University(清华大学人工智能学院) CCDS, Nanyang Technological University(南洋理工大学CCDS)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 54 pages, journal

详情

展开后加载摘要…

URL PDF HTML 收藏