arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 3363 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 3363 篇

2505.03581 2025-05-07 cs.CV 50%

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes

Sergey Linok, Vadim Semenov, Anastasia Trunova, Oleg Bulichev, Dmitry Yudin

机构 * Center for Cognitive Modeling(认知建模中心) Moscow Institute of Physics and Technology(莫斯科物理技术学院) Innopolis University(Innopolis大学) AIRI

专题命中 视觉空间推理 :reasoning(abstract)

Comments 8 pages, 5 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02836 2025-05-06 cs.CV 50%

Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation

Lu Ling, Chen-Hsuan Lin, Tsung-Yi Lin, Yifan Ding, Yu Zeng, Yichen Sheng, Yunhao Ge, Ming-Yu Liu, Aniket Bera, Zhaoshuo Li

机构 * NVIDIA Research(NVIDIA研究)

专题命中 视觉空间推理 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20091 2025-05-01 cs.CV cs.MA 50%

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

Noriyuki Kugo, Xiang Li, Zixin Li, Ashish Gupta, Arpandeep Khatua, Nidhish Jain, Chaitanya Patel, Yuta Kyuragi, Yasunori Ishii, Masamoto Tanabiki, Kazuki Kozuka, Ehsan Adeli

机构 * Panasonic Connect Co., Ltd.(松下电器(中国)有限公司) Stanford University(斯坦福大学) Panasonic R&D Company of America(松下美国研发公司) Panasonic Holdings Corporation(松下控股公司)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00114 2025-04-30 cs.RO cs.CV 50%

Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach

Aaron Hao Tan, Angus Fung, Haitong Wang, Goldie Nejat

专题命中 视觉空间推理 :planning(abstract)

Comments 8 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.18738 2025-04-29 cs.CV 50%

A Review of 3D Object Detection with Vision-Language Models

Ranjan Sapkota, Konstantinos I Roumeliotis, Rahul Harsha Cheppally, Marco Flores Calero, Manoj Karkee

机构 * Cornell University(康奈尔大学) University of Peloponnese(希腊皮洛斯大学) Kansas State University(堪萨斯州立大学) Universidad de las Fuerzas Armadas(武装力量大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.01540 2025-04-29 cs.CV 50%

DEVICE: Depth and Visual Concepts Aware Transformer for OCR-based Image Captioning

Dongsheng Xu, Qingbao Huang, Xingmao Zhang, Haonan Cheng, Feng Shuang, Yi Cai

机构 * School of Artificial Intelligence at Guangxi University(广西大学人工智能学院) College of General Education, Guangxi Arts University(广西艺术学院文学院) The State Key Laboratory of Media Convergence and Communication, Communication University of China(中国传媒大学媒体融合与传播国家重点实验室) Guangxi Key Laboratory of Intelligent Control and Maintenance of Power Equipment(广西智能控制与电力设备关键实验室) School of Software Engineering, South China University of Technology(华南理工大学软件学院)

专题命中 视觉空间推理 :reasoning(abstract)

Comments 13pages, 8figures. This work has been published in Pattern Recognition

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11748 2025-04-28 cs.CV 50%

Understanding Depth and Height Perception in Large Visual-Language Models

Shehreen Azad, Yash Jain, Rishit Garg, Yogesh S Rawat, Vibhav Vineet

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学) Microsoft Research(微软研究院) Indian Institute of Technology, Kharagpur(印度克里希纳布尔理工学院)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Accepted in CVPRW 2025. Project page: https://sacrcv.github.io/GeoMeter-website/

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10263 2025-04-28 cs.CV cs.MA 50%

PreGSU-A Generalized Traffic Scene Understanding Model for Autonomous Driving based on Pre-trained Graph Attention Network

Yuning Wang, Zhiyuan Liu, Haotian Lin, Junkai Jiang, Shaobing Xu, Jianqiang Wang

专题命中 视觉空间推理 :reasoning(abstract)

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17748 2025-04-25 cs.RO 50%

Robotic Task Ambiguity Resolution via Natural Language Interaction

Eugenio Chisari, Jan Ole von Hartz, Fabien Despinoy, Abhinav Valada

机构 * Department of Computer Science, University of Freiburg(弗赖堡大学计算机科学系) Toyota Motor Europe(丰田欧洲公司)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15305 2025-04-24 cs.RO cs.CV cs.SY eess.SY 50%

SLAM-Based Navigation and Fault Resilience in a Surveillance Quadcopter with Embedded Vision Systems

Abhishek Tyagi, Charu Gaur

专题命中 视觉空间推理 :planning(abstract)

Comments 18 pages, 21 figures, 15 tables. Onboard processing using Raspberry Pi 4 and Arduino Nano. Includes ORB-SLAM3-based navigation, LQR control, rotor fault recovery, object detection, and PCA face recognition. Real-world and simulation tests included. Designed for GPS-denied autonomous UAV surveillance

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14680 2025-04-24 cs.RO 50%

A Complete and Bounded-Suboptimal Algorithm for a Moving Target Traveling Salesman Problem with Obstacles in 3D

Anoop Bhat, Geordan Gutow, Bhaskar Vundurthy, Zhongqiang Ren, Sivakumar Rathinam, Howie Choset

机构 * Robotics Institute at Carnegie Mellon University(卡内基梅隆大学机器人研究所) UM-SJTU Joint Institute and Department of Automation at Shanghai Jiao Tong University(上海交通大学联合研究院和自动化系) Department of Mechanical Engineering and Department of Computer Science and Engineering at Texas A&M University(德克萨斯大学机械工程系和计算机科学与工程系)

专题命中 视觉空间推理 :planning(abstract)

Comments Accepted to ICRA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05938 2025-04-24 cs.RO 50%

Energy-Efficient Autonomous Aerial Navigation with Dynamic Vision Sensors: A Physics-Guided Neuromorphic Approach

Sourav Sanyal, Amogh Joshi, Manish Nagaraj, Rohan Kumar Manna, Kaushik Roy

机构 * School of Electrical and Computer Engineering, Purdue University(电气与计算机工程学院,普渡大学)

专题命中 视觉空间推理 :planning(abstract)

Comments This work has been accepted for presentation at the 2025 IEEE International Joint Conference on Neural Networks (IJCNN), June 30 - July 5, 2025, Rome, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15600 2025-04-23 cs.RO cs.SY eess.SY 50%

Research on Navigation Methods Based on LLMs

Anlong Zhang, Jianmin Ji

机构 * Institute of Advanced Technology University of Science and Technology of China(科学技术大学先进技术学院) School of Computer Science and Technology University of Science and Technology of China(科学技术大学计算机科学与技术学院)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15309 2025-04-23 cs.CV 50%

LLM-Enabled Style and Content Regularization for Personalized Text-to-Image Generation

Anran Yu, Wei Feng, Yaochen Zhang, Xiang Li, Lei Meng, Lei Wu, Xiangxu Meng

机构 * School of Software, Shandong University(软件学院,山东大学) Shandong Research Institute of Shandong University(山东大学山东研究院) Inspur Technology(Inspur技术公司)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13690 2025-04-22 cs.CV 50%

Analysing the Robustness of Vision-Language-Models to Common Corruptions

Muhammad Usama, Syeda Aishah Asim, Syed Bilal Ali, Syed Talal Wasim, Umair Bin Mansoor

机构 * Department of Electrical Engineering, DHA Suffa University Karachi(电气工程系,达哈萨夫大学卡拉奇) Department of Computer Science, University of Bonn(计算机科学系,波恩大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments arXiv admin note: text overlap with arXiv:2304.10592, arXiv:2301.12597 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.10727 2025-04-16 cs.CV 50%

Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization

Darryl Hannan, John Cooper, Dylan White, Timothy Doster, Henry Kvinge, Yijing Watkins

专题命中 视觉空间推理 :reasoning(abstract)

Comments 26 pages, CVPR MORSE Workshop 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03133 2025-04-10 cs.CV 50%

Joint Retrieval of Cloud properties using Attention-based Deep Learning Models

Zahid Hassan Tushar, Adeleke Ademakinwa, Jianwu Wang, Zhibo Zhang, Sanjay Purushotham

专题命中 视觉空间推理 :CoT(abstract)

Comments 6 Pages, 4 figures, to be published in 2025 IEEE International Geoscience and Remote Sensing Symposium (IGARSS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15185 2025-04-10 cs.RO 50%

Semantically Safe Robot Manipulation: From Semantic Scene Understanding to Motion Safeguards

Lukas Brunke, Yanni Zhang, Ralf Römer, Jack Naimer, Nikola Staykov, Siqi Zhou, Angela P. Schoellig

专题命中 视觉空间推理 :reasoning(abstract)

Comments 9 pages, 6 figures

Journal ref in IEEE Robotics and Automation Letters, vol. 10, no. 5, pp. 4810-4817, May 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05966 2025-04-09 eess.IV cs.CV 50%

AVP-AP: Self-supervised Automatic View Positioning in 3D cardiac CT via Atlas Prompting

Xiaolin Fan, Yan Wang, Yingying Zhang, Mingkun Bao, Bosen Jia, Dong Lu, Yifan Gu, Jian Cheng, Haogang Zhu

专题命中 视觉空间推理 :planning(abstract)

Comments 12 pages, 8 figures, published to TMI

Journal ref IEEE TRANSACTIONS ON MEDICAL IMAGING, March 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05463 2025-04-09 cs.CV 50%

REVEAL: Relation-based Video Representation Learning for Video-Question-Answering

Sofian Chaybouti, Walid Bousselham, Moritz Wolter, Hilde Kuehne

专题命中 视觉空间推理 :reasoning(abstract)

Comments 18 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05445 2025-04-09 cs.HC 50%

Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly

Lianghan Dong, Anamaria Crisan

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19889 2025-04-08 cond-mat.mtrl-sci cs.RO 50%

A Multi-Agent Framework Integrating Large Language Models and Generative AI for Accelerated Metamaterial Design

Jie Tian, Martin Taylor Sobczak, Dhanush Patil, Jixin Hou, Lin Pang, Arunachalam Ramanathan, Libin Yang, Xianyan Chen, Yuval Golan, Xiaoming Zhai, Hongyue Sun, Kenan Song, Xianqiao Wang

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15190 2025-04-08 cs.CV 50%

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues

Sagar Soni, Akshay Dudhane, Hiyam Debary, Mustansar Fiaz, Muhammad Akhtar Munir, Muhammad Sohail Danish, Paolo Fraccaro, Campbell D Watson, Levente J Klein, Fahad Shahbaz Khan, Salman Khan

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15717 2025-04-08 cs.RO cs.SY eess.SY 50%

Autonomous Wheel Loader Navigation Using Goal-Conditioned Actor-Critic MPC

Aleksi Mäki-Penttilä, Naeim Ebrahimi Toulkani, Reza Ghabcheloo

专题命中 视觉空间推理 :planning(abstract)

Comments Accepted to International Conference on Robotics and Automation (ICRA) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17703 2025-04-07 cs.RO 50%

RAIDER: Tool-Equipped Large Language Model Agent for Robotic Action Issue Detection, Explanation and Recovery

Silvia Izquierdo-Badiola, Carlos Rizzo, Guillem Alenyà

专题命中 视觉空间推理 :self-correction(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02197 2025-04-04 cs.ET cs.HC 50%

Design and Implementation of the Transparent, Interpretable, and Multimodal (TIM) AR Personal Assistant

Erin McGowan, Joao Rulff, Sonia Castelo, Guande Wu, Shaoyu Chen, Roque Lopez, Bea Steers, Iran R. Roman, Fabio F. Dias, Jing Qian, Parikshit Solunke, Michael Middleton, Ryan McKendrick, Claudio T. Silva

专题命中 视觉空间推理 :reasoning(abstract)

Comments Copyright 2025 IEEE. All rights reserved, including rights for text and data mining and training of artificial intelligence and similar technologies. Personal use is permitted, but republication/redistribution requires IEEE permission. Article accepted for publication in IEEE Computer Graphics and Applications. This is the author's version, content may change prior to final publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11402 2025-04-03 cs.RO 50%

M2Diffuser: Diffusion-based Trajectory Optimization for Mobile Manipulation in 3D Scenes

Sixu Yan, Zeyu Zhang, Muzhi Han, Zaijin Wang, Qi Xie, Zhitian Li, Zhehan Li, Hangxin Liu, Xinggang Wang, Song-Chun Zhu

专题命中 视觉空间推理 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03735 2025-04-02 cs.CV 50%

VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding

Chaoyu Li, Eun Woo Im, Pooyan Fazli

专题命中 视觉空间推理 :reasoning(abstract)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07135 2025-03-31 cs.RO cs.CV 50%

VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

Hanzhi Chen, Boyang Sun, Anran Zhang, Marc Pollefeys, Stefan Leutenegger

专题命中 视觉空间推理 :planning(abstract)

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20682 2025-03-27 cs.CV 50%

GLRD: Global-Local Collaborative Reason and Debate with PSL for 3D Open-Vocabulary Detection

Xingyu Peng, Si Liu, Chen Gao, Yan Bai, Beipeng Mu, Xiaofei Wang, Huaxia Xia

专题命中 视觉空间推理 :reasoning(abstract)

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏