arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 3363 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 3363 篇

2506.09469 2025-06-12 cs.CV 50%

Optimizing Cooperative Multi-Object Tracking using Graph Signal Processing

Maria Damanaki, Nikos Piperigkos, Alexandros Gkillas, Aris S. Lalos

机构 * Industrial Systems Institute, Athena Research Center, Patras Science Park, Greece(工业系统研究所,阿提卡研究中心,帕特拉斯科学公园,希腊) AviSense.AI, Patras Science Park, Greece(AviSense.AI,帕特拉斯科学公园,希腊) Dpt. of Informatics & Telecom., University of Ioannina, Arta, Greece(信息与电信系,伊奥安妮亚大学,阿塔,希腊)

专题命中 视觉空间推理 :planning(abstract)

Comments 2025 IEEE International Conference on Multimedia and Expo Workshops, 3DMM - 3D Multimedia Analytics, Search and Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08553 2025-06-11 cs.CV 50%

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge

Agnese Taluzzi, Davide Gesualdi, Riccardo Santambrogio, Chiara Plizzari, Francesca Palermo, Simone Mentasti, Matteo Matteucci

机构 * Politecnico di Milano(米兰理工大学) EssilorLuxottica(Essilor Luxottica)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Technical report for the HD-EPIC VQA Challenge 2025 (1st place)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00788 2025-06-11 cs.CV 50%

SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models

Wufei Ma, Luoxin Ye, Celso M de Melo, Jieneng Chen, Alan Yuille

机构 * Johns Hopkins University(约翰霍普金斯大学) DEVCOM Army Research Laboratory(国防部陆军研究实验室)

专题命中 视觉空间推理 :reasoning(abstract)

Comments CVPR 2025 highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07643 2025-06-10 cs.CV 50%

Synthetic Visual Genome

Jae Sung Park, Zixian Ma, Linjie Li, Chenhao Zheng, Cheng-Yu Hsieh, Ximing Lu, Khyathi Chandu, Quan Kong, Norimasa Kobori, Ali Farhadi, Yejin Choi, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Allen Institute for Artificial Intelligence(人工智能艾伦研究所) Stanford University(斯坦福大学) Woven by Toyota(丰田编织)

专题命中 视觉空间推理 :reasoning(abstract)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07091 2025-06-10 cs.CV 50%

SceneLCM: End-to-End Layout-Guided Interactive Indoor Scene Generation with Latent Consistency Model

Yangkai Lin, Jiabao Lei, Kui Jia

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05153 2025-06-10 cs.CV 50%

Hummingbird: High Fidelity Image Generation via Multimodal Context Alignment

Minh-Quan Le, Gaurav Mittal, Tianjian Meng, A S M Iftekhar, Vishwas Suryanarayanan, Barun Patra, Dimitris Samaras, Mei Chen

机构 * Microsoft(微软公司) Stony Brook University(史蒂文尼森布鲁克大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments Accepted to ICLR 2025. Project page with code release: https://roar-ai.github.io/hummingbird

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17641 2025-06-10 cs.RO 50%

Scene Exploration by Vision-Language Models

Venkatesh Sripada, Samuel Carter, Frank Guerin, Amir Ghalamzan

机构 * University of Surrey(萨里大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06199 2025-06-09 cs.RO cs.CV 50%

3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model

Hongyan Zhi, Peihao Chen, Siyuan Zhou, Yubo Dong, Quanxi Wu, Lei Han, Mingkui Tan

机构 * South China University of Technology(南方科技大学) Tencent Robotics X(腾讯机器人实验室) Hong Kong University of Science and Technology(香港科学与技术大学) Pazhou Laboratory(琶洲实验室)

专题命中 视觉空间推理 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04633 2025-06-06 cs.CV 50%

Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations

Linjie Li, Mahtab Bigverdi, Jiawei Gu, Zixian Ma, Yinuo Yang, Ziang Li, Yejin Choi, Ranjay Krishna

机构 * University of Washington(华盛顿大学) Sun Yat-sen University(中山大学) Stanford University(斯坦福大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments STARE is available at https://github.com/STARE-bench/STARE

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09586 2025-06-05 cs.CV 50%

Gaze-LLE: Gaze Target Estimation via Large-Scale Learned Encoders

Fiona Ryan, Ajay Bati, Sangmin Lee, Daniel Bolya, Judy Hoffman, James M. Rehg

机构 * Georgia Institute of Technology(佐治亚理工学院) Sungkyunkwan University(全南大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 视觉空间推理 :reasoning(abstract)

Comments CVPR 2025 Highlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02557 2025-06-04 cs.CV 50%

Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models

Shizhan Gong, Yankai Jiang, Qi Dou, Farzan Farnia

专题命中 视觉空间推理 :reasoning(abstract)

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01031 2025-06-03 cs.CV 50%

NavBench: Probing Multimodal Large Language Models for Embodied Navigation

Yanyuan Qiao, Haodong Hong, Wenqi Lyu, Dong An, Siqi Zhang, Yutong Xie, Xinyu Wang, Qi Wu

机构 * The University of Adelaide(阿德莱德大学) The University of Queensland(昆士兰大学) CSIRO Data61(澳大利亚联邦科学与工业研究组织数据61研究中心) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Tongji University(同济大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24315 2025-06-02 cs.CV 50%

InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing

Jinlu Zhang, Yixin Chen, Zan Wang, Jie Yang, Yizhou Wang, Siyuan Huang

机构 * Center on Frontiers of Computing Studies, School of Computer Science, Peking University(1 前沿计算研究中心,计算机科学学院,北京大学) State Key Laboratory of General Artificial Intelligence, BIGAI(2 通用人工智能国家重点实验室,BIGAI) Beijing Institute of Technology(3 北京理工大学) The Chinese University of Hong Kong, Shenzhen(4 香港中文大学(深圳)) Nat’l Eng. Research Center of Visual Technology, Peking University(5 视觉技术国家工程研究中心,北京大学) Institute for AI, Peking University(6 人工智能研究院,北京大学) State Key Laboratory of General Artificial Intelligence, Peking University(7 通用人工智能国家重点实验室,北京大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15962 2025-06-02 cs.RO 50%

Blimp-based Crime Scene Analysis

Martin Cooney, Fernando Alonso-Fernandez

机构 * School of Information Technology, Halmstad University(信息技术学院,哈姆斯塔德大学)

专题命中 视觉空间推理 :planning(abstract)

Comments 16 pages, 5 figures, 1 table; Accepted for SAIS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21351 2025-05-28 cs.RO 50%

EquAct: An SE(3)-Equivariant Multi-Task Transformer for Open-Loop Robotic Manipulation

Xupeng Zhu, Yu Qi, Yizhe Zhu, Robin Walters, Robert Platt

机构 * Northeastern University(东北大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20236 2025-05-27 cs.CV 50%

Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models

Weihao Xuan, Qingcheng Zeng, Heli Qi, Junjue Wang, Naoto Yokoya

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13617 2025-05-27 cs.CV 50%

Compile Scene Graphs with Reinforcement Learning

Zuyao Chen, Jinlin Wu, Zhen Lei, Marc Pollefeys, Chang Wen Chen

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.06397 2025-05-27 cs.CV 50%

PromptHMR: Promptable Human Mesh Recovery

Yufu Wang, Yu Sun, Priyanka Patel, Kostas Daniilidis, Michael J. Black, Muhammed Kocabas

专题命中 视觉空间推理 :reasoning(abstract)

Comments CVPR 2025. Project website: https://yufu-wang.github.io/phmr-page

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22679 2025-05-26 cs.CV 50%

Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

Weiqi Li, Xuanyu Zhang, Shijie Zhao, Yabin Zhang, Junlin Li, Li Zhang, Jian Zhang

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16602 2025-05-23 cs.CV 50%

MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation

Bohan Zhou, Yi Zhan, Zhongbin Zhang, Zongqing Lu

机构 * School of Computer Science, Peking University(北京大学计算机学院) Department of Automation, Tsinghua University(清华大学自动化系) BeingBeyond

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13959 2025-05-21 cs.RO 50%

MultiDrive: A Co-Simulation Framework Bridging 2D and 3D Driving Simulation for AV Software Validation

Marc Kaufeld, Korbinian Moller, Alessio Gambi, Paolo Arcaini, Johannes Betz

专题命中 视觉空间推理 :planning(abstract)

Comments 7 pages, Submitted to the IEEE International Conference on Intelligent Transportation Systems (ITSC 2025), Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15863 2025-05-21 cs.RO 50%

Task-oriented Robotic Manipulation with Vision Language Models

Nurhan Bulus Guran, Hanchi Ren, Jingjing Deng, Xianghua Xie

机构 * Department of Computer Science, Swansea University(计算机科学系,萨瑟兰大学) Department of Computer Science, Durham University(计算机科学系,杜ham大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12998 2025-05-20 cs.CV 50%

A Skull-Adaptive Framework for AI-Based 3D Transcranial Focused Ultrasound Simulation

Vinkle Srivastav, Juliette Puel, Jonathan Vappou, Elijah Van Houten, Paolo Cabras, Nicolas Padoy

机构 * University of Strasbourg, CNRS, INSERM, ICube, UMR7357(斯特拉斯堡大学,法国国家科学研究中心,法国国家医学研究院,ICube,UMR7357) IHU Strasbourg(斯特拉斯堡IHU) Image Guided Therapy(影像引导治疗)

专题命中 视觉空间推理 :planning(abstract)

Comments The project page is available at https://github.com/CAMMA-public/TFUScapes

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12194 2025-05-20 cs.RO 50%

Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding

Xuefei Sun, Doncey Albin, Cecilia Mauceri, Dusty Woods, Christoffer Heckman

机构 * Autonomous Robotics and Perception Group in the Computer Science Department at the University of Colorado Boulder(科罗拉多大学波德分校计算机科学系自主机器人与感知小组)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15891 2025-05-20 cs.GR cs.CV 50%

TexPro: Text-guided PBR Texturing with Procedural Material Modeling

Ziqiang Dang, Wenqi Dong, Zesong Yang, Bangbang Yang, Liang Li, Yuewen Ma, Zhaopeng Cui

专题命中 视觉空间推理 :reasoning(abstract)

Comments Accepted by CVM 2025 and CVMJ (Computational Visual Media Journal)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19090 2025-05-16 cs.CV 50%

EndoMamba: An Efficient Foundation Model for Endoscopic Videos via Hierarchical Pre-training

Qingyao Tian, Huai Liao, Xinyan Huang, Bingyu Yang, Dongdong Lei, Sebastien Ourselin, Hongbin Liu

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) The First Affiliated Hospital, Sun Yat-sen University(中山大学第一附属医院) Centre for Artificial Intelligence and Robotics, Chinese Academy of Sciences(中国科学院人工智能与机器人中心) School of Engineering and Imaging Sciences, King’s College London(伦敦大学国王学院工程与影像科学学院)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08725 2025-05-14 cs.CV 50%

Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving

Zongchuang Zhao, Haoyu Fu, Dingkang Liang, Xin Zhou, Dingyuan Zhang, Hongwei Xie, Bing Wang, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Xiaomi EV(小米电动车)

专题命中 视觉空间推理 :reasoning(abstract)

Comments The dataset and code will be released at https://github.com/zc-zhao/DriveMonkey

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06575 2025-05-13 cs.CV 50%

GRACE: Estimating Geometry-level 3D Human-Scene Contact from 2D Images

Chengfeng Wang, Wei Zhai, Yuhang Yang, Yang Cao, Zhengjun Zha

机构 * University of Science and Technology of China(中国科学技术大学)

专题命中 视觉空间推理 :reasoning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05495 2025-05-12 cs.CV cs.RO 50%

Learning 3D Persistent Embodied World Models

Siyuan Zhou, Yilun Du, Yuncong Yang, Lei Han, Peihao Chen, Dit-Yan Yeung, Chuang Gan

机构 * Institution1(机构1) Institution2(机构2)

专题命中 视觉空间推理 :planning(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05473 2025-05-09 cs.CV 50%

DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint Diffusion

Qitao Zhao, Amy Lin, Jeff Tan, Jason Y. Zhang, Deva Ramanan, Shubham Tulsiani

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉空间推理 :reasoning(abstract)

Comments CVPR 2025. Project website: https://qitaozhao.github.io/DiffusionSfM

详情

展开后加载摘要…

URL PDF HTML 收藏