arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 3363 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 3363 篇

1912.11684 2020-03-10 cs.CV cs.LG cs.RO cs.SD eess.AS 57%

Look, Listen, and Act: Towards Audio-Visual Embodied Navigation

Chuang Gan, Yiwei Zhang, Jiajun Wu, Boqing Gong, Joshua B. Tenenbaum

专题命中 视觉空间推理 :planning(abstract);分类 cs.LG

Comments Accepted by ICRA 2020. Project page: http://avn.csail.mit.edu

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.09589 2019-12-23 cs.HC cs.AI 57%

Smart Home Appliances: Chat with Your Fridge

Denis Gudovskiy, Gyuri Han, Takuya Yamaguchi, Sotaro Tsukizawa

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments NeurIPS 2019 demo track

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.00915 2019-12-03 cs.AI 57%

Just Ask:An Interactive Learning Framework for Vision and Language Navigation

Ta-Chung Chi, Mihail Eric, Seokhwan Kim, Minmin Shen, Dilek Hakkani-tur

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 8 pages, accepted to AAAI 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.09033 2019-11-21 cs.LG cs.CV stat.ML 57%

Exploiting Spatial Invariance for Scalable Unsupervised Object Tracking

Eric Crawford, Joelle Pineau

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

Comments Accepted at AAAI 2020. Code: https://github.com/e2crawfo/silot. Visualizations: https://sites.google.com/view/silot

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.13249 2019-10-30 cs.CV cs.HC cs.LG 57%

Navigation Agents for the Visually Impaired: A Sidewalk Simulator and Experiments

Martin Weiss, Simon Chamorro, Roger Girgis, Margaux Luck, Samira E. Kahou, Joseph P. Cohen, Derek Nowrouzezahrai, Doina Precup, Florian Golemo, Chris Pal

专题命中 视觉空间推理 :planning(abstract);分类 cs.LG

Comments Accepted at CoRL2019. Code & video available at https://mweiss17.github.io/SEVN/

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.07357 2019-09-18 cs.AI 57%

CHALET: Cornell House Agent Learning Environment

Claudia Yan, Dipendra Misra, Andrew Bennnett, Aaron Walsman, Yonatan Bisk, Yoav Artzi

专题命中 视觉空间推理 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.04306 2019-09-11 cs.CV cs.LG cs.RO 57%

Bayesian Relational Memory for Semantic Visual Navigation

Yi Wu, Yuxin Wu, Aviv Tamar, Stuart Russell, Georgia Gkioxari, Yuandong Tian

专题命中 视觉空间推理 :planning(abstract);分类 cs.LG

Comments Accepted at ICCV 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.04776 2019-07-30 cs.CV cs.LG 57%

Multi-Agent Tensor Fusion for Contextual Trajectory Prediction

Tianyang Zhao, Yifei Xu, Mathew Monfort, Wongun Choi, Chris Baker, Yibiao Zhao, Yizhou Wang, Ying Nian Wu

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.LG

Comments Presented in CVPR 19: http://openaccess.thecvf.com/content_CVPR_2019/html/Zhao_Multi-Agent_Tensor_Fusion_for_Contextual_Trajectory_Prediction_CVPR_2019_paper.html ; Architecture details available: https://github.com/programmingLearner/MATF-architecture-details

详情

展开后加载摘要…

URL PDF HTML 收藏
1906.00932 2019-06-05 cs.CV cs.LG stat.ML 57%

Y-GAN: A Generative Adversarial Network for Depthmap Estimation from Multi-camera Stereo Images

Miguel Alonso

专题命中 视觉空间推理 :planning(abstract);分类 cs.LG

Comments Accepted for Presentation at the ICML 2019 LatinX in AI Research Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.05757 2019-03-15 cs.HC cs.AI cs.CV 57%

VRKitchen: an Interactive 3D Virtual Environment for Task-oriented Learning

Xiaofeng Gao, Ran Gong, Tianmin Shu, Xu Xie, Shu Wang, Song-Chun Zhu

专题命中 视觉空间推理 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.02805 2019-01-04 cs.NE cs.AI cs.RO 57%

DeepTraffic: Crowdsourced Hyperparameter Tuning of Deep Reinforcement Learning Systems for Multi-Agent Dense Traffic Navigation

Lex Fridman, Jack Terwilliger, Benedikt Jenik

专题命中 视觉空间推理 :planning(abstract);分类 cs.AI

Comments Neural Information Processing Systems (NIPS 2018) Deep Reinforcement Learning Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.07488 2018-11-22 cs.CV cs.AI 57%

Quantifying Human Behavior on the Block Design Test Through Automated Multi-Level Analysis of Overhead Video

Seunghwan Cha, James Ainooson, Maithilee Kunda

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.07528 2018-10-18 cs.AI 57%

Machine Common Sense Concept Paper

David Gunning

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.08481 2017-02-08 cs.AI cs.CV 57%

GuessWhat?! Visual object discovery through multi-modal dialogue

Harm de Vries, Florian Strub, Sarath Chandar, Olivier Pietquin, Hugo Larochelle, Aaron Courville

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

Comments 23 pages; CVPR 2017 submission; see https://guesswhat.ai

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.00380 2016-12-02 cs.AI cs.CV stat.ML 57%

Playing Doom with SLAM-Augmented Deep Reinforcement Learning

Shehroze Bhatti, Alban Desmaison, Ondrej Miksik, Nantas Nardelli, N. Siddharth, Philip H. S. Torr

专题命中 视觉空间推理 :planning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1602.02261 2016-05-23 cs.AI 57%

End-to-End Goal-Driven Web Navigation

Rodrigo Nogueira, Kyunghyun Cho

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05336 2026-01-12 cs.RO 56%

Intent at a Glance: Gaze-Guided Robotic Manipulation via Foundation Models

意图一目了然:通过基础模型实现的注视引导的机器人操作

Tracey Yee Hsin Tay, Xu Yan, Jonathan Ouyang, Daniel Wu, William Jiang, Jonathan Kao, Yuchen Cui

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 视觉空间推理 :reasoning(abstract);planning(comments)

AI总结 GAMMA通过结合基础模型和注视技术,实现无需特定任务训练的机器人操作自主控制,提升人机交互的直观性和可扩展性。

Comments Accepted to 2025 RSS Robot Planning in the Era of Foundation Models (FM4RoboPlan) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03540 2024-11-07 cs.RO 56%

VLA-3D: A Dataset for 3D Semantic Scene Understanding and Navigation

Haochen Zhang, Nader Zantout, Pujith Kachana, Zongyuan Wu, Ji Zhang, Wenshan Wang

专题命中 视觉空间推理 :reasoning(abstract,comments)

Comments Accepted and presented at the 1st Workshop on Semantic Reasoning and Goal Understanding in Robotics (SemRob), Robotics Science and Systems Conference (RSS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.14257 2023-10-31 cs.CL cs.AI cs.LG 56%

Hierarchical Prompting Assists Large Language Model on Web Navigation

Abishek Sridhar, Robert Lo, Frank F. Xu, Hao Zhu, Shuyan Zhou

专题命中 视觉空间推理 :分类 cs.CL、cs.AI、cs.LG;reasoning(comments)

Comments EMNLP 2023 Findings; Natural Language Reasoning and Structured Explanations Workshop at ACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.06150 2018-09-20 cs.RO cs.AI cs.CL cs.LG 56%

FollowNet: Robot Navigation by Following Natural Language Directions with Deep Reinforcement Learning

Pararth Shah, Marek Fiser, Aleksandra Faust, J. Chase Kew, Dilek Hakkani-Tur

专题命中 视觉空间推理 :分类 cs.CL、cs.AI、cs.LG;planning(journal_ref)

Comments 7 pages, 8 figures

Journal ref Third Workshop in Machine Learning in the Planning and Control of Robot Motion at ICRA, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.03947 2017-09-13 cs.RO 56%

Constant Space Complexity Environment Representation for Vision-based Navigation

Jeffrey Kane Johnson

专题命中 视觉空间推理 :planning(abstract,comments)

Comments IROS 2017: 9th Workshop on Planning, Perception and Navigation for Intelligent Vehicles

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19298 2026-08-21 cs.CV 新提交 50%

SceneGTMM: A Conformal Mapping-based Scene-Aware Transferable GNN-Transformer Dual-Graph Interaction Framework for Map Matching

SceneGTMM:一种基于保角映射的场景感知可迁移GNN-Transformer双图交互地图匹配框架

Yongliang Zhang, Feng Song, Ji Chen, Lishuai Guo, Yong Deng, Yue Zheng, Tianyi Liu, Zhixiong Chen, Qixin Zhang

专题命中 视觉空间推理 :planning(abstract)

AI总结 本文提出SceneGTMM框架,通过保角映射场景策略、双图交互架构与CRF增强预测,提升地图匹配的噪声鲁棒性、跨区域迁移性与可解释性,在多源及跨城轨迹匹配中表现优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.14497 2026-08-21 cs.CV 版本更新 50%

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

通过自我场景增强在多模态大语言模型中强化自我中心空间感知

Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu, Xu Zheng, Zidong Cao, Lutao Jiang, Zixin Zhang, Huiyu Zhou, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Guangxi Zhuang Autonomous Region Information Center(广西壮族自治区信息中心) The Hong Kong University of Science and Technology(香港科技大学)

专题命中 视觉空间推理 :reasoning(abstract)

AI总结 研究如何强化多模态大语言模型的自我中心空间感知,提出自我场景增强框架ESA,利用自我元素图作为中间表示,通过视觉基础模型增强空间感知,在EgoTextVQA基准上取得显著性能提升。

Comments 14 pages, 8 figures. Chi Kit Wong and Ye Pan contributed equally. Code: https://github.com/Chikit-WONG/spatialGraph

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29786 2026-08-21 cs.CV cs.RO 版本更新 50%

OP3DSG: Open-Vocabulary Part-Aware 3D Scene Graph Generation for Real-World Environments

OP3DSG:面向真实环境的开放词汇部件感知3D场景图生成

Yirum Kim, Ue-Hwan Kim

机构 * Gwangju Institute of Science and Technology(全州科学技术院)

专题命中 视觉空间推理 :reasoning(abstract)

AI总结 提出OP3DSG框架,通过部件感知检测与融合、几何初始化先验图及LLM精炼,实现开放词汇下对象、交互部件、空间/功能关系及可供性的统一建模,在UniGraph3D基准上达到最优性能。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02846 2026-08-21 cs.CV 50%

Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?

瞬间动作预见:多模态线索能替代视频到何种程度?

Manuel Benavent-Lledo, Konstantinos Bacharidis, Victoria Manousaki, Konstantinos Papoutsakis, Antonis Argyros, Jose Garcia-Rodriguez

机构 * Universidad de Alicante(阿利坎特大学) Foundation for Research and Technology-Hellas(希腊基础研究与技术基金会) University of Crete(克里特大学) Hellenic Mediterranean University(希腊地中海大学)

专题命中 视觉空间推理 :reasoning(abstract)

AI总结 AAG通过结合单帧RGB特征与深度线索及先前动作信息,实现了多模态单帧动作预见,能与视频聚合基线和先进方法在教学活动数据集上竞争。

Comments Accepted in WACV 2026 - Applications Track

Journal ref 2026 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.19059 2026-08-20 cs.RO cs.CV 新提交 50%

LT-Mem: Volatility-Aware Spatio-Temporal Memory for Lifelong Scene Understanding

LT-Mem:面向终身场景理解的感知时间动态的时空记忆

Yumin Lee, Hyoseok Ju, Giseop Kim

机构 * DGIST(大邱庆北科学技术院)

专题命中 视觉空间推理 :reasoning(abstract)

AI总结 该研究针对机器人长期场景理解的时间遗忘问题,提出LT-Mem时空记忆框架,结合多会话SLAM与Tri-Memory结构,在LT-VQA数据集上性能优于基线且token消耗更少。

Comments 8 pages, 8 figures, 6 tables. Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18840 2026-08-20 cs.RO cs.CV 新提交 50%

Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction

超越布局与关节运动:面向具身交互的使用驱动型代码场景

Zijian Xiao, Zipeng Ye, Jinkun Hao, Xiong Yang, Yuchen Xie, Ran Yi

机构 * Meituan(美团)

专题命中 视觉空间推理 :reasoning(abstract)

AI总结 针对现有代码场景生成方法未建模场景功能使用的问题,提出RoomWright框架,通过使用驱动型物体推理与代码智能体实现可执行、可编辑的仿真就绪3D场景,为具身AI与策略学习提供交互式环境。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18498 2026-08-20 cs.CV 新提交 50%

DyG$^2$T: Modeling Object Dynamics with 3D Gaussian Temporal-Spatial Particle Graph Transformer

DyG²T:基于3D高斯时空粒子图Transformer的对象动力学建模

Yansong Wang, Zhaobo Qi, Xinyan Liu, Beichen Zhang, Shuhui Wang, Weigang Zhang, Qingming Huang

机构 * School of Computer Science and Technology, Harbin Institute of Technology at Weihai(哈尔滨工业大学(威海)计算机科学与技术学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)

专题命中 视觉空间推理 :reasoning(abstract)

AI总结 本文针对现有动力学建模方法丢失细粒度细节、轨迹漂移等问题,提出DyG²T框架,通过空间补全关键点、TDN增强时间判别性、粒子图Transformer建模交互,在合成与真实数据集上实现精准动力学建模,泛化能力强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21064 2026-08-20 cs.CV 版本更新 50%

2Xplat: Decoupling Geometry and Appearance Modeling for Feed-Forward 3D Gaussian Splatting

2Xplat:两个专家胜过一个通才

Hwasik Jeong, Seungryong Lee, Gyeongjin Kang, Seungkwon Yang, Xiangyu Sun, Seungtae Nam, Eunbyung Park

机构 * Yonsei University(延世大学) Sungkyunkwan University(成均馆大学)

专题命中 视觉空间推理 :reasoning(abstract)

AI总结 本文提出2Xplat框架,通过分离几何估计与高斯生成,改进3DGS生成效果,证明模块化设计在复杂3D建模中的优势。

Comments Project page: https://hwasikjeong.github.io/2Xplat

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17975 2026-08-19 cs.GR cs.CV 新提交 50%

aDSL: Agentic 3D Creation via Joint Agent-Program Design

aDSL:通过智能体-程序联合设计实现智能体驱动的3D内容创建

Rui-Huan Wang, Si-Tong Wei, Jia-Qi He, Heng-Yi Wei, Baoquan Chen, Peng-Shuai Wang

专题命中 视觉空间推理 :reasoning(abstract)

AI总结 该研究联合设计以智能体为中心的领域特定语言aDSL与角色专业化多智能体系统,通过“规划-执行-评审”循环提升3D内容创建的鲁棒性等,在文本到形状等任务上优于LLM基线。

详情

展开后加载摘要…

URL PDF HTML 收藏