arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

共收录 3352 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 3352 篇

1805.08565 2018-05-23 cs.LG stat.ML 74%

Global Navigation Using Predictable and Slow Feature Analysis in Multiroom Environments, Path Planning and Other Control Tasks

Stefan Richthofer, Laurenz Wiskott

专题命中 视觉空间推理 :planning(title);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16952 2025-06-23 cs.CY 73%

Modeling and Visualization Reasoning for Stakeholders in Education and Industry Integration Systems: Research on Structured Synthetic Dialogue Data Generation Based on NIST Standards

Wei Meng

专题命中 视觉空间推理 :reasoning(title,comments)

Comments This paper presents an innovative and rigorous framework for stakeholder modelling in education-industry integration, combining NIST-compliant synthetic data generation with interpretable visual reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05933 2026-06-25 cs.CL cs.AI 版本更新 73%

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs

强化学习改善LLM中参数化知识的遍历

Renfei Zhang, Manasa Kaniselvan, Rylan Schaeffer abd Niloofar Mireshghallah

机构 * Carnegie Mellon University(卡内基梅隆大学) MIT(麻省理工学院) Stanford University(斯坦福大学)

专题命中 视觉空间推理 :reasoning(abstract);self-correction(abstract);分类 cs.CL、cs.AI

AI总结 本文发现强化学习提升LLM知识回忆能力,原因在于改进了模型对参数化知识层次结构的遍历技能,而非获取新知识。

Comments `

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.00095 2026-06-02 cs.CV cs.AI cs.CL cs.RO 73%

Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation

弥合2D-3D鸿沟:面向视觉语言导航的分层语义几何地图

Kailing Li, Tianwen Qian, Lijin Yang, Yuqian Fu, Jingyu Gong, Xiaoling Wang, Liang He

机构 * School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院) Bosch Corporate Research(博世企业研究) King Abdullah University of Science and Technology(卡布斯大学)

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI

AI总结 提出分层语义几何地图(HSGM),将3D几何信息转化为VLM可理解的结构化表示,结合VLM高层语义规划与经典路径规划,实现零样本视觉语言导航,在R2R-CE和RxR-CE基准上达到最先进性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20287 2026-05-21 cs.LG cs.AI cs.CV 73%

FusionCell: Cross-Attentive Fusion of Layout Geometry and Netlist Topology for Standard-Cell Performance Prediction

FusionCell: 跨注意力融合布局几何与网络列表拓扑以实现标准单元性能预测

Haoyi Zhang, Kairong Guo, Bojie Zhang, Yibo Lin, Runsheng Wang

机构 * School of Integrated Circuits, Peking University, Beijing, China(集成电路学院,北京大学,北京,中国)

专题命中 视觉空间推理 :reasoning(abstract);logical reasoning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出FusionCell,通过跨注意力机制融合布局几何和网络列表拓扑,以提高标准单元性能预测的准确性,解决了传统方法忽略布局几何导致的耦合和布局依赖效应的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11509 2026-05-13 cs.AI cs.LG cs.MA cs.SY eess.SY 73%

Hierarchical LLM-Driven Control for HAPS-Assisted UAV Networks: Joint Optimization of Flight and Connectivity

分层LLM驱动控制用于HAPS辅助无人机网络:飞行与连接的联合优化

Zijiang Yan, Hao Zhou, Wael Jaafar, Jianhua Pei, Ping Wang, Halim Yanikomeroglu, Hina Tabassum

机构 * Department of Electrical Engineering and Computer Science, York University(约克大学电气工程与计算机科学系) Samsung Research America(三星美国研究院) Department of Software and IT Engineering, École de technologie supérieure (ÉTS), University of Quebec(魁北克大学软件与信息技术工程系,École de technologie supérieure) Non-Terrestrial Networks (Carleton-NTN) Lab and the Department of Systems and Computer Engineering, Carleton University(非地面网络(Carleton-NTN)实验室和系统与计算机工程系,卡尔顿大学)

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.AI、cs.LG

AI总结 本文研究了在整合陆地和非陆地网络中多无人机系统的联合优化问题,提出基于LLM的分层多速率控制框架,通过高精度仿真平台验证了其在提升运输效率、通信吞吐量和减少碰撞率方面的优势。

Comments Submission for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12597 2026-03-16 cs.LG cs.AI cs.HC cs.MA cs.SE 73%

Feynman: Knowledge-Infused Diagramming Agent for Scalable Visual Designs

Feynman: 一种融合知识的图表生成代理,用于可扩展的视觉设计

Zixin Wen, Yifu Cai, Kyle Lee, Sam Estep, Josh Sunshine, Aarti Singh, Yuejie Chi, Wode Ni

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出Feynman代理,通过知识组件枚举和代码规划生成高质量图表-描述对,构建了可扩展的图表生成流水线,并创建了视觉语言基准Diagramma。

Comments A previous version was submitted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19255 2026-03-06 cs.LG cs.AI 73%

VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use

VTool-R1: 通过在多模态工具使用上的强化学习使VLMs学会通过图像思考

Mingyuan Wu, Jingcheng Yang, Jize Jiang, Meitang Li, Kaizhuo Yan, Hanchao Yu, Minjia Zhang, Chengxiang Zhai, Klara Nahrstedt

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Michigan Ann Arbor(密歇根大学安娜堡分校) Independent Researcher(独立研究者)

专题命中 视觉空间推理 :reasoning(abstract);self-correction(abstract);分类 cs.AI、cs.LG

AI总结 VTool-R1通过强化学习训练视觉语言模型生成多模态思考链,提升其通过图像进行推理的能力。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01644 2026-02-03 cs.LG cs.AI cs.CV cs.MA cs.RO 73%

From Perception to Action: Spatial AI Agents and World Models

从感知到行动:空间AI代理与世界模型

Gloria Felicia, Nolan Bryant, Handi Putra, Ayaan Gazali, Eliel Lobo, Esteban Rojas

机构 * AtlasPro AI

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种统一的三轴分类法,将代理能力与空间任务联系起来,强调空间定位与符号定位的区别,并指出世界模型对跨尺度安全部署的重要性。

Comments 61 pages, 742 citations, 1 figure, 3 tables. Survey paper on spatial AI agents, embodied AI, graph neural networks, and world models

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13968 2026-01-27 cs.CV cs.AI cs.CL 73%

RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation

RotBench: 对多模态大语言模型识别图像旋转能力的评估

Tianyi Niu, Jaemin Cho, Elias Stengel-Eskin, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学夏洛特分校) Allen Institute for Artificial Intelligence(人工智能研究院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

AI总结 RotBench评估了多模态大语言模型在识别图像旋转角度方面的性能,发现大多数模型难以区分90°和270°旋转,但能识别0°和180°图像,揭示了模型空间推理能力与人类的差距。

Comments EACL 2026 Camera-Ready. Code and data: https://github.com/tianyiniu/RotBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09775 2026-01-16 cs.LG cs.CL 73%

The Geometry of Thought: Disclosing the Transformer as a Tropical Polynomial Circuit

思维的几何学:揭示Transformer为热带多项式电路

Faruk Alpay, Bilge Senturk

机构 * Bahçeşehir University(巴切希尔大学)

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG

AI总结 本研究揭示Transformer在高置信度下通过热带多项式电路实现动态规划,为链式思维提供几何解释。

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02382 2026-01-07 cs.NI cs.AI cs.IR cs.LG 73%

How to Discover Knowledge for FutureG: Contextual RAG and LLM Prompting for O-RAN

如何为未来G发现知识:面向O-RAN的上下文RAG和LLM提示

Nathan Conger, Nathan Scollar, Kemal Davaslioglu, Yalin E. Sagduyu, Sastry Kompella

专题命中 视觉空间推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.AI、cs.LG

AI总结 本文提出上下文RAG方法,通过引导文档检索和上下文增强LLM性能,提升ORAN领域问答的准确性和效率,同时保持低运行时间和碳排放。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23907 2025-12-02 cs.CV cs.AI cs.LG 73%

DynaStride: Dynamic Stride Windowing with MMCoT for Instructional Multi-Scene Captioning

DynaStride: 基于MMCoT的动态步长窗多场景描述生成

Eddison Pham, Prisha Priyadarshini, Adrian Maliackel, Kanishk Bandi, Cristian Meo, Kevin Zhu

机构 * Algoverse AI Research(Algoverse人工智能研究)

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

AI总结 DynaStride通过多模态链式思考和动态步长窗算法,实现无需手动分割的多场景教学视频连贯描述生成,提升描述的连贯性和信息量。

Comments 16 pages, 15 figures, 5 Tables, Accepted at NeurIPS 7HVU Workshop, Accepted at AAAI AI4ED Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19433 2025-10-13 cs.CV cs.AI cs.CL 73%

Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System

Lixuan He, Haoyu Dong, Zhenxing Chen, Yangcheng Yu, Jie Feng, Yong Li

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI

Comments The paper is currently under investigation regarding concerns of potential academic misconduct. While the investigation is ongoing, the authors have voluntarily requested to withdraw the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26539 2025-10-01 cs.CV cs.CL cs.LG 73%

Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents

Zhen Yang, Zi-Yi Dou, Di Feng, Forrest Huang, Anh Nguyen, Keen You, Omar Attia, Yuhao Yang, Michael Feng, Haotian Zhang, Ram Ramrakhya, Chao Jia, Jeffrey Nichols, Alexander Toshev, Yinfei Yang, Zhe Gan

机构 * Apple(苹果公司)

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11197 2025-09-16 cs.RO cs.AI cs.CL cs.CV 73%

DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation

Yunheng Wang, Yuetong Fang, Taowen Wang, Yixiao Feng, Yawen Tan, Shuning Zhang, Peiran Liu, Yiding Ji, Renjing Xu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Zhejiang Normal University(浙江师范大学)

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11247 2025-08-19 cs.AI cs.LG cs.RO 73%

LD-Scene: LLM-Guided Diffusion for Controllable Generation of Adversarial Safety-Critical Driving Scenarios

Mingxing Peng, Yuting Xie, Xusen Guo, Ruoyu Yao, Hai Yang, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)

专题命中 视觉空间推理 :chain-of-thought(abstract);CoT(abstract);分类 cs.AI、cs.LG

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15677 2025-07-31 cs.AI cs.CL cs.CV cs.MM cs.RO 73%

Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence

Yining Hong, Rui Sun, Bingxuan Li, Xingcheng Yao, Maxine Wu, Alexander Chien, Da Yin, Ying Nian Wu, Zhecan James Wang, Kai-Wei Chang

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18945 2025-07-29 cs.CV cs.AI cs.LG cs.RO 73%

Aether: Geometric-Aware Unified World Modeling

Aether Team, Haoyi Zhu, Yifan Wang, Jianjun Zhou, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Chunhua Shen, Jiangmiao Pang, Tong He

机构 * USTC Shanghai AI Lab(USTC上海人工智能实验室) SII SJTU(SJTU信息研究所) ZJU FDU(浙江大学福州大学)

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.AI、cs.LG

Comments Project Page: https://aether-world.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20897 2025-06-24 cs.CV cs.AI cs.CL cs.RO 73%

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

Pingrui Zhang, Yifei Su, Pengyuan Wu, Dong An, Li Zhang, Zhigang Wang, Dong Wang, Yan Ding, Bin Zhao, Xuelong Li

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MAIS, Institute of Automation of Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) University of Science and Technology of China(中国科学技术大学) TeleAI, China Telecom Corp Ltd(中国电信TeleAI)

专题命中 视觉空间推理 :reasoning(abstract);logical reasoning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09848 2025-04-15 cs.AI cs.CL 73%

A Survey of Large Language Model-Powered Spatial Intelligence Across Scales: Advances in Embodied Agents, Smart Cities, and Earth Science

Jie Feng, Jinwei Zeng, Qingyue Long, Hongyi Chen, Jie Zhao, Yanxin Xi, Zhilun Zhou, Yuan Yuan, Shengyuan Wang, Qingbin Zeng, Songwei Li, Yunke Zhang, Yuming Lin, Tong Li, Jingtao Ding, Chen Gao, Fengli Xu, Yong Li

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09819 2025-02-17 cs.CV cs.AI cs.GR cs.LG cs.PL 73%

A Solver-Aided Hierarchical Language for LLM-Driven CAD Design

Benjamin T. Jones, Felix Hähnlein, Zihan Zhang, Maaz Ahmad, Vladimir Kim, Adriana Schulz

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.10404 2024-07-30 cs.CV cs.AI cs.LG 73%

LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation

Kibum Kim, Kanghoon Yoon, Jaehyeong Jeon, Yeonjun In, Jinyoung Moon, Donghyun Kim, Chanyoung Park

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

Comments 8 pages; CVPR 2024

Journal ref CVPR (2024), 28306-28316

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.14562 2024-06-21 cs.CL cs.AI cs.CV 73%

Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities

Sachit Menon, Richard Zemel, Carl Vondrick

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI

Comments Project website: whiteboard.cs.columbia.edu/

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.02537 2024-06-05 cs.CL cs.CV cs.LG 73%

TopViewRS: Vision-Language Models as Top-View Spatial Reasoners

Chengzu Li, Caiqi Zhang, Han Zhou, Nigel Collier, Anna Korhonen, Ivan Vulić

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.LG

Comments 9 pages, 3 figures, 3 tables (21 pages, 4 figures, 15 tables including references and appendices)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02651 2024-05-24 cs.LG cs.AI cs.CV 73%

Vision-Language Models Provide Promptable Representations for Reinforcement Learning

William Chen, Oier Mees, Aviral Kumar, Sergey Levine

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09631 2024-03-15 cs.CV cs.AI cs.CL cs.RO 73%

3D-VLA: A 3D Vision-Language-Action Generative World Model

Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, Chuang Gan

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI

Comments Project page: https://vis-www.cs.umass.edu/3dvla/

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11854 2024-02-27 cs.LG cs.AI stat.ML 73%

Multimodal Web Navigation with Instruction-Finetuned Foundation Models

Hiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo, Aleksandra Faust, Shixiang Shane Gu, Izzeddin Gur

专题命中 视觉空间推理 :reasoning(abstract);planning(abstract);分类 cs.AI、cs.LG

Comments Accepted to ICLR 2024. Website: https://sites.google.com/view/mm-webnav/

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07872 2024-02-13 cs.RO cs.CL cs.CV cs.LG 73%

PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs

Soroush Nasiriany, Fei Xia, Wenhao Yu, Ted Xiao, Jacky Liang, Ishita Dasgupta, Annie Xie, Danny Driess, Ayzaan Wahid, Zhuo Xu, Quan Vuong, Tingnan Zhang, Tsang-Wei Edward Lee, Kuang-Huei Lee, Peng Xu, Sean Kirmani, Yuke Zhu, Andy Zeng, Karol Hausman, Nicolas Heess, Chelsea Finn, Sergey Levine, Brian Ichter

专题命中 视觉空间推理 :reasoning(abstract);logical reasoning(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07757 2024-02-13 cs.LG cs.AI 73%

Towards an Understanding of Stepwise Inference in Transformers: A Synthetic Graph Navigation Model

Mikail Khona, Maya Okawa, Jan Hula, Rahul Ramesh, Kento Nishi, Robert Dick, Ekdeep Singh Lubana, Hidenori Tanaka

专题命中 视觉空间推理 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏