arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型推理能力

大模型数学、逻辑、规划、多步推理和测试时计算能力。

2026-04-29 至 2026-04-29 共收录 8 信号源:cs.CL, cs.AI, cs.LG

1. 视觉空间推理 8 篇

2512.03043 2026-04-29 cs.CV 85%

OneThinker: All-in-one Reasoning Model for Image and Video

OneThinker:面向图像和视频的统一推理模型

Kaituo Feng, Manyuan Zhang, Hongyu Li, Kaixuan Fan, Shuang Chen, Yilei Jiang, Dian Zheng, Peiwen Sun, Yiyuan Zhang, Haoze Sun, Yan Feng, Peng Pei, Xunliang Cai, Xiangyu Yue

机构 * MMLab, CUHK(CUHK多媒体实验室) Meituan Home(美团家)

专题命中 视觉空间推理 :reasoning(title,abstract);CoT(abstract,abstract_cn)

AI总结 OneThinker提出一个统一的多模态推理模型,整合图像和视频理解,涵盖问答、描述生成、空间时间定位、跟踪和分割等任务,通过构建大规模训练语料和EMA-GRPO算法提升多任务强化学习效果。

Comments CVPR 2026, Project page: https://github.com/tulerfeng/OneThinker

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16918 2026-04-29 cs.CV 80%

AdaTooler-V: Adaptive Tool-Use for Images and Videos

AdaTooler-V:面向图像和视频的自适应工具使用

Chaoyang Wang, Kaituo Feng, Dongyang Chen, Zhongyu Wang, Zhixun Li, Sicheng Gao, Meng Meng, Xu Zhou, Manyuan Zhang, Yuzhang Shang, Xiangyu Yue

机构 * MMLab, CUHK(香港中文大学MML实验室) THU(清华大学) SJTU(上海交通大学) DB Group, CUHK(香港中文大学DB小组) UCF(佛罗里达大学) Sangfor(Sangfor公司) JMU(约翰·霍普金斯大学)

专题命中 视觉空间推理 :CoT(abstract,abstract_cn);reasoning(abstract);chain-of-thought(abstract)

AI总结 本文提出AdaTooler-V,通过自适应工具使用提升多模态大语言模型的视觉推理能力,通过强化学习算法和定制数据集优化工具调用策略,实验证明其在多种视觉任务中表现优异。

Comments ACL 2026 Findings, Project page: https://github.com/CYWang735/AdaTooler-V

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25318 2026-04-29 cs.GR cs.AI cs.CL 62%

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation

Cutscene Agent:一种用于自动化3D场景生成的LLM代理框架

Lanshan He, Haozhou Pang, Qi Gan, Xin Shen, Ziwei Zhang, Yibo Liu, Gang Fang, Bo Liu, Kai Sheng, Shengfeng Zeng, Chaofan Li, Zhen Hui, Keer Zhou, Lan Zhou, Shujun Dai

机构 * Kuaishou GameMind Lab(快手游戏大脑实验室)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Cutscene Agent框架,通过双向集成LLM代理与游戏引擎,实现多代理协作生成可编辑的电影化3D内容,并构建了针对长周期多步骤协作的评估基准。

Comments 27 pages excluding appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25268 2026-04-29 cs.CL cs.AI 62%

CRAFT: Grounded Multi-Agent Coordination Under Partial Information

CRAFT:在部分信息下的 grounded 多智能体协调

Abhijnan Nath, Hannah VanderHoeven, Nikhil Krishnaswamy

机构 * Department of Computer Science, Colorado State University(计算机科学系,科罗拉多州立大学)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.CL、cs.AI

AI总结 CRAFT 是评估大语言模型在严格部分信息下实用沟通能力的多智能体基准,通过自然语言构建共享3D结构。研究发现更强推理能力不必然带来更好协调,小模型常表现更优,表明多智能体协调仍是当前语言模型的挑战。

Comments Added revisions, corrected typos and additional analysis

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25526 2026-04-29 cs.SE cs.AI 57%

AI as Consumer and Participant: A Co-Design Agenda for MBSE Substrates and Methodology

AI作为消费者和参与者:MBSE子系统与方法论的协同设计议程

Siyuan Ji

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

AI总结 本文探讨了AI在MBSE模型中的应用问题,指出现有模型未为AI消费设计,提出需协同设计模型与方法论,使其成为可查询的知识子系统,而非仅为人导航的结构化 artifacts。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22875 2026-04-29 cs.CV cs.AI 57%

SketchVLM: Vision language models can annotate images to explain thoughts and guide users

SketchVLM:视觉语言模型可以注释图像以解释思路并引导用户

Brandon Collins, Logan Bolton, Hung Huy Nguyen, Mohammad Reza Taesiri, Trung Bui, Anh Totti Nguyen

机构 * Auburn University(阿伯拉罕大学) Adobe Research(Adobe研究)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

AI总结 SketchVLM通过生成可编辑的SVG叠加图提升视觉推理和绘图任务的准确性与注释质量,实现高达28.5%的精度提升和1.48倍的注释质量提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00376 2026-04-29 cs.AI 57%

NeuroHex: A Brain-Inspired Hex Coordinate System to Enable Highly Computationally-Efficient World Models for Continuous Online-Adaptive Learning

NeuroHex:一种脑启发的六边形坐标系统,用于构建高效计算的世界模型以支持连续在线自适应学习

Quinn Jacobson, Joe Luo, Jingfei Xu, Shanmuga Venkatachalam, Kevin Wang, Dingchao Rong, John Paul Shen

机构 * NeuroAI Computer Architecture Lab (NCAL)(神经人工智能计算机架构实验室)

专题命中 视觉空间推理 :reasoning(abstract);分类 cs.AI

AI总结 NeuroHex通过六边形坐标系统实现高效世界模型,利用环索引和量化角编码等方法,降低计算成本并支持空间匹配,同时提供OSM2Hex工具将地图数据转换为该坐标系统,提升自主系统空间推理效率。

Comments This is an expanded version of the paper titled "NeuroHex: Highly Efficient Hex Coordinate System for Creating World Models to Enable Adaptive AI" published in the proceedings of the 2026 Neuro Inspired Computational Elements (NICE) [1] conference. This is an archival version of the paper and is currently under review for an ACM journal publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08323 2026-04-29 cs.CV 50%

Detecting Dental Landmarks from Intraoral 3D Scans: the 3DTeethLand challenge

从口内3D扫描中检测牙齿标志:3DTeethLand挑战

Achraf Ben-Hamadou, Nour Neifar, Ahmed Rekik, Oussama Smaoui, Firas Bouzguenda, Sergi Pujades, Niels van Nistelrooij, Shankeeth Vinayahalingam, Kaibo Shi, Hairong Jin, Youyi Zheng, Tibor Kubík, Oldřich Kodym, Petr Šilling, Kateřina Trávníčková, Tomáš Mojžiš, Jan Matula, Jeffry Hartanto, Xiaoying Zhu, Kim-Ngan Nguyen, Tudor Dascalu, Huikai Wu, and Weijie Liu, Shaojie Zhuang, Guangshun Wei, Yuanfeng Zhou

机构 * Department of Oral Maxillofacial Surgery, Radboud University Medical Center, Geert Grooteplein Zuid 10, 6525 GA Nijmegen, Netherlands organization= State Key Lab of CAD\&CG, Zhejiang University, Hangzhou, 310058 , country= China organization= Department of Computer Graphics Multimedia, Brno University of Technology , city= Brno , country= Czech Republic text= National Dental Centre Singapore organization= Guangxi Colleges Universities Key Laboratory of Intelligent Software ,country= Wuzhou University text= National University of Singapore organization= Department of Computer Science , country = University of Copenhagen organization= School of Software, Shandong University , country= China Centre de Recherche en Num\' e rique de Sfax, Laboratory of Signals, Systems, Artificial Intelligence Inria, Univ. Grenoble Alpes, CNRS, Grenoble INP, LJK, France

专题命中 视觉空间推理 :planning(abstract)

AI总结 本文提出3DTeethLand挑战,通过公开数据集评估牙齿标志检测算法,推动临床应用。挑战引入340个口内3D扫描数据,49支队伍参与,最终6支进入决赛,展示高精度与召回率的平衡方法。

Comments MICCAI 2024, 3DTeethLand, Challenge report, under review

详情

展开后加载摘要…

URL PDF HTML 收藏