arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2773 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2773 篇

2012.01526 2020-12-04 cs.CV cs.AI cs.RO 62%

From Goals, Waypoints & Paths To Long Term Human Trajectory Forecasting

Karttikeya Mangalam, Yang An, Harshayu Girase, Jitendra Malik

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 7 figures (including 2 GIFs)

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.08744 2020-10-23 cs.CV cs.AI cs.RO 62%

PLOP: Probabilistic poLynomial Objects trajectory Planning for autonomous driving

Thibault Buhet, Emilie Wirbel, Andrei Bursuc, Xavier Perrotton

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at CorRL 2020 (matching camera-ready version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.01719 2020-10-15 cs.CL cs.AI 62%

Grounded Language Learning Fast and Slow

Felix Hill, Olivier Tieleman, Tamara von Glehn, Nathaniel Wong, Hamza Merzic, Stephen Clark

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.12394 2020-01-17 cs.CL cs.CV cs.LG 62%

All-in-One Image-Grounded Conversational Agents

Da Ju, Kurt Shuster, Y-Lan Boureau, Jason Weston

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
1910.06315 2019-10-15 cs.CV cs.CL cs.LG 62%

Dynamic Attention Networks for Task Oriented Grounding

Soumik Dasgupta, Badri N. Patro, Vinay P. Namboodiri

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted ICCV 2019 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1901.02701 2019-01-11 cs.CV cs.LG cs.MM 62%

Guess What's on my Screen? Clustering Smartphone Screenshots with Active Learning

Agnese Chiatti, Dolzodmaa Davaasuren, Nilam Ram, Prasenjit Mitra, Byron Reeves, Thomas Robinson

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV、cs.MM

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.06150 2018-09-20 cs.RO cs.AI cs.CL cs.LG 62%

FollowNet: Robot Navigation by Following Natural Language Directions with Deep Reinforcement Learning

Pararth Shah, Marek Fiser, Aleksandra Faust, J. Chase Kew, Dilek Hakkani-Tur

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI

Comments 7 pages, 8 figures

Journal ref Third Workshop in Machine Learning in the Planning and Control of Robot Motion at ICRA, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.04012 2018-06-12 cs.CV cs.MM 62%

Hierarchy of GANs for learning embodied self-awareness model

Mahdyar Ravanbakhsh, Mohamad Baydoun, Damian Campo, Pablo Marin, David Martin, Lucio Marcenaro, Carlo S. Regazzoni

专题命中 多模态Agent :cross-modal(abstract);分类 cs.CV、cs.MM

Comments 2018 IEEE International Conference on Image Processing - ICIP'18. arXiv admin note: text overlap with arXiv:1806.02609

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.06922 2018-04-12 cs.CL cs.AI 62%

Emergent Translation in Multi-Agent Communication

Jason Lee, Kyunghyun Cho, Jason Weston, Douwe Kiela

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL、cs.AI

Comments Accepted to ICLR 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1709.10423 2017-10-02 cs.CL cs.AI cs.LG cs.RO 62%

Learning how to learn: an adaptive dialogue agent for incrementally learning visually grounded word meanings

Yanchao Yu, Arash Eshghi, Oliver Lemon

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CL、cs.AI

Comments 10 pages, RoboNLP Workshop from ACL Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.03673 2017-01-16 cs.AI cs.CV cs.LG cs.RO 62%

Learning to Navigate in Complex Environments

Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J. Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, Dharshan Kumaran, Raia Hadsell

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

Comments 11 pages, 5 appendix pages, 11 figures, 3 tables, under review as a conference paper at ICLR 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.07133 2016-05-24 cs.CL cs.CV cs.LG 62%

Towards Multi-Agent Communication-Based Language Learning

Angeliki Lazaridou, Nghia The Pham, Marco Baroni

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.CL

Comments 9 pages, manuscript under submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13278 2026-06-08 cs.CL cs.LG 版本更新 61%

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning

AutoTool: 面向智能体推理的动态工具选择与集成

Jiaru Zou, Ling Yang, Yunzhe Qi, Sirui Chen, Mengting Ai, Ke Shen, Jingrui He, Mengdi Wang

机构 * Nanyang Technological University(南洋理工大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL;multi-modal(comments)

AI总结 提出AutoTool框架,通过双阶段优化(SFT+RL轨迹稳定化和KL正则化Plackett-Luce排序)使大语言模型具备动态工具选择能力,在数学、科学、代码和多模态推理等任务上平均提升6.4%-7.7%。

Comments ICML2026; Best Paper Award at ICCV 2025 Workshop on Multi-Modal Reasoning for Agentic Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.01714 2022-09-23 cs.AI cs.LG 61%

On the Horizon: Interactive and Compositional Deepfakes

Eric Horvitz

专题命中 多模态Agent :multimodal(abstract,comments);分类 cs.AI

Comments CCC Blue Sky Ideas paper, published at the ACM International Conference on Multimodal Interaction (ICMI '22), November 7-11, 2022

详情
URL PDF HTML 收藏
2105.04633 2021-05-12 cs.CL 61%

Language Acquisition is Embodied, Interactive, Emotive: a Research Proposal

Casey Kennington

专题命中 多模态Agent :multimodal(abstract,comments);分类 cs.CL

Comments 6 pages, ICLR 2021 Embodied Multimodal Learning Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.02040 2020-01-08 eess.IV cs.CV cs.LG 61%

Robust Semantic Segmentation of Brain Tumor Regions from 3D MRIs

Andriy Myronenko, Ali Hatamizadeh

专题命中 多模态Agent :multimodal(abstract,comments);分类 cs.CV

Comments Accepted to 2019 International MICCAI Brainlesion Workshop -- Multimodal Brain Tumor Segmentation Challenge (BraTS) 2019. arXiv admin note: substantial text overlap with arXiv:1810.11654

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.11094 2018-10-29 cs.HC cs.CV 61%

The Role of Emotion in Problem Solving: First Results from Observing Chess

Thomas Guntz, James Crowley, Dominique Vaufreydaz, Raffaella Balzarini, Philippe Dessus

专题命中 多模态Agent :multimodal(abstract,journal_ref);分类 cs.CV

Journal ref ICMI 2018 - Workshop at 20th ACM International Conference on Multimodal Interaction, Oct 2018, Boulder, Colorado, United States. pp.1-13

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15549 2026-08-18 cs.RO cs.AI 新提交 57%

MistyPilot: Enabling Social-Robot Control through Multi-Agent LLM Skill Orchestration

MistyPilot:通过多智能体大语言模型技能编排实现社交机器人控制

Xiao Wang, Lu Dong, Ifeoma Nwogu, Srirangaraj Setlur, Venu Govindaraju

机构 * State University of New York at Buffalo(纽约州立大学布法罗分校)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 该研究提出多智能体大语言模型框架MistyPilot,可解释自然语言指令并编排Misty社交机器人技能,经评估其在多项任务上准确率高、方差低,用户反馈积极,代码将公开。

Comments Accepted at the ECCV 2026 ACVR Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15424 2026-08-18 cs.MA cs.AI cs.LG 新提交 57%

ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems

ETHOS:面向临床多智能体系统的模块化伦理框架

Rakesh Sharma, Sydney Pugh, Cameron Beeche, Pankhuri Singhal, Rachel Wu, Margaret Eby, Jeffrey Duda, James Gee, Kyra O'Brien, Hersh Sagreiya, Marina Serper, Victoria Gershuni, Angela Bradbury, Anurag Verma, Eric Eaton, Kevin B. Johnson, Walter Witschey

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 ETHOS是可与现有临床多智能体系统集成的模块化伦理框架,通过分层治理提升决策可靠性,将AI伦理原则转化为可部署的安全保障。

Comments Preprint of an article submitted for consideration in Pacific Symposium on Biocomputing \textcopyright\ 2027 World Scientific Publishing Company. \url{https://psb.stanford.edu/}

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15009 2026-08-18 cs.RO cs.CV 新提交 57%

ForceU-VLA: A Force-Aware Vision-Language-Action Model for Embodied Ultrasound Scanning

ForceU-VLA:用于实体超声扫描的力感知视觉-语言-动作模型

Xingzheng Wu, Cheng Zhang, Guihao Yan, Xifeng Hu, Zhi Liu, Qing Cai

机构 * Faculty of Computer Science and Technology, Ocean University of China(中国海洋大学计算机科学与技术学院) School of Information Science and Engineering, Shandong University(山东大学信息科学与工程学院) Innovation School of Artificial Intelligence, Hefei University of Technology(合肥工业大学人工智能创新学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 本文针对现有实体超声扫描方法的不足,提出力感知视觉-语言-动作模型ForceU-VLA,设计FUSFM与SAMM模块,构建ForceU-VLA-Data数据集,实验证实其可提升超声扫描的接触稳定性与压力调节能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21576 2026-08-18 cs.CL 版本更新 57%

Vision Language Models Cannot Plan, but Can They Formalize?

视觉语言模型(VLMs)无法规划,但它们能进行形式化吗?

Muyu He, Yuxi Zheng, Yuchen Liu, Zijian An, Bill Cai, Jiani Huang, Lifeng Zhou, Feng Liu, Ziyang Li, Li Zhang

机构 * Drexel University(德雷塞尔大学) University of Pennsylvania(宾夕法尼亚大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

AI总结 本研究提出5种VLM作为形式化器的流水线,评估后发现其表现优于端到端规划生成,较弱VLMs的瓶颈为物体关系视觉grounding,较强模型已克服该限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14028 2026-08-17 cs.RO cs.AI 新提交 57%

AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning

AdvDex:通过关节对齐动作与对抗学习从人类演示中学习灵巧操作

Zhiyue Zhao, Jingyi Wu, Hairuo Liu, Mingyu Liu, Liyang Li, Hengdi Zhang, Tong He, Zhengxue Cheng

机构 * Zhejiang University(浙江大学) Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学) Shanghai Jiao Tong University(上海交通大学) Paxini Tech(帕西尼科技)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 AdvDex是一种统一视觉-语言-动作框架,通过OmniShare数据集、JAAS动作空间与领域对抗学习,实现从人类和机器人演示中学习灵巧操作,提升跨实体泛化与技能迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13552 2026-08-17 cs.CV 版本更新 57%

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

PlayWorld:基于智能体玩家的长程目标世界模型基准测试

Kaixin Ding, Xi Chen, Minghong Cai, Zhiyuan Xu, Yiyang Wang, Yuxiang Lu, Junyi Li, Shuyang Chen, Yuan Gao, Xin Tao, Pengfei Wan, Hengshuang Zhao

机构 * The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) Zhejiang University(浙江大学) Kuaishou Technology(快手科技)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

AI总结 该研究针对现有世界模型跨模型公平比较的难题,推出含171个场景的PlayWorld基准,通过多模态智能体玩家从多维度评估9种先进世界模型,发现其长程交互式目标表现仍不可靠。

Comments project page: https://kxding.github.io/project/PlayWorld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01236 2026-08-17 cs.RO cs.AI 版本更新 57%

LLM-Advisor: An LLM Advisor for Cost-efficient Path Planning across Multiple Terrains

LLM-Advisor: 一种用于多地形高效路径规划的LLM基准

Ling Xiao, Toshihiko Yamasaki

机构 * Graduate School of Information Science and Technology, Hokkaido University(弘前大学信息科学与技术研究生院) Department of Information and Communication Engineering, The University of Tokyo(东京大学信息与通信工程系)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 LLM-Advisor利用大型语言模型优化多地形路径规划,提升路径成本效率,适用于现实场景。

Comments This paper has been accepted by IEEE Transactions on Automation Science and Engineering

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12674 2026-08-14 cs.AI 新提交 57%

Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy

Lines and Ladders:面向大规模零售价格分类的上下文感知多智能体框架

Ravi Teja Chunduri, Srikaran Reddy Boya, Deep Narayan Mishra, Ajay Kumar B, Karthik Kumaran, Pranay Kona

机构 * Walmart Global Tech(沃尔玛全球科技)

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

AI总结 针对大规模零售商品定价管理难题,提出上下文感知多智能体框架自动化构建Lines and Ladders价格分类,3智能体系统在Lines任务F1达0.83,在多品类数据上表现优异且已投入生产。

Comments 8 pages. Accepted in the Main Conference of IEEE ICMLA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11458 2026-08-13 cs.CV 新提交 57%

Multi-Agent Target-Existence Verification and Learned Mask Geometry Refinement: Winning Report of the MeViS-Text Track at the 8th LSVOS Challenge 2026

多智能体目标存在性验证与学习型掩码几何优化:2026年第8届LSVOS挑战赛MeViS-Text赛道获胜报告

Jungyoon Lee, Gyuil Lim, Doeon Kim, Seong-heum Kim

机构 * Soongsil University(崇实大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 本研究提出SSUPER方案,通过多智能体审计解耦存在性验证、StyleRefiner优化掩码几何,获2026年LSVOS挑战赛MeViS-Text赛道冠军,最终得分0.9081339614。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10976 2026-08-12 cs.AI 新提交 57%

XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving

XCoT-VLA:面向视觉-语言-动作自动驾驶的可执行思维链

Foundation Model Team, XPeng Inc

机构 * XPeng Inc(小鹏汽车)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 该研究提出XCoT-VLA,用可执行CoT令牌替代自然语言CoT,结合Reason FFN与Control FFN实现轨迹生成,在自动驾驶任务中降低了误差并满足实时规划要求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23399 2026-08-12 cs.AI 版本更新 57%

GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning

GAM-Agent:面向复杂视觉推理的博弈论与感知不确定性感知协作框架

Jusheng Zhang, Yijia Fan, Wenjun Lin, Ruiqi Chen, Haoyi Jiang, Wenhao Chai, Jian Wang, Keze Wang

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 该研究提出GAM-Agent博弈论多智能体框架,通过基础感知智能体与关键验证智能体的非零和博弈及不确定性感知协作,在四个视觉推理基准上显著提升了中小及强规模VLM的性能,为可靠可解释多模态推理提供了新路径。

Comments Accepted at NeurIPS 2025. Code available at https://github.com/jushengzhang/Gam-Agent

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09876 2026-08-11 cs.RO cs.AI 新提交 57%

Energy-Structured Latent World Models with Neural Time Fields for Physically Constistent Open-World Motion Planning

用于物理一致性开放世界运动规划的能量结构化潜在世界模型与神经时间场

Yapeng Liu, Yuanzhao Zhai, Bo Ding, Huaimin Wang, Lin Wang

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 针对开放世界运动规划的物理一致性挑战,提出能量结构化潜在世界模型ELWM,结合物理条件神经时间场PC-NTF,显著提升了运动预测精度与导航成功率,降低了碰撞率与残差。

Comments 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08960 2026-08-11 cs.AI 新提交 57%

Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression

阅读并非推理:弥合视觉-文本压缩中的智能体策略差距

Cheng Fan, Junyi Zhou, Tingzhang Luo, RongJian Xu, Qiyanhui Lu, Mingjian Zhu, Hanting Chen, Jianyuan Guo

专题命中 多模态Agent :cross-modal(abstract);分类 cs.AI

AI总结 该研究针对视觉-文本压缩导致的智能体策略差距,提出CAPS跨模态智能体策略自蒸馏框架,在SearchQA、ALFWorld等数据集上显著提升性能并大幅降低上下文成本。

详情

展开后加载摘要…

URL PDF HTML 收藏