arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2759 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2759 篇

2606.24112 2026-08-12 cs.AI 版本更新 83%

ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection

ReMMD: 面向多模态虚假信息检测的现实多语言多图像智能体验证

Chenhao Dang, Dantong Zhu, Jun Yang, Conghui He, Weijia Li

机构 * Shanghai Jiaotong University(上海交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Tsinghua University(清华大学) Central South University(中南大学) China Electronics Technology Group Corporation 15th Research Institute(中国电子科技集团公司第十五研究所)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 提出ReMMD框架,包含多语言多图像基准ReMMDBench和持久记忆验证器ReMMD-Agent,通过原子点分解和可重用证据集实现高效准确的多模态虚假信息检测。

Comments The project is available at https://dang-ai.github.io/ReMMD

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14029 2026-07-30 cs.CV 版本更新 83%

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management

POINTS-Seeker:从零开始训练多模态代理搜索模型

Yikun Liu, Yuan Liu, Haicheng Wang, Zhongyin Zhao, Le Tian, Xiao Zhou, Jiangchao Yao, Yanfeng Wang, Weidi Xie

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) CMIC, Shanghai Jiao Tong University(上海交通大学CMIC) WeChat AI, Tencent(腾讯WeChat AI)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出POINTS-Seeker模型,通过引入代理播种机制和V-Fold压缩方案,解决长周期交互中的证据检索问题,并在六个基准测试中超越现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15231 2026-06-16 cs.AI 新提交 83%

Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning

Visual-Seeker:通过主动视觉推理实现视觉原生多模态智能搜索

Zhengbo Zhang, Changtao Miao, Jinbo Su, Zhaowen Zhou, Chunxia Zhang, Xukai Wang, Ruiqi Liu, Kaiyuan Zheng, Jiansheng Cai, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan

机构 * School of Artificial Intelligence UCAS(中国科学院大学人工智能学院) Institute of Automation CAS(中国科学院自动化研究所) Ant Digital Technologies Ant Group(蚂蚁数字科技蚂蚁集团) RUC(中国人民大学) BIT(北京理工大学)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 提出Visual-Seeker,一种通过主动视觉推理进行视觉原生多模态深度搜索的智能体,在五个基准上达到最先进性能,甚至超越专有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.07402 2026-06-08 cs.CL 新提交 83%

M$^3$Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions

M$^3$Exam: 面向真实用户-智能体交互的多模态记忆基准

Zhengjun Huang, Wenxuan Liu, Zhoujin Tian, Wei Chen, Junle Chen, Yuqian Wu, Fangyuan Zhang, Qintian Guo, Xiaofang Zhou

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) Beijing University of Chemical Technology(北京化工大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Beijing Institute of Technology (Zhuhai)(北京理工大学(珠海)) Tencent Hy(腾讯(深圳)) Peng Cheng Laboratory(鹏城实验室)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 提出M$^3$Exam基准,用于评估多模态大语言模型在真实用户-智能体交互中的跨模态推理和隐式信息推断能力,并设计M$^3$Proctor方法通过按需处理视觉源提升准确率13%,同时降低索引构建时间和检索token超70%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.03951 2026-06-03 cs.CV 83%

Demo2Tutorial: From Human Experience to Multimodal Software Tutorials

Demo2Tutorial:从人类经验到多模态软件教程

Zechen Bai, Zhiheng Chen, Yiqi Lin, Kevin Qinghong Lin, Difei Gao, Xiangwu Guo, Xin Wang, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室)

专题命中 多模态Agent :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 提出Demo2Tutorial框架,通过屏幕录制和交互日志将人类经验解析为结构化多模态教程,用于人类学习和GUI智能体训练,实验证明其生成质量超越人工教程并提升任务效率。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.18652 2026-05-19 cs.CV 83%

MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents

MementoGUI: 学习代理多模态记忆控制以实现长周期GUI代理

Ziyun Zeng, Hang Hua, Bocheng Zou, Mu Cai, Rogerio Feris, Jiebo Luo

机构 * University of Rochester(罗切斯特大学) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 多模态Agent :multimodal(title);MLLM(abstract,abstract_cn);分类 cs.CV

AI总结 本文提出MementoGUI,一种学习代理多模态记忆控制框架,用于提升长周期GUI代理的任务状态维持能力,通过模块化记忆控制和可扩展的数据管道提高记忆检索和决策效率。

Comments Preprint, 15 pages, 4 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17336 2026-05-19 cs.RO cs.CV eess.SP 83%

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms

基于触觉的多模态融合在具身智能中的应用:视觉、语言和接触驱动范式的综述

Zhixiang Cao, Di Tian, Runwei Guan, Yanzhou Mu, Xiaolou Sun, Shaofeng Liang, Daizong Liu, Tao Huang, Yutao Yue, Henghui Ding, Bin Fang, Alex Zhou, Qing-Long Han, Hui Xiong

机构 * School of Electronic Science and Engineering, Xi’an Jiaotong University, China(西安交通大学电子科学与技术学院) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou), China(香港科技大学(广州)人工智能研究所) State Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学新型软件技术国家重点实验室) Purple Mountain Laboratory, China(紫金山实验室) Institute for Math & AI, Wuhan University, China(武汉大学数学与人工智能学院) Centre for AI and Data Science Innovation and the School of Science and Engineering, James Cook University, Australia(詹姆斯库克大学人工智能与数据科学创新中心及科学与工程学院) School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China(北京邮电大学人工智能学院) Institute of Big Data, Fudan University, China(复旦大学大数据研究院) Linkerbot (Beijing) Technology Co., Ltd, China(北京链动科技有限公司) School of Engineering, Swinburne University of Technology, Melbourne(斯威本技术大学工程学院)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文综述了多模态触觉融合在具身智能中的研究,探讨了如何通过整合视觉、语言和触觉信息来提升物理交互与语义推理的结合,提出了一种分层的分类体系,并总结了当前的研究挑战和未来方向。

Comments 20 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06595 2026-05-08 cs.RO cs.AI cs.LG cs.MA 83%

Cross-Modal Navigation with Multi-Agent Reinforcement Learning

跨模态导航与多智能体强化学习

Shuo Liu, Xinzichen Li, Christopher Amato

机构 * Khoury College of Computer Sciences(计算机科学学院)

专题命中 多模态Agent :cross-modal(title,abstract);multi-modal(abstract);分类 cs.AI

AI总结 本文提出CRONA框架,通过多智能体强化学习实现跨模态导航,利用辅助信念和集中式多模态批评者提升协作效率,实验表明多智能体方法在视觉-听觉导航中优于单智能体基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20816 2026-04-22 cs.CL 83%

Rethinking Information Synthesis in Multimodal Question Answering A Multi-Agent Perspective

重新思考多模态问答中的信息合成:多智能体视角

Krishna Singh Rajput, Tejas Anvekar, Chitta Baral, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文提出多智能体框架MAMMQA,通过分解查询、跨模态推理和整合答案,提升多模态问答的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12213 2026-04-16 cs.AI cs.MA cs.SE 83%

Modality-Native Routing in Agent-to-Agent Networks: A Multimodal A2A Protocol Extension

代理到代理网络中的模态本原路由:一种多模态A2A协议扩展

Vasundra Srinivasan

机构 * AI Architect, Author—Data Engineering for Multimodal AI (O’Reilly)(人工智能架构师,作者—多模态AI的数据工程(O’Reilly)) Stanford School of Engineering (April 2026)(斯坦福大学工程学院(2026年4月))

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 本文提出MMA2A架构,通过多模态本原路由提升任务准确率,其在跨模态任务中表现优于文本瓶颈基线,尤其在视觉依赖任务中效果显著,但增加了1.8倍的延迟。

Comments 14 pages, 4 figures (TikZ). PDFLaTeX. Supplementary code and experiment artifacts: https://github.com/vasundras/modality-native-routing-a2a-protocol

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07956 2026-04-13 cs.AI 83%

MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems

MONETA:通过地理信息和多智能体系统进行多模态行业分类

Arda Yüksel, Gabriel Thiem, Susanne Walter, Patrick Felka, Gabriela Alves Werb, Ivan Habernal

机构 * Trustworthy Human Language Technologies(可信人类语言技术实验室) Technical University of Darmstadt, Germany(德国达姆施塔特工业大学) Deutsche Bundesbank(德国联邦银行) Frankfurt University of Applied Sciences, Germany(德国法兰克福应用技术大学) Research Center for Trustworthy Data Science and Security, Ruhr University Bochum, Germany(德国波鸿鲁尔大学可信数据科学与安全研究中心)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 本文提出MONETA,首个基于文本和地理空间数据的多模态行业分类基准,利用多智能体系统提升分类精度,实现62.10%和74.10%的分类性能。

Comments Accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19420 2026-04-09 cs.AI 83%

Commander-GPT: Dividing and Routing for Multimodal Sarcasm Detection

Commander-GPT:多模态讽刺检测的分工与路由

Yazhou Zhang, Chunwang Zou, Bo Wang, Jing Qin, Prayag Tiwari

专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

AI总结 本文提出Commander-GPT框架,通过分工协作的LLM代理团队,结合三种指挥官类型,提升多模态讽刺检测性能,实验结果显示在F1分数上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03586 2026-04-07 cs.CL 83%

MultiPress: A Multi-Agent Framework for Interpretable Multimodal News Classification

MultiPress:一种用于可解释多模态新闻分类的多智能体框架

Tailong Luo, Hao Li, Rong Fu, Xinyue Jiang, Huaxuan Ding, Yiduo Zhang, Zilin Zhao, Simon Fong, Guangyin Jin, Jianyuan Ni

机构 * New York Institute of Technology(纽约理工学院) University of Arizona(亚利桑那大学) University of Macau(澳门大学) Peking University(北京大学) Juniata College(朱尼亚塔学院)

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL

AI总结 本文提出MultiPress多智能体框架,通过多模块协作和检索增强推理提升多模态新闻分类的准确性和可解释性。

Comments Accepted in International Joint Conference on Neural Networks (IJCNN) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29902 2026-04-01 cs.AI 83%

ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation

ATP-Bench: 向多模态大语言模型交错生成的代理工具规划迈进

Yinuo Liu, Zi Qian, Heng Zhou, Jiahao Zhang, Yajie Zhang, Zhihang Li, Mengyu Zhou, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang

机构 * Qwen Large Model Application Team, Alibaba(阿里巴巴通义千问大模型应用团队) Huazhong University of Science and Technology(华中科技大学) Zhejiang University(浙江大学)

专题命中 多模态Agent :MLLM(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 针对多模态大语言模型交错生成中事实性与创造性难以统一的问题,提出ATP-Bench基准,包含7702个问题-答案对,评估代理工具规划能力,揭示模型在交错规划中的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29161 2026-04-01 cs.AI 83%

Webscraper: Leverage Multimodal Large Language Models for Index-Content Web Scraping

Webscraper:利用多模态大语言模型进行索引-内容网页抓取

Guan-Lun Huang, Yuh-Jzer Joung

机构 * Dept. of Information Management, National Taiwan University, Taipei, Taiwan(国立台湾大学资讯管理学系,台北,台湾)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 本文提出Webscraper框架,利用多模态大语言模型自动导航交互界面并提取结构化数据,通过五阶段提示和定制工具提升动态网站抓取准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26908 2026-03-31 cs.CV 83%

FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition

FusionAgent:一种具有动态模型选择的多模态代理用于人体识别

Jie Zhu, Xiao Guo, Yiyang Su, Anil Jain, Xiaoming Liu

机构 * Michigan State University(密歇根州立大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 FusionAgent通过动态模型选择提升人体识别鲁棒性,利用多模态大语言模型进行样本特定的模型选择,结合ACT分数融合方法,实验表明其在效率和性能上优于现有方法。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22460 2026-03-26 cs.AI 83%

GeoSketch: A Neural-Symbolic Approach to Geometric Multimodal Reasoning with Auxiliary Line Construction and Affine Transformation

GeoSketch: 一种基于神经符号的方法,用于几何多模态推理中的辅助线构造与仿射变换

Shichao Weng, Zhiqiang Wang, Yuhua Zhou, Rui Lu, Ting Liu, Zhiyang Teng, Xiaozhang Liu, Hanmeng Liu

机构 * Fudan University(复旦大学) IFLYTEK CO.LTD(若lytek有限公司) Zhejiang University(浙江大学) The Hong Kong University of Science and Technology(香港科学与技术大学) National University of Defense Technology(国防科技大学) Hainan University(海南大学)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 GeoSketch通过整合感知、符号推理和绘图动作模块,实现动态几何推理,提升多模态推理的准确性和解决问题的成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18722 2026-03-11 cs.AI 83%

VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft

VistaWise: 构建低成本代理的跨模态知识图谱用于Minecraft

Honghao Fu, Junlong Ren, Qi Chai, Deheng Ye, Yujun Cai, Hao Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Queensland(昆士兰大学) Tencent(腾讯)

专题命中 多模态Agent :cross-modal(title,abstract);multimodal(abstract);分类 cs.AI

AI总结 VistaWise通过整合跨模态知识图谱和专用模型,实现低成本、高效率的Minecraft代理构建。

Comments Accepted by EMNLP 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08240 2026-03-10 cs.CV 83%

SiMO: Single-Modality-Operable Multimodal Collaborative Perception

SiMO: 单模态可操作的多模态协作感知

Jiageng Wen, Shengjie Zhao, Bing Li, Jiafeng Huang, Kenan Ye, Hao Deng

机构 * Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学智能自主系统上海研究院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院) School of Mechatronic Engineering and Automation, Shanghai University(上海大学机械电子工程与自动化学院)

专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 SiMO通过单模态可操作的多模态协作感知方法,解决多模态特征融合中的语义不匹配问题,提升协作感知性能。

Comments Accepted to ICLR 2026. This arXiv version includes an additional appendix (Appendix 15) containing further philosophical discussion not included in the official ICLR peer-reviewed version

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19255 2026-03-06 cs.LG cs.AI 83%

VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use

VTool-R1: 通过在多模态工具使用上的强化学习使VLMs学会通过图像思考

Mingyuan Wu, Jingcheng Yang, Jize Jiang, Meitang Li, Kaizhuo Yan, Hanchao Yu, Minjia Zhang, Chengxiang Zhai, Klara Nahrstedt

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Michigan Ann Arbor(密歇根大学安娜堡分校) Independent Researcher(独立研究者)

专题命中 多模态Agent :multimodal(title,abstract);multi-modal(abstract);分类 cs.AI

AI总结 VTool-R1通过强化学习训练视觉语言模型生成多模态思考链,提升其通过图像进行推理的能力。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08211 2026-02-10 cs.CV 83%

Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension

Chain-of-Caption: 无需训练提升多模态大语言模型的指称表达理解

Yik Lung Pang, Changjae Oh

机构 * Queen Mary University of London(伦敦女王学院)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出无需训练的Chain-of-Caption框架,通过结合多种上下文提升多模态大语言模型在指称表达理解任务中的性能。

Comments 4 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01854 2026-02-03 cs.CV 83%

Fact or Fake? Assessing the Role of Deepfake Detectors in Multimodal Misinformation Detection

事实还是假象?评估深度伪造检测器在多模态虚假信息检测中的作用

A S M Sharifuzzaman Sagar, Mohammed Bennamoun, Farid Boussaid, Naeha Sharif, Lian Xu, Shaaban Sahmoud, Ali Kishk

机构 * The University of Western Australia(西澳大学) Fatih Sultan Mehmet Vakif University(法提赫·苏丹·梅赫梅特·瓦基夫大学) Aljazeera Media Network Investigative Department(半岛电视台调查部门)

专题命中 多模态Agent :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 研究评估了深度伪造检测器在多模态虚假信息检测中的作用,发现其独立价值有限,而证据驱动的事实核查系统表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03215 2026-01-14 cs.AI 83%

COSINT-Agent: A Knowledge-Driven Multimodal Agent for Chinese Open Source Intelligence

COSINT-Agent: 一种面向中文开源情报的知识驱动多模态智能体

Wentao Li, Congcong Wang, Xiaoxiao Cui, Zhi Liu, Wei Guo, Lizhen Cui

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

AI总结 COSINT-Agent通过整合多模态大语言模型与实体-事件-场景知识图谱,提升中文开源情报的多模态推理能力与情报生成效率。

Comments This manuscript (arXiv:2503.03215) is being withdrawn at the supervisor's request. The content is preliminary and needs further internal revision and approval before public release. We will resubmit a revised version after completion. Apologies for the inconvenience

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09245 2025-12-18 cs.CV 83%

DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving

DriveMLM: 通过行为规划状态对齐多模态大语言模型以实现自动驾驶

Erfei Cui, Wenhai Wang, Zhiqi Li, Jiangwei Xie, Haoming Zou, Hanming Deng, Gen Luo, Lewei Lu, Xizhou Zhu, Jifeng Dai

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) Beijing National Research Center for Information Science and Technology(北京信息科学与技术国家研究中心)

专题命中 多模态Agent :multi-modal(title);multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 DriveMLM通过多模态大语言模型对齐行为规划状态,提升自动驾驶系统的决策能力与闭环性能。

Comments Accepted to Visual Intelligence

Journal ref Visual Intelligence, Volume 3, article number 22, (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12799 2025-12-16 cs.CV 83%

DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning

DrivePI: 基于空间感知的4D MLLM用于统一自动驾驶理解、感知、预测与规划

Zhe Liu, Runhui Huang, Rui Yang, Siming Yan, Zining Wang, Lu Hou, Di Lin, Xiang Bai, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) Yinwang Intelligent Technology Co. Ltd.(英维智能科技有限公司) Tianjin University(天津大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 多模态Agent :MLLM(title,abstract);multi-modal(abstract);分类 cs.CV

AI总结 DrivePI是一种基于空间感知的4D MLLM,用于统一自动驾驶的理解、感知、预测和规划,通过端到端优化实现多任务并行,提升性能并减少碰撞率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06396 2025-12-09 cs.CR cs.AI 83%

AgenticCyber: A GenAI-Powered Multi-Agent System for Multimodal Threat Detection and Adaptive Response in Cybersecurity

AgenticCyber: 一种基于生成式AI的多智能体系统,用于多模态威胁检测与自适应响应在网络安全中

Shovan Roy

机构 * TNTech

专题命中 多模态Agent :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI

AI总结 AgenticCyber通过生成式AI驱动的多智能体系统,实现多模态威胁检测与自适应响应,提升网络安全性能和态势感知能力。

Comments 6 pages for IEEE conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02981 2025-12-03 cs.CV 83%

InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration

InEx:通过内省与跨模态多智能体协作缓解幻觉

Zhongyu Yang, Yingfang Yuan, Xuanming Jiang, Baoyi An, Wei Pang

专题命中 多模态Agent :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 InEx通过内省推理和跨模态多智能体协作,自主缓解大型语言模型的幻觉问题,实验表明其在多个基准上表现优异。

Comments Published in AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10017 2025-11-14 cs.CV 83%

AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models

Xinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu, Zhen Li, Na Zhao

机构 * University of Science and Technology of China(科学技术大学) Singapore University of Technology and Design(新加坡科技设计大学) Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00940 2025-11-04 cs.RO cs.AI 83%

URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model

Zhe Li, Xiang Bai, Jieyu Zhang, Zhuangzhe Wu, Che Xu, Ying Li, Chengkai Hou, Shanghang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) University of Washington(华盛顿大学)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.AI

Comments Accepted to the 39th Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11498 2025-10-14 cs.LG cs.CL 83%

ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding

Yuhang Li, Chenchen Zhang, Ruilin Lv, Ao Liu, Ken Deng, Yuanxing Zhang, Jiaheng Liu, Wiggin Zhou, Bo Zhou

机构 * LLM Department, Tencent(腾讯大语言模型部门) Peking University(北京大学) Nanjing University(南京大学)

专题命中 多模态Agent :multimodal(title,abstract);MLLM(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏