arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 2777 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态Agent 2777 篇

1802.03116 2018-02-12 cs.CL 57%

Zero-Resource Neural Machine Translation with Multi-Agent Communication Game

Yun Chen, Yang Liu, Victor O. K. Li

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

Comments Published at AAAI-18

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.05376 2017-02-20 cs.AI cs.DM stat.ML 57%

Towards a Unified Taxonomy of Biclustering Methods

Dmitry I. Ignatov, Bruce W. Watson

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

Comments http://ceur-ws.org/Vol-1552/

Journal ref Russian and South African Workshop on Knowledge Discovery Techniques Based on Formal Concept Analysis (RuZA 2015), November 30 - December 5, 2015, Stellenbosch, South Africa, In CEUR Workshop Proceedings, Vol. 1552, p. 23-39

详情

展开后加载摘要…

URL PDF HTML 收藏
1609.01472 2016-09-07 cs.CY cs.AI 57%

OpenTripPlanner, OpenStreetMap, General Transit Feed Specification: Tools for Disaster Relief and Recovery

Chelcie Narboneta, Kardi Teknomo

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments 6 pages, Narboneta, C. G. and Teknomo, K. (2014) OpenTripPlanner, OpenStreetMap, General Transit Feed Specification: Tools for Disaster Relief and Recovery, Proceeding of the 7th IEEE International Conference Humanoid, Nanotechnology, Information Technology Communication and Control, Environment and Management (HNICEM) 12-16 November 2014 Hotel Centro, Puerto Princesa, Palawan, Philippines

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.03845 2016-08-15 cs.RO cs.AI 57%

Traversing Environments Using Possibility Graphs for Humanoid Robots

Michael X. Grey, Aaron D. Ames, C. Karen Liu

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments Submitted to the International Workshop on the Algorithmic Foundations of Robotics (2016)

详情

展开后加载摘要…

URL PDF HTML 收藏
1405.5164 2014-05-21 cs.CV 57%

Multi-ellipses detection on images inspired by collective animal behavior

Erik Cuevas, Maurici Gonzalez, Daniel Zaldivar, Marco Perez

专题命中 多模态Agent :multi-modal(abstract);分类 cs.CV

Comments 21 pages

Journal ref Neural Computing and Applications, 24(5), (2014), 1019-1033

详情

展开后加载摘要…

URL PDF HTML 收藏
1301.0216 2013-01-03 cs.AI 57%

Applying Strategic Multiagent Planning to Real-World Travel Sharing Problems

Jan Hrnčíř, Michael Rovatsos

专题命中 多模态Agent :multi-modal(abstract);分类 cs.AI

Comments 7th International Workshop on Agents in Traffic and Transportation, AAMAS, 2012

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.13357 2026-07-24 cs.HC 版本更新 56%

TANDE: Disentangling Verbal and Nonverbal Backchannels in Emotional AI-Avatar Conversations with Young Adults

TANDE:在与年轻人的情感人工智能-虚拟化身对话中区分言语和非言语反馈渠道

Ann-Kareen Gedeus, Jack Good, Nadine Wagener, Angelique Taylor

专题命中 多模态Agent :multimodal(abstract,comments)

AI总结 研究在与年轻人的情感人工智能-虚拟化身对话中反馈渠道模式的影响,引入TANDE这个由LLM驱动的ECA,通过实验探讨其对融洽关系、同理心和参与度的作用及性别差异,得出相关设计启示。

Comments This paper has been accepted for publication at the 28th ACM International Conference on Multimodal Interaction (ICMI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16843 2026-08-18 cs.RO 新提交 50%

Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation

基于基础模型的具身智能体的安全性:攻击面、攻击、防御与评估

Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao

机构 * Wuhan University(武汉大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 该研究以信任边界为核心,针对基于基础模型的具身智能体安全,划分了五个层级与十二个攻击面,分析了58种攻击、61种防御的现状,指出部分领域研究不足并提出开放挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15198 2026-08-18 stat.ML cs.LG physics.comp-ph 新提交 50%

Identifying parameter couplings and uncertainties of mixed-noise stochastic systems via full-covariance Gaussian mixture network

通过全协方差高斯混合网络识别混合噪声随机系统的参数耦合与不确定性

Xiaolong Wang, Xiangwen Hao, Jing Feng, Yuanyuan Liu, Yong Xu

专题命中 多模态Agent :multi-modal(abstract)

AI总结 研究针对混合噪声随机系统参数识别的难点,提出PENN-GMD神经网络,采用全协方差高斯混合分布,经五个数值示例验证可准确恢复似然分布、捕捉参数耦合,为复杂随机系统参数识别提供实用工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18518 2026-08-18 cs.LG stat.ME stat.ML 版本更新 50%

Measuring the Prevalence of Policy Violating Content with ML Assisted Sampling and LLM Labeling

利用机器学习辅助抽样和大语言模型标注测量违规内容的普及率

Attila Dobi, Aravindh Manickavasagam, Benjamin Thompson, Xiaohan Yang, Faisal Farooq

机构 * Pinterest

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出了一种基于机器学习和大语言模型的系统,用于高效测量违反政策内容的普及率,通过概率抽样和多模态标注提升测量准确性。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22983 2026-08-18 cs.RO 版本更新 50%

Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives

基础模型时代中的具身机器人操作:规划与学习视角

Shuanghao Bai, Wenxuan Song, Jiayi Chen, Yuheng Ji, Zhide Zhong, Jin Yang, Han Zhao, Wanqi Zhou, Zhe Li, Pengxiang Ding, Cheng Chi, Chang Xu, Xiaolong Zheng, Donglin Wang, Haoang Li, Shanghang Zhang, Badong Chen

机构 * Xi’an Jiaotong Univeristy(西安交通大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Chinese Academy of Sciences(中国科学院) Westlake University(西湖大学) Zhejiang University(浙江大学) University of Sydney(悉尼大学) BAAI(百度人工智能研究院) Peking University(北京大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文探讨了基础模型时代机器人操作的规划与学习方法,分析了高层推理与低层控制的统一框架,并提出了未来研究方向。

Comments This work is a re-architected core derived from the full survey (arXiv:2510.10903), refined to highlight the most central themes and representative studies

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10903 2026-08-18 cs.RO 版本更新 50%

Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey

面向机器人操作的统一理解:一项综合调查

Shuanghao Bai, Wenxuan Song, Jiayi Chen, Yuheng Ji, Zhide Zhong, Jin Yang, Han Zhao, Wanqi Zhou, Wei Zhao, Zhe Li, Pengxiang Ding, Cheng Chi, Haoang Li, Chang Xu, Xiaolong Zheng, Donglin Wang, Shanghang Zhang, Badong Chen

专题命中 多模态Agent :multimodal(abstract)

AI总结 本调查针对机器人操作这一具身智能核心挑战,提出方法的统一分类与瓶颈分类,为机器人操作研究提供了全面的路线图与结构化参考。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14131 2026-08-17 cs.SE 新提交 50%

LegacyWorld: Atomicity-Aware Evaluation of GUI Agents for Legacy Workflows

LegacyWorld:面向遗留工作流的GUI智能体的原子性感知评估

Thilo Reintjes, Sivajeet Chand, Derui Zhu, Sushant Kumar Pandey, Alexander Pretschner

专题命中 多模态Agent :multimodal(abstract)

AI总结 该研究开发LegacyUse框架以自动化遗留工作流,构建含28个Windows GUI工作流的基准,采用原子性评估6款计算机使用智能体,发现不同操作特征并提出相关首要要求。

Comments Accepted for publication in the Industry Track of the 42nd IEEE International Conference on Software Maintenance and Evolution (ICSME 2026), 14-18 September 2026, Benevento, Italy

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12763 2026-08-14 cs.CE 新提交 50%

ARIES-Mission2: A Zero-Shot Vision-Language-Action Framework for Fast Large-Scale Aerial Mission Generation

ARIES-Mission2:用于快速大规模空中任务生成的零样本视觉-语言-动作框架

Junhao Wei, Yanxiao Li, Haochen Li, Yifu Zhao, Dexing Yao, Baili Lu, Zikun Li, Yapeng Wang, Sio-Kei Im, Dingcheng Yang, Xu Yang

专题命中 多模态Agent :multimodal(abstract)

AI总结 ARIES-Mission2是解耦视觉语义感知与路径优化的零样本VLA框架,在UAV基准上,其飞行距离更短、任务生成速度更快,且TSP模块可扩展性佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10252 2026-08-12 stat.ME stat.AP 新提交 50%

Joint return levels of maximum temperature and minimum relative humidity by combining copulas with an extreme value framework for bimodal data

结合Copula与适用于双峰数据的极值框架,分析最高气温与最小相对湿度的联合重现水平

Beatriz G da Cruz Albernaz, Cira E G Otiniano, Carolyne Soares de Brito, Fidel E C Morales, Enzo Porto Brasil

专题命中 多模态Agent :multimodal(abstract)

AI总结 本研究结合Copula与适用于双峰数据的极值框架,分析巴西利亚最高气温与最小相对湿度的联合重现水平,为当地气候风险应对提供关键依据。

Comments 19 pages, 13 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09771 2026-08-11 cs.RO 新提交 50%

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

SLIM-0.5B:学习机器人操作的动作基础预测隐变量

Jingkai Wang, Zihan Tang, Gu Zhang, Mingyu Cao, Jiapeng Chen, Jingjiao Zhao, Xiansheng Chen, Pengwei Wang, Lemao Liu, Dejing Dou

机构 * Fudan University(复旦大学) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Tsinghua University(清华大学) Renmin University of China(中国人民大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 SLIM-0.5B是0.5B参数的紧凑机器人操作策略,通过自监督掩码轨迹预测学习动作基础预测隐变量,性能优于或匹配大型基线,参数量少、推理延迟低、内存占用小。

Comments 18 pages, 11 figures. Project page: https://kzz1031.github.io/slim-project-page/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09410 2026-08-11 cs.RO 新提交 50%

Skills in Weights, Memory in Code: Hybrid Learning for Memory-Dependent Robot Manipulation

权重中的技能,代码中的记忆:面向依赖记忆的机器人操作的混合学习

Yunhao Zhao, Zhenyang Ni, Haoyang Chen, Ruohan Zhang, Qi Zhu

专题命中 多模态Agent :multimodal(abstract)

AI总结 针对现实机器人操作的非马尔可夫性,提出混合学习框架 HyMeS,结合编码智能体与马尔可夫 VLA,在 RoboMemArena 上显著提升操作成功率,实现数据高效的组合泛化。

Comments 9 pages, 4 figures, and 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.09098 2026-08-11 cs.RO 新提交 50%

UnsDrive: Towards Robust End-to-End Autonomous Driving in Unstructured Scenes

UnsDrive:面向非结构化场景的鲁棒端到端自动驾驶

Nanxin Zeng, Ruiqi Song, Xiangyu Guo, Baiyong Ding, Yunfeng Ai

机构 * University of Chinese Academy of Sciences(中国科学院大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Waytous Inc.(文远知行公司)

专题命中 多模态Agent :multimodal(abstract)

AI总结 针对非结构化采矿环境自动驾驶泛化差的问题,提出UnsDrive规划器,结合未知感知占用表示与流匹配规划器,引入专用损失和评分器,辅以MineLoop模拟器验证,性能优于基线。

Comments 9 pages, 4 figures, conference

Journal ref the 34th ACM International Conference on Multimedia, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07002 2026-08-10 cs.RO 新提交 50%

A Haptic Robot Finger Designed for Guqin Instrument Playing

用于古琴演奏的触觉机器人手指

Tianwei Zhang, Hanming Yan, Yang Yang. Ziya Wang

机构 * The Shenzhen Institute of Artificial Intelligence and Robotics for Society(深圳市人工智能与机器人研究院) The Chinese University of Hong Kong - Shenzhen(香港中文大学(深圳)) School of Physics and Optoelectronic Engineering, Shenzhen University(深圳大学物理与光电工程学院) Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

专题命中 多模态Agent :multimodal(abstract)

AI总结 该研究设计了一款模仿人类手指的高精度触觉传感机器人指尖,以古琴为验证场景,开展了古琴弦接触相关任务的验证,将触觉传感与机器人技术结合,助力世界遗产保护与文化传播。

Comments Accepted by IEEE Transactions on Haptics

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04755 2026-08-06 cs.CR 新提交 50%

"Allow" to Achieve, Over-Privileged Inadvertently: The Unintended Cost of Task-Completion-Driven Pop-up Decisions in Mobile GUI Agents

“允许”达成目标:移动GUI智能体中任务完成驱动的弹窗决策的意外代价

Dongsheng Chen, Yuxuan Li, Guanhua Chen, Jiaxin Zhang, Xiangyu Zhao, Lei Ma, Xin Yao, Xuetao Wei

专题命中 多模态Agent :multimodal(abstract)

AI总结 研究移动GUI智能体的权限素养,发现其存在应用信任偏差、任务优先级覆盖等问题,提示干预效果不一,建议将任务执行与权限授权分离。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03556 2026-08-05 cs.RO 新提交 50%

Human Centric Embodied Intelligence for Soft Wearable Robotics

面向软质可穿戴机器人的以人为中心的具身智能

Rainier Natividad, Raye Chen-Hua Yeow

专题命中 多模态Agent :multimodal(abstract)

AI总结 本综述提出以人为中心的具身智能(HCEI)概念,引入感知-认知-执行-增强(PCAA)框架,整合多领域进展以指导软质可穿戴机器人向个性化、可预测的以人为中心方向发展。

Comments Review article; 48 pages, 6 figures, 4 tables, and 2 supplementary tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02959 2026-08-05 cs.MA 新提交 50%

SABRE: A Multi-Agent Approach for Selecting Out-of-Distribution Detectors Under a Budget

SABRE:预算约束下选择分布外检测器的多智能体方法

Mary Wisell, Salimeh Sekeh

专题命中 多模态Agent :multimodal(abstract)

AI总结 SABRE是一种多智能体方法,可在预算约束下针对视觉-语言模型的分布外检测,自动选择各领域最优检测器,解决固定检测器跨领域失效的问题,实现可靠检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25459 2026-08-05 cs.RO 版本更新 50%

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning

GS-Playground:一种高吞吐量的视觉真实模拟器用于基于视觉的机器人学习

Yufei Jia, Heng Zhang, Ziheng Zhang, Junzhe Wu, Mingrui Yu, Zifan Wang, Dixuan Jiang, Zheng Li, Chenyu Cao, Zhuoyuan Yu, Xun Yang, Haizhou Ge, Yuchi Zhang, Jiayuan Zhang, Zhenbiao Huang, Tianle Liu, Shenyu Chen, Jiacheng Wang, Bin Xie, Xuran Yao, Xiwa Deng, Guangyu Wang, Jinzhi Zhang, Lei Hao, Zhixing Chen, Yuxiang Chen, Anqi Wang, Hongyun Tian, Yiyi Yan, Zhanxiang Cao, Yizhou Jiang, Hanyang Shao, Yue Li, Lu Shi, Bokui Chen, Wei Sui, Hanqing Cui, Yusen Qin, Ruqi Huang, Lei Han, Tiancai Wang, Guyue Zhou

机构 * THU(清华大学) Motphys Dexmal DISCOVER Robotics HKUST(GZ)(香港科技大学(广州)) BIT(北京理工大学) NUS(新加坡国立大学) HITSZ(哈尔滨工业大学) XJTU(西安交通大学) NJU(南京大学) SJTU(上海交通大学) D-Robotics

专题命中 多模态Agent :multi-modal(abstract)

AI总结 GS-Playground通过高吞吐量的多模态模拟框架,结合高效物理引擎和3DGS渲染管道,提升视觉强化学习的效率,降低大规模视觉RL的门槛,实现感知与物理的无缝衔接。

Comments Robotics: Science and Systems 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29113 2026-08-03 cs.HC 新提交 50%

Designing a digital word-learning intervention with neurodiverse children: Experiences and ideas from children with developmental language disorder

面向神经多样性儿童的数字词汇学习干预方案设计:发育性语言障碍儿童的经验与思路

Rafiah Ansari, Rina R. Wehbe, Lizbeth Olivia Escobedo

专题命中 多模态Agent :multimodal(abstract)

AI总结 该研究针对发育性语言障碍儿童,采用定制化参与式设计方法,设计词汇学习干预方案,明确了多模态等关键设计要求,为相关干预开发提供了宝贵见解。

Comments 56 pages, 7 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07792 2026-08-03 quant-ph 50%

Wigner and friends, a map is not the territory! Contextuality in multi-agent paradoxes

Sidiney B. Montanhano

专题命中 多模态Agent :multi-modal(abstract)

Comments Minor corrections. 12 pages, 2 figures, 3 tables; 20 pages with appendix and references

Journal ref Foundations of Science, 30(4), 971-1002 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28399 2026-07-31 cs.LG 新提交 50%

Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees

为什么GUI智能体正确但延迟?基于决策时间关键路径的解码,用预编译策略树测试

Zihan Dong, Rui Qian, Qishi Zhan, Dongshen Peng, Kaixin Li, Yu Li

专题命中 多模态Agent :multimodal(abstract)

AI总结 针对GUI智能体因决策时间关键路径的自回归解码延迟导致瞬态事件失败的问题,提出AAPT策略树方法,提升了决策窗口内的动作成功率,验证了分支路由为关键瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28055 2026-07-31 cs.IR 新提交 50%

VIG-RL: Learning to Search and Insert for Verified Image Grounding

VIG-RL:学习搜索与插入以实现可验证的图像定位

Qinhan Yu, Jun Guang, Chong Chen, Wentao Zhang

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文针对现有检索增强框架无法动态推理视觉证据插入时机与位置的问题,提出自主智能体框架VIG-RL,将相关工作流建模为主动决策过程,经强化学习优化后实现SOTA性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26460 2026-07-30 cs.RO 新提交 50%

RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning

RLMM-Flow:一种基于流的隐空间强化学习移动操作框架

Shuhang Wang, Ziming Li, Hui Cheng

专题命中 多模态Agent :multimodal(abstract)

AI总结 该研究提出RLMM-Flow框架,结合流策略预训练与隐空间强化学习,在移动操作基准上提升任务成功率等指标,同时保留流的快速推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18060 2026-07-29 cs.RO 版本更新 50%

RoboHarness: Memory-Driven Orchestration of Heterogeneous Robot Policies for Long-Horizon Planning

RoboHarness:用于长期规划的异构机器人策略的内存驱动编排

Jinbang Huang, Yuanzhao Hu, Zhiyuan Li, Ran Qi, Yixin Xiao, Zhanguang Zhang, Mark Coates, Tongtong Cao, Yingxue Zhang

机构 * Huawei Noah’s Ark Lab(华为诺亚方舟实验室) University of British Columbia(英属哥伦比亚大学) University of Toronto(多伦多大学) McGill University(麦吉尔大学) Labs(2012实验室)

专题命中 多模态Agent :multi-modal(abstract)

AI总结 针对长期机器人任务需多种能力、异构策略编排难的问题,提出RoboHarness框架,通过多模态执行内存等表征策略能力边界,经内存桥接稳定策略交接,实验验证其在长期规划和分布外鲁棒性上有显著提升。

Comments 21 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.23176 2026-07-28 eess.SY cs.SY 新提交 50%

An Adjoint-Based Differentiable Physics Framework for Online Parameter Inversion in Closed-Brayton Gas-Cooled Reactor Digital Twins

用于闭式布雷顿气冷反应堆数字孪生体在线参数反演的基于伴随的可微物理框架

Chengyuan Li, Shanfang Huang, Jian Deng

专题命中 多模态Agent :multi-modal(abstract)

AI总结 研究先进反应堆数字孪生体在线参数反演问题,提出基于伴随的可微物理框架,通过隐式微分代数模型传播反向模式自动微分,驱动AD - 海森增量4D - Var估计器,在多工况下与多种基线对比,展现出不同优势。

详情

展开后加载摘要…

URL PDF HTML 收藏