arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-08-18 至 2026-08-18 共收录 178 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态评测 36 篇

2608.14724 2026-08-18 cs.CV cs.AI cs.LG 新提交 62%

Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering

面向吉隆坡城市交通的隐私保护数据集构建:结合空间车辆上下文过滤的接地视觉-语言检测

Mohammed Abdul Al Arafat Tanzin, Rudzidatul Akmam Dziyauddin

机构 * Faculty of Artificial Intelligence, Universiti Teknologi Malaysia(马来西亚理工大学人工智能学院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 针对吉隆坡热带城市交通场景的隐私保护数据集构建难题,提出结合Grounding DINO与空间车辆ROI包含引擎的自动化匿名化框架,在1266帧图像上实现约95%的匿名化成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16594 2026-08-18 cs.AI 新提交 57%

CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction

CACSurv:基于大型语言模型的一致性对齐比较学习用于癌症生存预测

Tianqi Xiang, Qixiang Zhang, Xinpeng Ding, Yi Li, Xiaomeng Li

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 该研究针对癌症生存预测中LLM时间回归的公式与监督不匹配问题,提出CACSurv框架,构建TCGA-SurvReport基准,在6个TCGA癌症队列上的平均C指数达0.722,显著优于现有模型与基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16015 2026-08-18 cs.CV 新提交 57%

Multi-scale Decomposed Convolution Refinement Network for Visible-Infrared Person Re-Identification

用于可见光-红外行人重识别的多尺度分解卷积细化网络

Mingsheng Zheng, Zirui Jiang, Bo Liu, Yupeng Chen, Jun Zhang, Kai Zhao

机构 * School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing, Xinjiang University(新疆大学丝绸之路多语言认知计算联合国际研究实验室)

专题命中 多模态评测 :cross-modal(abstract);分类 cs.CV

AI总结 针对可见光-红外行人重识别的跨模态差异与判别能力不足问题,提出MDCRNet网络,结合HDCA模块与JDML损失,在SYSU-MM01和RegDB数据集上取得最优性能。

Comments 15 pages, 4 figures. Accepted for publication in the LNCS proceedings of ICONIP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15877 2026-08-18 cs.AI 新提交 57%

Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation

Dear Algo:面向统一搜索与推荐的精准优先智能体意图层

Rui Wang, Jiazhou Wang, Zheng Wei, Chenglin Lu, Fangcheng Sun, Ivy Sun, Jin Sun, Hui Geng, Lillian Zhang, Chao Yang, Lei Chen, Shahin Sefati, Reem Helou, Joe Zhou, Babak Shakibi, Yiyi Pan, Bi Xue, Hong Yan, Shujian Bu

机构 * Meta Platforms Inc.(元平台公司) Google DeepMind(谷歌DeepMind)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 该研究提出面向统一搜索与推荐的精准优先智能体意图层Dear Algo,将多类型意图编译为可执行计划,实验显示其在精度、候选数量及用户相关性上均优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15539 2026-08-18 cs.CV 新提交 57%

CrossView: Can Vision-Language Models Reason Across Cameras?

CrossView:视觉-语言模型能否跨相机进行推理?

Sahil Shah, S P Sharan, Harsh Goel, Manvik Pasula, Adithya Hebbalae, Minkyu Choi, Sandeep P. Chinchali

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出CrossView多相机视频问答基准,评估发现GPT-5.2等模型跨相机推理准确率低,开源模型表现更差,该基准可用于测试模型联合处理多视角的能力。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14947 2026-08-18 cs.AI 新提交 57%

RETRACE: Resilience-Guided Trait-Conditioned Craving Estimation from Wearable Physiology in Opioid Use Disorder

RETRACE:基于可穿戴生理信号的阿片类药物使用障碍中韧性引导的特质条件渴求估计

Yi Xiao, Harshit Sharma, Dessa Bergen-Cico, Asif Salekin

机构 * Arizona State University(亚利桑那州立大学) Syracuse University(雪城大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 本研究针对阿片类药物使用障碍跨受试者渴求检测难题,提出RETRACE框架,通过双编码器结合韧性上下文实现轻量个性化,在多模态数据集上较基线提升7%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14868 2026-08-18 cs.CV cs.RO 新提交 57%

Beam-Wise Statistical Background Subtraction for Static Roadside LiDAR: A Cross-Sensor Benchmark Study

面向静态路侧激光雷达的波束级统计背景减除:跨传感器基准研究

Alexander Baumann, Marcel Vosshans, Thao Dang

机构 * University of Applied Sciences Esslingen(埃斯林根应用科学大学) Faculty of Computer Science and Engineering(计算机科学与工程学院) Institute for Intelligent Systems(智能系统研究所)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

AI总结 本文针对静态路侧激光雷达缺失系统性跨传感器评估的问题,构建波束级统计背景减除基准,引入新数据集并结合空间滤波,实现了鲁棒可迁移的背景减除方案,提升了精度并保持高召回率与实时性,相关资源已公开。

Comments Accepted for publication at the 2026 IEEE 29th International Conference on Intelligent Transportation Systems (ITSC), Naples, Italy, September 15-18, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14721 2026-08-18 cs.CV 新提交 57%

AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning

AeroGround:用于空-地协同推理的综合基准

Shenghong Yi, Lin Zhang, Muzian Li, Jiakang Yuan, Haoyu Zhang, Peng Ye, Jiayuan Fan, Huafeng Qin, Tao Chen

机构 * Shanghai Innovation Institute(上海创新研究院) College of Future Information Technology, Fudan University(复旦大学未来信息技术学院) College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院) Chongqing Technology and Business University(重庆工商大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出AeroGround基准,针对现有VLMs在空-地协同场景推理能力的研究空白,构建含29000组多模态观测与2250个问答实例的数据集,实验发现模型与人类表现差距显著,为相关系统开发提供基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14655 2026-08-18 cs.LG cs.CV 新提交 57%

Diagnosing and Mitigating Perception-Decision Misalignment in Omni-LLMs via Modality Subspace Activation

通过模态子空间激活诊断并缓解全模态大语言模型(Omni-LLMs)中的感知-决策失配问题

Hongbo Jiang, Jie Li, Yunhang Shen, Tianyu Xie, Pingyang Dai

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

AI总结 针对Omni-LLMs存在的感知-决策失配问题,本文提出因果模态敏感性及对应诊断方法,构建CausalMSBench数据集,提出无需训练的MSA框架以恢复模型的因果模态敏感性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11891 2026-08-18 cs.CY cs.AI cs.HC 版本更新 57%

Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework

基于基准的公开基准印度基础模型对比评估:能力与评估成熟度框架

Avinash Agarwal, Vridhi Jain

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 本文提出能力与评估成熟度框架,对公开基准的印度基础模型与全球同类模型对比评估,发现其在传统基准表现尚可但参与新评估不足,还提出基准成熟度指数,指出能力差距或源于评估生态问题。

Comments 19 pages, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09005 2026-08-18 cs.CV cs.HC 版本更新 57%

A Survey of Body and Face Motion: Datasets, Performance Evaluation Metrics and Generative Techniques

人体与面部运动的综述:数据集、性能评估指标和生成技术

Lownish Rai Sookha, Nikhil Pakhale, Mudasir Ganaie, Abhinav Dhall

机构 * Indian Institute of Technology Ropar(印度理工学院罗帕尔) Monash University(莫纳什大学)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV

AI总结 本文综述了人体和面部运动生成,涵盖数据集、评估指标和生成技术,首次全面探讨了该领域的发展与未来方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16337 2026-08-18 stat.ME 新提交 50%

High-Dimensional Assisted Learning for Vertically Distributed Data with Blockwise Missingness

面向含分块缺失的垂直分布式数据的高维辅助学习

Yuwen Long, Shuyuan Wu, Yin Xia

专题命中 多模态评测 :multimodal(abstract)

AI总结 针对含分块缺失的垂直分布式数据,提出 ALB 算法,无需合并记录即可实现高维线性估计与推理,性能优于完整样本 Lasso,可近似集中式基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26431 2026-08-18 math.OC 版本更新 50%

Adjoint-Compatible Surrogates of the Expected Information Gain for Optimal Experimental Design

伴随兼容的期望信息增益代理用于最优实验设计

Luc de Montella, Sebastian Sager

专题命中 多模态评测 :multimodal(abstract)

AI总结 本文提出两种伴随兼容的期望信息增益代理,用于解决动态系统参数估计中的最优实验设计问题,在非高斯或多模态先验不确定性下表现更优。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态Agent 14 篇

2507.20776 2026-08-18 cs.CV 版本更新 85%

RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning

RingMo-Agent: 多平台多模态遥感基础模型

Huiyang Hu, Peijin Wang, Yingchao Feng, Kaiwen Wei, Wenxin Yin, Wenhui Diao, Mengyu Wang, Hanbo Bi, Kaiyue Kang, Tong Ling, Kun Fu, Xian Sun

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院 aerospace information research institute) School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院) University of Chinese Academy of Sciences(中国科学院大学) Key Laboratory of Target Cognition and Application Technology (TCAT)(目标认知与应用技术重点实验室)

专题命中 多模态Agent :multi-modal(title,abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV

AI总结 RingMo-Agent是一种多平台多模态遥感基础模型,通过大规模数据集和任务特定标记提升遥感图像的感知与推理能力。

Comments 24 pages, 5 figures, 20 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08759 2026-08-18 cs.CV cs.RO 版本更新 79%

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

通过技能级评估与诊断解构多模态语言模型的具身能力

Yu Qi, Haibo Zhao, Ziyu Guo, Siyuan Ma, Ziyan Chen, Yaokun Han, Renrui Zhang, Zitiantao Lin, Yizhe Zhu, Shiji Xin, Yijian Huang, Boce Hu, Kai Cheng, Peiheng Wang, Jiazheng Liu, Jiayi Zhang, Yizhe Zhu, Wenqing Wang, Yiran Qin, Haojie Huang, Lawson L. S. Wong

机构 * Northeastern University, Boston, MA, USA The Chinese University of Hong Kong, Hong Kong, China Peking University, Beijing, China Westlake University, Hangzhou, China Harvard University, Cambridge, MA, USA Purdue University, West Lafayette, IN, USA University of Oxford, Oxford, United Kingdom

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出BEAR基准,通过分解具身任务为14个原子技能进行细粒度评估,发现感知能力是推理失败的主要瓶颈,并提出BEAR-Agent多模态对话代理,显著提升具身技能性能。

Comments Accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22447 2026-08-18 cs.SE 版本更新 78%

Latent Reuse in Agent Skills: Multi-modal Clone Detection at Ecosystem Scale

SkillClone: 多模态克隆检测与克隆传播分析在智能体技能生态系统中

Jiaying Zhu, Lyuye Zhang, Wenbo Guo, Yang Liu

专题命中 多模态Agent :multi-modal(title,abstract)

AI总结 本文提出SkillClone,首个多模态智能体技能克隆检测方法,通过融合TF-IDF与通道分解实现高精度检测,揭示技能生态系统中大量克隆关系及传播问题。

Comments 13 pages, ASE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14720 2026-08-18 physics.chem-ph cs.AI 新提交 74%

Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spectra

基于多智能体闭环推理的多模态光谱有机结构解析

Bingsen Xue, Zhuojun Jiang, Jianhao Zhang, Mingcheng Gu, Yizhe Yuan, Yongtai Zhuo, Yifan Zhang, Li Wang, Ya Su, Yue Yuan, Jiang Liu, Xueqian Kong, Cheng Jin

专题命中 多模态Agent :multimodal(title);分类 cs.AI

AI总结 本文提出多智能体系统MACROS,经海量光谱数据训练后可实现零样本泛化,提升结构解析速度与准确性,为全自动结构解析及自主实验室发展奠定基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15930 2026-08-18 cs.AI cs.CV 新提交 62%

UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

UI-Mate:利用上下文内演示推进开放权重基础图形用户界面智能体

Zihan Ding, Longxu Dou, Qi Gao, Xiangwu Guo, Shengchao Hu, Zilong Huang, Zihang Jiang, Lei Ke, Mengcheng Lan, Weixian Lei, Hanxuan Li, Honglin Li, Xiyun Li, Zaitang Li, Leowei Liang, Xin Luo, Haozhe Ma, Jiayi Mao, Zhoujie Pan, Can Qin, Tianyuan Qu, Weiqi Wang, Wenkai Wang, Yonglin Wang, Yuxin Wang, Chenxu Wu, Yingchen Yu, Chenyu Zhang, Yuhao Zheng

机构 * Tencent Hy Frontier Team(腾讯 Hy 前沿团队)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 UI-Mate是集成环境基训练栈与上下文内演示学习的GUI智能体,在多计算机使用基准上创开放权重SOTA,提升了长周期办公任务的执行可靠性

Comments UI-Mate Technical Report. Project page: https://ui-mate.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15549 2026-08-18 cs.RO cs.AI 新提交 57%

MistyPilot: Enabling Social-Robot Control through Multi-Agent LLM Skill Orchestration

MistyPilot:通过多智能体大语言模型技能编排实现社交机器人控制

Xiao Wang, Lu Dong, Ifeoma Nwogu, Srirangaraj Setlur, Venu Govindaraju

机构 * State University of New York at Buffalo(纽约州立大学布法罗分校)

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 该研究提出多智能体大语言模型框架MistyPilot,可解释自然语言指令并编排Misty社交机器人技能,经评估其在多项任务上准确率高、方差低,用户反馈积极,代码将公开。

Comments Accepted at the ECCV 2026 ACVR Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15424 2026-08-18 cs.MA cs.AI cs.LG 新提交 57%

ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems

ETHOS:面向临床多智能体系统的模块化伦理框架

Rakesh Sharma, Sydney Pugh, Cameron Beeche, Pankhuri Singhal, Rachel Wu, Margaret Eby, Jeffrey Duda, James Gee, Kyra O'Brien, Hersh Sagreiya, Marina Serper, Victoria Gershuni, Angela Bradbury, Anurag Verma, Eric Eaton, Kevin B. Johnson, Walter Witschey

专题命中 多模态Agent :multimodal(abstract);分类 cs.AI

AI总结 ETHOS是可与现有临床多智能体系统集成的模块化伦理框架,通过分层治理提升决策可靠性,将AI伦理原则转化为可部署的安全保障。

Comments Preprint of an article submitted for consideration in Pacific Symposium on Biocomputing \textcopyright\ 2027 World Scientific Publishing Company. \url{https://psb.stanford.edu/}

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15009 2026-08-18 cs.RO cs.CV 新提交 57%

ForceU-VLA: A Force-Aware Vision-Language-Action Model for Embodied Ultrasound Scanning

ForceU-VLA:用于实体超声扫描的力感知视觉-语言-动作模型

Xingzheng Wu, Cheng Zhang, Guihao Yan, Xifeng Hu, Zhi Liu, Qing Cai

机构 * Faculty of Computer Science and Technology, Ocean University of China(中国海洋大学计算机科学与技术学院) School of Information Science and Engineering, Shandong University(山东大学信息科学与工程学院) Innovation School of Artificial Intelligence, Hefei University of Technology(合肥工业大学人工智能创新学院)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CV

AI总结 本文针对现有实体超声扫描方法的不足,提出力感知视觉-语言-动作模型ForceU-VLA,设计FUSFM与SAMM模块,构建ForceU-VLA-Data数据集,实验证实其可提升超声扫描的接触稳定性与压力调节能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21576 2026-08-18 cs.CL 版本更新 57%

Vision Language Models Cannot Plan, but Can They Formalize?

视觉语言模型(VLMs)无法规划,但它们能进行形式化吗?

Muyu He, Yuxi Zheng, Yuchen Liu, Zijian An, Bill Cai, Jiani Huang, Lifeng Zhou, Feng Liu, Ziyang Li, Li Zhang

机构 * Drexel University(德雷塞尔大学) University of Pennsylvania(宾夕法尼亚大学) Johns Hopkins University(约翰霍普金斯大学)

专题命中 多模态Agent :multimodal(abstract);分类 cs.CL

AI总结 本研究提出5种VLM作为形式化器的流水线,评估后发现其表现优于端到端规划生成,较弱VLMs的瓶颈为物体关系视觉grounding,较强模型已克服该限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16843 2026-08-18 cs.RO 新提交 50%

Security of Foundation-Model-Powered Embodied Agents: Attack Surfaces, Attacks, Defenses, and Evaluation

基于基础模型的具身智能体的安全性:攻击面、攻击、防御与评估

Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao

机构 * Wuhan University(武汉大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 该研究以信任边界为核心,针对基于基础模型的具身智能体安全,划分了五个层级与十二个攻击面,分析了58种攻击、61种防御的现状,指出部分领域研究不足并提出开放挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15198 2026-08-18 stat.ML cs.LG physics.comp-ph 新提交 50%

Identifying parameter couplings and uncertainties of mixed-noise stochastic systems via full-covariance Gaussian mixture network

通过全协方差高斯混合网络识别混合噪声随机系统的参数耦合与不确定性

Xiaolong Wang, Xiangwen Hao, Jing Feng, Yuanyuan Liu, Yong Xu

专题命中 多模态Agent :multi-modal(abstract)

AI总结 研究针对混合噪声随机系统参数识别的难点,提出PENN-GMD神经网络,采用全协方差高斯混合分布,经五个数值示例验证可准确恢复似然分布、捕捉参数耦合,为复杂随机系统参数识别提供实用工具。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18518 2026-08-18 cs.LG stat.ME stat.ML 版本更新 50%

Measuring the Prevalence of Policy Violating Content with ML Assisted Sampling and LLM Labeling

利用机器学习辅助抽样和大语言模型标注测量违规内容的普及率

Attila Dobi, Aravindh Manickavasagam, Benjamin Thompson, Xiaohan Yang, Faisal Farooq

机构 * Pinterest

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文提出了一种基于机器学习和大语言模型的系统,用于高效测量违反政策内容的普及率,通过概率抽样和多模态标注提升测量准确性。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22983 2026-08-18 cs.RO 版本更新 50%

Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives

基础模型时代中的具身机器人操作:规划与学习视角

Shuanghao Bai, Wenxuan Song, Jiayi Chen, Yuheng Ji, Zhide Zhong, Jin Yang, Han Zhao, Wanqi Zhou, Zhe Li, Pengxiang Ding, Cheng Chi, Chang Xu, Xiaolong Zheng, Donglin Wang, Haoang Li, Shanghang Zhang, Badong Chen

机构 * Xi’an Jiaotong Univeristy(西安交通大学) Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Chinese Academy of Sciences(中国科学院) Westlake University(西湖大学) Zhejiang University(浙江大学) University of Sydney(悉尼大学) BAAI(百度人工智能研究院) Peking University(北京大学)

专题命中 多模态Agent :multimodal(abstract)

AI总结 本文探讨了基础模型时代机器人操作的规划与学习方法,分析了高层推理与低层控制的统一框架,并提出了未来研究方向。

Comments This work is a re-architected core derived from the full survey (arXiv:2510.10903), refined to highlight the most central themes and representative studies

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10903 2026-08-18 cs.RO 版本更新 50%

Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey

面向机器人操作的统一理解:一项综合调查

Shuanghao Bai, Wenxuan Song, Jiayi Chen, Yuheng Ji, Zhide Zhong, Jin Yang, Han Zhao, Wanqi Zhou, Wei Zhao, Zhe Li, Pengxiang Ding, Cheng Chi, Haoang Li, Chang Xu, Xiaolong Zheng, Donglin Wang, Shanghang Zhang, Badong Chen

专题命中 多模态Agent :multimodal(abstract)

AI总结 本调查针对机器人操作这一具身智能核心挑战,提出方法的统一分类与瓶颈分类,为机器人操作研究提供了全面的路线图与结构化参考。

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 多模态训练与对齐 23 篇

2608.15614 2026-08-18 cs.CV cs.AI 新提交 84%

EgoGazeLite: On-Device Egocentric Gaze Prediction for Token-Efficient Multimodal LLM Video Input

EgoGazeLite:面向 token 高效多模态大语言模型视频输入的设备内自我中心注视预测

Matteo Stoiber, Niels Buus Lassen

机构 * Copenhagen Business School(哥本哈根商学院)

专题命中 多模态训练与对齐 :multimodal(title,abstract);MLLM(abstract_cn);分类 cs.CV、cs.AI

AI总结 EgoGazeLite 是轻量双进程注视预测器,可在消费级硬件上实时运行,无需眼动追踪硬件即可实现 token 高效的自我中心视频理解,性能与真实注视裁剪无显著差异。

Comments 16 pages. Accepted at the WearableAI Workshop, ECCV 2026 (Archival Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13741 2026-08-18 cs.CL cs.LG 版本更新 83%

GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis

GALA:面向文本到时间序列合成的生成感知跨模态对齐

Haochen Zhang, Gengwei Zhang, Laura Yao, Nicholas Konz, Tianlong Chen

专题命中 多模态训练与对齐 :cross-modal(title,abstract);分类 cs.CL

AI总结 本研究针对文本到时间序列合成中条件表示与信号模态不匹配的问题,提出GALA两阶段跨模态对齐方法,在TSFragment-600K数据集上实现SOTA,打破了生成器内部文本编码器的保真度与贴合度权衡。

Comments 21 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.16201 2026-08-18 cs.LG 新提交 82%

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

面向基于大语言模型(LLM)的多模态情感分析的多粒度情感集成

Shanshan Lin, Yuesheng Wu, Chao Chen, Yizhe Yang, Zhihao Chen, Zexian Yang, Xiangwen Liao

机构 * Fuzhou University(福州大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Jiangxia University(江夏大学)

专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract)

AI总结 该研究提出MGSI多粒度情感集成框架,通过多尺度编码、文本引导对齐等优化,提升基于LLM的多模态情感分析性能,在四个公开基准上效果优于冻结LLM基线。

Comments Accepted to NLPCC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏