arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4882 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4882 篇

1411.1784 2014-11-10 cs.LG cs.AI cs.CV stat.ML 62%

Conditional Generative Adversarial Nets

Mehdi Mirza, Simon Osindero

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1205.1365 2012-05-08 cs.MM cs.CV 62%

Image Enhancement with Statistical Estimation

Aroop Mukherjee, Soumen Kanrar

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments 9 pages,6 figures; ISSN:0975-5578 (Online); 0975-5934 (Print)

Journal ref The International Journal of Multimedia & Its Applications (IJMA) April 2012, Volume 4, Number 2, page 59-67

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0105003 2009-11-30 cs.CL cs.AI 62%

Rule Writing or Annotation: Cost-efficient Resource Usage for Base Noun Phrase Chunking

Grace Ngai, David Yarowsky

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 4 figures, appeared in ACL2000

Journal ref Proceedings of the 38th Annual Meeting of the Association for Computational Linguistics, pages 117-125, Hong Kong (2000)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07878 2025-10-21 cs.LG cs.AI eess.SP q-bio.NC 61%

Comparative Analysis of Deep Learning Approaches for Harmful Brain Activity Detection Using EEG

Shivraj Singh Bhatti, Aryan Yadav, Mitali Monga, Neeraj Kumar

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Thapar Institute of Engineering and Technology(泰帕尔工程与技术学院)

专题命中 其他多模态 :multimodal(abstract,comments);分类 cs.AI

Comments 6 pages, 5 figures. Presented at IEEE CICT 2024. The paper discusses the application of multimodal data and training strategies in EEG-based brain activity classification

Journal ref 2024 IEEE 8th Int. Conf. on Info. and Comm. Tech. (CICT), pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01542 2025-05-06 cs.HC cs.AI 61%

Emotions in the Loop: A Survey of Affective Computing for Emotional Support

Karishma Hegde, Hemadri Jayalath

机构 * School of Computing University of Georgia(计算学院 佐治亚大学)

专题命中 其他多模态 :multimodal(abstract,comments);分类 cs.AI

Comments 20 pages, 7 tables, 96 references. Survey paper on affective computing applications using large language models, multimodal AI, and therapeutic chatbots

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.01536 2021-09-21 cs.CL 61%

BERT meets LIWC: Exploring State-of-the-Art Language Models for Predicting Communication Behavior in Couples' Conflict Interactions

Jacopo Biggiogera, George Boateng, Peter Hilpert, Matthew Vowels, Guy Bodenmann, Mona Neysari, Fridtjof Nussbeck, Tobias Kowatsch

专题命中 其他多模态 :multimodal(abstract,comments);分类 cs.CL

Comments 5 pages. Accepted at the 2nd Workshop on Social Affective Multimodal Interaction for Health (SAMIH) at ICMI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.01466 2018-07-05 cs.CL 61%

Polarity and Intensity: the Two Aspects of Sentiment Analysis

Leimin Tian, Catherine Lai, Johanna D. Moore

专题命中 其他多模态 :multimodal(abstract,comments);分类 cs.CL

Comments Published at the First Grand Challenge and Workshop on Human Multimodal Language (Challenge-HML) of ACL 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17447 2026-08-19 cs.CV 新提交 57%

NGS-Marker: Robust Native Watermarking for 3D Gaussian Splatting

NGS-Marker:面向3D高斯溅射(3DGS)的鲁棒原生水印

Hao Qin, Yukai Sun, Luyuan Chen, Mengxu Lu, Feng Zhang, Ming Kong, Zhenhong Du, Qiang Zhu

机构 * Zhejiang University(浙江大学) Zhejiang Key Laboratory of Geographic Information Science(浙江省地理信息科学重点实验室)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 该研究针对现有3D高斯溅射水印技术无法抵御部分侵权的问题,提出NGS-Marker原生水印框架,通过联合训练的注入器与解码器及梯度渐进注入策略实现全场景覆盖,可抵御部分侵权并支持混合与多模态水印,具备实际部署灵活性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15115 2026-08-18 cs.CV 新提交 57%

Perspective-Invariant Attack with Enhanced Transferability of Adversarial Examples

具有增强对抗样本迁移性的视角不变攻击

Kaisheng Liang, Yiming Cao, Bin Xiao

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 针对对抗样本跨模型迁移性带来的安全威胁,提出视角不变攻击(PIA)及其扩展PIA-Mix,通过多自由度顶点采样策略提升对抗样本迁移性,实验显示其性能优于当前最优基于迁移的攻击方法。

Journal ref IEEE Transactions on Information Forensics and Security, vol. 21, pp. 6818-6831, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14306 2026-08-17 cs.AI cs.SY eess.SY 新提交 57%

Sensor-Driven Mission Synthesis for UAV/UGV Swarms: A TB-CSPN Coordination Architecture with Hardware-Enforced Safety

传感器驱动的无人机/无人车集群任务合成:具备硬件强制安全保障的TB-CSPN协同架构

Uwe M. Borghoff, Paolo Bottoni, Remo Pareschi

机构 * Institute for Software Technology, University of the Bundeswehr Munich(慕尼黑联邦国防军大学软件技术研究所) Department of Computer Science, Sapienza University of Rome(罗马大学计算机科学系) SofTware And Knowledge Engineering Lab(软件与知识工程实验室)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出一种具备硬件强制安全保障的TB-CSPN协同架构,用于异构无人机/无人车集群,结合多模态传感器观测,通过顾问与监督智能体实现可审计的决策路径,提升对抗环境下的韧性,经沿海监视案例验证其有效性。

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10434 2026-08-12 cs.AI 新提交 57%

Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance

无人机入侵检测中对话式与仪表盘式可解释人工智能:操作员信任与依赖的实证研究

Cong Chi Nguyen, Trang Mai Xuan, Vu-Duc Ngo, Kim-Ngan Thi Nguyen, Trong-Nghia Nguyen, Thien Van Luong

机构 * Phenikaa University(菲卡大学) Phenikaa School of Computing(菲卡计算机学院) MobiFone Corporation(MobiFone集团) MobiFone HighTech Center(MobiFone高科技中心) National Economics University(国民经济大学) College of Technology(技术学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本研究对比对话式与仪表盘式XAI界面对无人机入侵检测操作员信任和依赖的影响,发现对话式界面可用性更高但易引发过度依赖,为未来XAI系统设计提供启示。

Comments 12 pages, 3 figures, EIDT conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10203 2026-08-12 cs.CV 新提交 57%

A Convolutional Layer Activation Dimensionality Reduction for Out-of-Distribution and Adversarial Attack Detection Methods

用于分布外样本与对抗攻击检测方法的卷积层激活降维技术

Leandro de Souza Rosa, Lorenzo Capelli, Clara Nunes Barrancos, Mauro Mangia, Riccardo Rovatti

机构 * Alma Mater Studiorum Università di Bologna(博洛尼亚大学) Advanced Research Center on Electronic Systems “Ercole De Castro” (ARCES) - Alma Mater Studiorum Università di Bologna(博洛尼亚大学“埃尔科莱·德·卡斯特罗”电子系统高级研究中心)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文针对卷积层激活降维的不足,提出一种可控高压缩率的新型降维方法,扩展两种最先进的检测方法并在OOD与对抗攻击检测任务上验证,其性能更优且计算内存占用更低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19125 2026-08-11 cs.LG cs.AI 版本更新 57%

Transformer Circuits Can Realize Clustering Algorithms

Transformer电路可实现聚类算法

Kenneth L. Clarkson, Lior Horesh, Takuya Ito, Charlotte Park, Parikshit Ram

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 研究表明Transformer电路可实现Lloyd算法等精确聚类算法,提出的k均值Transformer聚类算法泛化性强且质量优于Lloyd算法,改动后可衍生多种新型聚类算法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.10961 2026-08-11 cs.LG cs.AI 57%

Bike Sharing Demand Prediction based on Knowledge Sharing across Modes: A Graph-based Deep Learning Approach

Yuebing Liang, Guan Huang, Zhan Zhao

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06758 2026-08-10 cs.CL 新提交 57%

Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control

Stockmark-Nemotron-3-Nano-Omni-JapanDocReader:基于能力注入与遗忘控制的结构化文档解析模型

Shi Chen, Hayato Aida, Makoto Morinaga, Shohei Tanaka, Kosuke Arima

机构 * Stockmark Inc.(斯托克马克公司)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

AI总结 本研究提出Stockmark-Nemotron-3-Nano-Omni-JapanDocReader模型,通过能力注入与遗忘控制优化结构化文档解析性能,结合混合SFT与DAPO式RL取得优于SFT的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05684 2026-08-07 cs.RO cs.AI cs.LG cs.SY eess.SY 新提交 57%

Nonvisual Classification of Ground-Condition by Artificial Proprioception in an Amoeba-Inspired Autonomous Walking Robot

变形虫启发的自主步行机器人中基于人工本体感觉的地面状况非视觉分类

Hyoto Yamaguchi, Zenji Yatabe, Seiya Kasai

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 该研究为变形虫启发的自主步行机器人,整合传感器与储备池计算实现人工本体感觉,完成地面状况非视觉分类,可依地面切换步态并分析传感器贡献。

Comments 5 pages, 7 figures, The paper has been submitted to IEEE SCIS ISIS 2026 for consideration

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04930 2026-08-06 cs.LG cs.AI 新提交 57%

SVI-DAG: A Structured Variational Inference Approach to Bayesian Causal Discovery

SVI-DAG:一种用于贝叶斯因果发现的结构化变分推断方法

Shrenik Zinage

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 SVI-DAG是一种贝叶斯因果发现的结构化变分推断方法,用归一化流建模边依赖、斯坦变分梯度下降优化,在不确定性量化上优于5种现有方法,结构准确性具竞争力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02955 2026-08-05 cs.HC cs.AI cs.CY eess.IV 新提交 57%

Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits

聊天调试:人机协作调试模拟电路的探索性研究

John Hu, Andrew Ash

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 该研究通过分析本科生与开源大型语言模型(LLMs)协作调试模拟电路的聊天记录,揭示了人机协作调试的模式、LLMs的优势与不足及学生的技能短板,为优化人机协作调试提供了依据。

Comments This is the accepted version of a paper to be presented at the 2026 IEEE Frontiers in Education Conference (FIE). The final version will be available via IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01112 2026-08-04 cs.AI 新提交 57%

Fighting Fire with Fire: On the Feasibility of Protecting Exercises Against AI Cheating

以毒攻毒:利用对抗机器学习保护练习题抵御AI作弊的可行性研究

Tobias Braun, Jonas Grebe, Louis Rethfeld, Marcus Rohrbach

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本研究探究利用对抗机器学习,通过在多模态选择题视觉组件添加扰动引导AI作弊者给出固定错误答案,再以统计检验检测作弊,以保护教育练习题抵御AI作弊的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00743 2026-08-04 cs.CV 新提交 57%

LUT: Latent Utility Training for Visual Reasoning

LUT:用于视觉推理的隐式效用训练

Jiaxuan Kang, Siyu Chen, Mingda Li, Mingjie Liu, Tianyue Wang, Zhaoyang Wei, Yongheng Zhang, Yanchao Hao, Zheng Wei

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出仅用标准VQA对训练的LUT框架,通过轨迹和步骤层面的隐式效用优化,在感知密集型视觉推理基准上性能优于现有隐式推理方法,且标注成本更低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15622 2026-08-04 cs.CV cs.LG 版本更新 57%

AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference

AdaVFM:通过LLM引导执行实现边缘智能的自适应视觉基础模型

Yiwei Zhao, Yi Zheng, Huapeng Su, Jieyu Lin, Stefano Ambrogio, Cijo Jose, Michael Ramamonjisoa, Patrick Labatut, Barbara De Salvo, Chiao Liu, Phillip B. Gibbons, Ziyun Li

机构 * Carnegie Mellon University(卡内基梅隆大学) Meta

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出AdaVFM,一种通过LLM引导执行实现边缘设备上语言对齐视觉基础模型高效推理的自适应框架,通过动态调整计算实现性能与效率的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29278 2026-08-03 cs.CV 新提交 57%

Training-Free Entity-Level Few-Shot Segmentation of Remote Sensing Images with Advection Refinement

基于平流优化的遥感图像无训练实体级小样本分割

Xueting Bai, Huan Ni

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 该研究针对现有跨域小样本分割方法训练成本高、预测结果碎片化的问题,提出一种基于平流优化的无训练实体级遥感图像小样本分割框架,可提升 SAM3 的相关适应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00177 2026-08-03 cs.GR cs.CV 版本更新 57%

FieryGS: In-the-Wild Fire Synthesis with Physics-Integrated Gaussian Splatting

FieryGS: 在真实世界中实现火灾合成的物理集成高斯点云方法

Qianfan Shen, Ningxiao Tao, Qiyu Dai, Tianle Chen, Minghan Qin, Yongjie Zhang, Mengyu Chu, Wenzheng Chen, Baoquan Chen

机构 * School of EECS, Peking University(电子工程系,北京大学) School of Intelligence Science and Technology, Peking University(智能科学与技术学院,北京大学) Yuanpei College, Peking University(元培学院,北京大学) ByteDance Seed(字节跳动种子) Wangxuan Institute of Computer Technology, Peking University(王璇计算机技术研究所,北京大学) Beijing Academy of Artificial Intelligence, Beijing, China(北京人工智能研究院,北京,中国)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 FieryGS通过整合物理准确的燃烧模拟与渲染,实现真实世界3D场景中逼真的火灾合成,结合多模态大语言模型进行物理材料推理,提升火灾动态的可控性和真实性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27428 2026-07-31 physics.med-ph cs.AI 新提交 57%

Rethinking Artificial Intelligence in Medical Imaging: Assumptions, Reality, and Reframing

重新思考医学影像中的人工智能:假设、现实与重构

Arman Rahmim, Nourhan Bayasi, Xiaoxiao Li, Babak Saboury, Fereshteh Yousefirizi

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文指出医学影像AI研究转化不足源于结构性错位,明确六大错位维度并提出重构路径,最终愿景是开发与医生对齐、扩展而非替代临床判断的智能体AI。

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13174 2026-07-31 eess.IV cs.CV cs.LG 版本更新 57%

Scalable Drift Monitoring in Medical Imaging AI

医学影像AI中的可扩展漂移监测

Jameson Merkow, Felix J. Dorfner, Xiyu Yang, Alexander Ersoy, Giridhar Dasegowda, Mannudeep Kalra, Matthew P. Lungren, Christopher P. Bridge, Ivan Tarapov

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本研究针对医学影像AI的模型漂移与可靠性问题,开发了基于CheXstray框架的增强型可扩展漂移监测框架MMC+,经真实世界数据验证可有效检测数据偏移并预警性能偏差,助力AI在临床场景的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24806 2026-07-29 q-bio.NC cs.AI 新提交 57%

Decoding Error-Related Potentials under Multisensory Feedback with Varying Congruency

在具有不同一致性的多感官反馈下解码错误相关电位

Yixin Liu, Kang Yin, Hye-Bin Shin, Seong-Whan Lee

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 研究在多模态视觉、听觉和触觉反馈及可控感官一致性下的ErrP解码挑战,采用基于多分支EEGNet的架构及辅助监督学习策略,经迷宫观察任务实验,该方法在异质感官条件下分类性能一致且准确度提高,提升了ErrP解码稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.24113 2026-07-28 cs.RO cs.AI cs.HC 新提交 57%

A Case Study on the Acceptance of a Humanoid Robotic Head Employed in Three Public Spaces

关于在三个公共场所使用的仿人机器人头部接受度的案例研究

Marcel Heisler, Luca Randecker, Christian Becker-Asano

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 研究仿人机器人头部在公共场所的接受度,通过情感模拟后端处理自然语言生成多模态响应,邀请访客交流。结果显示用户有使用意愿,公共场所更适配,多语言响应受欢迎,但响应时间待改进。

Comments accepted at the 35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN 2026) Kitakyushu, Japan

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22734 2026-07-28 cs.CV eess.IV physics.geo-ph 新提交 57%

Fast Fourier Convolutional GAN for 30 m Clear-Sky Land Surface Temperature Gap-Free Reconstruction

用于30米晴空陆地表面温度无间隙重建的快速傅里叶卷积生成对抗网络

Marwa Alfouly, Smajil Halilovic, Nils Bochow, Thomas Hamacher, Niklas Boers, Konrad Schindler

机构 * Technical University of Munich(慕尼黑工业大学) Helmholtz Centre for Polar and Marine Research, Alfred Wegener Institute(亥姆霍兹极地与海洋研究中心阿尔弗雷德·韦格纳研究所) Swiss Federal Institute of Technology Zurich (ETH Zurich)(瑞士联邦理工学院苏黎世分校)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 针对卫星LST数据因云层有间隙、重建难问题,提出多模态快速傅里叶卷积生成对抗网络,利用快速傅里叶卷积实现全局感受野,由卫星观测和SAR数据引导,能恢复大量缺失区域,重建效果好。

Comments 35 pages, 9 figures, Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15442 2026-07-20 cs.AI 新提交 57%

Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes

超越玩笑:用于检测和解释表情包中有害幽默的多角度推理

Shanhong Liu, Pai Chet Ng, De Wen Soh, Malika Meghjani, Konstantinos N. Plataniotis

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 研究如何检测和解释表情包中有害幽默,提出MAR-12框架,利用视觉语言模型,从十二个角度解读表情包,经注意力机制和原型分类器预测,在多数据集上准确率超现有方法,且能给出有说服力的解释。

Comments Accepted for Publication at AAAI-ICWSM 2027

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.15374 2026-07-20 cs.CV 新提交 57%

Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning

通过强化学习进行推理引导的部件级视觉定位

Kazi Sajeed Mehrab, Hani Alomari, Najibul Haque Sarker, Chia-Wei Tang, Zaber Ibn Abdul Hakim, Anuj Karpatne, Chris Thomas

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 研究部件级视觉定位难题,提出OP - HRG粗到细推理引导定位策略,先定位父物体再定位部件,经自我检查反思结果,引入部件感知GRPO框架训练,训练的4B模型性能优异且可迁移到推理分割。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏