arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-04-29 至 2026-04-29 共收录 67 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2604.07802 2026-04-29 cs.CV cs.AI 62%

Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models

潜在异常知识挖掘:揭示视觉-语言模型中的稀疏敏感神经元

Shaotian Li, Shangze Li, Chuancheng Shi, Wenhua Wu, Yanqiu Wu, Xiaohan Yu, Fei Shen, Tat-Seng Chua

机构 * Macquarie University(麦考瑞大学) Nanjing University of Science and Technology(南京理工大学) The University of Sydney(悉尼大学) National University of Singapore(新加坡国立大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出LAKE框架,通过挖掘视觉-语言模型中稀疏敏感神经元,实现异常检测的内在可解释性,实验表明其在工业基准上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25884 2026-04-29 quant-ph cs.CV 57%

QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding

QCalEval:用于量子校准图理解的视觉-语言模型基准测试

Shuxiang Cao, Zijian Zhang, Abhishek Agarwal, Grace Bratrud, Niyaz R. Beysengulov, Daniel C. Cole, Alejandro Gómez Frieiro, Elena O. Glen, Hao Hsu, Gang Huang, Raymond Jow, Greshma Shaji, Tom Lubowe, Ligeng Zhu, Luis Mantilla Calderón, Nicola Pancotti, Joel Pendleton, Brandon Severin, Charles Etienne Staub, Sara Sussman, Antti Vepsäläinen, Neel Rajeshbhai Vora, Yilun Xu, Varinia Bernales, Daniel Bowring, Elica Kyoseva, Ivan Rungger, Giulia Semeghini, Sam Stanwyck, Timothy Costa, Alán Aspuru-Guzik, Krysta Svore

机构 * NVIDIA University of Toronto(多伦多大学) IQM Quantum Computers(IQM量子计算机) Lawrence Berkeley National Laboratory(伯克利国家实验室) Conductor Quantum(Conductor量子) National Physical Laboratory(国家物理实验室) Infleqtion Harvard University(哈佛大学) Fermi National Accelerator Laboratory(费米国家加速器实验室) Northwestern University(西北大学) EeroQ Corporation(EeroQ公司) Royal Holloway University of London(伦敦皇家霍洛威大学) Vector Institute for Artificial Intelligence(人工智能向量研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出QCalEval,首个用于评估视觉-语言模型理解量子校准图能力的基准测试,包含243个样本和87种场景类型,测试零样本和上下文学习下的六种问题类型,展示了不同模型的性能差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25562 2026-04-29 cs.CR cs.AI 57%

SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents

SnapGuard: 基于截图的轻量级提示注入检测方法

Mengyao Du, Han Fang, Haokai Ma, Jiahao Chen, Kai Xu, Quanjun Yin, Ee-Chien Chang

机构 * National University of Defense Technology(国防科技大学) University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Zhejiang University(浙江大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

AI总结 针对基于截图的网络代理面临的提示注入攻击问题,提出SnapGuard方法,通过视觉稳定指标和文本信号分析实现高效检测,达到F1得分0.75,速度提升8倍。

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23214 2026-04-29 cs.CL 57%

DARC-CLIP: Dynamic Adaptive Refinement with Cross-Attention for Meme Understanding

DARC-CLIP:动态自适应细化与跨注意力机制用于表情包理解

Qiyuan Jin

机构 * The Hong Kong University of Science(香港科技大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

AI总结 本文提出DARC-CLIP框架,通过层次化细化栈实现自适应多模态融合,提升表情包中多模态线索的建模精度,尤其在仇恨检测任务中取得显著提升。

Comments Accepted to IEEE ICASSP 2026. 5 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29080 2026-04-29 cs.CV cs.LG 57%

Is the Modality Gap a Bug or a Feature? A Robustness Perspective

模态间隙是bug还是feature?从鲁棒性视角

Rhea Chowers, Oshri Naparstek, Udi Barzelay, Yair Weiss

机构 * Hebrew University(希伯来大学) IBM Research(IBM研究院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文从鲁棒性角度探讨模态间隙的存在原因,发现减少间隙可提升模型鲁棒性而不影响清洁准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13302 2026-04-29 cs.CL 57%

Images Amplify Misinformation Sharing in Vision-Language Models

图像在视觉-语言模型中放大虚假信息的传播

Alice Plebe, Timothy Douglas, Diana Riazi, R. Maria del Rio-Chanona

机构 * Department of Industrial Engineering, University of Trento(特伦托大学工业工程系) Computer Science Department, University College London(伦敦大学学院计算机科学系)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

AI总结 研究探讨了图像如何影响视觉-语言模型分享新闻内容的倾向,发现图像能提高虚假新闻的分享率,且不同模型对图像的反应存在差异。

Comments Accepted for oral presentation at ICWSM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 7 篇

2512.06757 2026-04-29 cs.SD cs.CV 79%

XM-ALIGN: Unified Cross-Modal Embedding Alignment for Face-Voice Association

XM-ALIGN:面向人脸-语音关联的统一跨模态嵌入对齐框架

Zhihua Fang, Shumei Tao, Junxu Wang, Liang He

机构 * School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) Xinjiang Multimodal Information Technology Engineering Research Center(新疆多模态信息处理工程技术研究中心) Urumqi Branch, China Mobile Group Xinjiang Co., Ltd(中国移动新疆乌鲁木齐分公司) School of Intelligence Science and Technology, Xinjiang University(新疆大学智能科学与技术学院) Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出XM-ALIGN框架,通过显式与隐式对齐机制提升跨模态验证性能,采用共享分类器联合优化人脸与语音嵌入,并通过数据增强提升泛化能力。

Comments FAME 2026 Technical Report

Journal ref ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25255 2026-04-29 cs.CV 79%

Personalized Cross-Modal Emotional Correlation Learning for Speech-Preserving Facial Expression Manipulation

面向语音保留的面部表情操控的个性化跨模态情感相关学习

Tianshui Chen, Yujie Zhu, Jianman Lin, Zhijing Yang, Chunmei Qing, Feng Gao, Liang Lin

机构 * Guangdong University of Technology(广东工业大学) South China University of Technology(华南理工大学) Peking University(北京大学) Sun Yat-Sen University(中山大学)

专题命中 音频语音多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出PCMECL算法,通过个性化提示和特征差分提升跨模态对齐,解决面部表情操控中情感操控的监督问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18612 2026-04-29 cs.CR cs.CV 79%

Multimodal Privacy-Preserving Entity Resolution with Fully Homomorphic Encryption

多模态隐私保护实体解析与全同态加密

Susim Roy, Nalini Ratha

机构 * University at Buffalo, The State University of New York Department of Computer Science(布法罗大学计算机科学与工程系)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出多模态框架,利用政府和金融机构的数据集,解决数据量、匹配精度和隐私问题,通过全同态加密保障隐私,实现低误判率和高效计算。

Comments 5 pages, 3 figures, IEEE ICASSP'26

Journal ref ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25591 2026-04-29 eess.AS cs.AI cs.CL cs.LG cs.SD 67%

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

穿越不确定性:音频感知大语言模型不确定性估计的实证研究

Chun-Yi Kuan, Wei-Ping Huang, Hung-yi Lee

机构 * Graduate Institute of Communication Engineering, National Taiwan University, Taiwan(台湾大学通讯工程研究所) Artificial Intelligence Center of Research Excellence (AI-CoRE), National Taiwan University, Taiwan(台湾大学人工智能卓越研究中心)

专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、cs.AI、eess.AS

AI总结 本文研究了音频感知大语言模型的不确定性估计,通过多种方法对比发现语义层面方法在通用音频推理中表现更优,且在可靠性导向任务中效果依赖模型和基准。

Comments Manuscript in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11110 2026-04-29 cs.SD 67%

Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan

Ti-Audio:首个多方言端到端藏语语音大模型

Jialing Wang, Yue Zhao, Yuhao Zhang, Jing Yu, Shaosai Li, Zhanchen Dai, Benyou Wang, Haizhou Li

机构 * Minzu University of China(中国民族大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract)

AI总结 本文提出Ti-Audio,首个多方言端到端藏语语音大模型,通过动态Q-Former适配器和温度采样策略解决低资源多方言环境下的语音识别与翻译问题,实验显示其在藏语基准测试中达到最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24179 2026-04-29 cs.CL cs.AI 62%

MemeScouts@LT-EDI 2026: Asking the Right Questions -- Prompted Weak Supervision for Meme Hate Speech Detection

MemeScouts@LT-EDI 2026:问对问题——针对弱监督的提示方法用于表情包仇恨言论检测

Ivo Bueno, Lea Hirlimann, Enkelejda Kasneci

机构 * Technical University of Munich(慕尼黑技术大学) LMU Munich(慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心(MCML))

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出一种提示弱监督方法,通过分解表情包理解为基于问题的标注函数,提升多语言多模态仇恨言论检测效果,尤其在中文和印地语中表现突出。

Comments Accepted at Sixth Workshop on Language Technology for Equality, Diversity and Inclusion at ACL2026 (LT-EDI@ACL26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25383 2026-04-29 cs.SD cs.AI eess.AS 62%

ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations

ML-SAN:多级说话人自适应网络用于对话中的情绪识别

Kexue Wang, Yinfeng Yu, Liejun Wang

机构 * Joint Research Laboratory for Embodied Intelligence, Xinjiang University(新疆大学具身智能联合研究实验室) Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing, Xinjiang University(丝绸之路多语种认知计算国际联合研究实验室) School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI、eess.AS

AI总结 本文提出ML-SAN多级说话人自适应网络,通过三级适应过程解决说话人身份信息混淆问题,提升多轮对话中情绪识别的准确性和鲁棒性。

Comments Main paper (12 pages). Accepted for publication by International Conference on Intelligent Computing 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 7 篇

2412.07584 2026-04-29 cs.CV cs.AI 81%

Multimodal Contextualized Support for Enhancing Video Retrieval System

多模态上下文支持用于增强视频检索系统

Quoc-Bao Nguyen-Le, Thanh-Huy Le-Nguyen

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种多模态视频检索系统,通过整合多帧信息提取更高层次的抽象信息,提升视频检索的准确性与深度。

Comments This paper has been withdrawn by the author. After further review, the author believes that the current version does not meet the desired standards and plans to revise the work before any potential resubmission

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25584 2026-04-29 cs.AI 79%

DualFact+: A Multimodal Fact Verification Framework for Procedural Video Understanding

DualFact+: 一种用于过程视频理解的多模态事实验证框架

Cennet Oguz, Yasser Hamidullah, Josef van Genabith, Simon Ostermann

机构 * German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心(DFKI)) Saarland Informatics Campus(萨尔兰信息学校区)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 DualFact+通过双层多模事实验证框架,针对过程视频描述中的概念事实和上下文事实进行评估,揭示了多模态事实 grounding 的挑战。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05959 2026-04-29 cs.CV cs.LG 74%

Multi-Modal Landslide Detection from Sentinel-1 SAR and Sentinel-2 Optical Imagery Using Multi-Encoder Vision Transformers and Ensemble Learning

基于Sentinel-1 SAR和Sentinel-2光学影像的多模态滑坡检测:使用多编码器视觉Transformer和集成学习

Ioannis Nasios

机构 * NodalPoint

专题命中 视频多模态 :multi-modal(title);分类 cs.CV

AI总结 本文提出融合Sentinel-2光学影像与Sentinel-1 SAR数据的多模型框架,利用多编码器视觉Transformer和集成学习提升滑坡检测精度,实现91.9的F1分数。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25276 2026-04-29 cs.CV 70%

OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding

OmniVTG:一种大规模数据集和开放世界视频时间定位的训练范式

Minghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng, Yang Liu

机构 * Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) Central Media Technology Institute, Huawei Technologies Ltd.(华为技术有限公司中央媒体技术研究所) PKU-WUHAN Institute for Artificial Intelligence, Peking University(北京大学武汉人工智能研究所)

专题命中 视频多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出OmniVTG数据集和Self-Correction Chain-of-Thought训练范式,通过语义覆盖迭代扩展管道构建大规模数据集,并利用多模态大语言模型的密集描述能力提升视频时间定位性能。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24842 2026-04-29 cs.AI cs.MA cs.MM 62%

Co-Director: Agentic Generative Video Storytelling

Co-Director: 基于代理的生成视频叙事

Yale Song, Yiwen Song, Nick Losier, Nathan Hodson, Ye Jin, Rhyard Zhu, Yan Xu, Daniel Vlasic, Carina Claassen, Jasmine Leon, Khanh G. LeViet, Zack Chomyn, Joe Timmons, Brett Slatkin, Scott Penberthy, Tomas Pfister

机构 * Google(谷歌)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI、cs.MM

AI总结 本文提出Co-Director框架,通过分层多代理方法解决视频生成的语义一致性问题,引入分层参数化和多模态自优化循环,实现叙事策略探索与有效配置的平衡,通过GenAD-Bench验证其在个性化广告中的优越性。

Comments Project Page: https://co-director-agent.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29844 2026-04-29 cs.RO cs.AI cs.CV cs.LG 62%

DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA

DIAL: 通过潜在世界建模解耦意图与动作以实现端到端VLA

Yi Chen, Yuying Ge, Hui Zhou, Mingyu Ding, Yixiao Ge, Xihui Liu

机构 * The University of Hong Kong(香港大学) XPENG Robotics(小鹏机器人) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 DIAL通过潜在意图瓶颈解耦意图与动作,利用VLM进行潜在世界建模并结合轻量策略实现端到端VLA,实验表明其在RoboCasa GR1任务中优于现有方法,且在真实世界部署中表现稳健。

Comments Project page: https://xpeng-robotics.github.io/dial

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03043 2026-04-29 cs.CV 57%

OneThinker: All-in-one Reasoning Model for Image and Video

OneThinker:面向图像和视频的统一推理模型

Kaituo Feng, Manyuan Zhang, Hongyu Li, Kaixuan Fan, Shuang Chen, Yilei Jiang, Dian Zheng, Peiwen Sun, Yiyuan Zhang, Haoze Sun, Yan Feng, Peng Pei, Xunliang Cai, Xiangyu Yue

机构 * MMLab, CUHK(CUHK多媒体实验室) Meituan Home(美团家)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 OneThinker提出一个统一的多模态推理模型,整合图像和视频理解,涵盖问答、描述生成、空间时间定位、跟踪和分割等任务,通过构建大规模训练语料和EMA-GRPO算法提升多任务强化学习效果。

Comments CVPR 2026, Project page: https://github.com/tulerfeng/OneThinker

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 9 篇

2604.25273 2026-04-29 cs.CV 90%

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval

对抗大多模态模型中的视觉忽视与语义漂移以提升跨模态检索

Guosheng Zhang, Linkai Liu, Keyao Wang, Haixiao Yue, Zhiwen Tan, Xiao Tan

机构 * Baidu Inc(百度公司)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 本文提出SSA-ME框架,通过显式建模显著视觉主体,提升细粒度表示学习,解决多模态检索中的语义漂移和视觉模态忽视问题,实验表明其在MMEB基准上达到最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25533 2026-04-29 cs.CV 57%

DualGeo: A Dual-View Framework for Worldwide Image Geo-localization

DualGeo: 一个用于全球图像地理定位的双视角框架

Junchao Cui, Wenqi Shi, Shaoyong Du, Hang He, Xuanzi Ma, Hao Tang, Xiangyang Luo

机构 * Henan Key Laboratory of Cyberspace Situation Awareness, Zhengzhou, China(河南空域态势感知重点实验室,郑州,中国) Information Engineering University, Zhengzhou, China(信息工程大学,郑州,中国)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 DualGeo通过融合图像与语义分割特征构建地理表示,并利用双视角对比学习与地理聚类提升全球图像定位精度,实验表明其在街道和城市级别定位准确率提升显著。

Comments ICME2026 Accept

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25390 2026-04-29 cs.IR cs.CV 57%

GeoSearch: Augmenting Worldwide Geolocalization with Web-Scale Reverse Image Search and Image Matching

GeoSearch:通过网络级反向图像搜索和图像匹配增强全球地理定位

Tung-Duong Le-Duc, Hoang-Quoc Nguyen-Son, Minh-Son Dao

机构 * University of Science, VNU-HCM(越南国家科学大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 本文提出GeoSearch框架,通过整合网络级反向图像搜索与图像匹配技术,提升全球图像地理定位的准确性,实验表明其在考虑泄漏因素下的优越性。

Comments Accepted to SIGIR 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25296 2026-04-29 cs.CL 57%

Learning from Medical Entity Trees: An Entity-Centric Medical Data Engineering Framework for MLLMs

从医学实体树学习:一种以实体为中心的医疗数据工程框架用于多模态大语言模型

Jianghang Lin, Haihua Yang, Deli Yu, Kai Wu, Kai Ye, Jinghao Lin, Zihan Wang, Yuhang Wu, Liujuan Cao

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, China(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学,中国) ByteDance(字节跳动) Northeastern University(东北大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 本文提出以实体为中心的医疗数据工程框架,通过构建医学实体树,提升多模态大语言模型在医疗领域的表现,通过实体引导检索、双重过滤和知识感知数据合成等方法,增强模型处理复杂临床问题的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25102 2026-04-29 cs.CV 57%

One Perturbation, Two Failure Modes: Probing VLM Safety via Embedding-Guided Typographic Perturbations

一个扰动,两种失效模式:通过嵌入引导的字形扰动探测VLM安全性

Ravikumar Balakrishnan, Sanket Mendapara

机构 * Cisco Systems(思科系统)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 本文通过实证研究揭示多模态嵌入距离对VLM攻击成功率的预测作用,并提出基于嵌入引导的字形扰动方法,验证了可读性与安全对齐的交互影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25732 2026-04-29 cs.IR 50%

Personalized Multi-Interest Modeling for Cross-Domain Recommendation to Cold-Start Users

面向冷启动用户的跨域推荐的个性化多兴趣建模

Xiaodong Li, Jiawei Sheng, Jiangxia Cao, Xinghua Zhang, Wenyuan Zhang, Yong Sun, Shirui Pan, Zhihong Tian, Tingwen Liu

专题命中 跨模态检索 :multimodal(abstract)

AI总结 本文提出NF-NPCDR框架,通过个性化偏好编码器和共同偏好编码器捕捉用户多兴趣偏好,结合随机自适应解码器提升冷启动用户推荐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25489 2026-04-29 physics.acc-ph cs.LG 50%

Adaptable phase retrieval for coherent transition radiation spectroscopy based on differentiable physics information

基于可微物理信息的适应性相位恢复用于相干跃迁辐射光谱学

Ritz Ann Aguilar, Maxwell LaBerge, Andreas Doepp, Alexander Debus, Zewu Bi, Michael Bussmann, Arie Irman, Ulrich Schramm, Jeffrey Kelling

机构 * Institute of Radiation Physics, Helmholtz-Zentrum Dresden-Rossendorf(辐射物理研究所,德累斯顿-罗斯托克研究中心) Centre for Advanced Laser Applications, Ludwig-Maximilians-Universität München(先进激光应用中心,慕尼黑路德维希-马克西米利安大学) Center for Advanced Systems Understanding, Görlitz(先进系统理解中心,戈尔茨) Technische Universität Dresden(德累斯顿技术大学) Technische Universität Chemnitz(切恩茨技术大学)

专题命中 跨模态检索 :multimodal(abstract)

AI总结 本文提出了一种基于可微物理信息的相位恢复方法,用于解决相干跃迁辐射光谱学中束流剖面恢复问题,通过梯度下降优化傅里叶相位并结合物理先验约束,提升重建精度和鲁棒性。

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21745 2026-04-29 cs.DL 50%

AI-Augmented Bibliometric Framework: A Paradigm Shift with Agentic AI for Dynamic, Snippet-Based Research Analysis

增强型文献计量框架:基于代理AI的范式转变用于动态、片段式研究分析

Adela Bara, Simona-Vasilica Oprea

专题命中 跨模态检索 :multimodal(abstract)

AI总结 本文提出一个生成式多代理AI框架,通过自然语言指令实现动态代码基文献计量分析,无需专业编程技能,支持多模态全文检索、代理探索和动态指标创建,突破传统工具的限制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01553 2026-04-29 cs.IR 50%

IoDResearch: Deep Research on Private Heterogeneous Data via the Internet of Data

IoDResearch:通过互联网的数据进行私有异构数据的深度研究

Zhuofan Shi, Zijie Guo, Xinjian Ma, Gang Huang, Yun Ma, Xiang Jing

专题命中 跨模态检索 :multimodal(abstract)

AI总结 本文提出IoDResearch框架,通过将异构资源转化为FAIR合规的数字对象,提升私有数据的检索效率和可重用性,实验表明其在检索、问答和报告生成任务中优于现有基线方法。

Comments Accepted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 8 篇

2603.12118 2026-04-29 cs.LG cs.DC 89%

Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models

Cornserve:一种用于任意到任意多模态模型的分布式服务系统

Jae-Won Chung, Jeff J. Ma, Jisang Ahn, Yizhuo Liang, Akshay Jajoo, Myungjin Lee, Mosharaf Chowdhury

机构 * University of Michigan(密歇根大学) University of Southern California(南加州大学) Cisco Research(思科研究)

专题命中 多模态生成 :any-to-any(title,abstract);multimodal(title,abstract)

AI总结 本文提出Cornserve,一种支持任意到任意多模态模型的分布式服务系统,通过灵活的任务抽象和组件解耦实现高效部署,提升了吞吐量和延迟性能。

Comments CAIS 2026 Demo track | Open source at https://github.com/cornserve-ai/cornserve | Demo video at https://www.youtube.com/watch?v=nb8R-vztLRg

详情

展开后加载摘要…

URL PDF HTML 收藏