EsurvFusion: An evidential multimodal survival fusion model based on Gaussian random fuzzy numbers
专题命中 多模态训练与对齐 :multimodal(title,abstract)
Comments Multimodal survival analysis, Epistemic random fuzzy sets theory, Uncertainty
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态训练与对齐 :multimodal(title,abstract)
Comments Multimodal survival analysis, Epistemic random fuzzy sets theory, Uncertainty
专题命中 多模态训练与对齐 :multi-modal(title,abstract)
Comments Appendix for the IEEE FUSION 2019 submission on multi-modal variational Autoencoders for sensor fusion
仅训练投影器就足够
机构 * Ramen VR(拉面VR公司) ; University of California, Berkeley(加州大学伯克利分校)
专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL
AI总结 该研究探究多模态大语言模型适配新模态是否需微调主干,经实验发现仅训练投影器即可实现强多模态性能,还能避免联合训练导致的语言模型能力漂移,且训练样本吞吐量约为联合训练的两倍,通过多类基准验证了结论。
真相留在家族中:通过模型谱系中继承的真相头增强上下文基础
机构 * University of Science and Technology of China(中国科学技术大学)
专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CL、cs.AI
AI总结 研究发现基础LLM与下游变体间存在上下文真相分数的强继承性,提出TruthProbe软门控策略放大真相头以提升上下文真实性并减少多模态幻觉。
Comments Accepted at ICML 2026
OPD-V:结合模态平衡的视觉在线策略自蒸馏
机构 * National University of Singapore(新加坡国立大学) ; Ludwig Maximilian University of Munich(慕尼黑大学) ; Munich Center for Machine Learning(慕尼黑机器学习中心) ; Sun Yat-sen University(中山大学)
专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 该研究针对多模态大语言模型的模态不平衡问题,提出视觉在线策略自蒸馏范式OPD-V,通过正、负教师模型实现模态平衡,在多基准与骨干上提升推理性能并降低训练成本。
Comments Corrected the uploaded manuscript. Project Page:https://github.com/aniri15/OPD-V
Omni-Prune:用于高效全模态大语言模型的查询感知统一令牌剪枝
机构 * Nanjing University(南京大学)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV、cs.CL
AI总结 研究针对全模态大语言模型推理时音频-视频令牌序列长、预填充延迟高和GPU内存使用量大的问题,提出无需训练的查询感知Omni-Prune框架,联合去除冗余并保留跨模态证据,实验证明其性能优于基线方法,能加速预填充并减少内存。
Comments 14 pages, 7 figures. Code: https://github.com/kimberlyii/Omni-Prune
看一次就够了?用于3D问答的在线几何感知令牌剪枝
机构 * National Tsing Hua University(国立清华大学) ; Industrial Technology Research Institute ITRI(工业技术研究院) ; University of Toronto(多伦多大学) ; Atmanity Inc(Atmanity公司)
专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multi-modal(abstract);分类 cs.CV、cs.AI
AI总结 针对3D问答中多模态大语言模型推理成本高的问题,提出在线令牌剪枝方法,利用深度和相机姿态投影到体素空间,识别重叠区域并剪枝冗余令牌,减少令牌使用,提升效率和性能。
Comments published at ICLR 2026 Workshop on Efficient Spatial Reasoning
BioMatrix:迈向覆盖序列、结构和语言模态矩阵的综合性生物基础模型
机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院) ; OpenDataLab, Shanghai Artificial Intelligence Laboratory(上海人工智能实验室 OpenDataLab) ; Zhejiang University(浙江大学) ; Shanghai Innovation Institute(上海创新研究院) ; East China Normal University(华东师范大学) ; Zhongguancun Academy(中关村学院) ; School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CL、cs.AI
AI总结 提出首个原生多模态生物基础模型BioMatrix,通过统一分词方案将分子序列、结构、蛋白质序列、结构和自然语言映射到共享离散标记空间,在单一解码器架构下实现所有模态的统一生成,在80项任务中77项达到最优或竞争力水平。
VITAL: 视觉-语义双重监督增强可解释的医学多模态大语言模型潜在推理
机构 * Zhejiang University(浙江大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Tencent(腾讯) ; Ningbo Global Innovation Center, Zhejiang University(宁波全球创新中心,浙江大学) ; Zhejiang Key Laboratory of Digital-Intelligence Service Technology(浙江省数字智能服务技术重点实验室)
专题命中 多模态训练与对齐 :MLLM(summary_cn,abstract_cn);分类 cs.CV、cs.AI
AI总结 提出VITAL框架,通过视觉-语义双重监督(文本解码器重构推理链、视觉投影器回归ROI特征)实现医学MLLM的可解释潜在推理,在7个基准上达到SOTA。
GA-VLN: 用于高效视觉-语言导航的几何感知鸟瞰图表示
机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院) ; University of Chinese Academy of Sciences(中国科学院大学) ; Robbyant ; School of Computing, National University of Singapore(新加坡国立大学计算机学院) ; The Hong Kong University of Science and Technology(香港科技大学)
专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 本文提出GA-VLN框架,通过引入几何感知的鸟瞰图表示(GA-BEV),整合显式和隐式几何信息,提升视觉-语言导航的效率和性能,实验表明其在仅使用导航数据的情况下取得了最先进的结果。
通过强化学习为深度不平衡回归注入分布意识
机构 * The Hong Kong University of Science(香港科学与技术大学)
专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL
AI总结 本文提出基于组相对策略优化的分布感知强化学习框架,通过一致性相关系数奖励实现跨样本关系监督,提升长尾回归任务的分布对齐能力,在中等和少样本场景下表现优异。
Comments Accepted by ICML 2026
各向异性模态对齐
机构 * HKUST(GZ)(香港科技大学(广州)) ; NUS(国立新加坡大学) ; UCSD(加州大学圣地亚哥分校) ; Stanford(斯坦福大学) ; PKU(北京大学) ; THU(清华大学)
专题命中 多模态训练与对齐 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.MM
AI总结 本文研究了多模态模型中模态间转换的可行性,提出各向异性模态对齐方法,通过几何修正框架提升单模态数据的多模态训练效果。
通过协同多模态对齐和训练时融合实现图像与文本一体化:ITO
机构 * School of Computer Science(计算机科学学院) ; Huazhong University of Science and Technology(华中科技大学) ; Li Auto Inc.(力汽车公司) ; Institute of AI for Industries, Chinese Academy of Sciences(产业人工智能研究院,中国科学院)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI
AI总结 ITO通过协同多模态对齐与训练时融合机制,提升图像与文本表示的一致性,有效解决模态间结构化交互问题。
跨模态冲突下的全方位安全:漏洞、动态机制和高效对齐
机构 * Nanyang Technological University(南洋理工大学) ; Beijing University of Posts and Telecommunications(北京邮电大学) ; Tsinghua University(清华大学) ; Fudan University(复旦大学) ; University of Science and Technology of China(中国科学技术大学) ; Renmin University of China(中国人民大学)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);omni-modal(abstract);分类 cs.CL、cs.AI
AI总结 本文提出OmniSteer方法,通过提取黄金拒绝向量和轻量级适配器提升多模态模型的安全性与通用能力。
机构 * Department of Data Science(数据科学系) ; New York University Shanghai(纽约大学上海分校)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI
Comments Capstone Paper
机构 * School of Software, Northwestern Polytechnical University, Shaanxi, China(西北工业大学软件学院) ; Yangtze River Delta Research Institute of Northwestern Polytechnical University, Taicang, China(西北工业大学长江三角研究 institute) ; Department of Computer Science and Information Technology, La Trobe University, Melbourne, Australia(拉筹伯大学计算机科学与信息技术系) ; CSIRO Data61, Sydney, Australia(CSIRO Data61) ; School of Intelligence Science and Technology, Nanjing University, Suzhou, China(南京大学智能科学与技术学院) ; Australian Institute of Health Innovation (AIHI), Macquarie University, Australia(麦考瑞大学健康创新研究所) ; School of Computer Software, College of Intelligence and Computing, Tianjin University, Tianjin, China(天津大学计算机软件学院) ; Centre for Artificial Intelligence driven Drug Discovery, Faculty of Applied Science, Macao Polytechnic University, Macao Special Administrative Region of China(澳门理工学院人工智能驱动药物发现中心) ; Department of Computer Science, School of Engineering, Shantou University, Guangdong, China(汕头大学计算机科学系)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI
Comments The paper has been accepted by the 33rd Pacific Conference on Computer Graphics and Applications (Pacific Graphics 2025)
Journal ref PG2025 Conference Papers, Posters, and Demos, 2025
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CL、cs.AI
机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) ; School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.MM
机构 * Research Center for Social Computing and Information Retrieval, Harbin Institute of Technology(社会计算与信息检索研究中心,哈尔滨工业大学) ; School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学) ; School of Computing, National University of Singapore(计算学院,新加坡国立大学)
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.CL
Comments Accepted by IEEE Transactions on Circuits and Systems for Video Technology (TCSVT). June 2025. DOI: https://doi.org/10.1109/TCSVT.2025.3578266
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI
专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
Comments 6 pages, 3 figures, accepted by IEEE ISBI 2025
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);omni-modal(abstract);分类 cs.CL、eess.AS
专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.MM
Comments 10 pages, 3 figures
专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.CL
Comments 17 pages
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI
超越多模态对齐:通过响应替换与有序执行验证物理语言
机构 * New York University(纽约大学) ; Carnegie Mellon University(卡内基梅隆大学) ; Columbia University(哥伦比亚大学)
专题命中 多模态训练与对齐 :multimodal(title,abstract)
AI总结 本文提出DBOSC方法,在Cluster Haptic数据集与弹塑性系统中验证了多模态物理表示的可执行含义,明确了执行器、图表等因素对动作组合的影响,区分了多项可测试的操作能力成就。
Malformer:一种基于Transformer的多模态恶意软件检测器
专题命中 多模态训练与对齐 :multi-modal(title);multimodal(abstract)
AI总结 本研究提出四模态恶意软件检测模型Malformer,融合文本、图像、图形、音频表示,采用多模态Transformer融合,在201549个样本数据集上准确率98.3%,优于单/双模态检测器,为恶意软件防御提供稳健基础。
室内60GHz网络中多模态传感辅助非视距波束搜索的激光雷达衍生表面先验
专题命中 多模态训练与对齐 :multimodal(title);cross-modal(abstract)
AI总结 该研究提出用激光雷达衍生的局部表面结构作为先验,结合射频测量,可将60GHz网络的波束对探测减少72%,同时保持74.5%位置的波束在详尽搜索的3dB范围内,降低毫米波波束搜索不确定性。
临床路径对早期阿尔茨海默病检测中的多模态深度学习至关重要
专题命中 多模态训练与对齐 :multimodal(title,abstract)
AI总结 该研究提出基于SigLIP的零样本多模态框架,结合结构MRI与常规临床变量,在ADNI队列中实现早期阿尔茨海默病风险分层,性能优于传统模型,且可扩展至纵向数据。
Comments Published