arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 3450 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 3450 篇

2603.29259 2026-04-01 cs.IR cs.CL 79%

Aligning Multimodal Sequential Recommendations via Robust Direct Preference Optimization with Sparse MoE

通过鲁棒直接偏好优化与稀疏MoE对齐多模态序列推荐

Hejin Huang, Jusheng Zhang, Kaitong Cai, Jian Wang, Rong Pan

机构 * Sun Yat-sen University(中山大学) Snap Inc(Snap公司)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 本文研究了在隐式反馈下直接偏好优化的行为,通过系统实验比较了常见负样本选择策略及其与DPO训练的交互,发现用动态top-K候选池进行随机采样可提升排名性能,原因在于减少错误抑制梯度和保留信息硬信号。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26695 2026-03-31 eess.SP cs.AI cs.LG eess.IV quant-ph 79%

Complementarity-Preserving Generative Theory for Multimodal ECG Synthesis: A Quantum-Inspired Approach

保持互补性的多模态ECG生成理论:一种受量子启发的方法

Timothy Oladunni, Farouk Ganiyu-Adewumi, Clyde Baidoo, Kyndal Maclin

机构 * Morgan State University(摩根州立大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出保持互补性的生成理论,通过量子启发框架Q-CFD-GAN生成多模态ECG数据,减少潜在嵌入方差并恢复三领域互补性,提升临床应用的生理意义。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.26669 2026-03-31 cs.IR cs.AI cs.LG 79%

ReCQR: Incorporating conversational query rewriting to improve Multimodal Image Retrieval

ReCQR:将对话查询重写纳入多模态图像检索以提高性能

Yuan Hu, ZhiYu Cao, PeiFeng Li, QiaoMing Zhu

机构 * School of Computer Science and Technology, Soochow University(苏州大学计算机科学与技术学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出ReCQR任务,通过构建多轮对话查询重写数据集,提升多模态图像检索的准确性和用户查询建模能力。

Comments 4 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.24736 2026-03-27 cs.AI cs.LG 79%

AutoSAM: an Agentic Framework for Automating Input File Generation for the SAM Code with Multi-Modal Retrieval-Augmented Generation

AutoSAM:一种用于自动化生成SAM代码输入文件的代理框架,结合多模态检索增强生成

Zaid Abulawi, Zavier Ndum Ndum, Eric Cervi, Rui Hu, Yang Liu

机构 * Department of Nuclear Engineering, Texas A\&M University. Engineering Division, Argonne National Laboratory

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.AI

AI总结 AutoSAM通过多模态检索增强生成技术,自动化生成SAM代码输入文件,解决异构工程文档中提取设计数据并转换为求解器语法的难题,实现100%结构化输入利用和88%PDF文本提取。

Comments 34 Pages, 14 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22821 2026-03-25 cs.CV 79%

Cross-Slice Knowledge Transfer via Masked Multi-Modal Heterogeneous Graph Contrastive Learning for Spatial Gene Expression Inference

跨切片知识迁移 via 遮蔽多模态异质图对比学习用于空间基因表达推断

Zhiceng Shi, Changmiao Wang, Jun Wan, Wenwen Min

机构 * Yunnan University(云南大学) Shenzhen Research Institute of Big Data(深圳大数据研究院) Zhongnan University of Economics and Law(中南财经政法大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出SpaHGC模型,通过多模态异质图对比学习,结合局部空间上下文和跨切片相似性,提升空间基因表达预测精度,优于现有九种方法。

Comments Accepted by CVPR-2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04846 2026-03-24 cs.CV 79%

Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models

多范式协同对抗攻击多模态大语言模型

Yuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou, Xiao-Jun Wu, Josef Kittler

机构 * School of Artificial Intelligence and Computer Science, Jiangnan University(江南大学人工智能与计算机科学学院) Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(Surrey 大学视觉、语音和信号处理中心)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 针对多模态大语言模型的多范式协同对抗攻击方法,通过聚合视觉和语言特征进行联合优化,提升对抗示例的可转移性,实验表明优于现有方法。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02258 2026-03-24 cs.CV 79%

Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning

病理代理RAG:通过强化学习实现多模态代理检索增强生成用于病理学视觉语言模型

Wenchuan Zhang, Jingru Guo, Hengzhe Zhang, Penghao Zhang, Jie Chen, Shuwan Zhang, Zhang Zhang, Yuhao Yi, Hong Bu

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出Patho-AgenticRAG,通过强化学习实现多模态代理检索增强生成,解决病理学视觉语言模型在高分辨率、复杂组织结构和临床语义上的挑战,提升诊断准确性。

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(35): 29921-29929, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20970 2026-03-24 cs.CV 79%

GraPHFormer: A Multimodal Graph Persistent Homology Transformer for the Analysis of Neuroscience Morphologies

GraPHFormer:一种多模态图持久同调变换器,用于神经科学形态学分析

Uzair Shah, Marco Agus, Mahmoud Gamal, Mahmood Alzubaidi, Corrado Cali, Pierre J. Magistretti, Abdesselam Bouzerdoum, Mowafa Househ

机构 * Hamad Bin Khalifa University(哈马德·本·卡西姆大学) University of Turin(都灵大学) BESE, King Abdullah University of Science and Technology(贝赛,国王阿卜杜勒阿齐兹大学科学与技术学院) University of Wollongong(沃林根大学) Neuroscience Institute Cavalieri Ottolenghi(卡瓦利埃-奥托伦奇神经科学研究所) Université Grenoble-Alpes(格勒诺布尔阿尔卑斯大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 GraPHFormer通过CLIP式对比学习统一拓扑和图结构分析,利用持久图像编码和树状LSTM编码器,实现对神经形态的高精度识别与分类,优于传统方法。

Comments Accepted to IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20181 2026-03-23 cs.CR cs.AI 79%

Improving Generalization on Cybersecurity Tasks with Multi-Modal Contrastive Learning

通过多模态对比学习提升网络安全任务的泛化能力

Jianan Huang, Rodolfo V. Valentim, Luca Vassio, Matteo Boffa, Marco Mellia, Idilio Drago, Dario Rossi

机构 * Huawei Paris Research Center, France(华为巴黎研究中心,法国)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出多模态对比学习框架,通过文本指导 payloads 分类,提升网络安全任务的泛化能力,并在合成基准和真实数据集上验证效果。

Comments Submitted to Euro S&P - 5th International Workshop on Designing and Measuring Security in Systems with AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20016 2026-03-23 cs.CV 79%

CFCML: A Coarse-to-Fine Crossmodal Learning Framework For Disease Diagnosis Using Multimodal Images and Tabular Data

CFCML:一种用于多模态图像和表格数据疾病诊断的粗到细跨模态学习框架

Tianling Liu, Hongying Liu, Fanhua Shang, Lequan Yu, Tong Han, Liang Wan

机构 * College of Intelligence and Computing(智能与计算学院) Tianjin University(天津大学) Medical School of Tianjin University(天津大学医学院) Peng Cheng Lab(鹏城实验室) Department of Statistics and Actuarial Science, School of Computing and Data Science, The University of Hong Kong(统计与精算系,计算与数据科学学院,香港大学) Department of Radiology, Tianjin Huanhu Hospital(天津华医院放射科) Tianjin Key Laboratory of Cerebral Vascular and Neurodegenerative Diseases(天津脑血管与神经退行性疾病重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出CFCML框架,通过粗到细的跨模态学习逐步缩小多模态数据间的模态差距,提升诊断准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15623 2026-03-18 cs.IR cs.AI 79%

Finder: A Multimodal AI-Powered Search Framework for Pharmaceutical Data Retrieval

Finder:一种多模态AI驱动的制药数据检索框架

Suyash Mishra, Srikanth Patil, Satyanarayan Pati, Sagar Sahu, Baddu Narendra

机构 * Researcher, Global Product Strategy F. Hoffmann-La Roche Ltd. Basel, Switzerland(全球产品战略研究员 罗氏有限公司 巴塞尔,瑞士) Associate Vice President (Gen AI) Involead Services Pvt Ltd. Pune, India(高级副总裁(生成式人工智能) Involead服务私人有限公司 普纳,印度) Lead Data Scientist Involead Services Pvt Ltd. Delhi, India(首席数据科学家 Involead服务私人有限公司 德里,印度) Data Scientist Involead Services Pvt Ltd. Bhubaneswar, India(数据科学家 Involead服务私人有限公司 奇塔拉,印度)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 Finder利用混合向量搜索统一文本、图像、音频和视频的检索,通过稀疏词汇和密集语义模型提升多模态内容处理能力,支持自然语言推理搜索,已处理超过29万文档、3.1万视频和1192音频文件。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15341 2026-03-17 cs.AI cs.HC cs.MA 79%

Intelligent Co-Design: An Interactive LLM Framework for Interior Spatial Design via Multi-Modal Agents

智能协同设计:一种基于多模态代理的交互式LLM框架用于室内空间设计

Ren Jian Lim, Rushi Dai

机构 * Hong Kong Center for Construction Robotics(香港建设机器人中心)

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.AI

AI总结 本文提出一种基于LLM的多模态多代理框架,通过自然语言描述和图像动态生成3D设计,提升用户参与度和设计效率。

Comments 25 pages, 20 figures; accepted for publication in the Proceedings of ACADIA 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19554 2026-03-17 cs.LG cs.AI 79%

CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning

CARE:对比锚定反射用于可验证多模态推理

Yongxin Wang, Zhicheng Yang, Meng Cao, Mingfei Han, Haokun Lin, Yingying Zhu, Xiaojun Chang, Xiaodan Liang

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 CARE通过对比锚定反射框架,将错误转化为监督信号,提升多模态推理的准确性和训练平滑度,在六个可验证视觉推理基准上提升4.6个百分点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04268 2026-03-12 cs.CV 79%

KVSmooth: Mitigating Hallucination in Multi-modal Large Language Models through Key-Value Smoothing

KVSmooth: 通过键值平滑缓解多模态大语言模型中的幻觉

Siyu Jiang, Feiyang Chen, Xiaojin Zhang, Kun He

专题命中 跨模态检索 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 KVSmooth通过键值平滑技术有效缓解多模态大语言模型中的幻觉问题,提升生成精度和召回率。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07997 2026-03-10 cs.AI 79%

CMMR-VLN: Vision-and-Language Navigation via Continual Multimodal Memory Retrieval

CMMR-VLN:通过持续多模态记忆检索实现视觉与语言导航

Haozhou Li, Xiangyu Dong, Huiyan Jiang, Yaoming Zhou, Xiaoguang Ma

机构 * Foshan Graduate School of Innovation at Northeastern University(东北大学创新研究生院) Faculty of Robot Science and Engineering at Northeastern University(东北大学机器人科学与工程学院) College of Software at Northeastern University(东北大学软件学院) School of Aeronautic Science and Engineering at Beihang University(北航航空科学与工程学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 CMMR-VLN通过引入持续多模态记忆检索机制,提升视觉与语言导航任务中对先前经验的选择性利用能力,显著提高导航成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07039 2026-03-10 cs.AI 79%

Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding

具有4D空间-时间嵌入的自监督多模态世界模型

Lance Legel, Qin Huang, Brandon Voelker, Daniel Neamati, Patrick Alan Johnson, Favyen Bastani, Jeff Rose, James Ryan Hennessy, Robert Guralnick, Douglas Soltis, Pamela Soltis, Shaowen Wang

机构 * Ecological Intelligence Lab(生态智能实验室) School of Complex Adaptive Systems(复杂适应系统学院) University of Houston(休斯顿大学) Geosensing Systems Engineering & Sciences Lab(传感系统工程与科学实验室) Stanford University(斯坦福大学) Allen Institute for Artificial Intelligence(人工智能研究院) Spatial Intelligence Lab(空间智能实验室) Department of Computer Science(计算机科学系) Georgia Institute of Technology(佐治亚理工学院) Florida Museum of Natural History(佛罗里达自然历史博物馆) University of Florida(佛罗里达大学) NSF Institute for Geospatial Understanding(国家科学基金会地理理解研究所) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

AI总结 DeepEarth通过4D空间-时间嵌入实现自监督多模态世界模型,在生态预测中取得最佳性能。

Comments 8 pages, 5 figures, 1 table. Presented at 2026 World Modeling Workshop, Mila Quebec

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06982 2026-03-10 cs.CV cs.IR 79%

Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning

优化多模态模型用于基于图像的形状检索:预对齐和硬对比学习的作用

Paul Julius Kühn, Cedric Spengler, Michael Weinmann, Arjan Kuijper, Saptarshi Neil Sinha

机构 * Fraunhofer IGD(弗劳恩霍夫图像研究中心) Delft University of Technology(代尔夫特理工大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出通过预对齐和硬对比学习优化多模态模型,提升基于图像的形状检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06503 2026-03-09 cs.CL 79%

Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editing

超越行到推理:面向多模态电子表格理解与编辑的智能检索

Anmol Gulati, Sahil Sen, Waqar Sarguroh, Kevin Paul

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 BRTR通过迭代工具调用框架实现多模态电子表格的端到端理解与编辑,取得领先性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01511 2026-03-09 cs.AI 79%

Multimodal Mixture-of-Experts with Retrieval Augmentation for Protein Active Site Identification

多模态专家混合模型与检索增强用于蛋白质活性位点识别

Jiayang Wu, Jiale Zhou, Rubo Wang, Xingyi Zhang, Xun Lin, Tianxu Lv, Leong Hou U, Yefeng Zheng

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 MERA通过检索增强和可靠性引导的多专家模型,实现了蛋白质活性位点识别的高精度预测和性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05446 2026-03-06 cs.CV 79%

NaiLIA: Multimodal Nail Design Retrieval Based on Dense Intent Descriptions and Palette Queries

NaiLIA:基于密集意图描述和调色板查询的多模态美甲设计检索

Kanon Amemiya, Daichi Yashima, Kei Katsumata, Takumi Komatsu, Ryosuke Korekata, Seitaro Otsuki, Komei Sugiura

机构 * Keio University(庆应大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 NaiLIA通过密集意图描述和调色板查询实现美甲设计图像的多模态检索,提升检索精度。

Comments Accepted to CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02929 2026-03-05 cs.CV 79%

TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval

TRACE:面向通用多模态检索的任务自适应推理与表征学习

Xiangzhao Hao, Shijie Wang, Tianyu Yang, Tianyue Wang, Haiyun Guo, Jinqiao Wang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 TRACE通过结合生成推理与判别表征学习,实现通用多模态检索中的任务自适应推理,提升检索准确率与推理效率,同时具备强零样本迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02047 2026-03-03 cs.CV 79%

NICO-RAG: Multimodal Hypergraph Retrieval-Augmented Generation for Understanding the Nicotine Public Health Crisis

NICO-RAG:多模态超图检索增强生成用于理解尼古丁公共卫生危机

Manuel Serna-Aguilera, Raegan Anderes, Page Dobbs, Khoa Luu

机构 * University of Arkansas(亚拉巴马大学) University of Arkansas for Medical Sciences(亚拉巴马大学医学科学分校)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 NICO-RAG通过多模态超图检索增强生成,解决尼古丁公共卫生危机中的信息检索与生成问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00511 2026-03-03 cs.CV cs.LG 79%

Multimodal Adaptive Retrieval Augmented Generation through Internal Representation Learning

多模态自适应检索增强生成通过内部表示学习

Ruoshuang Du, Xin Sun, Qiang Liu, Bowen Song, Zhongqi Chen, Weiqiang Wang, Liang Wang

机构 * Shanghaitech University School of Information Science(上海科技大学信息科学学院) Chinese Academy of Sciences Institute of Automation(中国科学院自动化研究所)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出MMA-RAG,通过动态评估模型内部知识置信度,提升视觉问答系统在多模态场景下的检索增强生成性能。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13104 2026-03-03 cs.LG cs.CL 79%

Equitable Electronic Health Record Prediction with FAME: Fairness-Aware Multimodal Embedding

基于FAME的公平电子健康记录预测:具有公平性的多模态嵌入

Nikkie Hooman, Zhongjie Wu, Eric C. Larson, Mehak Gupta

机构 * Department of Computer Science(计算机科学系)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CL

AI总结 FAME通过公平性意识的多模态嵌入框架,在电子健康记录预测中实现性能与公平性的双重优化。

Comments 21 pages, 3 figures

Journal ref Proceedings of Machine Learning Research 2025 (Machine Learning for Healthcare, ML4H)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10030 2026-03-03 cs.CR cs.AI 79%

Safeguarding Multimodal Knowledge Copyright in the RAG-as-a-Service Environment

在RAG-as-a-Service环境中保护多模态知识版权

Tianyu Chen, Jian Lou, Wenjie Wang

机构 * ShanghaiTech University(上海科技大学) Sun Yat-sen University(中山大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 AQUA是一种用于多模态RAG系统中图像知识保护的首个水印框架,通过嵌入语义信号实现高效、隐蔽的版权追溯。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22217 2026-02-27 cs.IR cs.AI 79%

RAGdb: A Zero-Dependency, Embeddable Architecture for Multimodal Retrieval-Augmented Generation on the Edge

RAGdb: 一种无依赖、可嵌入的多模态检索增强生成边缘计算架构

Ahmed Bin Khalid

机构 * SKAS IT

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 RAGdb通过单文件SQLite容器实现多模态检索增强生成,无需GPU推理,提升边缘计算效率与数据主权

Comments 6 pages, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20877 2026-02-25 cs.IR cs.AI 79%

E-MMKGR: A Unified Multimodal Knowledge Graph Framework for E-commerce Applications

E-MMKGR:面向电子商务应用的统一多模态知识图谱框架

Jiwoo Kang, Yeon-Chang Lee

机构 * UNIST(全南国立科学技术院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 E-MMKGR提出了一种统一多模态知识图谱框架,通过GNN和图谱优化学习统一物品表示,提升电子商务推荐和搜索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19735 2026-02-24 cs.CV 79%

VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments

VGGT-MPR:基于VGGT的多模态地点识别在自动驾驶环境中的应用

Jingyi Xu, Zhangshuo Qi, Zhongmiao Yan, Xuyu Gao, Qianyun Jiao, Songpengcheng Xia, Xieyuanli Chen, Ling Pei

机构 * Shanghai Jiao Tong University(上海交通大学) Beijing Institute of Technology(北京理工大学) National University of Defense Technology(国防科技大学)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 VGGT-MPR通过视觉几何基础的Transformer实现多模态地点识别,在自动驾驶中表现出强鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12774 2026-02-16 cs.CV 79%

Bootstrapping MLLM for Weakly-Supervised Class-Agnostic Object Counting

基于MLLM的弱监督类无关物体计数

Xiaowen Zhang, Zijie Yue, Yong Luo, Cairong Zhao, Qijun Chen, Miaojing Shi

机构 * College of Electronic and Information Engineering, Tongji University(电子信息工程学院,同济大学) College of Computer Science, Wuhan University(计算机科学学院,武汉大学) College of Computer Science, Tongji University(计算机科学学院,同济大学) State Key Laboratory of Autonomous Intelligent Unmanned Systems(自主智能无人系统国家重点实验室)

专题命中 跨模态检索 :MLLM(title,abstract);分类 cs.CV

AI总结 本文提出 WS-COC,基于 MLLM 的弱监督框架,通过三种策略提升类无关物体计数性能,有效降低标注成本。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05096 2026-02-16 cs.CV cs.LG 79%

Visual concept ranking uncovers medical shortcuts used by large multimodal models

视觉概念排名揭示大型多模态模型使用的医学捷径

Joseph D. Janizek, Sonnet Xu, Junayd Lateef, Roxana Daneshjou

机构 * Stanford University(斯坦福大学) University of California Berkeley(加州大学伯克利分校)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出视觉概念排名方法,用于揭示大型多模态模型在医疗任务中表现的潜在视觉特征依赖性。

详情

展开后加载摘要…

URL PDF HTML 收藏