Hi-BEHRT: Hierarchical Transformer-based model for accurate prediction of clinical events using multimodal longitudinal electronic health records
专题命中 跨模态检索 :multimodal(title,abstract)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 跨模态检索 :multimodal(title,abstract)
专题命中 跨模态检索 :multimodal(title,abstract)
专题命中 跨模态检索 :cross-modal(title,abstract)
Comments For further details, please visit https://arg-nctu.github.io/projects/deeprl-mmWave.html
Journal ref IEEE Robotics and Automation Letters, 2021
专题命中 跨模态检索 :multi-modal(title,abstract)
专题命中 跨模态检索 :multi-modal(title,abstract)
Comments 10 pages, 7 figures
Journal ref KDD 2020
专题命中 跨模态检索 :multimodal(title,abstract)
Comments Accepted to International Conference on Robotics and Automation (ICRA) 2020, IEEE copyright
专题命中 跨模态检索 :multi-modal(title,abstract)
专题命中 跨模态检索 :cross-modal(title,abstract)
专题命中 跨模态检索 :multimodal(title,abstract)
Comments 15 pages, 10 figures
专题命中 跨模态检索 :cross-modal(title,abstract)
专题命中 跨模态检索 :multimodal(title,abstract)
专题命中 跨模态检索 :multi-modal(title,abstract)
Comments IEEE Transactions on Software Engineering
专题命中 跨模态检索 :multi-modal(title,abstract)
Comments 6 pages, 4 figures
专题命中 跨模态检索 :cross-modal(title,abstract)
Comments 7 pages, 11 figures
专题命中 跨模态检索 :multi-modal(title,abstract)
专题命中 跨模态检索 :cross-modal(title,abstract)
专题命中 跨模态检索 :multimodal(title,abstract)
Comments We need to correct certain errors both in the software description as well as in the algorithms
专题命中 跨模态检索 :cross-modal(title,abstract)
Comments 12 pages
专题命中 跨模态检索 :multimodal(title,abstract)
专题命中 跨模态检索 :multimodal(title,abstract)
Comments 11 pages, 5 figures
作为辅助监督的生成:通过解耦嵌入预测以零推理开销增强视觉理解
机构 * ByteDance(字节跳动)
专题命中 跨模态检索 :multimodal(abstract,abstract_cn);cross-modal(abstract);分类 cs.CV
AI总结 本研究提出GAS框架,将视觉生成作为辅助监督,通过解耦MoT架构的NEP实现零推理开销,提升了多模态理解尤其是感知与空间理解能力。
多样化意图的多轮时尚图像检索
专题命中 跨模态检索 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV
AI总结 研究现实世界时尚多轮图像检索问题,提出FashionAM框架直接对齐多模态对话查询与时尚图库嵌入空间,避免文本化,引入DIM-Fashion数据集,实验证明该框架优于现有方法,数据集和代码将公开。
通过优势加权排序揭示二元抑郁检测中的潜在抑郁严重程度
机构 * South China Normal University(华南师范大学)
专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.AI
AI总结 研究利用视听数据检测抑郁的挑战,提出含时间编码器和相互变压器的多模态框架,核心是二元优势加权排序损失,通过两种机制优化潜在空间分布,实验表明该模型重建潜在有序结构,性能达最优。
协同感知-推理治理:用可验证解剖证据为医学多模态大语言模型提供基础
机构 * Huazhong University of Science and Technology(华中科技大学) ; Imperial Global Singapore, Imperial College London(帝国理工学院新加坡全球中心) ; Nanyang Technological University(南洋理工大学) ; National University of Singapore(新加坡国立大学) ; Anhui University(安徽大学)
专题命中 跨模态检索 :MLLM(summary_cn);multimodal(abstract);分类 cs.CV
AI总结 提出一种无需训练的证据注入框架,通过ROI引导的视觉激活调制和解剖坐标语义标记,协同校准视觉感知与文本推理,动态路由任务特定干预,有效减少医学MLLM的幻觉。
Comments Accepted by MICCAI 2026 (Early Accept, Top 9%)
非结构化度假租赁图像集合中的房间场景发现与分组
机构 * Expedia Group(Expedia集团)
专题命中 跨模态检索 :MLLM(abstract,abstract_cn);multi-modal(abstract);分类 cs.CV
AI总结 针对度假租赁平台图片缺乏结构化分类的问题,提出一种低延迟、样本高效的机器学习流水线,通过监督式房间类型检测、重叠检测和聚类算法实现房间分组,并利用多模态大语言模型识别床型,显著优于对比学习和预训练嵌入聚类方法。
Comments Presented at the Two-sided Marketplace Optimization Workshop, KDD 2025
LightSTAR: 通过视觉自适应精炼的轻量级选择实现高效视觉文档检索
机构 * MoE Key Lab of Artificial Intelligence, AI Institute, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学与工程学院人工智能研究院教育部人工智能重点实验室)
专题命中 跨模态检索 :MLLM(abstract,abstract_cn);multi-modal(abstract);分类 cs.CV
AI总结 提出LightSTAR框架,通过免LLM的视觉选择快速筛选候选页,再经视觉自适应语义精炼进行细粒度匹配,在保持高检索精度的同时大幅降低延迟。
Comments Accpeted by ECCV 2026
MIRAGE:基于层次分解的多向量图像检索运行时调度
机构 * School of Computer Science, Peking University(北京大学计算机科学学院) ; School of Electronics Engineering and Computer Science, Peking University(北京大学电子工程与计算机科学学院) ; School of Information, Renmin University of China(中国人民大学信息学院) ; School of Integrated Circuit Science and Engineering, Beihang University(北京航空航天大学集成电路科学与工程学院)
专题命中 跨模态检索 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV
AI总结 提出MIRAGE框架,通过层次化分解和跨层次相似性一致性减少冗余计算,实现多向量图像检索的精度提升和3.5倍计算加速。
Comments Will appear in DAC'2026, camera ready
弥合法医图像检索中的模态差距
机构 * Advanced Technologies Application Center (CENATAV)(先进技术应用中心(CENATAV)) ; Centro de Sistemas Complejos, Facultad de Física, Universidad de La Habana(哈瓦那大学物理学院复杂系统中心)
专题命中 跨模态检索 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV
AI总结 提出统一检索框架,利用多模态大语言模型生成文本描述并结合视觉与文本特征融合,提升纹身、人脸素描等法医任务的检索精度与鲁棒性。
Comments 23 pages, 5 figures, paper submitted to Elsevier journal
混合模态双人脸-发型检索
机构 * Vietnam National University, Ho Chi Minh City, Vietnam(越南国家大学,胡志明市,越南) ; University of Information Technology, VNU-HCM, Ho Chi Minh City, Vietnam(信息技术大学,VNU-HCM,胡志明市,越南)
专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);image-text(abstract);分类 cs.CV
AI总结 提出混合模态双参考检索任务DFHR,通过解耦身份与发型特征并融合多模态嵌入,实现跨模态的身份感知与属性可控检索。
ValueGround: 评估多模态大语言模型中文化条件化的视觉价值基础
机构 * University of Technology Nuremberg (UTN)(图恩大学)
专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract_cn);cross-modal(abstract);分类 cs.CL
AI总结 提出ValueGround基准,通过最小对比图像对评估多模态大语言模型在文化条件化视觉价值判断中的表现,发现模型在可视化选项下准确率显著低于文本选项。
Comments Updated preprint