A Keypoint Detection and Description Network Based on the Vessel Structure for Multi-Modal Retinal Image Registration
专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments 6 pages, 4 figures, 1 table, accepted to BVM 2022
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV
Comments 6 pages, 4 figures, 1 table, accepted to BVM 2022
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
Comments Accepted for publication in AAAI-2022
专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV
Comments In Submission
专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
Comments Accepted in Neurocomputing
专题命中 多模态评测 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV
Comments 8 pages
专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV
Comments Accepted at Perception Beyond Visible Spectrum Workshop, CVPR 2019
专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV
Journal ref IEEE Transactions on Medical Imaging, vol. 39, no. 5, pp. 1703-1711, May 2020
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV
Comments Accepted at CVPR 2020. For a demo video, see http://tiny.cc/xmuda
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL
专题命中 多模态评测 :multimodal(title,abstract);multi-modal(abstract);分类 cs.CL
Comments 6 pages, 3 figures
专题命中 多模态评测 :cross-modal(title,abstract);multi-modal(abstract);分类 cs.CV
Comments Accepted to CVPR 2019. Project Page: http://visgel.csail.mit.edu/
专题命中 多模态评测 :multi-modal(title,abstract);multimodal(abstract);分类 cs.CV
Comments Working Draft
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV
M$^3$Eval: 通过认知基础视频任务的多模态记忆评估
机构 * School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) ; State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室) ; Yuanpei College, Peking University(北京大学元培学院) ; Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) ; School of Psychological and Cognitive Sciences, Peking University(北京大学心理学与认知科学学院) ; University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 提出首个多模态模型记忆评估框架M$^3$Eval,通过认知心理学设计的视频任务系统评估模型在记忆保持、忠实性和鲁棒性上的表现,发现模型在并行视频流处理、干扰模式、时空记忆和符号记忆方面的显著缺陷。
Comments We present an evaluation designed for multi-modal memory in multi-modal models
测试时匹配:在多模态模型中解锁组合推理
机构 * University of California, Riverside(加州大学河滨分校)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 本文提出测试时匹配算法,通过改进评估指标提升多模态模型的组合推理能力,使SigLIP-B16和GPT-4.1在Winoground等基准上取得新突破。
Comments To appear at ICLR 2026; extended results to generative multimodal models
机构 * Fontys University of Applied Sciences(Fontys应用科学大学) ; Technical University of Eindhoven(埃因霍温技术大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
Comments We present CrypticBio, the largest publicly available multimodal dataset of visually confusing species, specifically curated to support the development of AI models for biodiversity identification using images, language and spatiotemporal data
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments LogicVista benchmarks the logical reasoning of multimodal large language models in visual tasks
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Multimodal Benchmark, Project Url: https://zeyofu.github.io/blink/, ECCV 2024
PanDent:面向牙科放射学中全面的牙级结构-语言一致性
机构 * Faculty of Dentistry, The University of Hong Kong(香港大学牙医学院) ; Imperial College London(帝国理工学院) ; University of Science and Technology of China(中国科学技术大学) ; Department of Data and Systems Engineering, The University of Hong Kong(香港大学数据与系统工程系) ; Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系)
专题命中 多模态评测 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV、cs.MM
AI总结 本研究推出PanDent牙科OPG基准,经实验发现现有MLLM生成的牙科报告流畅但临床一致性差,在PanDent上微调可提升其结构-语言一致性,该基准可用于评估MLLM的牙级临床推理能力。
十年视觉语言人工智能模型中准确性和视觉认知错误的演变
机构 * Psychological & Brain Sciences, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校心理与脑科学系) ; Department of Computer Science, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校计算机科学系) ; Department of Electrical and Computer Engineering, University of California, Santa Barbara(加利福尼亚大学圣巴巴拉分校电气与计算机工程系)
专题命中 多模态评测 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 研究十年间视觉语言模型进展,引入CSB数据集,评估模型在其上及MS-COCO样本中的准确性与视觉认知错误类型,发现MLLM消除简单与复杂场景描述准确性差距,几乎消除多数错误类型,为模型发展提供全面评估。
解释比单独预测更难:评估基于概念的MLLM解释作为ICL视觉分类器
专题命中 多模态评测 :MLLM(title_cn,abstract_cn);multimodal(abstract);分类 cs.CL、cs.AI
AI总结 本文通过五种形式化程度递增的条件,系统评估多模态大语言模型在少样本上下文学习中的基于概念的可解释性,发现解释比预测更难,且强制生成形式化解释会降低预测准确性。
Comments Accepted to the CompLearn Workshop at ICML 2026
VLRS-Bench: 一种面向遥感的视觉-语言推理基准
机构 * School of Computer Science, Wuhan University(武汉大学计算机学院)
专题命中 多模态评测 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 本文提出VLRS-Bench,首个专注于复杂遥感推理的基准,包含2000个问题-答案对,涵盖14项任务和八个时间阶段,揭示现有MLLM在遥感任务中的瓶颈。
PPU-Bench: 用于视觉语言模型个性化部分遗忘的现实世界基准
机构 * Harbin Institute of Technology(哈尔滨工业大学) ; Pengcheng Laboratory(鹏城实验室) ; The Hong Kong Polytechnic University(香港理工大学) ; Sichuan University(四川大学) ; Zhejiang Normal University(浙江师范大学)
专题命中 多模态评测 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 本文提出PPU-Bench,一个无需微调的现实世界基准,用于评估视觉语言模型中个性化部分遗忘的效果,通过24K多模态和单模态样本测试遗忘与保留的平衡及跨模态一致性。
专题命中 多模态评测 :multimodal(abstract,comments);multi-modal(abstract);MLLM(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Total 120 pages. See our project at https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models
MetaReason:通过编辑元信息实现精确的交错多模态推理以解决几何问题
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
AI总结 本研究提出MetaReason框架,构建TutorGeo与ExamGeo数据集,结合监督微调与强化学习,实现几何问题的精确交错多模态推理,性能优于现有开源模型。
NARRATE:用于自动驾驶中以人为本解释的多模态真实世界澳大利亚驾驶数据集
机构 * ARC Training Centre for Automated Vehicles in Rural and Remote Regions (AVR3)(澳大利亚农村及偏远地区自动驾驶车辆ARC培训中心) ; Université Gustave Eiffel(Gustave Eiffel大学) ; Queensland University of Technology (QUT)(昆士兰科技大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
AI总结 研究人员推出多模态真实世界澳大利亚驾驶数据集NARRATE,含2050个带注释事件,提供多类标签与情境意识注释,开展四项基准任务验证其价值,为开发以人为本的自动驾驶解释模型奠定基础。
Comments Accepted at The 19th European Conference on Computer Vision (ECCV 2026) DriveX Workshop (Foundation Models for Autonomous Driving)
DepressionAgent:用于抑郁风险评估的阅读、倾听、观察与审议多模态证据框架
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)
AI总结 研究针对现有多模态抑郁评估隐式特征融合的不足,提出以证据为中心的DepressionAgent框架,通过显式证据审议等机制在多基准上取得竞争力性能,且有效性与可检查性获多维度验证。
CONFER:面向多模态情感识别中规则校准弱监督的冲突感知证据协商框架
专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract)
AI总结 本文提出CONFER框架,针对多模态情感识别中跨模态冲突与自我报告标签不可靠问题,通过图结构的冲突感知证据协商实现弱标签校准,在多数据集严格LOSO协议下取得具竞争力的情感识别准确率,提升了对弱标签损坏的鲁棒性。