How to Configure Good In-Context Sequence for Visual Question Answering
专题命中 视觉问答 :visual question answering(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
Comments 8 pages, 6 figures
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉问答 :visual question answering(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
Comments 8 pages, 6 figures
专题命中 视觉问答 :visual question answering(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments 16 pages; see https://toloka.ai/challenges/wsdm2023/ for more details
专题命中 视觉问答 :visual question answering(title,abstract);visual reasoning(abstract);分类 cs.CV、cs.AI
专题命中 视觉问答 :visual question answering(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI
专题命中 视觉问答 :vision language model(title);vision-language model(abstract);visual question answering(abstract);分类 cs.CV、cs.AI
专题命中 视觉问答 :visual language model(title,abstract);visual question answering(abstract);分类 cs.CV、cs.LG
Comments Published at ICLR 2022
专题命中 视觉问答 :visual question answering(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
专题命中 视觉问答 :visual question answering(title,abstract);grounding(abstract);分类 cs.CV、cs.LG
Comments Paper at CVPR 2021. 14 pages, 7 figures
专题命中 视觉问答 :visual question answering(title,abstract);grounding(abstract);分类 cs.CV、cs.AI
Comments Submitted for review at CIKM 2019
专题命中 视觉问答 :visual question answering(title,abstract);grounding(abstract);分类 cs.CV
Comments Second Place of WSDM2023 Toloka Visual Question Answering Challenge
通过隐式文化对齐奖励建模消除文本到图像评估中的偏差
机构 * National Tsing Hua University(国立清华大学) ; National Yang Ming Chiao Tung University(国立阳明交通大学)
专题命中 视觉问答 :MLLM(abstract,abstract_cn);visual question answering(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 研究文本到图像评估中文化真实性问题,提出基于轻量级多模态大语言模型构建的隐式文化对齐奖励模型,集成隐式文化探测器与跳跃连接交叉注意力机制,实验证明该模型准确率高、速度快,能为偏好优化管道提供有效信号。
Comments 16 pages, 2 figures, ECCV 2026 Workshop FAILED
CHaystack:中文文档检索与视觉问答基准测试
专题命中 视觉问答 :VLM(summary_cn,abstract);visual question answering(abstract)
AI总结 本文介绍了CHaystack中文文档检索与视觉问答基准测试,涵盖四类文档。提出CDocRAG系统,用VLM相关性过滤器验证文档图像。评估开源模型发现Qwen系列在文本丰富文档表现佳,中文大规模DocumentVQA在文本编码方面有挑战,仍需改进。
HueManity: 探索 MLLMs 中的细粒度视觉感知
机构 * Google(谷歌) ; Waymo
专题命中 视觉问答 :visual reasoning(abstract);visual question answering(abstract);grounding(abstract);multimodal large language model(abstract)
AI总结 HueManity 通过细粒度视觉感知基准测试揭示 MLLMs 在捕捉细粒度视觉细节方面的显著缺陷。
Journal ref ICML 2025 Workshop on Assessing World Models
社会视听问答的推理:我们处于什么阶段?
机构 * Inria(法国国家信息与自动化研究所) ; Univ. Grenoble Alpes(格勒诺布尔大学) ; CNRS(法国国家科学研究中心) ; University of Twente(特文特大学)
专题命中 视觉问答 :visual question answering(title);multimodal large language model(abstract,abstract_cn);分类 cs.CV
AI总结 本文针对社会视听问答研究,发现IntentBench存在高噪声,Vanilla SFT基线性能优于现有推理方法,仅用文本模态即可实现与视频相当的性能,并发布了IntentBench-Prime等资源。
Comments Accepted at HCMIW ECCV workshop. Code available here: https://github.com/koenv759/VanillaSFT
LHSDet:基于视觉问答的高分辨率AI生成图像检测方法
机构 * College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机学院) ; Academy of Military Science(军事科学院)
专题命中 视觉问答 :visual question answering(title,abstract);vision-language model(abstract);分类 cs.CV
AI总结 针对现有AI生成图像检测方法忽略高分辨率细节、难以应对未知生成模型的问题,提出LHSDet,将检测任务转为视觉问答,采用三分支架构,在各类生成模型上实现高检测准确率与稳健性能。
提升大型视觉-语言模型对流场数据的理解能力
机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院)
专题命中 视觉问答 :vision-language model(title,abstract);visual question answering(abstract);分类 cs.CV
AI总结 FieldLVLM通过场感知语言生成策略和数据压缩多模态模型调优,提升大型视觉-语言模型对流场数据的理解能力。
Comments Accepted by Machine Intelligence Research
多模态大语言模型能理解光学相干断层扫描(OCT)吗?
专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV
AI总结 研究探讨多模态大语言模型能否理解OCT,引入OCT - Bench基准,包含多维度细粒度任务,基于多个数据集构建大量选择题,评估20个代表性模型,发现当前模型理解OCT能力不足,为评估模型和推动OCT理解提供基础。
通过自我场景增强在多模态大语言模型中强化自我中心空间感知
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Guangxi Zhuang Autonomous Region Information Center(广西壮族自治区信息中心) ; The Hong Kong University of Science and Technology(香港科技大学)
专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV
AI总结 研究如何强化多模态大语言模型的自我中心空间感知,提出自我场景增强框架ESA,利用自我元素图作为中间表示,通过视觉基础模型增强空间感知,在EgoTextVQA基准上取得显著性能提升。
Comments 14 pages, 8 figures. Chi Kit Wong and Ye Pan contributed equally. Code: https://github.com/Chikit-WONG/spatialGraph
InfraQR:对红外视觉语言模型的边缘放置的受QR码启发的结构化补丁攻击
机构 * China University of Petroleum-Beijing at Karamay(中国石油大学(北京)克拉玛依校区) ; Guizhou University(贵州大学) ; Shenzhen Research Institute of Big Data(深圳大数据研究院) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 视觉问答 :vision-language model(title,abstract);visual question answering(abstract);分类 cs.CV
AI总结 研究红外视觉语言模型对局部结构化扰动的鲁棒性,提出InfraQR攻击方法,沿图像边界放置结构化补丁并通过替代编码器优化网格单元,实验表明该方法能大幅降低分类器准确率,使对抗图像影响黑盒模型,凸显模型易受此类扰动影响。
Robobench:作为具身大脑的多模态大语言模型的综合评估基准
机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(信息多媒体国家重点实验室,计算机学院,北京大学) ; Beijing Academy of Artificial Intelligence(北京人工智能研究院) ; Institute for Brain and Intelligence, Fudan University(脑与智能研究院,复旦大学) ; University of Science and Technology Beijing(北京科技大学) ; Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
专题命中 视觉问答 :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
AI总结 研究针对在动态非结构化环境中构建能感知、推理和行动的机器人的挑战,介绍RoboBench基准,涵盖多维度、多任务等,能评估多模态大语言模型,揭示其局限性并分析与下游控制关系,为量化高级认知提供支架。
Comments ECCV 2026 Camera Ready
4DP-QA:面向视觉语言模型中4D感知的可扩展问答
机构 * NVIDIA(英伟达) ; Yale University(耶鲁大学) ; KAIST AI(韩国科学技术院人工智能学院)
专题命中 视觉问答 :vision language model(title,abstract);VLM(abstract_cn);分类 cs.CV
AI总结 针对视觉语言模型难以理解动态场景的问题,提出一种关注运动场景理解的问答生成流水线,通过真运动追踪解耦物体与相机运动,生成大规模数据集4DP-QA和基准4DP-QA-Bench,训练现有模型在外部基准上取得性能提升。
Comments Project page: https://research.nvidia.com/labs/lpr/4dpqa
建模复杂行为:视觉语言模型中的多人格组合与动态切换
机构 * Xi'an Jiaotong University(西安交通大学) ; Beihang University(北京航空航天大学)
专题命中 视觉问答 :vision-language model(title);visual question answering(abstract);multimodal large language model(abstract);分类 cs.AI
AI总结 本研究在视觉语言模型中引入显式人格条件,建立包括单人格、多人格和人格切换的系统评估框架,发现人格提示可提升图像描述但损害精确推理任务,并观察到多特质组合与动态切换中的平衡与残留效应。
Comments 16 pages, 4 figures, 10 tables
ChinaHeritaQA:面向中国世界遗产地的文化基础视觉问答数据集
机构 * LMU Munich(慕尼黑大学) ; FAU Erlangen-Nuremberg(埃尔朗根-纽伦堡大学) ; Munich Center for Machine Learning(慕尼黑机器学习中心) ; University of Tübingen & Tübingen AI Center(图宾根大学与图宾根人工智能中心) ; Sun Yat-sen University(中山大学) ; University of Copenhagen(哥本哈根大学) ; University of Maryland, College Park(马里兰大学帕克分校)
专题命中 视觉问答 :visual question answering(title);vision-language model(abstract);VLM(abstract_cn);分类 cs.CV
AI总结 提出ChinaHeritaQA多模态基准数据集,包含2279张图像和14133个双语多项选择题,覆盖七个认知维度,评估视觉语言模型在中国世界遗产上的文化推理能力。
MAOAM: 基于视觉语言模型的统一对象与材质选择
机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ; Adobe Research(Adobe研究)
专题命中 视觉问答 :vision-language model(title);VLM(abstract,abstract_cn);分类 cs.CV
AI总结 提出MAOAM框架,利用视觉语言模型和分割头,通过文本或点击交互实现对象和材质的精确选择,并设计数据生成流水线解决材质选择数据缺乏问题。
Comments Accepted to SIGGRAPH 2026 Conference. Project page: \href{https://jadenpark0.github.io/project_pages/maoam/}{here}
NextMotionQA: 使用视觉语言模型基准测试和评判人体运动理解
机构 * University of Tübingen(图宾根大学) ; Tübingen AI Center(图宾根人工智能中心) ; Max Planck Institute for Informatics(马克斯·普朗克信息学院) ; Saarland Informatics Campus(萨尔兰州信息学院)
专题命中 视觉问答 :vision-language model(title,abstract);VLM(abstract_cn);分类 cs.CV
AI总结 提出NextMotionQA基准,通过三项互补任务和多粒度难度分层,系统评估视觉语言模型在人体运动理解中的能力,并揭示其在细粒度评判中的局限性。
Comments 23 pages, 8 figures, 9 tables
VIHD: 基于视觉干预的医学视觉问答幻觉检测
机构 * Department of Data Science \& AI, Faculty of Information Technology, Monash University, Melbourne, VIC 3800, Australia Alfred Health Radiology, Alfred Health, Melbourne, VIC 3004, Australia School of Translational Medicine, Faculty of Medicine, Nursing ; Health Sciences, Monash University, Melbourne, VIC 3800, Australia Hong Kong Polytechnic University, Hong Kong SAR, China
专题命中 视觉问答 :visual question answering(title);multimodal large language model(abstract);MLLM(abstract_cn);分类 cs.CV
AI总结 提出VIHD方法,通过视觉依赖探测和视觉干预解码校准语义熵,有效检测医学多模态大语言模型中的幻觉响应。
Comments Early accepted by MICCAI 2026. This version of the contribution has been accepted for publication, after peer review (when applicable) but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections
测量多模态大语言模型中的认知谦逊
机构 * Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学) ; Hong Kong Baptist University(香港 Baptist大学)
专题命中 视觉问答 :multimodal large language model(title,abstract);visual question answering(abstract);分类 cs.CV
AI总结 提出HumbleBench基准,通过强制选择多项选择中引入“以上皆非”选项,评估多模态大语言模型拒绝错误选项的谦逊行为。
DisasterVQA: 一个用于灾难场景的视觉问答基准数据集
机构 * Qatar Computing Research Institute(卡塔尔计算研究所) ; Hamad Bin Khalifa University(哈马德·本·卡伊夫大学) ; College of Science & Engineering(科学与工程学院) ; Qatar University(卡塔尔大学)
专题命中 视觉问答 :visual question answering(title,abstract);vision-language model(abstract);分类 cs.CV
AI总结 本文提出DisasterVQA数据集,用于灾难场景中的感知与推理任务,通过1395张真实图像和4405对专家 curated 的问答对,评估了七种最先进的视觉-语言模型在灾难响应中的性能,发现模型在细粒度定量推理、物体计数和上下文敏感解释方面存在不足。
Comments Accepted at ICWSM 2026
在大型视觉语言模型中检测和评估医学幻觉
机构 * Academy for Engineering and Technology, Fudan University(复旦大学工程与技术学院) ; Tencent Youtu Lab(腾讯优图实验室) ; Cognition and Intelligent Technology Laboratory(认知与智能技术实验室)
专题命中 视觉问答 :vision language model(title,abstract);visual question answering(abstract);分类 cs.CV
AI总结 本文提出Med-HallMark基准,用于医疗多模态领域中的幻觉检测与评估,引入MediHall Score和MediHallDetector,通过多任务训练提升模型可靠性,实验表明其在医疗应用中更有效。
无需指令的大型视觉语言模型在医疗指令遵循中的调优
机构 * Department of Electrical and Computer Engineering, The University of British Columbia(英属哥伦比亚大学电气与计算机工程系) ; Department of Robotics and Mechatronics Engineering, Daegu Gyeongbuk Institute of Science and Technology (DGIST)(大邱庆北科学技术院机器人与机电工程系) ; Division of Intelligent Robot, Daegu Gyeongbuk Institute of Science and Technology (DGIST)(大邱庆北科学技术院智能机器人系) ; Vector Institute(向量研究所) ; Department of Computer Science and Engineering, Pohang University of Science and Technology (POSTECH)(釜山科学技术大学计算机科学与工程系)
专题命中 视觉问答 :vision language model(title,abstract);visual question answering(abstract);分类 cs.CV
AI总结 本文提出无需指令的调优方法,利用图像描述对训练模型,提升医疗领域视觉语言模型的指令遵循能力,实现多选视觉问答任务的最优性能。