arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-03-11 至 2026-03-11 共收录 7 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 7 篇

2510.13756 2026-03-11 cs.CV cs.AI cs.LG 87%

RECODE: Reasoning Through Code Generation for Visual Question Answering

RECODE: 通过代码生成进行视觉问答的推理

Junhong Shen, Mu Cai, Bo Hu, Ameet Talwalkar, David A Ross, Cordelia Schmid, Alireza Fathi

专题命中 视觉问答 :visual question answering(title);visual reasoning(abstract);grounding(abstract);multimodal large language model(abstract)

AI总结 RECODE通过代码生成实现视觉问答的可验证推理,优于传统方法。

Comments The authors are withdrawing this manuscript temporarily to conduct additional checks of the experimental setup and implementation. We plan to post an updated version after completing these checks

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08927 2026-03-11 cs.CV cs.MM 85%

MEGC2026: Micro-Expression Grand Challenge on Visual Question Answering

MEGC2026:面向视觉问答的微表情大奖挑战

Xinqi Fan, Jingting Li, John See, Moi Hoon Yap, Su-Jing Wang, Adrian K. Davison

机构 * Department of Computing and Mathematics, Manchester Metropolitan University(曼彻斯特 Metropolitan 大学计算与数学系) State Key Laboratory of Cognitive Science and Mental Health, Institute of Psychology, CAS & Department of Psychology, University of the Chinese Academy of Sciences(中国科学院心理研究所认知科学与心理健康国家重点实验室及中国科学院大学心理学系) School of Mathematical and Computer Sciences, Heriot-Watt University Malaysia(赫瑞-沃森大学马来西亚分校数学与计算机科学学院)

专题命中 视觉问答 :visual question answering(title,abstract);vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 MEGC2026通过视觉问答任务探索微表情识别,利用多模态大语言模型和大视觉-语言模型提升微表情分析能力。

Comments MEGC 2026 at IEEE FG 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10729 2026-03-11 cs.CV cs.AI 81%

EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering

EgoCross:多模态大语言模型跨领域眼动视频问答基准测试

Yanjun Li, Yuqian Fu, Tianwen Qian, Qi'ao Xu, Silong Dai, Danda Pani Paudel, Luc Van Gool, Xiaoling Wang

专题命中 视觉问答 :multimodal large language model(title,abstract);分类 cs.CV、cs.AI

AI总结 EgoCross提出一个跨领域眼动视频问答基准,评估多模态大语言模型在不同领域中的泛化能力,揭示现有模型在非日常领域任务中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09155 2026-03-11 cs.CL cs.AI 79%

OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering

OPENXRD:一个全面的基准框架用于LLM/MLLM XRD问答

Ali Vosoughi, Ayoub Shahnazari, Yufeng Xi, Zeliang Zhang, Griffin Hess, Chenliang Xu, Niaz Abdolrahim

专题命中 视觉问答 :MLLM(title);LLaVA(abstract);分类 cs.AI

AI总结 OPENXRD是一个用于评估LLM和MLLM在晶体学问答中性能的综合基准框架,通过专家审核材料验证内容质量对模型性能的关键影响。

Comments Accepted at Digital Discovery (Royal Society of Chemistry)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09896 2026-03-11 cs.CV 57%

Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports

迈向赛场:评估体育中的空间智能

Yuchen Yang, Yuqing Shao, Duxiu Huang, Linfeng Dong, Yifei Liu, Suixin Tang, Xiang Zhou, Yuanyuan Gao, Wei Wang, Yue Zhou, Xue Yang, Yanfeng Wang, Xiao Sun, Zhihang Zhong

专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV

AI总结 CourtSI通过大规模体育场景数据集评估VLMs的空间智能,发现现有基准存在局限,并验证了模型在体育场景中的提升效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09696 2026-03-11 cs.CV 57%

TemporalDoRA: Temporal PEFT for Robust Surgical Video Question Answering

TemporalDoRA: 用于鲁棒手术视频问答的时序PEFT

Luca Carlini, Chiara Lena, Cesare Hassan, Danail Stoyanov, Elena De Momi, Sophia Bano, Mobarak I. Hoque

机构 * Dipartimento di Elettronica, Informazione e Bioingegneria (DEIB), Politecnico di Milano(电子信息与生物工程系(DEIB),米兰理工学院) IRCCS Humanitas Research Hospital(人类学研究所医院) UCL Hawkes Institute and Department of Computer Science, University College London(伦敦大学学院 Hawkes 研究所和计算机科学系) Division of Informatics, Imaging and Data Science, University of Manchester(信息学、影像学与数据科学系,曼彻斯特大学)

专题命中 视觉问答 :grounding(abstract);分类 cs.CV

AI总结 TemporalDoRA通过引入时序多头注意力和选择性权重分解,提升手术视频问答的鲁棒性和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09173 2026-03-11 cs.CV 57%

Point Cloud as a Foreign Language for Multi-modal Large Language Model

点云作为多模态大语言模型的外语

Sneha Paul, Zachary Patterson, Nizar Bouguila

机构 * Concordia University(康科迪亚大学)

专题命中 视觉问答 :MLLM(abstract);分类 cs.CV

AI总结 SAGE是首个端到端3D多模态大语言模型,通过轻量级3D分词器直接处理点云,提升3D任务的推理能力与鲁棒性。

Comments Accepted in The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2026

详情

展开后加载摘要…

URL PDF HTML 收藏