RECODE: Reasoning Through Code Generation for Visual Question Answering
RECODE: 通过代码生成进行视觉问答的推理
Junhong Shen, Mu Cai, Bo Hu, Ameet Talwalkar, David A Ross, Cordelia Schmid, Alireza Fathi
专题命中
视觉问答
:visual question answering(title);visual reasoning(abstract);grounding(abstract);multimodal large language model(abstract)
AI总结
RECODE通过代码生成实现视觉问答的可验证推理,优于传统方法。
CommentsThe authors are withdrawing this manuscript temporarily to conduct additional checks of the experimental setup and implementation. We plan to post an updated version after completing these checks
MEGC2026: Micro-Expression Grand Challenge on Visual Question Answering
MEGC2026:面向视觉问答的微表情大奖挑战
Xinqi Fan, Jingting Li, John See, Moi Hoon Yap, Su-Jing Wang, Adrian K. Davison
机构
*
Department of Computing and Mathematics, Manchester Metropolitan University(曼彻斯特 Metropolitan 大学计算与数学系)
;
State Key Laboratory of Cognitive Science and Mental Health, Institute of Psychology, CAS & Department of Psychology, University of the Chinese Academy of Sciences(中国科学院心理研究所认知科学与心理健康国家重点实验室及中国科学院大学心理学系)
;
School of Mathematical and Computer Sciences, Heriot-Watt University Malaysia(赫瑞-沃森大学马来西亚分校数学与计算机科学学院)
专题命中
视觉问答
:visual question answering(title,abstract);vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV
TemporalDoRA: Temporal PEFT for Robust Surgical Video Question Answering
TemporalDoRA: 用于鲁棒手术视频问答的时序PEFT
Luca Carlini, Chiara Lena, Cesare Hassan, Danail Stoyanov, Elena De Momi, Sophia Bano, Mobarak I. Hoque
机构
*
Dipartimento di Elettronica, Informazione e Bioingegneria (DEIB), Politecnico di Milano(电子信息与生物工程系(DEIB),米兰理工学院)
;
IRCCS Humanitas Research Hospital(人类学研究所医院)
;
UCL Hawkes Institute and Department of Computer Science, University College London(伦敦大学学院 Hawkes 研究所和计算机科学系)
;
Division of Informatics, Imaging and Data Science, University of Manchester(信息学、影像学与数据科学系,曼彻斯特大学)
机构
*
University of Maryland(马里兰大学)
;
Brown University(布朗大学)
;
Washington University in St. Louis(圣路易斯华盛顿大学)
;
Adobe
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Southern California(南加州大学)
;
NVIDIA
专题命中
视觉推理
:vision language model(title,abstract);VLM(abstract);分类 cs.CV、cs.LG
World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models
World2Mind: 用于基础模型的方位认知工具包
Shouwei Ruan, Bin Wang, Zhenyu Wu, Qihui Zhu, Yuxiang Zhang, Hang Su, Yubin Wang
机构
*
Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能研究院)
;
Huawei Noah’s Ark Lab(华为诺亚方舟实验室)
;
Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(清华大学人工智能院计算机科学与技术系、清华大学-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学)
MGCR-Net:Multimodal Graph-Conditioned Vision-Language Reconstruction Network for Remote Sensing Change Detection
MGCR-Net:多模态图条件视觉-语言重建网络用于遥感变化检测
Chengming Wang, Guodong Fan, Jinjiang Li, Min Gan, C. L. Philip Chen
机构
*
School of Computer Science and Technology, Shandong Technology and Business University(山东科技职业大学计算机科学与技术学院)
;
School of Computer Science and Technology, Qingdao University(青岛大学计算机科学与技术学院)
;
School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
AI总结
MGCR-Net通过多模态图条件视觉-语言重建机制提升遥感变化检测的语义交互能力。
Journal refIEEE Transactions on Geoscience and Remote Sensing, vol. 64, pp. 1-15, 2026, Art no. 4701515
机构
*
The State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, China(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院)
;
Spatiotemporal AI, China(时空人工智能,中国)
;
Hangzhou International Innovation Institute, Beihang University, China(杭州国际创新研究院,北航,中国)
;
Georgia Institute of Technology, China(佐治亚理工学院,中国)
;
Key Laboratory of Computing Power Network and Information Security, Ministry of Education(计算功率网络与信息安全重点实验室,教育部;山东省计算机科学中心,齐鲁工业大学(山东省科学院),中国)
;
Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences), China
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract)
机构
*
Tsinghua University(清华大学)
;
University of California San Diego(加州大学圣地亚哥分校)
;
West China Second University Hospital, Sichuan University(四川大学华西第二医院)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);分类 cs.CV
CommentsThis is an accepted workshop paper at CHI '26, "W37: Human-AI Interaction Alignment: Designing, Evaluating, and Evolving Value-Centered AI For Reciprocal Human-AI Futures", or https://bialign-workshop.github.io/2026/cfp