机构
*
Peking University Shanghai AI Laboratory(北京大学上海人工智能实验室)
;
Nanjing University(南京大学)
;
Peking University(北京大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Zhongguancun Academy Beijing Key Laboratory of Data Intelligence and Security(中关村北京数据智能与安全重点实验室)
专题命中
视觉问答
:LLaVA(abstract,abstract_cn);vision language model(abstract);分类 cs.CV
Journal refAdvances in Information Retrieval: 48th European Conference on Information Retrieval, ECIR 2026, Delft, The Netherlands, March 29 - April 2, 2026, Proceedings, Part IV
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Tsinghua University(清华大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Wuhan AI Research(武汉人工智能研究院)
机构
*
Shenyang Institute of Computing Technology, Chinese Academy of Sciences(中国科学院沈阳计算技术研究所)
;
University of Chinese Academy of Sciences 3 ByteDance 4 Westlake University(中国科学院大学 3 字节跳动 4 西湖大学)
;
Key Laboratory of Computing Power Network(计算功率网络重点实验室)
;
Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(信息安全,教育部,山东计算机科学中心(济南国家超级计算中心),齐鲁工业大学(山东省科学院))
专题命中
视觉推理
:visual reasoning(abstract);grounding(abstract);multimodal large language model(abstract)
机构
*
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(人工智能研究院,合肥国家科学中心)
;
Anhui Polytechnic University(安徽理工大学)
;
Hefei University of Technology(合肥工业大学)
;
Anhui University(安徽大学)
;
IGS, Imperial College London(帝国理工学院伦敦分校)
Comments4 pages, 2 figures, and 1 table. This is a methodology paper for the DataCV 2026 Challenge (CVPR Workshops), Task 1, where our method ranked 2nd
SurgiSR4K: A High-Resolution Endoscopic Video Dataset for Robotic-Assisted Minimally Invasive Procedures
SurgiSR4K:一种用于机器人辅助微创手术的高分辨率内窥镜视频数据集
Fengyi Jiang, Xiaorui Zhang, Lingbo Jin, Ruixing Liang, Yuxin Chen, Adi Chola Venkatesh, Jason Culman, Tiantian Wu, Lirong Shao, Wenqing Sun, Cong Gao, Hallie McNamara, Jingpei Lu, Omid Mohareri
机构
*
Intuitive Surgical, Inc.(Intuitive Surgical公司)
;
Johns Hopkins Medicine Neurosurgery(约翰霍普金斯医学神经外科)
;
Johns Hopkins University Electrical and Computer Engineering(约翰霍普金斯大学电气与计算机工程)
;
University of British Columbia Electrical and Computer Engineering(不列颠哥伦比亚大学电气与计算机工程)
;
Wilford & Kate Bailey Small Animal Teaching Hospital(威尔福德与凯蒂·贝利小动物教学医院)
Retrieval-Augmented Multimodal Model for Fake News Detection
增强检索的多模态模型用于虚假新闻检测
Yiheng Li, Weihai Lu, Hanyi Yu, Yue Wang
机构
*
University of International Business and Economics(国际商务经济大学)
;
Peking University(北京大学)
;
University of Southern California(南加州大学)
;
Upstart Holdings, Inc.(Upstart Holdings公司)
专题命中
视觉定位与Grounding
:MLLM(abstract,abstract_cn);multimodal large language model(abstract)
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
对比语义投影:利用对比示例实现忠实的神经元标注
Oussama Bouanani, Jim Berend, Wojciech Samek, Sebastian Lapuschkin, Maximilian Dreyer
机构
*
Fraunhofer Heinrich Hertz Institute(弗劳恩霍夫海因里希·赫兹研究所)
;
Technische Universität Berlin(柏林技术大学)
;
BIFOLD – Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究所)
;
Centre of eXplainable Artificial Intelligence, Technological University Dublin(可解释人工智能中心,都柏林技术大学)
专题命中
视觉定位与Grounding
:vision language model(abstract);分类 cs.CV、cs.LG
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
告诉模型该看哪里:通过视觉引导注意力缓解大语言模型的幻觉
Jianfei Zhao, Feng Zhang, Xin Sun, Chong Feng, Zhixing Tan
机构
*
School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院)
;
Zhongguancun Academy(中关村学院)
;
Southeast Academy of Information Technology, Beijing Institute of Technology(北京理工大学信息科技东南学院)
;
Zhongguancun Laboratory(中关村实验室)
Journal refProceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), July 20--24, 2026, Melbourne, VIC, Australia