机构
*
Hubei Key Laboratory of Transportation Internet of Things, Wuhan University of Technology(武汉理工大学交通物联网湖北省重点实验室)
;
State Key Laboratory for Multimedia Information Processing, Peking University(北京大学多媒体信息处理全国重点实验室)
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
理解并缓解多模态推理链模型中的幻觉
Ji Ma, Wei Suo, Peng Wang, Yanning Zhang
机构
*
School of Computer Science and Ningbo Institute, Northwestern Polytechnical University, China(西北工业大学计算机学院和宁波研究院)
;
National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology, China(国家空天地海一体化大数据应用技术国家工程实验室)
Echoes of ownership: Adversarial-guided dual injection for copyright protection in MLLMs
所有权的回声:对抗引导的双注入用于MLLMs中的版权保护
Chengwei Xia, Fan Ma, Ruijie Quan, Yunqiu Xu, Kun Zhan, Yi Yang
机构
*
School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection
通过因果发现和双模态安全子空间投影诊断和修复视觉-语言模型中的不安全通道
Jinhu Fu, Yihang Lou, Qingyi Si, Shudong Zhang, Yan Bai, Sen Su
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Huawei Technologies Ltd.(华为技术有限公司)
;
Peking University(北京大学)
;
Chongqing University of Posts and Telecommunications(重庆邮电大学)
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
针对LAION-400M的人为中心标注:审查偏见及其在模型中的转移
Leander Girrbach, Stephan Alaniz, Genevieve Smith, Trevor Darrell, Zeynep Akata
机构
*
Technical University of Munich, Munich Center for Machine Learning (MCML), MDSI(慕尼黑工业大学,慕尼黑机器学习中心(MCML),MDSI)
;
LTCI, Télécom Paris, Institut Polytechnique de Paris, France(法国巴黎理工学院,巴黎电信学院,LTCI)
;
Helmholtz Munich(亥姆霍兹慕尼黑中心)
;
University of California, Berkeley(加州大学伯克利分校)
Explaining CLIP Zero-shot Predictions Through Concepts
通过概念解释CLIP零样本预测
Onat Ozdemir, Anders Christensen, Stephan Alaniz, Zeynep Akata, Emre Akbas
机构
*
School of Informatics, University of Edinburgh(爱丁堡大学信息学院)
;
Dept. of Computer Eng., Middle East Technical University (METU)(中东技术大学计算机工程系)
;
Orbital
;
DTU Compute, Technical University of Denmark(丹麦技术大学DTU计算学院)
;
Dept. of Biology, University of Copenhagen(哥本哈根大学生物系)
;
LTCI, Télécom Paris, Institut Polytechnique de Paris(巴黎理工学院巴黎电信学院LTCI实验室)
;
Technical University of Munich (TUM)(慕尼黑工业大学)
;
Helmholtz Munich(亥姆霍兹慕尼黑中心)
;
MCML
;
MDSI
;
Robotics & AI Center (ROMER), METU(中东技术大学机器人与人工智能中心)
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
ShotBench:视觉语言模型中的专家级电影叙事理解
Hongbo Liu, Jingwen He, Yi Jin, Dian Zheng, Yuhao Dong, Fan Zhang, Ziqi Huang, Yinan He, Yangguang Li, Weichao Chen, Yu Qiao, Wanli Ouyang, Shengjie Zhao, Ziwei Liu
机构
*
Tongji University(同济大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
S-Lab, Nanyang Technological University(南洋理工大学S-Lab)
From Prediction to Diagnosis: Reasoning-Aware AI for Photovoltaic Defect Inspection
从预测到诊断:面向光伏缺陷检测的推理感知AI
Dev Mistry, Feng Qiu, Bo Chen, Feng Liu, Can Chen, Mohammad Shahidehpour, Ren Wang
机构
*
Illinois Institute of Technology(伊利诺伊理工学院)
;
Argonne National Laboratory(阿贡国家实验室)
;
Commonwealth Edison(联邦爱迪生公司)
;
Drexel University(德雷塞尔大学)
;
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
Cross-Modal Urban Sensing: Evaluating Sound-Vision Alignment Across Street-Level and Aerial Imagery
跨模态城市感知:评估声音与视觉在街道级和航空影像中的对齐情况
Pengyu Chen, Xiao Huang, Teng Fei, Sicheng Wang
机构
*
Department of Geography, University of South Carolina(南卡罗来纳大学地理系)
;
Department of Environmental Sciences, Emory University(埃默里大学环境科学系)
;
University of Canterbury(坎特伯雷大学)
;
School of Resource and Environmental Sciences, Wuhan University(武汉大学资源与环境科学学院)