CommentsAn earlier version of this work appeared at the NeurIPS 2025 Workshop on Symmetry and Geometry in Neural Representations (NeurReps). Workshop version: https://openreview.net/forum?id=jOmZsvXoK5
VOPE: Revisiting Hallucination of Vision-Language Models in Voluntary Imagination Task
VOPE:重新审视视觉语言模型在自愿想象任务中的幻觉现象
Xingming Long, Jie Zhang, Shiguang Shan, Xilin Chen
机构
*
Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences (CAS)(中国科学院人工智能安全重点实验室,计算技术研究所,中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Zhongguancun Academy(中关村学院)
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
VideoRAE:通过表示自动编码器驯服用于生成建模的视频基础模型
Zhihao Xie, Junfeng Wu, Xinting Hu, Junchao Huang, Li Jiang
机构
*
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Huazhong University of Science and Technology(华中科技大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
University of Science and Technology of China(中国科学技术大学)
AgentFoX: LLM Agent-Guided Fusion with eXplainability for AI-Generated Image Detection
AgentFoX: 基于可解释性的大语言模型引导融合的AI生成图像检测
Yangxin Yu, Yue Zhou, Bin Li, Kaiqing Lin, Haodong Li, Jiangqun Ni, Bo Cao
机构
*
Guangdong Provincial Key Laboratory of Intelligent Information Processing(广东省智能信息处理重点实验室)
;
Shenzhen Key Laboratory of Media Security(深圳媒体安全重点实验室)
;
SZU-AFS Joint Innovation Center for AI Technology(深圳大学-AFS人工智能技术联合创新中心)
;
Shenzhen University(深圳大学)
;
School of Cyber Science and Technology(网络安全科学与技术学院)
;
The Smart City Research Institute of China Electronics Technology Group Corporation(中国电子科技集团有限公司智慧城市研究院)
SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling
SLAP: 用于变分视频-语言建模的语义最小作用原理
Xiang Fang, Wanlong Fang
机构
*
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院)
;
Nanyang Technological University, Singapore(新加坡南洋理工大学)
Discriminative-Generative Target Speaker Extraction with Decoder-Only Language Models
判别-生成目标说话人提取与解码器-only语言模型
Bang Zeng, Beilong Tang, Wang Xiang, Ming Li
机构
*
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Suzhou Municipal Key Laboratory of Multimodal Intelligent Systems, Digital Innovation Research Center, Duke Kunshan University(多模态智能系统苏州市重点实验室、数字创新研究中心、杜克昆山大学)
;
North Carolina State University(北卡罗来纳州立大学)