CommentsICLR 2026Camera Ready Version. TL;DR: The first multi-person dialogue video generation method from pairs of reference image and audio via explicit layout-aligned condition injection. Project page https://zhenzhiwang.github.io/interacthuman/
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
对自监督语音表示中说话人特定属性的大规模探测分析
Aemon Yat Fei Chiu, Kei Ching Fung, Roger Tsz Yeung Li, Jingyu Li, Tan Lee
机构
*
Department of Electronic Engineering, The Chinese University of Hong Kong, Hong Kong(香港中文大学电子工程系)
;
Independent Researcher(独立研究者)
;
School of Data Science, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学深圳学校数据科学系)
Accurate and Efficient Hybrid-Ensemble Atmospheric Data Assimilation in Latent Space with Uncertainty Quantification
在潜在空间中实现准确且高效的混合-集成大气数据同化与不确定性量化
Hang Fan, Juan Nathaniel, Yi Xiao, Ce Bian, Fenghua Ling, Ben Fei, Lei Bai, Pierre Gentine
机构
*
Department of Earth and Environmental Engineering, School of Engineering and Applied Sciences, Climate School, Columbia University(地球与环境工程系,工程与应用科学学院,气候学院,哥伦比亚大学)
;
Learning the Earth with Artificial Intelligence and Physics (LEAP) Center, Columbia University(人工智能与物理学习地球中心,哥伦比亚大学)
;
Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国)
;
The Chinese University of Hong Kong, Hong Kong, China(香港中文大学,香港,中国)
;
Department of Computer Science and Technology, Tsinghua University, Beijing, China(计算机科学与技术系,清华大学,北京,中国)
机构
*
Department of Information Engineering(信息工程系)
;
The Chinese University of Hong Kong(香港中文大学)
;
National Key Laboratory on Wireless Communications(无线通信国家重点实验室)
;
University of Electronic Science and Technology of China(电子科学与技术大学)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
鲁棒的基于大语言模型的音频视觉语音识别与稀疏模态对齐和视觉单元引导的细化
Fei Su, Cancan Li, Juan Liu, Wei Ju, Hongbin Suo, Ming Li
机构
*
School of Computer Science, Wuhan University, China(武汉大学计算机学院)
;
School of Artificial Intelligence, Wuhan University, China(武汉大学人工智能学院)
;
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)人工智能学院)
;
AI Center, OPPO, China(OPPO人工智能中心)
;
Digital Innovation Research Center, Duke Kunshan University, China(杜克大学昆山数字创新研究中心)
Toward Early Quality Assessment of Text-to-Image Diffusion Models
面向文本到图像扩散模型早期质量评估
Huanlei Guo, Hongxin Wei, Bingyi Jing
机构
*
Southern University of Science and Technology(南方科技大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Shenzhen Loop Area Institute(深圳河套学院)
Factuality Matters: When Image Generation and Editing Meet Structured Visuals
事实性至关重要:当图像生成与编辑遇见结构化视觉
Le Zhuo, Songhao Han, Yuandong Pu, Boxiang Qiu, Sayak Paul, Yue Liao, Yihao Liu, Jie Shao, Xi Chen, Si Liu, Hongsheng Li
机构
*
CUHK MMLab(香港大学多模态实验室)
;
Beihang University(北京航空航天大学)
;
Krea AI(Krea人工智能)
;
Shanghai Jiao Tong University(上海交通大学)
;
Shanghai AI Lab(上海人工智能实验室)
;
Hugging Face
;
National University of Singapore(新加坡国立大学)
;
ByteDance(字节跳动)
;
The University of Hong Kong(香港大学)
机构
*
Dalian Maritime University(大连海事大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Tsinghua University(清华大学)
;
Nanyang Technological University(南洋理工大学)
;
Xi'an Jiaotong University(西安交通大学)
;
Renmin University of China(中国人民大学)
;
Wuhan University(武汉大学)
ACE-Brain-0: Spatial Intelligence as a Shared Scaffold for Universal Embodiments
ACE-Brain-0:空间智能作为通用具身化体系的共享框架
Ziyang Gong, Zehang Luo, Anke Tang, Zhe Liu, Shi Fu, Zhi Hou, Ganlin Yang, Weiyun Wang, Xiaofeng Wang, Jianbo Liu, Gen Luo, Haolan Kang, Shuang Luo, Yue Zhou, Yong Luo, Li Shen, Xiaosong Jia, Yao Mu, Xue Yang, Chunxiao Liu, Junchi Yan, Hengshuang Zhao, Dacheng Tao, Xiaogang Wang
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Nanyang Technological University(南洋理工大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
The University of Hong Kong(香港大学)
;
University of Science(科学技术大学)
;
Fudan University(复旦大学)
;
Xiamen University(厦门大学)
;
East China Normal University(华东师范大学)
;
Wuhan University(武汉大学)
;
Sun Yat-sen University(中山大学)
AI总结
ACE-Brain-0 通过空间智能作为共享框架,统一了自动驾驶、机器人和 UAVs 的具身化任务,采用 SSR 范式和 GRPO 方法实现跨领域泛化和领域精通的平衡。
The Dresden Dataset for 4D Reconstruction of Non-Rigid Abdominal Surgical Scenes
德里森数据集用于非刚性腹部手术场景的4D重建
Reuben Docea, Rayan Younis, Yonghao Long, Maxime Fleury, Jinjing Xu, Chenyang Li, André Schulze, Ann Wierick, Johannes Bender, Micha Pfeiffer, Qi Dou, Martin Wagner, Stefanie Speidel
机构
*
Department of Translational Surgical Oncology, National Center for Tumor Diseases (NCT), NCT/UCC Dresden, a partnership between DKFZ, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology, and Helmholtz-Zentrum Dresden-Rossendorf (HZDR), Dresden, Germany(转化外科肿瘤中心部门,肿瘤疾病国家中心(NCT),NCT/UCC德累斯顿,由DKFZ、医学院和卡尔·戈斯瓦尔德·卡尔医院、德累斯顿技术大学以及德累斯顿-罗斯恩多夫赫尔姆霍兹中心(HZDR)组成的合作机构,德累斯顿,德国)
;
Department of Translational Surgical Oncology, NCT/UCC Dresden, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology(转化外科肿瘤中心部门,NCT/UCC德累斯顿,医学院和卡尔·戈斯瓦尔德·卡尔医院,德累斯顿技术大学)
;
Department for Visceral, Thoracic and Vascular Surgery, Faculty of Medicine and University Hospital Carl Gustav Carus, Technische Universität Dresden(visceral、胸腔和血管外科部门,医学院和卡尔·戈斯瓦尔德·卡尔医院,德累斯顿技术大学)
;
Centre for Tactile Internet with Human-in-the-Loop (CeTI), Technical University Dresden(人机交互触觉互联网中心(CeTI),德累斯顿技术大学)
;
Department of Computer Science and Engineering, The Chinese University of Hong Kong(计算机科学与工程部门,香港中文大学)
Value Gradient Guidance for Flow Matching Alignment
价值梯度指导下的流匹配对齐
Zhen Liu, Tim Z. Xiao, Carles Domingo-Enrich, Weiyang Liu, Dinghuai Zhang
机构
*
The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳))
;
University of Tübingen(图宾根大学)
;
Microsoft Research(微软研究院)
;
The Chinese University of Hong Kong(香港中文大学)
;
Mila – Quebec AI Institute(魁北克AI研究所)
AI总结
VGG-Flow通过价值梯度指导实现流匹配模型的高效微调与先验保留对齐。
CommentsAccepted at NeurIPS 2025; 26 pages, 20 figures
机构
*
Zhiyuan College, Shanghai Jiao Tong University, Shanghai 200240, China(上海交通大学紫阳学院)
;
School of Data Science, The Chinese University of Hong Kong, Shenzhen, Guangdong, China(香港中文大学(深圳)数据科学学院)
;
McCombs School of Business, The University of Texas at Austin, Austin, TX, USA(德克萨斯大学奥斯汀分校麦克拉姆商学院)
机构
*
Hong Kong University of Science and Technology (HKUST)(香港科技大学)
;
Tongyi Fun Team, Alibaba Group(通义Fun团队,阿里巴巴集团)
;
The Chinese University of Hong Kong (CUHK)(香港中文大学)
机构
*
Tsinghua University(清华大学)
;
Infinigence AI
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
SLAI
;
Shanghai AI Laboratory(上海人工智能实验室)
Exploiting Low-Dimensional Manifold of Features for Few-Shot Whole Slide Image Classification
利用特征的低维流形进行少样本全滑动图像分类
Conghao Xiong, Zhengrui Guo, Zhe Xu, Yifei Zhang, Raymond Kai-Yu Tong, Si Yong Yeo, Hao Chen, Joseph J. Y. Sung, Irwin King
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Centre of AI in Medicine, Singapore(新加坡人工智能医学中心)
;
The Hong Kong University of Science and Technology(香港科学大学)
;
Nanyang Technological University(南洋理工大学)
;
Lee Kong Chian School of Medicine, Nanyang Technological University(南洋理工大学Lee Kong Chian医学学院)
;
MedVisAI Lab, Singapore(新加坡MedVisAI实验室)
机构
*
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
The Chinese University of Hong Kong, Hong Kong SAR(香港中文大学)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳研究院)
;
Shenzhen Loop Area Institute, China(深圳环园院)
;
McGill University(麦吉尔大学)
;
The Hong Kong University of Science and Technology, Guangzhou(香港科学与技术大学)
Words & Weights: Streamlining Multi-Turn Interactions via Co-Adaptation
词语与权重:通过协同适应流式多轮交互
Chenxing Wei, Hong Wang, Ying He, Zhongxiang Dai, Bo Jiang, F. Richard Yu, Yao Shu
机构
*
Shenzhen University(深圳大学)
;
Hong Kong University of Science(香港科学大学)
;
Guangdong Laboratory of Artificial Intelligence(广东人工智能与数字经济实验室)
;
University of Science(科学大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Carleton University(卡尔顿大学)
Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction
人类还是机器?一种语音到语音交互的初步图灵测试
Xiang Li, Jiabao Gao, Sipei Lin, Xuan Zhou, Chi Zhang, Bo Cheng, Jiale Han, Benyou Wang
机构
*
State Key Laboratory of Networking and Switching Technology(网络与交换技术国家重点实验室)
;
Shenzhen Research Institute of Big Data(深圳大数据研究 institute)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Shenzhen Loop Area Institute(深圳河套学院)
;
The Hong Kong University of Science and Technology(香港科技大学)
机构
*
Shanghai AI Laboratory(上海人工智能实验室)
;
Shanghai Jiao Tong University(上海交通大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Nanjing University(南京大学)
;
Wuhan University(武汉大学)
Token-Importance Guided Direct Preference Optimization
基于令牌重要性的直接偏好优化
Ning Yang, Hai Lin, Yibo Liu, Baoliang Tian, Guoqing Liu, Haijun Zhang
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
ByteDance(字节跳动)
;
Microsoft Research AI4Science(微软研究院AI4Science)
;
University of Science and Technology Beijing(北京科技大学)