CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning
CapRL++:基于可验证奖励的统一强化学习用于密集图像和视频描述生成
Penghui Yang, Long Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Yibin Wang, Yujie Zhou, Jiazi Bu, Jianze Liang, Qidong Huang, Jiaqi Wang, Feng Wu, Dahua Lin
机构
*
Tsinghua University(清华大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Microsoft(微软)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Shanghai Innovation Institute(上海创新研究院)
;
Alibaba Cloud(阿里云)
;
The Chinese University of Hong Kong(香港中文大学)
机构
*
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
University of Waterloo(滑铁卢大学)
;
The University of Hong Kong(香港大学)
;
The University of Sydney(悉尼大学)
;
Université Lyon 1(里昂第一大学)
Zero-Parameter Geometric Gating for Temporally Stable Low-Altitude UAV Video Semantic Segmentation
用于低空无人机视频语义分割的零参数几何门控以实现时间稳定性
Jingpu Yang, Fengxian Ji, Zhengzhao Lai, Juanfan Wu, Mingxuan Cui, Yufeng Wang
机构
*
Beihang University(北京航空航天大学)
;
Northeastern University(东北大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Beijing Institute of Technology(北京理工大学)
机构
*
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
New Jersey Institute of Technology(新泽西理工学院)
;
Institute of Computing Technology, CAS(中国科学院计算技术研究所)
Bridging the Agent-World Gap: Text World Models for LLM-based Agents
弥合智能体-世界鸿沟:面向基于LLM的智能体的文本世界模型
Yixia Li, Hongru Wang, Peng Lai, Zhiwen Ruan, He Zhu, Youxin Zhu, Ganlong Zhao, Minda Hu, Yun Chen, Sibei Yang, Peng Li, Jeff Z. Pan, Jia Pan, Guanhua Chen, Yang Liu, Guanbin Li
机构
*
Southern University of Science and Technology(南方科技大学)
;
University of Edinburgh(爱丁堡大学)
;
Peking University(北京大学)
;
Sun Yat-sen University(中山大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai University of Finance and Economics(上海财经大学)
;
Tsinghua University(清华大学)
;
The University of Hong Kong(香港大学)
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Beijing Normal-Hong Kong Baptist University(北京师范大学-香港 Baptist大学)
;
Guangzhou Nanfang College(广州南方学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Xiangtan University(湘潭大学)
;
University of Utah(犹他大学)
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
NJIT(新泽西理工学院)
;
Jilin University(吉林大学)
;
Institute of Computing Technology, CAS(中国科学院计算技术研究所)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Guangdong Technion–Israel Institute of Technology(广东以色列理工学院)
;
Technion–Israel Institute of Technology(以色列理工学院)
;
Multiscale Medical Robotics Centre(多尺度医疗机器人中心)
A Multi-modal Agentic Co-pilot for Evidence Grounded Computational Pathology
面向证据基础计算病理学的多模态智能体协同助手
Zhe Xu, Zhengyu Zhang, Zhiyuan Cai, Jiahao Xu, Yijie Lin, Ziyi Liu, Junlin Hou, Hongyi Wang, Yuxiang Nie, Ling Liang, Yihui Wang, Yingxue Xu, Ronald Cheong Kin Chan, Li Liang, Hao Chen
机构
*
Department of Computer Science and Engineering, Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)
;
Department of Pathology, Nanfang Hospital, Southern Medical University(南方医科大学南芳医院病理科)
;
Department of Pathology, School of Basic Medical Sciences, Southern Medical University(南方医科大学基础医学学院病理科)
;
Department of Anatomical and Cellular Pathology, Chinese University of Hong Kong(香港中文大学解剖与细胞病理学系)
;
Guangdong Provincial Key Laboratory of Molecular Tumor Pathology(广东省分子肿瘤病理学重点实验室)
;
Jinfeng Laboratory(锦风实验室)
;
Department of Chemical and Biological Engineering, Hong Kong University of Science and Technology(香港科技大学化学与生物工程系)
;
Division of Life Science, Hong Kong University of Science and Technology(香港科技大学生命科学系)
;
State Key Laboratory of Nervous System Disorders, The Hong Kong University of Science and Technology(香港科技大学神经系统疾病国家重点实验室)
;
HKUST Shenzhen-Hong Kong Collaborative Innovation Research Institute, The Hong Kong University of Science and Technology(香港科技大学深圳-香港协同创新研究院)
机构
*
Nanyang Technological University(南洋理工大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
A*STAR Institute for Infocomm Research (I2R)(新加坡科技研究局资讯通信研究院)
BiEAR: A Human Auditory-Inspired Adaptive Binaural Front-end for Multi-Speaker Localisation and Distance Estimation
BiEAR: 一种受人类听觉启发的自适应双耳前端,用于多说话人定位和距离估计
Hanyu Meng, Eliathamby Ambikairajah, Vidhyasaharan Sethu, Qiquan Zhang, Haizhou Li
机构
*
The University of New South Wales(新南威尔士大学)
;
Tongyi Speech Lab, Alibaba Group(通义语音实验室,阿里巴巴集团)
;
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen(人工智能学院,香港中文大学(深圳))
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Xinjiang University(新疆大学)
;
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Independent Researcher(独立研究者)
;
The Chinese University of Hong Kong(香港中文大学)
Phonetic Error Analysis of Raw Waveform Acoustic Models
原始波形声学模型的音素错误分析
Erfan Loweimi, Zhengjun Yue, Andrea Carmantini, Zoran Cvetkovic, Steve Renals, Peter Bell
机构
*
Centre for Speech Technology Research (CSTR), University of Edinburgh, UK(语音技术研究中心(CSTR),爱丁堡大学,英国)
;
Cisco, UK(思科公司,英国)
;
SLAI & CUHK-SZ, China(SLAI与CUHK-SZ,中国)
;
King's College London, UK(伦敦国王学院,英国)
RASFT: Rollout-Adaptive Supervised Fine-Tuning for Reasoning
RASFT: 用于推理的滚动自适应监督微调
Yongliang Miao, Fengyuan Liu, Wei Shi, Yanguang Liu, Fei Sun, Na Zou, Mengnan Du
机构
*
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
New Jersey Institute of Technology(新泽西理工学院)
;
Institute of Computing Technology, CAS(中国科学院计算技术研究所)
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
S-Lab, Nanyang Technological University(南洋理工大学S实验室)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
University of Science and Technology of China(中国科学技术大学)
;
Stanford University(斯坦福大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
The Chinese University of Hong Kong(香港中文大学)
;
Fudan University(复旦大学)
;
CPII under InnoHK(InnoHK下的CPII)
;
Adobe Research(Adobe研究)
IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems
IRAF:面向噪声鲁棒的端到端全双工口语对话系统的抗干扰自适应融合
Tao Zhong, Jiajun Deng, Nikita Kuzmin, Yinke Zhu, Tianxiang Cao, Tristan Tsoi, Zhili Tan, Simon Lui, Xunying Liu
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
AudioLab Hong Kong, Huawei Leibniz Research Center(香港AudioLab,华为Leibniz研究中心)
;
Nanyang Technological University(南洋理工大学)