机构
*
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
WINDY Lab, Department of Artificial Intelligence, Westlake University(西湖大学人工智能系WINDY实验室)
;
CVGL Lab(CVGL实验室)
;
The Chinese University of Hong Kong(香港中文大学)
Comments8 pages, 7 figures. This paper is accepted by IEEE Robotics and Automation Letters
Journal refIEEE Robotics and Automation Letters, 2025
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
Long Xing, Qidong Huang, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Jinsong Li, Shuangrui Ding, Weiming Zhang, Nenghai Yu, Jiaqi Wang, Feng Wu, Dahua Lin
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong(香港中文大学)
Position: Intelligent Science Laboratory Requires the Integration of Cognitive and Embodied AI
Sha Zhang, Suorong Yang, Tong Xie, Xiangyuan Xue, Zixuan Hu, Rui Li, Wenxi Qu, Zhenfei Yin, Tianfan Fu, Di Hu, Andres M Bran, Nian Ran, Bram Hoex, Wangmeng Zuo, Philippe Schwaller, Wanli Ouyang, Lei Bai, Yanyong Zhang, Lingyu Duan, Shixiang Tang, Dongzhan Zhou
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Nanjing University(南京大学)
;
University of New South Wales(新南威尔士大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Peking University(北京大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Tongji University(同济大学)
;
The University of Sydney, Oxford(悉尼大学,牛津)
;
Renmin University of China(中国人民大学)
;
Swiss Federal Institute of Technology Lausanne(苏黎世联邦理工学院)
;
Shanghai Institute of Ceramics, Chinese Academy of Sciences(中国科学院上海硅酸盐研究所)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Shanghai AI Laboratory & The Chinese University of Hong Kong(上海人工智能实验室 & 香港中文大学)
机构
*
University of Electronic Science and Technology of China(电子科技大学)
;
CISPA Helmholtz Center for Information Security(信息安全赫尔姆霍茨中心)
;
The Chinese University of Hong Kong(香港中文大学)
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
Nianchen Deng, Lixin Gu, Shenglong Ye, Yinan He, Zhe Chen, Songze Li, Haomin Wang, Xingguang Wei, Tianshuo Yang, Min Dou, Tong He, Wenqi Shao, Kaipeng Zhang, Yi Wang, Botian Shi, Yanting Zhang, Jifeng Dai, Yu Qiao, Hongjie Zhang, Wenhai Wang
机构
*
Shanghai AI Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong(香港中文大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Donghua University(东华大学)
;
Nanjing University(南京大学)
;
Tsinghua University(清华大学)
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
The Chinese University of Hong Kong(香港中文大学)
;
Bytedance(字节跳动)
;
National University of Singapore(新加坡国立大学)
;
Tsinghua University(清华大学)
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学)
;
Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)
;
CUHK(香港中文大学)
机构
*
The University of Sydney, Sydney, Australia(悉尼大学)
;
The University of Hong Kong, Hong Kong SAR, China(香港大学)
;
The Chinese University of Hong Kong, Hong Kong SAR, China(香港中文大学)
Hierarchical Control of Emotion Rendering in Speech Synthesis
Sho Inoue, Kun Zhou, Shuai Wang, Haizhou Li
机构
*
School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), China(数据科学学院,香港中文大学(深圳))
;
School of Intelligence Science and Technology, Nanjing University, Suzhou, China(智能科学与技术学院,南京大学)
;
Shenzhen Research Institute of Big Data, Shenzhen, China(大数据研究 institute,深圳,中国)
;
Tongyi Speech Lab, Alibaba Group, Singapore(通义语音实验室,阿里巴巴集团,新加坡)
;
Department of ECE, National University of Singapore, Singapore(电子工程系,新加坡国立大学)
CommentsAccepted to IEEE Transactions on Affective Computing
机构
*
University of Science and Technology of China(中国科学技术大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Institute of High Performance Computing, A*STAR, Singapore(新加坡A*STAR高性能计算研究所)
A CLIP-Powered Framework for Robust and Generalizable Data Selection
Suorong Yang, Peng Ye, Wanli Ouyang, Dongzhan Zhou, Furao Shen
机构
*
National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong(香港中文大学)
Ziran Zhu, Tongda Xu, Minye Huang, Dailan He, Xingtong Ge, Xinjie Zhang, Ling Li, Yan Wang
机构
*
Institute for AI Industry Research, Tsinghua University(清华人工智能产业研究院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
;
SenseTime Research(商汤科技研究院)
;
The Chinese University of Hong Kong(香港中文大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Hong Kong University of Science and Technology(香港科技大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
National University of Singapore(新加坡国立大学)
;
Nanyang Technological University(南洋理工大学)
;
Macquarie University(麦考瑞大学)
机构
*
State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室)
;
School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
;
University of Chicago(芝加哥大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
East China Normal University(华东师范大学)
Cellular Traffic Prediction via Deep State Space Models with Attention Mechanism
Hui Ma, Kai Yang, Man-On Pun
机构
*
College of Electronic and Information Engineering, Tongji University(同济大学电子与信息学院)
;
School of Science and Engineering, The Chinese University of Hong Kong at Shenzhen(香港中文大学(深圳)理工学院)
;
Shenzhen Key Laboratory of IoT Intelligent Systems and Wireless Network Technology, Chinese University of Hong Kong at Shenzhen(深圳物联网智能系统与无线网络技术重点实验室,香港中文大学(深圳))
;
Shenzhen Research Institute of Big Data, Chinese University of Hong Kong at Shenzhen(深圳大数据研究 institute,香港中文大学(深圳))
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
Junfei Wu, Jian Guan, Kaituo Feng, Qiang Liu, Shu Wu, Liang Wang, Wei Wu, Tieniu Tan
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Ant Group(蚂蚁集团)
;
CUHK MMLab(香港中文大学MMLab)
;
Nanjing University(南京大学)
Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation
Ning Wang, Zihan Yan, Weiyang Li, Chuan Ma, He Chen, Tao Xiang
机构
*
College of computer science, Chongqing University(重庆大学计算机学院)
;
Department of Information Engineering, The Chinese University of Hong Kong(香港中文大学信息工程系)
Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-Task Learning
Ziyu Zhao, Yixiao Zhou, Zhi Zhang, Didi Zhu, Tao Shen, Zexi Li, Jinluan Yang, Xuwu Wang, Jing Su, Kun Kuang, Zhongyu Wei, Fei Wu, Yu Cheng
机构
*
Zhejiang University(浙江大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
ByteDance Inc.(字节跳动公司)
;
Fudan University(复旦大学)
;
The Chinese University of Hong Kong(香港中文大学)
Demystify Transformers & Convolutions in Modern Image Deep Networks
Xiaowei Hu, Min Shi, Weiyun Wang, Sitong Wu, Linjie Xing, Wenhai Wang, Xizhou Zhu, Lewei Lu, Jie Zhou, Xiaogang Wang, Yu Qiao, Jifeng Dai
机构
*
South China University of Technology(南方科技大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Huazhong University of Science and Technology(华中科技大学)
;
Fudan University(复旦大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Tsinghua University(清华大学)
;
SenseTime Research(商汤科技研究院)
CommentsThis paper was accepted to IEEE Transactions on Pattern Analysis and Machine Intelligence (IEEE TPAMI). All models and codes used in this study are publicly available at https://github.com/OpenGVLab/STM-Evaluation
Journal refIEEE Trans Pattern Anal Mach Intell, 47, 2025, pp. 2416-2428