机构
*
School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院)
;
National Institute of Health Data Science, Peking University(北京大学健康数据科学国家研究院)
;
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
;
Tianjin Institute of Cardiology, the Second Hospital of Tianjin Medical University(天津医科大学第二医院心内科)
;
National University of Singapore(新加坡国立大学)
;
Jarvis Lab, Tencent(腾讯 Jarvis实验室)
;
HeartVoice Medical Technology(HeartVoice医疗科技)
专题命中
视觉定位与Grounding
:MLLM(title,abstract);grounding(abstract);multimodal large language model(abstract)
Language Movement Primitives: Grounding Language Models in Robot Motion
语言运动基元:将语言模型锚定在机器人运动中
Yinlong Dai, Benjamin A. Christie, Daniel J. Evans, Dylan P. Losey, Simon Stepputtis
机构
*
Collab , Dept. of Mechanical Engineering, Virginia Tech, Blacksburg, VA 24061(合作组,机械工程系,弗吉尼亚理工学院,黑斯堡,VA 24061)
;
TEA Lab , Dept. of Mechanical Engineering, Virginia Tech, Blacksburg, VA 24061(TEA实验室,机械工程系,弗吉尼亚理工学院,黑斯堡,VA 24061)
You Can Ground Earlier than See: An Effective and Efficient Pipeline for Temporal Sentence Grounding in Compressed Videos
你可以比看见更早定位:一种用于压缩视频中时序句子定位的高效流程
Xiang Fang, Daizong Liu, Pan Zhou, Guoshun Nan
机构
*
The Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology(大数据安全湖北工程研究中心,网络安全科学与工程学院,华中科技大学)
;
Peking University(北京大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
Hierarchical Local-Global Transformer for Temporal Sentence Grounding
层次化局部-全局Transformer用于时间语句定位
Xiang Fang, Daizong Liu, Pan Zhou, Zichuan Xu, Ruixuan Li
机构
*
Hubei Engineering Research Center on Big Data Security, School of Cyber Science and Engineering, Huazhong University of Science and Technology(大数据安全湖北工程研究中心,华中科技大学网络安全科学与工程学院)
;
Wangxuan Institute of Computer fTechnology, Peking University(王宣计算机技术研究院,北京大学)
;
School of software, Dalian University of Technology(软件学院,大连理工大学)
;
School of Computer Science, and Technology, Huazhong University of Science, and Technology(计算机科学与技术学院,华中科技大学)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Third Affiliated Hospital of Sun Yat-Sen University(中山大学第三附属医院)
;
Hong Kong Metropolitan University(香港 Metropolitan 大学)
PathMem: Toward Cognition-Aligned Memory Transformation for Pathology MLLMs
PathMem: 面向病理学多模态大模型的认知对齐记忆转换
Jinyue Li, Yuci Liang, Qiankun Li, Xinheng Lyu, Jiayu Qian, Huabao Chen, Kun Wang, Zhigang Zeng, Anil Anthony Bharath, Yang Liu
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shenzhen University(深圳大学)
;
Nanyang Technological University(南洋理工大学)
;
Imperial College London(伦敦帝国学院)
;
Huazhong University of Science and Technology(华中科技大学)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);分类 cs.AI
机构
*
National College for Excellent Engineers, Beihang University, Beijing, China(北京航空航天大学优秀工程师学院)
;
School of Artificial Intelligence, Beihang University, Beijing, China(北京航空航天大学人工智能学院)
;
School of Electronic and Information Engineering, Beihang University, Beijing, China(北京航空航天大学电子与信息工程学院)
;
King Abdullah University of Science and Technology, Saudi Arabia(沙特国王 Abdullah 科学技术大学)
;
Huawei Noah’s Ark Lab, China(华为诺亚实验室)
Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs
通过先验增强的音频大语言模型统一语音编辑检测与内容定位
Jun Xue, Yi Chai, Yanzhen Ren, Jinshen He, Zhiqiang Tang, Zhuolin Yi, Yihuan Huang, Yuankun Xie, Yujie Chen
机构
*
Key Laboratory of Aerospace Information Security(航空信息安全与可信计算重点实验室)
;
School of Cyber Science and Engineering(网络安全工程学院)
;
Wuhan University(武汉大学)
;
Independent Researcher(独立研究员)
;
School of Computer Science and Technology(计算机科学与技术学院)
;
Anhui University(安徽大学)
;
Communication University of China(中国通信大学)
;
Beihang University(北京航空航天大学)