机构
*
University of Virginia(弗吉尼亚大学)
;
J. Crayton Pruitt Family Department of Biomedical Engineering, Herbert Wertheim College of Engineering, University of Florida(佛罗里达大学赫伯特·韦特海姆工程学院J. Crayton Pruitt家庭生物医学工程系)
Co-policy: Responsive Human-Robot Co-Creation for Musical Performances
Co-policy: 响应式人机音乐共创框架
Xuetao Li, Wenke Huang, Mang Ye, Zijian Liu, Jinhua Xie, Jifeng Xuan, Miao Li
机构
*
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
School of Automation, Wuhan University of Technology(武汉理工大学自动化学院)
;
School of Geodesy and Geomatics, Wuhan University(武汉大学测绘学院)
;
School of Robotics, Wuhan University(武汉大学机器人学院)
LLM-Based Synthetic Ground Truth Generation for Audio-Based Emotion Classification via In-Context Learning
基于上下文学习的音频情感分类的LLM合成真实标签生成
Qing Huang, Pooja Pol, Jianing Zhang
机构
*
School of Business, Technical University of Applied Sciences Augsburg(应用技术大学阿沙芬堡商学院)
;
Data Science und Autonome Systeme Technologietransferzentrum (TTZ)(数据科学与自主系统技术转移中心(TTZ))
NEST: Narrative Event Structures in Time for Long Video Understanding
NEST:面向长视频理解的时间叙事事件结构
Ali Asgarov, Kaushik Narasimhan, Najibul Haque Sarker, Hani Alomari, Chia-Wei Tang, Anushka Sivakumar, Zaber Ibn Abdul Hakim, Shaurya Mallampati, Chris Thomas
机构
*
Department of Computer Science, Virginia Tech(弗吉尼亚理工大学计算机科学系)
Reliability-Aware Prototype Calibration for Frozen Pose-Flow Video Anomaly Detection
面向冻结姿态流视频异常检测的可靠性感知原型校准
Ning Dong, Yingna Su, Xin Dong, Ziyun Jiao, Xinnian Guo, Zhuangzhuang Pan
机构
*
School of Information Engineering(信息工程学院)
;
Suqian University(宿迁大学)
;
School of Electronic & Information Engineering(电子与信息工程学院)
;
Nanjing University of Information Science and Technology(南京信息科学技术大学)
;
University of Electronic Science and Technology of China(电子科学与技术大学)
;
Institute for Advanced Studies(高级研究机构)
;
Universiti Malaya(马来亚大学)
Frequency-Aware Flow Matching for Continuous and Consistent Robotic Action Generation
频率感知流匹配用于连续且一致的机器人动作生成
Jianing Guo, Fangzheng Chen, Zihao Mao, Wong Lik Hang Kenny, Zhenhong Wu, Yu Li, Yishuai Cai, Yuanpei Chen, Yikun Ban, Kai Chen, Qi Dou, Yaodong Yang, Xianglong Liu, Huijie Zhao, Simin Li
机构
*
Beihang University(北京航空航天大学)
;
Peking University(北京大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
PKU-Psibot Lab(北大-智源机器人实验室)
;
Zhongguancun Laboratory(中关村实验室)
;
Hefei Comprehensive National Science Center(合肥综合性国家科学中心)
CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs
CARE: 面向视频多模态大语言模型的自适应推理长度的能力感知奖励塑形
Chengwen Liu, Hao Peng, Jisheng Dang, Hong Peng, Bin Hu, Tat-Seng Chua
机构
*
School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院)
;
School of Medical Technology, Beijing Institute of Technology(北京理工大学医学技术学院)
;
School of Computing, National University of Singapore(新加坡国立大学计算机学院)
SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm
SARLO-80:全球斜距SAR语言光学数据集80cm
Solène Debuysère, Nicolas Trouvé, Nathan Letheule, Elise Colin, Georgia Channing
机构
*
DEMR-ONERA – The French Aerospace Lab, Université Paris-Saclay(法国航空航天实验室DEMR-ONERA,巴黎-萨克雷大学)
;
DTIS-ONERA – The French Aerospace Lab, Université Paris-Saclay(法国航空航天实验室DTIS-ONERA,巴黎-萨克雷大学)
;
Hugging Face
专题命中
跨模态检索
:multimodal(abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI
Multi-Modal Contrastive Learning for Implicit Earth Embeddings via Location Tying
多模态对比学习用于基于位置绑定的隐式地球嵌入
Jonathan Hecht, Lukas Arzoumanidis, Ziyue Li, Youness Dehbi
机构
*
Computational Methods Lab, HafenCity University Hamburg(汉堡港城大学计算方法实验室)
;
Dept. of Operations & Technology, Technical University of Munich(慕尼黑工业大学运营与技术系;海尔布隆数据科学中心;慕尼黑数据科学研究所)
;
Heilbronn Data Science Center(波恩大学大地测量与地理信息研究所)
;
Munich Data Science Institute
;
Institute of Geodesy and Geoinformation, University of Bonn
Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA
多模态大语言模型的置信度校准:基于医学视觉问答的实证研究
Yuetian Du, Yucheng Wang, Ming Kong, Tian Liang, Qiang Long, Bingdi Chen, Qiang Zhu
机构
*
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院)
;
Zhihui Medical Technology (Shanghai) Co., Ltd.(智汇医疗科技(上海)有限公司)
HEad and neCK TumOR (HECKTOR) 2025: Benchmark of Segmentation, Diagnosis, and Prognosis in Multimodal PET/CT
头颈肿瘤 (HECKTOR) 2025 挑战赛:多模态 PET/CT 中的分割、诊断与预后基准
Numan Saeed, Salma Hassan, Shahad Hardan, Lishan Cai, Xinglong Liang, Moona Mazher, Abdul Qayyum, Yansong Bu, Mengye Lyu, Yue Lin, Mingyuan Meng, Chuanyi Huang, Lisheng Wang, Dalal Chamseddine, Shamimeh Ahrari, Beining Wu, Yifei Chen, Fuyou Mao, Hao Zhang, Baixiang Zhao, Surajit Ray, Muzi Guo, Lei Xiang, Jakob Dexl, Michael Ingrisch, Adrien Depeursinge, Arman Rahmim, Mathieu Hatt, Vincent Andrearczyk, Mohammad Yaqub
机构
*
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)(穆罕默德·本·扎耶德人工智能大学)
;
Amsterdam UMC(阿姆斯特丹大学医学中心)
;
The Netherlands Cancer Institute(荷兰癌症研究所)
;
Radboud University Medical Centre(拉德堡德大学医学中心)
;
University College London(伦敦大学学院)
;
Imperial College London(帝国理工学院)
;
Shenzhen Technology University(深圳技术大学)
;
Shenzhen University(深圳大学)
;
Newland Digital Technology(新大陆数字技术)
;
The University of Sydney(悉尼大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
University Hospital, Nantes(南特大学医院)
;
Nantes Université, Centrale Nantes, CNRS, LS2N(南特大学、南特中央理工学院、法国国家科学研究中心、LS2N实验室)
;
Hangzhou Dianzi University(杭州电子科技大学)
;
Tsinghua University(清华大学)
;
Central South University(中南大学)
;
University of Glasgow(格拉斯哥大学)
;
China Mobile System Integration Co., Ltd.(中移系统集成有限公司)
;
Subtle Medical Inc.(Subtle Medical公司)
;
University Hospital, LMU Munich(慕尼黑大学医院)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
BC Cancer Research Institute(不列颠哥伦比亚癌症研究所)
;
HES-SO Valais-Wallis University of Applied Sciences and Arts(HES-SO瓦莱州应用科学与艺术大学)
;
Lausanne University Hospital (CHUV)(洛桑大学医院)
;
LaTIM, INSERM, UMR 1101, Univ Brest(LaTIM实验室、法国国家健康与医学研究院、UMR 1101、布雷斯特大学)
Comments17 pages, 4 figures, 4 tables. Overview paper for the HECKTOR 2025 challenge, held as a satellite event at MICCAI 2025. Challenge website: https://hecktor.grand-challenge.org/
StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs
StylisticBias: 少数人类视觉线索驱动多模态大语言模型中的大部分社会偏见
Shaghayegh Kolli, Timo Cavelius, Nafiseh Nikeghbal, Samantha Dalal, Jana Diesner
机构
*
Technical University of Munich(慕尼黑工业大学)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
Princeton Center for Information and Technology Policy(普林斯顿信息与技术政策中心)
NRITYAM: Language Models Meet Art and Heritage of Dance
NRITYAM:语言模型遇见舞蹈的艺术与遗产
Punit Kumar Singh, Niladri Ghosh, Advait Joshiınst, Shailee Choudhary, Michael Färber, Haiqin Yang
机构
*
Shenzhen Technology University(深圳技术大学)
;
New Delhi Institute of Management(新德里管理学院)
;
Technische Universität Dresden(德累斯顿工业大学)
;
Ramakrishna Mission Vivekananda Educational and Research Institute(罗摩克里希纳传道会维韦卡南达教育与研究学院)
;
Indian Institute of Technology(印度理工学院)
;
Swami Vivekananda Institute of Technology(斯瓦米·维韦卡南达技术学院)
;
GuangDong Engineering Technology Research Center of Edge Intelligence(广东省边缘智能工程技术研究中心)