Does the Question Really Matter? Training-Free Data Selection for Vision-Language SFT
问题真的重要吗?视觉-语言SFT的无训练数据选择
Peng Sun, Yi Yang, Huawen Shen, Yi Ban, Tianfan Fu, Yanbo Wang, Yuqiang Li
机构
*
Nanjing University(南京大学)
;
Institute of Information Engineering(信息工程研究所)
;
North University of China(中国北方大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
Xiaobao Guo, Zitong Yu, Nithish Muthuchamy Selvaraj, Bingquan Shen, Adams Wai-Kin Kong, Alex C. Kot
机构
*
Rapid-Rich Object Search (ROSE) Lab and the College of Computing and Data Science, Nanyang Technological University (NTU)(快速丰富对象搜索(ROSE)实验室和南洋理工大学计算与数据科学学院)
;
School of Computing and Information Technology and Dongguan Key Laboratory for Intelligence and Information Technology, Great Bay University(计算与信息科技学院和东莞智能与信息技术重点实验室,大湾大学)
;
DSO National Laboratories(国防科学实验室)
;
College of Computing and Data Science, Nanyang Technological University (NTU)(计算与数据科学学院,南洋理工大学)
;
SMBU, Shenzhen 518172, China(深圳SMBU,越南河内VinUniversity,和新加坡NTU)
;
VinUniversity, Hanoi 100000, Vietnam
;
and NTU, Singapore
On Aligning Hierarchical Standardized Embedding for Audio-visual Generalized Zero-shot Learning
面向音视频广义零样本学习的层次化标准化嵌入对齐
Zihan Zhang, Jie Hong, Siyuan Fan, Yanghao Zhou, Pengfei Fang
机构
*
Southeast University(东南大学)
;
The University of Hong Kong(香港大学)
;
Beijing Institute of Technology(北京理工大学)
;
Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education(新一代人工智能技术及其跨学科应用重点实验室(东南大学),教育部)
;
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models
跨模态一致性引导用于自动回归TTS模型中的鲁棒情绪控制
Yizhou Peng, Yukun Ma, Chong Zhang, Yi-Wen Chao, Chongjia Ni, Bin Ma, Eng Siong Chng
机构
*
Alibaba-NTU Global e-Sustainability CorpLab(阿里-国立大学全球可持续发展公司实验室)
;
Nanyang Technological University(南洋理工大学)
;
College of Computing and Data Science(计算与数据科学学院)
;
Alibaba(阿里)
;
Alibaba Inc.(阿里公司)
机构
*
LMU Munich(慕尼黑大学)
;
Harvard University(哈佛大学)
;
University of Cambridge(剑桥大学)
;
Mina AI
;
Konrad Zuse School of Excellence in Reliable AI (relAI)(康拉德·楚泽可靠人工智能卓越学校(relAI))
Decoding Multimodal Cues: Unveiling the Implicit Meaning Behind Hateful Videos
解码多模态线索:揭示仇恨视频背后的隐含意义
Junyu Lu, Deyi Ji, Liqun Liu, Xiaokun Zhang, Youlin Wu, Roy Ka-Wei Lee, Peng Shu, Huan Yu, Jie Jiang, Bo Xu, Liang Yang, Hongfei Lin
机构
*
Dalian University of Technology(大连理工大学)
;
Tencent(腾讯)
;
City University of Hong Kong(香港城市大学)
;
Singapore University of Technology and Design(新加坡科技设计大学)
机构
*
Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)
;
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China(电子科技大学深圳高等研究院)
机构
*
Mohamed bin Zayed University of AI(莫扎德·本·扎耶德人工智能大学)
;
University of Chicago(芝加哥大学)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
;
Linköping University(林奈大学)
机构
*
Advanced Technologies Application Center (CENATAV)(先进技术应用中心(CENATAV))
;
Centro de Sistemas Complejos, Facultad de Física, Universidad de La Habana(哈瓦那大学物理学院复杂系统中心)
GLACIER: A Multimodal Student-Teacher Foundation Model for Molecular Property Prediction
GLACIER:用于分子性质预测的多模态师生基础模型
Emily Nguyen, Yongchan Hong, Harsh Toshniwal, Yan Liu, Andreas Luttens
机构
*
Department of Computer Science, University of Southern California(南加州大学计算机科学系)
;
Department of Quantitative and Computational Biology, University of Southern California(南加州大学定量与计算生物学系)
;
Amazon(亚马逊)
;
Department of Medical Biochemistry and Biophysics, Science for Life Laboratory, Karolinska Institutet(卡罗林斯卡学院医学生物化学与生物物理系,生命科学实验室)
Alexander Martin, Dengjia Zhang, Joel Brogan, Francis Ferraro, Jeremy Gwinnup, Reno Kriz, Teng Long, Kenton Murray, Andrew Yates, Xiang Xiang
机构
*
Johns Hopkins University(约翰霍普金斯大学)
;
OpenAI
;
University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)
;
Air Force Research Laboratory(空军研究实验室)
;
Human Language Technology Center of Excellence, Johns Hopkins University(约翰霍普金斯大学人类语言技术卓越中心)
;
University of Amsterdam(阿姆斯特丹大学)
;
Huazhong University of Science and Technology(华中科技大学)
CommentsFindings of the 2nd workshop on Multimodal Augmented Generation via Multimodal Retrieval (MAGMaR); Resources at this url: https://github.com/rekriz11/MAGMAR_2026