机构
*
Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)
;
Zhongguancun Academy(中关村学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Nanyang Technological University(南洋理工大学)
;
Department of Mechanical Engineering, Imperial College London(伦敦帝国理工学院机械工程系)
机构
*
Department of Electronic Engineering, Tsinghua University(清华大学电子工程系)
;
Institute for Embodied Intelligence and Robotics, Tsinghua University(清华大学智能感知与机器人研究院)
;
Department of Computer Science and Engineering, Shanghai Jiao Tong University(上海交通大学计算机科学与工程系)
;
Huakong AI Plus Company Limited(华冠AIplus有限公司)
;
Didi International Business Group(滴滴国际商务集团)
Attention, not scale, drives human-AI alignment in multimodal language prediction
注意力,而非规模,驱动多模态语言预测中的人机对齐
Viktor Kewenig, Andrew Lampinen, Samuel A. Nastase, Christopher Edwards, Quitterie Lacome D'Elascombe, Akilles Rechardt, Jeremy I Skipper, Gabriella Vigliocco
机构
*
Psychology and Language Science, Experimental Psychology, University College London, London, UK(心理学与语言科学、实验心理学,伦敦大学学院,伦敦,英国)
;
Google Deepmind, Mountain View, US(谷歌DeepMind,山景城,美国)
;
Princeton Neuroscience Institute, Princeton University, Princeton, NJ, USA(普林斯顿神经科学研究所,普林斯顿大学,普林斯顿,新泽西州,美国)
;
Computer Science Department, Exeter University(计算机科学系,埃克塞特大学)
机构
*
Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院,清华大学)
;
Pengcheng Laboratory(鹏城实验室)
;
Ant Group(蚂蚁集团)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳))
;
University of Pennsylvania(宾夕法尼亚大学)
IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models
IoUPD:用于多模态大语言模型视觉定位的IoU感知特权蒸馏
Xiuyuan Zhu, Ke Lu, Hao Wu, Siwen Jiao, Zijin Du, Dongming Zhang, Jian Xue
机构
*
University of Chinese Academy of Sciences(中国科学院大学)
;
State Key Laboratory of Communication Content Cognition(通信内容认知技术国家重点实验室)
;
Peng Cheng Laboratory(鹏城实验室)
RankByGene: Gene-Guided Histopathology Representation Learning Through Cross-Modal Ranking Consistency
RankByGene: 通过跨模态排序一致性实现基因引导的组织病理学表示学习
Wentao Huang, Meilong Xu, Xiaoling Hu, Shahira Abousamra, Aniruddha Ganguly, Saarthak Kapse, Alisa Yurovsky, Prateek Prasanna, Tahsin Kurc, Joel Saltz, Michael L. Miller, Chao Chen
机构
*
Stony Brook University(石英溪大学)
;
Athinoula A. Martinos Center for Biomedical Imaging, Massachusetts General Hospital and Harvard Medical School(阿提诺拉A.马丁努斯生物医学影像中心,麻省总医院和哈佛医学院)
;
Department of Biomedical Data Science, Stanford University(生物医学数据科学系,斯坦福大学)
;
Department of Pathology and Cell Biology, Columbia University(病理学与细胞生物学系,哥伦比亚大学)
Yan Zhou, Suncheng Xiang, Zhen Huang, Yue Ouyang, Yingqiu Li, Zehua Wang
机构
*
School of Mathematics and Statistics, Changsha University of Science and Technology(数学与统计学学院,长沙理工大学)
;
School of Biomedical Engineering, Shanghai Jiao Tong University(生物医学工程学院,上海交通大学)
;
Shanghai Chest Hospital, Shanghai Jiao Tong University School of Medicine(上海胸科医院,上海交通大学医学院)
From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States
从方向到大小:多模态指令微调如何在Transformer隐藏状态中重新组织身份指定提示的几何编码
Jorge A. Castillo, Marco Torres Yévenes, Juan Carlos Lanas
EHR2Path: Comprehensive Pathway-Level Modeling of Longitudinal Patient Trajectories from Multimodal Electronic Health Records
EHR2Path:从多模态电子健康记录中可扩展地建模纵向患者路径
Chantal Pellegrini, Ege Özsoy, David Bani-Harouni, Matthias Keicher, Nassir Navab
机构
*
Technical University of Munich(慕尼黑技术大学)
;
TUM School of Computation, Information and Technology(慕尼黑技术大学计算、信息与技术学院)
;
Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
机构
*
1 School of Medical Science \& Technology, IIT Kharagpur, India 2 Department of Electrical Engineering, IIT Kharagpur, India 3 Dr. B.C. Roy Multispeciality Medical Research Centre, IIT Kharagpur, India
LP-SFT: Local-Preserving Supervised Fine-Tuning via Multimodal Entropy Structure
LP-SFT:通过多模态熵结构进行局部保持监督微调
Yueyang Wang, Baolong Bi, Shuo Lu, Jingyuan Zhang, Jiajun Shi
机构
*
School of Mathematical Sciences, Peking University(北京大学数学科学学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
College of Computing, Georgia Institute of Technology(佐治亚理工学院计算学院)
CommentsThis preprint is withdrawn for unauthorized posting and incorrect author metadata.It was uploaded without full consent of all co-authors, with wrong name and affiliation information. We withdraw it to avoid copyright disputes. This corrects submission irregularities only, not academic content or conclusions
Where Will They Go? Modelling Multimodal Pedestrian Manoeuvres from Ego-centric Videos
他们将去哪里?从自我中心视频建模多模态行人机动
Yuxuan Xie, Nicolas Pugeault, Chongfeng Wei, Hubert P. H. Shum, Edmond S. L. Ho
机构
*
School of Computing Science, University of Glasgow(格拉斯哥大学计算机科学学院)
;
James Watt School of Engineering, University of Glasgow(格拉斯哥大学詹姆斯·瓦特工程学院)
;
Department of Computer Science, Durham University(杜伦大学计算机科学系)
机构
*
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统实验室)
;
Institute of Microelectronics, University of Chinese Academy of Sciences(中国科学院大学微电子学院)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学前沿交叉科学学院)