MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
MedPMC:一种用于为基础模型扩展高保真医学多模态数据的系统框架
Hyunjae Kim, Dain Kim, Pan Xiao, Serina S. Applebaum, Younjoon Chung, Xuguang Ai, Yu Yin, Roy Jiang, Yuexi Du, Yawen Wei, Yiming Kong, Tuo Guo, Zhiyuan Cao, Mengmeng Du, Yuelei Fu, Yan Hu, Rui Shi, Gui Yang, Kevin W. Jin, Yuntian Liu, Yuxuan Tian, Jonathan Marquez, Zhen Chen, Sheng Zhang, Hoifung Poon, Hua Xu, Jaewoo Kang, Qingyu Chen
机构
*
Yale University(耶鲁大学)
;
Korea University(韩国大学)
;
Washington University in St. Louis(圣路易斯华盛顿大学)
;
The University of Queensland(昆士兰大学)
;
The University of Texas Health Science Center at Houston(德克萨斯大学休斯顿健康科学中心)
;
University of Washington(华盛顿大学)
;
Microsoft Research(微软研究院)
专题命中
图文多模态
:multimodal(title,abstract);image-text(abstract);multimodal foundation model(abstract);分类 cs.CV
Creativity from Friction: Human-AI Interaction for Exploratory Structural Design
摩擦产生创造力:用于探索性结构设计的人机交互
Ricardo Maia Avelino, Rita Sevastjanova, Tom Van Mele, Philippe Block, Mennatallah El-Assady
机构
*
Block Research Group, Department of Architecture, ETH Zurich(建筑系,苏黎世联邦理工学院)
;
IVIA Lab, Department of Computer Science, ETH Zurich(计算机科学系,苏黎世联邦理工学院)
LEMUR 2: Unlocking Neural Network Diversity for AI
LEMUR 2:释放人工智能的神经网络多样性
Tolgay Atinc Uzun, Waleed Khalid, Saif U Din, Sai Revanth Mulukuledu, Akashdeep Singh, Chandini Vysyaraju, Raghuvir Duvvuri, Avi Goyal, Yashkumar Rajeshbhai Lukhi, Muhammad A. Hussain, Krunal Jesani, Usha Shrestha, Yash Mittal, Roman Kochnev, Pritam Kadam, Mohsin Ikram, Harsh R. Moradiya, Alice Arslanian, Dmitry Ignatov, Radu Timofte
A Good Initialization is All You Need for Faithful Visual Attribution
忠实视觉归因只需一个良好的初始化
Zihan Gu, Jiayu Wang, Hua Zhang, Yue Hu
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院)
Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders
用于社交机器人对话轮次转换的多模态语音活动投影及与语音活动相关的预训练编码器
Antonio Cano, Guillermo Pérez, Luis Merino, Randy Gomez
机构
*
i Intelligent Insights(4i智能洞察公司)
;
Universidad de Sevilla(塞维利亚大学)
;
Universidad Pablo de Olavide(巴勃罗·德奥拉维德大学)
;
Honda Research Institute Japan(本田日本研究所)
CommentsAccepted for presentation at the 35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN 2026). Acceptance notification date: 30 May 2026. Final published version pending
CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training
CarbonCLIP:通过集成街景语义和时间上下文训练增强卫星图像的碳排放预测
Zeru Yang, Fang-Ying Gong, Steve H. L. Yim, Chau Yuen
机构
*
Energy Research Institute at NTU(国立理工学院能源研究所)
;
Interdisciplinary Graduate Programme(跨学科研究生项目)
;
School of Electrical and Electronic Engineering(电气电子工程学院)
;
School of Public Administration and Policy(公共管理与政策学院)
;
Asian School of the Environment(亚洲环境学院)
;
Center for Climate Change and Environmental Health(气候变化与环境健康中心)
MMAgent-R$^2$: Learning to Rerank and Reject for Agentic mRAG
MMAgent-R$^2$:用于智能mRAG的重排与拒绝学习
Tao Zhang, Ziqi Zhang, Zongyang Ma, Yuxin Yang, Bing Li, Chunfeng Yuan, Kang Rong, Fengyun Rao, Jing Lyu, Weiming Hu
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(中国科学院复杂系统管理与控制国家重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information(北京多模态信息超智能安全重点实验室)
;
WeChat Vision, Tencent Inc.(腾讯微信视觉团队)
;
PeopleAI Inc.(人智公司)
;
School of Information Science and Technology, ShanghaiTech University(上海科技大学信息科学与技术学院)
机构
*
NepAl Applied Mathematics and Informatics Institute(尼泊尔应用数学与信息学研究所)
;
Fogsphere (Redev.AI Ltd)(福格球(Redev.AI有限公司))
;
University of Lausanne(洛桑大学)
;
West Virginia University(西弗吉尼亚大学)
;
University of Verona(维罗纳大学)
;
University College London(伦敦大学学院)
;
University of Aberdeen(阿伯丁大学)
Triple-Phase Multimodal Knowledge Aggregation Framework for Microbial Keratitis Subtype Diagnosis on Slit-Lamp Photography
基于裂隙灯摄影的微生物性角膜炎亚型诊断三相多模态知识聚合框架
Yiqing Wang, Maria A. Woodward, Ziyun Yang, N. Venkatesh Prajna, Chunming He, Leslie M. Niziol, Mercy Pawar, Ming-Chen Lu, Guillermo Amescua, Rachel Wozniak, Sejal Amin, Abinaya Krishnan, Prabhleen Kochar, Sina Farsiu
机构
*
Department of Biomedical Engineering, Duke University(杜克大学生物医学工程系)
;
Kellogg Eye Center, Department of Ophthalmology and Visual Sciences, University of Michigan(密歇根大学凯洛格眼科中心,眼科学与视觉科学系)
;
Department of Cornea and Refractive Surgery Services, Aravind Eye Care System(阿瓦因眼科医疗系统角膜与屈光手术部)
;
Bascom Palmer Eye Institute, Department of Ophthalmology, University of Miami Miller School of Medicine(迈阿密大学米勒医学院巴斯科姆·帕勒眼科研究所,眼科学系)
;
Flaum Eye Institute, Department of Ophthalmology, University of Rochester Medical Center(罗切斯特大学医学中心弗劳姆眼科研究所,眼科学系)
;
Department of Ophthalmology, Henry Ford Hospital(亨利福特医院眼科部)
;
Duke Eye Center, Duke University School of Medicine(杜克大学医学院杜克眼科中心)
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
通过深度原生结构推理实现准确、跨学科和透明的结构-属性理解
Chen Tang, Yizhou Wang, Jianyu Wu, Lintao Wang, Shixiang Tang, Pengze Li, Encheng Su, Jun Yao, Jiabei Xiao, Yuqi Shi, Jielan Li, Hongxia Hao, Zhangyang Gao, Fang Wu, Ben Fei, Xiangyu Yue, Pan Tan, Bozitao Zhong, Jinouwen Zhang, Aoran Wang, Yan Lu, Jiaheng Liu, Xinzhu Ma, Liang Hong, Mingyue Zheng, Phil Torr, Bowen Zhou, Wanli Ouyang, Lei Bai
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Fudan University(复旦大学)
;
University of Sydney(悉尼大学)
;
Nanjing University(南京大学)
;
University of Oxford(牛津大学)
;
The University of Science and Technology of China(中国科学技术大学)
;
Drug Discovery and Design Center, State Key Laboratory of Drug Research, Shanghai Institute of Materia Medica, Chinese Academy of Sciences(药物发现与设计中心、国家药物研究重点实验室、上海中医药材料医学研究所、中国科学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Stanford University(斯坦福大学)
Comparative Study of Domain-adapted VLMs for General Document Visual Question Answering
用于通用文档视觉问答的领域适应视觉语言模型的比较研究
Miguel Lopez-Duran, Elena Marrero, Julian Fierrez, Marta Robledo-Moreno, Ruben Vera-Rodriguez, Daniel DeAlcala, Aythami Morales, Ruben Tolosana, Oscar Delgado, Alvaro Ortigosa, Javier Ortega-Garcia
机构
*
Universidad Autónoma de Madrid (UAM)(马德里自治大学)
;
BiometricsAI(生物识别人工智能)