Med-K2N: Flexible K-to-N Modality Translation for Medical Image Synthesis
Feng Yuan, Yifan Gao, Yuehua Ye, Haoyue Li, Xin Gao
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Suzhou Institute of Biomedical Engineering and Technology(苏州生物医学工程与技术研究所)
;
Chinese Academy of Sciences(中国科学院)
;
The Third Affiliated Hospital of Sun Yat-sen University(中山大学第三附属医院)
CommentsThis is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland, https://doi.org/10.1145/3746027.3758297
Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
Punit Kumar Singh, Nishant Kumar, Akash Ghosh, Kunal Pasad, Khushi Soni, Manisha Jaishwal, Sriparna Saha, Syukron Abu Ishaq Alfarozi, Asres Temam Abagissa, Kitsuchart Pasupa, Haiqin Yang, Jose G Moreno
机构
*
Indian Institute of Technology Patna(印度理工学院帕纳布分校)
;
Sardar Patel Institute of Technology(萨达尔·帕特尔技术学院)
;
Universitas Gadjah Mada(加查马大学)
;
King Mongkut’s Institute of Technology Ladkrabang(拉差班国王技术学院)
;
Shenzhen Technology University(深圳技术大学)
;
Université de Toulouse(图卢兹大学)
专题命中
其他VLM
:multimodal large language model(abstract);分类 cs.AI
Comments52 pages, 56 figures; appearing at EMNLP'25
From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation
Fan Yang, Zhiyang Chen, Yousong Zhu, Xin Li, Jinqiao Wang
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(基础模型研究中心、自动化研究所、中国科学院)
;
Peng Cheng Laboratory, Shenzhen, China(鹏城实验室、深圳中国)
;
School of Artificial Intelligence, University of Chinese Academy of Science, Beijing, China(人工智能学院、中国科学院大学、北京中国)
;
Wuhan AI Research, Wuhan, China(武汉人工智能研究、武汉中国)
;
MAPLE Lab, Westlake University(MAPLE实验室、西湖大学)
专题命中
其他VLM
:vision language model(abstract);分类 cs.CV
机构
*
The College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(南京航空航天大学人工智能学院)
;
The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)
Resolving Ambiguity in Gaze-Facilitated Visual Assistant Interaction Paradigm
Zeyu Wang, Baiyu Chen, Kun Yan, Hongjing Piao, Hao Xue, Flora D. Salim, Yuanchun Shi, Yuntao Wang
机构
*
Key Laboratory of Pervasive Computing, Tsinghua University(清华大学普适计算重点实验室)
;
The University of New South Wales(新南威尔士大学)
;
SKLSDE Lab, Beihang University(北航SKLSDE实验室)
机构
*
Department of Pathology and Laboratory Medicine, University of Pennsylvania(病理学与实验室医学系,宾夕法尼亚大学)
;
Department of Electrical and System Engineering, University of Pennsylvania(电气与系统工程系,宾夕法尼亚大学)
;
The Wharton School, University of Pennsylvania(沃顿商学院,宾夕法尼亚大学)
;
Department of Bioengineering, University of Pennsylvania(生物工程系,宾夕法尼亚大学)
;
Department of Computer and Information Science, University of Pennsylvania(计算机与信息科学系,宾夕法尼亚大学)
;
Department of Biostatistics, Epidemiology & Informatics, University of Pennsylvania(生物统计学、流行病学与信息学系,宾夕法尼亚大学)
Compositional Concept Generalization with Variational Quantum Circuits
Hala Hawashin, Mina Abbaszadeh, Nicholas Joseph, Beth Pearson, Martha Lewis, Mehrnoosh sadrzadeh
机构
*
School of Computer Science
;
Engineering University of New South Wales Sydney, Australia
;
Stanford University California, USA
;
Computer Science University College London London, UK
;
School of Eng. Maths. \& Tech University of Bristol Bristol, UK
;
Inst. Logic Language \& Computation University of Amsterdam Amsterdam, NL
CommentsAccepted to: 2025 IEEE International Conference on Quantum Artificial Intelligence (QAI), Naples, Italy, Nov 2-5, 2025. This is the authors' accepted manuscript (AAM). An IEEE copyright notice appears on page 1. The final published version will appear in IEEE Xplore; DOI to be added when available
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
Muhammad Huzaifah, Geyu Lin, Tianchi Liu, Hardik B. Sailor, Kye Min Tan, Tarun K. Vangani, Qiongqiong Wang, Jeremy H. M. Wong, Jinyang Wu, Nancy F. Chen, Ai Ti Aw
机构
*
MERaLiON Team(MERaLiON团队)
;
Institute for Infocomm Research (I 2 R), A*STAR, Singapore(信息通信研究所(I2R),A*STAR,新加坡)
专题命中
其他VLM
:multimodal large language model(abstract);分类 cs.AI
机构
*
The Hong Kong University of Science and Technology(Guangzhou)(香港科技大学(广州))
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
专题命中
其他VLM
:multimodal large language model(abstract);分类 cs.LG