Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation
超越补丁:面向多模态少样本字体生成的全局感知自回归模型
Haonan Cai, Yuxuan Luo, Zhouhui Lian
机构
*
Wangxuan Institute of Computer Technology, Peking University(北京大学王学轩计算机技术研究所)
;
School of Electronics Engineering and Computer Science, Peking University(北京大学电子工程与计算机科学学院)
机构
*
Laboratory of Intelligent Collaborative Computing, University of Electronic Science(智能协同计算实验室,电子科学科技大学)
;
School of Computer Science(计算机科学学院)
;
Technology (School of Artificial Intelligence), Yibin University(技术(人工智能学院),宜宾大学)
;
College of Humanities(人文学院)
;
General Education, Chengdu Textile College(通识教育,成都纺织学院)
Multimodal synthesis of MRI and tabular data with diffusion in a joint latent space via cross-attention
通过交叉注意力在联合潜在空间中利用扩散模型进行MRI和表格数据的多模态合成
Daniel Mensing, Jan Kapar, Jochen G. Hirsch, Matthias Günther, Horst Hahn, Marvin N. Wright
机构
*
Fraunhofer Institute for Digital Medicine MEVIS(弗劳恩霍夫数字医学研究所MEVIS)
;
Leibniz Institute for Prevention Research and Epidemiology – BIPS(莱比锡预防研究与流行病学研究所 – BIPS)
;
Faculty of Mathematics and Computer Science, University of Bremen(不莱梅大学数学与计算机科学学院)
;
Faculty of Physics and Electrical Engineering, University of Bremen(不莱梅大学物理与电气工程学院)
Uncovering Entity Identity Confusion in Multimodal Knowledge Editing
揭示多模态知识编辑中的实体身份混淆
Shu Wu, Xiaotian Ye, Xinyu Mou, Dongsheng Liu, Xiaohan Wang, Mengqi Zhang
机构
*
New Laboratory of Pattern Recognition (NLPR)(模式识别新实验室)
;
State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS)(多模态人工智能系统国家重点实验室)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院)
;
Huazhong University of Science and Technology(华中科技大学)
;
Shandong University(山东大学)
Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding
统一多模态模型的免费午餐:通过内在理解的反思校正增强生成
Yibo Jiang, Tao Wu, Rui Jiang, Yehao Lu, Chaoxiang Cai, Zequn Qin, Xi Li
机构
*
School of Software Technology, Zhejiang University(浙江大学软件技术学院)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
College of Computer Science(计算机科学学院)
机构
*
Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身智能研究所)
;
Shanghai Innovation Institute(上海创新研究院)
;
Shanghai Key Laboratory of Multimodal Embodied AI(上海市多模态具身人工智能重点实验室)
;
College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)
;
The Chinese University of Hong Kong(香港中文大学)
;
Central South University(中南大学)
;
Fudan University(复旦大学)
Understanding vs. Generation: Navigating Optimization Dilemma in Multimodal Models
理解与生成:多模态模型中的优化困境导航
Sen Ye, Mengde Xu, Shuyang Gu, Di He, Liwei Wang, Han Hu
机构
*
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
;
Tencent(腾讯)
;
Center for Data Science, Peking University(北京大学数据科学中心)
;
Center for Machine Learning Research, Peking University(北京大学机器学习研究中心)
Reflect to Inform: Boosting Multimodal Reasoning via Information-Gain-Driven Verification
反思以获取信息:通过信息增益驱动的验证提升多模态推理
Shuai Lv, Chang Liu, Feng Tang, Yujie Yuan, Aojun Zhou, Kui Zhang, Xi Yang, Yangqiu Song
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Huawei Foundation Model Department(华为基础模型部)
;
The Chinese University of Hong Kong(香港中文大学)
;
Hong Kong University of Science and Technology(香港科技大学)
机构
*
College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)