AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
AR-Omni:一种统一的自回归模型用于任意到任意生成
Dongjie Cheng, Ruifeng Yuan, Yongqi Li, Runyang You, Wenjie Wang, Liqiang Nie, Lei Zhang, Wenjie Li
机构
*
The Hong Kong Polytechnic University(香港理工大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
Beyond Visual Safety: Jailbreaking Multimodal Large Language Models for Harmful Image Generation via Semantic-Agnostic Inputs
超越视觉安全:通过语义无关输入对多模态大语言模型进行有害图像生成的劫持
Mingyu Yu, Lana Liu, Zhehao Zhao, Wei Wang, Sujuan Qin
机构
*
State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学)
;
School of Cyberspace Security, Beijing University of Posts and Telecommunications(网络安全学院,北京邮电大学)
机构
*
School of Xingzhi College, South China Normal University(星智学院,华南师范大学)
;
College of Big Data and Internet, Shenzhen Technology University(大数据与互联网学院,深圳科技大学)
;
School of Data Science and Engineering, Xingzhi College, South China Normal University(数据科学与工程学院,星智学院,华南师范大学)
;
Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(电子与计算机工程系,香港科技大学)
;
College of Intelligence Science and Technology, National University of Defense Technology(智能科学与技术学院,国防科技大学)
CMIE: Combining MLLM Insights with External Evidence for Explainable Out-of-Context Misinformation Detection
Fanxiao Li, Jiaying Wu, Canyuan He, Wei Zhou
机构
*
School of Information Science and Engineering, Yunnan University(云南大学信息科学与工程学院)
;
National University of Singapore(新加坡国立大学)
;
Engineering Research Center of Cyberspace, Yunnan University(云南大学网络空间研究院)
Zekun Wang, King Zhu, Chunpu Xu, Wangchunshu Zhou, Jiaheng Liu, Yibo Zhang, Jiashuo Wang, Ning Shi, Siyu Li, Yizhi Li, Haoran Que, Zhaoxiang Zhang, Yuanxing Zhang, Ge Zhang, Ke Xu, Jie Fu, Wenhao Huang
机构
*
Beihang University(北航)
;
M-A-P
;
The Hong Kong Polytechnic University(香港理工大学)
;
AIWaves
;
University of Alberta(阿尔伯塔大学)
;
University of Waterloo(滑铁卢大学)
;
University of Manchester(曼彻斯特大学)
;
Chinese Academy of Sciences(中国科学院)
;
Peking University(北京大学)
;
Shanghai AI Lab(上海AI实验室)
;
Nanjing University(南京大学)
;
Kuaishou Technology(快手科技)
What is the Visual Cognition Gap between Humans and Multimodal LLMs?
Xu Cao, Yifan Shen, Bolin Lai, Wenqian Ye, Yunsheng Ma, Joerg Heintz, Jintai Chen, Meihuan Huang, Jianguo Cao, Aidong Zhang, James M. Rehg
机构
*
Department of Computer Science, University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校计算机科学系)
;
College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院)
;
Department of Computer Science, University of Virginia(弗吉尼亚大学计算机科学系)
;
Digital Twin Lab, Purdue University(普渡大学数字孪生实验室)
;
HKUST (Guangzhou)(香港科技大学(广州))
;
Department of Rehabilitation Medicine, Shenzhen Children’s Hospital(深圳儿童医院康复医学系)
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
Shaojun E, Yuchen Yang, Jiaheng Wu, Yan Zhang, Tiejun Zhao, Ziyan Chen
机构
*
Global Tone Communication Technology Co., Ltd.(全球 tone 通信技术有限公司)
;
Faculty of computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院)
;
School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)
机构
*
National Centre for Computer Animation, Bournemouth University(伯恩茅斯大学计算机动画国家中心)
;
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Hong Kong University of Science and Technology(香港科技大学)
;
Department of Computer Science and Information Engineering, National Cheng Kung University(国立成功大学计算机科学与信息工程系)
Generate, Not Recommend: Personalized Multimodal Content Generation
Jiongnan Liu, Zhicheng Dou, Ning Hu, Chenyan Xiong
机构
*
School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学系)
;
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学agate智能学院)
;
Serendipity One Inc.(Serendipity One公司)