机构
*
Institute of Image Communication and Network Engineering, Shanghai Jiao Tong University(上海交通大学图像通信与网络工程研究所)
;
University of Electronic and Science Technology of China(电子科技大学)
;
Tianjin University(天津大学)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Latent Reconstruction from Generated Data for Multimodal Misinformation Detection
从生成数据中进行潜在重建用于多模态虚假信息检测
Stefanos-Iordanis Papadopoulos, Christos Koutlis, Symeon Papadopoulos, Panagiotis C. Petrantonakis
机构
*
Information Technology Institute, Centre for Research & Technology, Hellas(信息科技研究所,研究中心,希腊)
;
Department of Electrical & Computer Engineering, Aristotle University of Thessaloniki(电气与计算机工程系,亚里士多德大学)
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Department of Pulmonary and Critical Care Medicine, The First Affiliated Hospital of Sun Yat-sen University(中山大学附属第一医院呼吸与危重症医学科)
;
Centre of AI and Robotics, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences(香港科学院人工智能与机器人中心)
;
School of Biomedical Engineering and Imaging Sciences, King’s College London(伦敦国王学院生物医学工程与影像科学学院)
机构
*
Indian Institute of Technology, Ropar, India(印度理工学院罗帕尔分校)
;
Machine Intelligence Group, Birla Institute of Technology and Science, Pilani, Hyderabad Campus, India(比拉理工科学院帕利尼 Hyderabad 分校机器智能小组)
;
Monash University, Melbourne, Australia(墨尔本大学)
专题命中
视觉定位与Grounding
:vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
VCIP, CS, Nankai University(南开大学计算机科学与技术学院)
;
NKIARI, Shenzhen Futian(深圳福田国家信息研究院)
;
OpenGVLab, Shanghai AI Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong(香港中文大学)
机构
*
College of Computer Science, Sichuan University(四川大学计算机学院)
;
Columbia University(哥伦比亚大学)
;
School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)
;
Apon AI and Brain-Computer Engineering Research Institute(Apon人工智能与脑机工程研究院)
;
Faculty of Science and Technology, University of Macau(澳门大学科学与技术学院)
;
Faculty of Applied Science and Engineering, University of Toronto(多伦多大学应用科学与工程学院)
;
Faculty of Computer Science and Information Technology, University of Malaya(马来亚大学计算机科学与信息技术学院)
;
Zhaolong Technology(智龙科技)
;
Purdue University(普渡大学)
HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
HalluShift++: 通过内部表示转移弥合语言与视觉,解决多模态大语言模型中的层级幻觉
Sujoy Nath, Arkaprabha Basu, Sharanya Dasgupta, Swagatam Das
机构
*
Netaji Subhash Engineering College (NSEC)(奈尔贾伊·萨布哈工程学院)
;
TCG Crest
;
Electronics and Communication Sciences Unit (ECSU)(电子与通信科学单位)
;
Indian Statistical Institute(印度统计研究所)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV