机构
*
School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院)
;
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室)
机构
*
Human-Robot Interfaces and Interaction Lab, Istituto Italiano di Tecnologia, Genova, Italy(人类-机器人接口与交互实验室,意大利技术研究院,热那亚,意大利)
;
Ph.D. program of national interest in Robotics and Intelligent Machines (DRIM) and Università di Genova, Genoa, Italy(机器人与智能机器国家利益博士项目(DRIM)和热那亚大学,热那亚,意大利)
;
Edwardson School of Industrial Engineering, Purdue University, West Lafayette, IN, USA(工业工程埃德华森学校,普渡大学,西拉法伊斯,美国)
3D Smoke Scene Reconstruction Guided by Vision Priors from Multimodal Large Language Models
由多模态大语言模型视觉先验引导的3D烟雾场景重建
Xinye Zheng, Fei Wang, Yiqi Nie, Kun Li, Junjie Chen, Jiaqi Zhao, Yanyan Wei, Zhiliang Wu
机构
*
Hefei University of Technology(合肥工业大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
;
Anhui University(安徽大学)
;
United Arab Emirates University(阿联酋大学)
;
Nanyang Technological University(南洋理工大学)
专题命中
幻觉与鲁棒性
:multimodal large language model(title);分类 cs.CV
机构
*
School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院)
;
School of Computer Science, University of Wisconsin–Madison(威斯康星大学麦迪逊分校计算机科学系)
;
School of Engineering Science, University of the Chinese Academy of Sciences(中国科学院工程科学学院)
;
NExT++ Research Centre, National University of Singapore(新加坡国立大学NExT++研究中心)
AI总结
本文提出部分重校正softmax损失,通过限制top K softmax输出提升预训练多模态模型的对抗鲁棒性,实验显示细调后模型对主流攻击更具鲁棒性。
CommentsThe study described in Section 4 was conducted without required institutional review board approval. The paper is withdrawn pending completion of the approval process
Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
面向视觉语言模型的层次化通用多模态攻击
Peng-Fei Zhang, Zi Huang
机构
*
School of Electrical Engineering and Computer Science, the University of Queensland(电气工程与计算机科学学院,昆士兰大学)
;
Department of Computer Science, City University of Hong Kong(计算机科学系,香港城市大学)
Robust Prompt Tuning for Vision-Language Models with Mild Semantic Noise
Yansheng Gao, Yufei Zheng, Shengsheng Wang
机构
*
College of Computer Science and Technology, Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(计算机科学与技术学院、教育部符号计算与知识工程重点实验室、吉林大学)
;
College of Software, Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University(软件学院、教育部符号计算与知识工程重点实验室、吉林大学)
机构
*
National Technical University of Athens(希腊雅典国家技术大学)
;
Université Grenoble Alpes(格勒诺布尔阿尔卑斯大学)
;
CSX-AI(CSX-AI公司)
;
Carl von Ossietzky University of Oldenburg(奥尔登堡卡尔·冯·奥西特齐大学)
;
Chalmers University of Technology(查尔姆斯理工大学)
Distribution-Based Masked Medical Vision-Language Model Using Structured Reports
Shreyank N Gowda, Ruichi Zhang, Xiao Gu, Ying Weng, Lu Yang
机构
*
School of Computer Science, University of Nottingham(计算机科学学院,诺丁汉大学)
;
Department of Computer Science and Technology, School of Informatics, Xiamen University(计算机科学与技术系,信息学院,厦门大学)
;
CHI Lab, University of Oxford(CHI实验室,牛津大学)
;
School of Computer Science, University of Nottingham Ningbo China(计算机科学学院,宁波大学中国)
ELEMENTAL: Interactive Learning from Demonstrations and Vision-Language Models for Reward Design in Robotics
Letian Chen, Nina Moorman, Matthew Gombolay
机构
*
School of Interactive Computing, Georgia Institute of Technology, Atlanta, United States(交互计算学院,佐治亚理工学院,美国亚特兰大)
;
Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,Location Country)
;
School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,Location Country)