A Provable Energy-Guided Test-Time Defense Boosting Adversarial Robustness of Large Vision-Language Models
一种可证明的能量引导测试时间防御提升大视觉-语言模型的对抗鲁棒性
Mujtaba Hussain Mirza, Antonio D'Orazio, Odelia Melamed, Iacopo Masi
机构
*
OmnAI Lab, Computer Science Department, Sapienza University of Rome, Italy(罗马大学计算机科学系OmnAI实验室,意大利)
;
Weizmann Institute of Science, Israel(魏茨曼科学研究所,以色列)
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Southeast University(东南大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
National University of Singapore(新加坡国立大学)
;
Wuhan AI Research(武汉人工智能研究院)
;
Guangdong Provincial Key Laboratory of Intellectual Property and Big Data, Guangdong Polytechnic Normal University(广东技术师范大学广东省知识产权大数据重点实验室)
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
通过视觉记忆机制扩展多模态大语言模型的长视频理解
Tao Chen, Kun Zhang, Qiong Wu, Xiao Chen, Chao Chang, Xiaoshuai Sun, Yiyi Zhou, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
;
National University of Defense Technology(国防科技大学)
Multi-Modal Representation Learning via Semi-Supervised Rate Reduction for Generalized Category Discovery
通过半监督率减少实现多模态表示学习的通用类别发现
Wei He, Xianghan Meng, Zhiyuan Huang, Xianbiao Qi, Rong Xiao, Chun-Guang Li
机构
*
School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院)
;
Intellifusion Inc., Shenzhen, P.R. China(深圳云天励飞技术股份有限公司)
MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
MVGGT:多模态视觉几何 grounded 变换器用于多视角3D指称表达分割
Changli Wu, Haodong Wang, Jiayi Ji, Yutian Yao, Chunsai Du, Jihua Kang, Yanwei Fu, Liujuan Cao
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
;
Shanghai Innovation Institute(上海创新研究院)
;
Fudan University(复旦大学)
;
ByteDance(字节跳动)
;
Tianjin University of Science and Technology(天津科技大学)
Guiding a Diffusion Transformer with the Internal Dynamics of Itself
通过自身内部动态引导扩散变换器
Xingyu Zhou, Qifan Li, Xiaobin Hu, Hai Chen, Shuhang Gu
机构
*
University of Electronic Science and Technology of China(电子科技大学)
;
National University of Singapore(新加坡国立大学)
;
Sun Yat-sen University(中山大学)
;
North China Institute of Computer Systems Engineering(华北计算机系统工程研究所)