A Provable Energy-Guided Test-Time Defense Boosting Adversarial Robustness of Large Vision-Language Models
一种可证明的能量引导测试时间防御提升大视觉-语言模型的对抗鲁棒性
Mujtaba Hussain Mirza, Antonio D'Orazio, Odelia Melamed, Iacopo Masi
机构
*
OmnAI Lab, Computer Science Department, Sapienza University of Rome, Italy(罗马大学计算机科学系OmnAI实验室,意大利)
;
Weizmann Institute of Science, Israel(魏茨曼科学研究所,以色列)
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
通过视觉记忆机制扩展多模态大语言模型的长视频理解
Tao Chen, Kun Zhang, Qiong Wu, Xiao Chen, Chao Chang, Xiaoshuai Sun, Yiyi Zhou, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
;
National University of Defense Technology(国防科技大学)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
Know-Show:在时空 grounded 推理上评估视频语言模型的基准测试
Chinthani Sugandhika, Chen Li, Deepu Rajan, Basura Fernando
机构
*
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
Institute of High-Performance Computing, Agency for Science, Technology and Research(新加坡科技研究局高性能计算研究所)
;
Centre for Frontier AI Research, Agency for Science, Technology and Research(新加坡科技研究局前沿人工智能研究中心)
机构
*
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Southeast University(东南大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
National University of Singapore(新加坡国立大学)
;
Wuhan AI Research(武汉人工智能研究院)
;
Guangdong Provincial Key Laboratory of Intellectual Property and Big Data, Guangdong Polytechnic Normal University(广东技术师范大学广东省知识产权大数据重点实验室)