CapCLIP: A Vision-Language Representation Alignment Approach for Wireless Capsule Endoscopy Analysis
CapCLIP:一种用于无线胶囊内镜分析的视觉-语言表示对齐方法
Haroon Wahab, Irfan Mehmood, Hassan Ugail
机构
*
School of Computer Science, AI and Electronics Faculty of Engineering and Digital Technologies(计算机科学与电子工程学院,工程与数字技术学院)
;
School of Management Faculty of Mgmt, Law & Social Sciences(管理学院,管理、法律与社会科学学院)
;
Centre for Visual Computing and Intelligent Systems(视觉计算与智能系统中心)
机构
*
State Key Laboratory of Physical Oceanography and the Faculty of Information Science and Engineering, Ocean University of China(物理海洋学国家重点实验室和中国海洋大学信息科学与工程学院)
;
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen university(中山大学深圳校区计算机科学与技术学院)
Precision Synthesis of Multi-Tracer PET via VLM-Modulated Rectified Flow for Stratifying Mild Cognitive Impairment
多示踪PET的高精度合成通过VLM调制的校正流用于区分轻度认知障碍
Tuo Liu, Shuijin Lin, Shaozhen Yan, Haifeng Wang, Jie Lu, Jianhua Ma, Chunfeng Lian
机构
*
School of Mathematics and Statistics, Xi'an Jiaotong University(西安交通大学数学与统计学学院)
;
Key Laboratory of Biomedical Information Engineering of Ministry of Education, School of Life Science and Technology, Xi'an Jiaotong University(教育部生物医学信息工程重点实验室,西安交通大学生命科学与技术学院)
;
Department of Radiology and Nuclear Medicine, Xuanwu Hospital, Capital Medical University(首都医科大学宣武医院放射科与核医学科)
;
Research Center for Intelligent Medical Equipment and Devices (IMED), Xi'an Jiaotong University(智能医疗设备与器件研究中心(IMED),西安交通大学)
Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility
风险感知注入:为安全校准视觉-语言模型而不牺牲实用性
Mengxuan Wang, Yuxin Chen, Gang Xu, Tao He, Hongjie Jiang, Ming Li
机构
*
Shien-Ming Wu School of Intelligent Engineering, South China University of Technology(华南理工大学吴贤铭智能工程学院)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳))
;
Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)
;
University of Electronic Science and Technology of China(电子科技大学)
Echoes of ownership: Adversarial-guided dual injection for copyright protection in MLLMs
所有权的回声:对抗引导的双注入用于MLLMs中的版权保护
Chengwei Xia, Fan Ma, Ruijie Quan, Yunqiu Xu, Kun Zhan, Yi Yang
机构
*
School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
Towards Highly Transferable Vision-Language Attack via Semantic-Augmented Dynamic Contrastive Interaction
通过语义增强的动态对比交互实现高度可迁移的视觉-语言攻击
Yuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou, Xiao-Jun Wu, Josef Kittler
机构
*
School of Artificial Intelligence and Computer Science, Jiangnan University(江南大学人工智能与计算机科学学院)
;
Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(Surrey 大学视觉、语音和信号处理中心)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
PixelVLA:推进视觉-语言-动作模型中的像素级理解
Wenqi Liang, Gan Sun, Yao He, Jiahua Dong, Suyan Dai, Ivan Laptev, Salman Khan, Yang Cong
机构
*
University of Trento(特伦托大学)
;
School of Automation Science and Engineering, South China University of Technology(华南理工大学自动化科学与工程学院)
;
Mohamed bin Zayed University of Artificial Intelligence(马尔代夫人工智能大学)
;
Australian National University(澳大利亚国立大学)
机构
*
State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology, Anhui University, Hefei, China(光电信息采集与防护技术国家重点实验室,安徽大学,合肥,中国)
;
School of Artificial Intelligence, Anhui University, Hefei, China(人工智能学院,安徽大学,合肥,中国)
;
College of Intelligence Science and Technology, National University of Defense Technology, Changsha, China(智能科学与技术学院,国防科技大学,长沙,中国)
机构
*
University of Amsterdam(阿姆斯特丹大学)
;
The Netherlands Cancer Institute(荷兰癌症研究所)
;
Vrije Universiteit Amsterdam(阿姆斯特丹自由大学)
;
The Arctic University of Norway(挪威北极大学)
;
Singapore Management University(新加坡管理大学)