机构
*
Harbin Institute of Technology(哈尔滨工业大学)
;
Pengcheng Laboratory(鹏城实验室)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Sichuan University(四川大学)
;
Zhejiang Normal University(浙江师范大学)
专题命中
VLM训练与架构
:vision language model(title);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
From Pixels to Explanations: Interpretable Diabetic Retinopathy Grading with CNN-Transformer Ensembles, Visual Explainability and Vision-Language Models
机构
*
Department of Informatics, University of Salerno(萨勒诺大学信息学院)
;
Department of Information Security and Communication Technology (IIK), Norwegian University of Science and Technology(挪威科技大学信息安全部)
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
MLLMEraser: 通过激活引导在多模态大语言模型中实现测试时遗忘
Chenlu Ding, Jiancan Wu, Leheng Sheng, Fan Zhang, Yancheng Yuan, Xiang Wang, Xiangnan He
机构
*
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Huawei(华为)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);LLaVA(abstract);MLLM(abstract);分类 cs.AI、cs.LG
Speculative Decoding Reimagined for Multimodal Large Language Models
Luxi Lin, Zhihang Lin, Zhanpeng Zeng, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(多媒体可信感知与高效计算重点实验室、中华人民共和国教育部、厦门大学)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);LLaVA(abstract);MLLM(abstract);分类 cs.CV、cs.AI
Towards Interpreting Visual Information Processing in Vision-Language Models
Clement Neo, Luke Ong, Philip Torr, Mor Geva, David Krueger, Fazl Barez
机构
*
Nanyang Technological University(南洋理工大学)
;
University of Oxford(牛津大学)
;
Tel Aviv University(特拉维夫大学)
;
MILA(蒙特利尔人工智能研究院)
;
ERA-Krueger AI Safety Lab(ERA-Krueger人工智能安全实验室)