Vision-Language Based Expert Reporting for Painting Authentication and Defect Detection
基于视觉-语言的专家报告用于绘画认证和缺陷检测
Eman Ouda, Mohammed Salah, Arsenii O. Chulkov, Gianfranco Gargiulo, Gian Luca Tartaglia, Stefano Sfarra, Yusra Abdulrahman
机构
*
organization= Khalifa University of Science
;
Technology, Department of Aerospace Engineering , city= Abu Dhabi , country= United Arab Emirates
;
organization= National Research Tomsk Polytechnic University , country= Russia
;
organization= Academy of Fine Arts of Naples , city= Naples , country= Italy
;
organization= Academy of Fine Arts of Sassari , city= Sassari , country= Italy
;
organization= Department of Industrial
;
Economics (DIIIE), University of L’Aquila , city= L'Aquila I-67100 , country= Italy
;
organization= Advanced Research
;
Innovation Center (ARIC), Khalifa University of Science \& Technology , city= Abu Dhabi , country= UAE
MGCR-Net:Multimodal Graph-Conditioned Vision-Language Reconstruction Network for Remote Sensing Change Detection
MGCR-Net:多模态图条件视觉-语言重建网络用于遥感变化检测
Chengming Wang, Guodong Fan, Jinjiang Li, Min Gan, C. L. Philip Chen
机构
*
School of Computer Science and Technology, Shandong Technology and Business University(山东科技职业大学计算机科学与技术学院)
;
School of Computer Science and Technology, Qingdao University(青岛大学计算机科学与技术学院)
;
School of Computer Science and Engineering, South China University of Technology(华南理工大学计算机科学与工程学院)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
AI总结
MGCR-Net通过多模态图条件视觉-语言重建机制提升遥感变化检测的语义交互能力。
Journal refIEEE Transactions on Geoscience and Remote Sensing, vol. 64, pp. 1-15, 2026, Art no. 4701515
机构
*
Shanghai AI Lab(上海人工智能实验室)
;
Northwestern Polytechnical University(西北工业大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Peking University(北京大学)
;
Nanyang Technological University(南洋理工大学)
;
Beihang University(北京航空航天大学)
;
Sichuan University(四川大学)
;
Tsinghua University(清华大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Fudan University(复旦大学)
;
Hong Kong University of Science and Technology(香港科技大学)
机构
*
Hong Kong Polytechnic University, HK SAR(香港理工大学)
;
Fondazione Bruno Kessler, Italy(布鲁诺·凯斯勒基金会)
;
University of Technology Sydney, Australia(悉尼科技大学)
专题命中
视觉定位与Grounding
:vision language model(abstract);multimodal large language model(abstract);分类 cs.CV
GroundedSurg: A Multi-Procedure Benchmark for Language-Conditioned Surgical Tool Segmentation
GroundedSurg: 一种多手术流程的语言条件手术工具分割基准
Tajamul Ashraf, Abrar Ul Riyaz, Wasif Tak, Tavaheed Tariq, Sonia Yadav, Moloud Abdar, Janibul Bashir
机构
*
King Abdullah University of Science and Technology (KAUST)(卡奥尔大学科学与技术学院)
;
Thapar Institute of Engineering and Technology(塔帕尔工程与技术学院)
;
The University of Queensland(昆士兰大学)
;
Gaash Research Lab, National Institute of Technology Srinagar(加什研究实验室,锡纳加尔国家理工学院)
Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
为MLLMs定制视觉情绪评估:一种开放词汇、多维且可扩展的方法
Daiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma, Yu Zhou
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
VCIP & TMCC & DISSec, College of Computer Science, Nankai University(南开大学计算机学院)
;
Department of Psychological and Cognitive Sciences, Tsinghua University(清华大学心理学与认知科学系)
;
University of Chinese Academy of Sciences(中国科学院大学)
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Hidden in the Metadata: Stealth Poisoning Attacks on Multimodal Retrieval-Augmented Generation
元数据中的隐藏攻击:多模态检索增强生成中的隐秘污染攻击
Kennedy Edemacu, Mohammad Mahdi Shokri
机构
*
The City University of New York, CSI, Staten Island, NY 10314, USA(纽约城市大学,CSI,史泰登岛分校)
;
The City University of New York, Graduate Center, New York, NY 10016, USA(纽约城市大学研究生中心)
专题命中
视觉定位与Grounding
:grounding(abstract);multimodal large language model(abstract);分类 cs.AI
Leveraging Multimodal LLM Descriptions of Activity for Explainable Semi-Supervised Video Anomaly Detection
利用多模态大语言模型对活动的描述进行可解释的半监督视频异常检测
Furkan Mumcu, Michael J. Jones, Anoop Cherian, Yasin Yilmaz
机构
*
Department of Electrical Engineering University of South Florida(电气工程系 佛罗里达州立大学)
;
Mitsubishi Electric Research Laboratories (MERL)(三菱电机研究实验室(MERL))
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
CommentsThis is the extended version of the paper accepted in ICASSP'26, which will be publicly available in May. Authors' contributions may vary among the versions
Minfeng Zhu, Zi Wang, Sizhe Ji, Zhengtong Du, Shengqiang Tai, Junming Ke, Xiao Deng, Zanlang Yin, Xiuqi Huang, Heyu Wang, Wei Chen
机构
*
State Key Lab of CAD&CG, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室)
;
Polytechnic Institute, Zhejiang University(浙江大学多科大学院)
;
Hangzhou Research Institute of AI and Holographic Technology(杭州人工智能与全息技术研究 institutes)
;
Volkswagen Group Innovation(大众集团创新)
;
School of Mathematical Science, Zhejiang University(浙江大学数学科学学院)
专题命中
视觉定位与Grounding
:visual language model(abstract);grounding(abstract);分类 cs.AI