Yuezhe Yang, Hao Wang, Yige Peng, Jinman Kim, Lei Bi
机构
*
Institute of Translational Medicine, Shanghai Jiao Tong University(翻译医学研究院,上海交通大学)
;
School of Computer Science, University of Sydney(计算机科学学院,悉尼大学)
TGC-Net: A Structure-Aware and Semantically-Aligned Framework for Text-Guided Medical Image Segmentation
TGC-Net:一种结构感知且语义对齐的文本引导医学图像分割框架
Gaoren Lin, Huangxuan Zhao, Yuan Xiong, Lefei Zhang, Bo Du, Wentao Zhu
机构
*
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Department of Orthopedics, Tongji Hospital, Tongji Medical College, Huazhong University of Science(同济大学同济医学学院骨科部,华中科技大学)
Multi-Aspect Knowledge-Enhanced Medical Vision-Language Pretraining with Multi-Agent Data Generation
多方面知识增强的医学视觉-语言预训练与多代理数据生成
Xieji Li, Siyuan Yan, Yingsheng Liu, H. Peter Soyer, Monika Janda, Victoria Mar, Zongyuan Ge
机构
*
Department of Data Science and AI, Faculty of Information Technology, Monash University(数据科学与人工智能系,信息科技学院,墨尔本大学)
;
Victorian Melanoma Service, Alfred Health(维多利亚黑色素瘤服务,阿尔弗雷德健康)
;
Frazer Institute, The University of Queensland, Dermatology Research Centre(弗雷泽研究所,昆士兰大学,皮肤科研究中心)
LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
LLaVA-UHD v3:渐进式视觉压缩用于多模态大语言模型中的高效原分辨率编码
Shichu Sun, Yichen Zhang, Haolin Song, Zonghao Guo, Chi Chen, Yidan Zhang, Yuan Yao, Zhiyuan Liu, Maosong Sun
机构
*
Tsinghua University(清华大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航天信息研究所)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院)
机构
*
School of Cyber Science and Technology, University of Science and Technology of China(中国科学技术大学网络安全学院)
;
Anhui Province Key Laboratory of Digital Security(安徽省数字安全重点实验室)
;
The University of Hong Kong(香港大学)
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
Xiangyang Wu, Liu Liu, Baosheng Yu, Jiayan Qiu, Zhenwei Shi
机构
*
Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院)
;
School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)
;
Nanyang Technological University(南洋理工大学)
;
University of Leicester(莱斯特大学)
;
School of Astronautics, Beihang University(北京航空航天大学航天学院)
专题命中
图文多模态
:multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.AI
机构
*
The Hong Kong Polytechnic University(香港理工大学)
;
Tsinghua University(清华大学)
;
InspireOmni AI
;
Alibaba Group(阿里巴巴集团)
;
Case Western Reserve University(凯斯西储大学)
机构
*
School of Engineering, Westlake University, Hangzhou, China(西湖大学工程学院)
;
School of Cyberspace Security, Nanjing University of Science and Technology, Nanjing, China(南京理工大学网络安全学院)
Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs
Yudong Zhang, Ruobing Xie, Yiqing Huang, Jiansheng Chen, Xingwu Sun, Zhanhui Kang, Di Wang, Yu Wang
机构
*
Tsinghua University, Tencent(清华大学,腾讯)
;
Tencent(腾讯)
;
University of Science and Technology Beijing(北京科技大学)
;
Tencent, University of Macau(腾讯,澳门大学)
;
Tsinghua University(清华大学)
SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models
Zheng Liu, Hao Liang, Bozhou Li, Wentao Xiong, Chong Chen, Conghui He, Wentao Zhang, Bin Cui
机构
*
Peking University Beijing China
;
Huawei Technologies Ltd. Beijing China
;
Shanghai AI Laboratory Shanghai China
;
Peking University Center for Machine Learning Research Beijing China
;
Peking University School of Computer Science \& Key Lab of High Confidence Software Technologies (MOE) Beijing China
;
Peking University
;
Huawei Technologies Ltd.
;
Shanghai AI Laboratory