GTA-Net: Cooperative Game Theory for Vision-Language Alignment in Chest X-Ray Report Generation
GTA-Net:合作博弈论在胸部X光报告生成中的视觉-语言对齐
Saif ur Rehman Khan, Imad Ahmed Waqar, Sebastian Vollmer, Andreas Dengel, Muhammad Nabeel Asim
机构
*
Department of Computer Science, Rhineland-Palatinate Technical University of Kaiserslautern-Landau(莱茵兰-普法尔茨凯泽斯劳滕-兰道工业大学计算机科学系)
;
German Research Center for Artificial Intelligence (DFKI)(德国人工智能研究中心)
;
IntelligentX GmbH
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey
视觉与视觉-语言应用中的模态感知特征匹配:全面综述
Weide Liu, Wei Zhou, Jun Liu, Ping Hu, Jun Cheng, Jungong Han, Weisi Lin
机构
*
School of Computing and Artificial Intelligence, Jiangxi University of Finance and Economics(江西财经大学计算机与人工智能学院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院)
;
School of Computer Science and Informatics, Cardiff University(卡迪夫大学计算机科学与信息学院)
;
School of Computing and Communications, Lancaster University(兰卡斯特大学计算机与通讯学院)
;
School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)
;
Institute for Infocomm Research, Agency for Science, Technology and Research (A*STAR)(新加坡资讯研究院,科技研究局(A*STAR))
;
Department of Automation, Tsinghua University(清华大学自动化系)
CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning
CapRL++:基于可验证奖励的统一强化学习用于密集图像和视频描述生成
Penghui Yang, Long Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Yibin Wang, Yujie Zhou, Jiazi Bu, Jianze Liang, Qidong Huang, Jiaqi Wang, Feng Wu, Dahua Lin
机构
*
Tsinghua University(清华大学)
;
University of Science and Technology of China(中国科学技术大学)
;
Microsoft(微软)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Shanghai Innovation Institute(上海创新研究院)
;
Alibaba Cloud(阿里云)
;
The Chinese University of Hong Kong(香港中文大学)
机构
*
Dwarkadas Jivanlal Sanghvi College of Engineering(达沃拉斯·吉万拉尔·桑格维工程学院)
;
King’s College London(伦敦国王学院)
;
Indian Institute of Technology Jodhpur(印度理工学院朱罗普尔)
TraversalBench: Challenging Paths to Follow for Vision Language Models
TraversalBench: 为视觉语言模型设计的复杂路径挑战测试集
Clara Petrova, Zhuo Chen, Marin Soljačić
机构
*
Massachusetts Institute of Technology, Department of Physics(麻省理工学院物理系)
;
Massachusetts Institute of Technology, Institute for Data, Systems, and Society(麻省理工学院数据、系统与社会研究所)
;
NSF AI Institute for Artificial Intelligence and Fundamental Interactions(国家科学基金会人工智能与基本相互作用AI研究所)
Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization
基于LLM增强优化的无人机低空经济网络高效机载视觉-语言推理
Yang Li, Ruichen Zhang, Yinqiu Liu, Guangyuan Liu, Abbas Jamalipour, Xianbin Wang, Dong In Kim
机构
*
College of Computing and Data Science, Nanyang Technological University, Singapore(计算与数据科学学院、新加坡国立科技大学)
;
The University of Sydney, Sydney, Australia(悉尼大学、澳大利亚悉尼)
;
Department of Electrical and Computer Engineering, Western University, London, Canada(电气与计算机工程系、西方大学、加拿大伦敦)
;
Department of Electrical and Computer Engineering, Sungkyunkwan University, South Korea(电气与计算机工程系、全州大学、韩国)
CL-CLIP: CLIP-Based Continual Learning Framework with Cost-Volume Category Decoupling for Object Detection
CL-CLIP: 基于CLIP的持续学习框架与代价体积类别解耦用于目标检测
Zihan Liu, Yuguang Yang, Shengjie Su, Jianing Pang, Linlin Yang, Chunyu Xie, Nikolai Yu. Zolotykh, Baochang Zhang
机构
*
National College for Excellent Engineers, Beihang University(卓越工程师学院,北京航空航天大学)
;
AI Research, Qihoo 360(360人工智能研究院,奇虎360)
;
School of Electronic Information Engineering, Beihang University(电子信息学院,北京航空航天大学)
;
School of Cyber Science and Technology, Beihang University(网络安全科学与技术学院,北京航空航天大学)
;
School of Computer Science and Engineering, Beihang University(计算机科学与工程学院,北京航空航天大学)
;
State Key Laboratory of Media Convergence and Communication, Communication University of China(媒体融合与传播国家重点实验室,中国传媒大学)
;
Institute of Information Technology, Mathematics and Mechanics, Lobachebsky University(信息技术、数学与力学学院,洛瓦茨基大学)
;
School of Artificial Intelligence, Beihang University(人工智能学院,北京航空航天大学)