Improving Adversarial Transferability on Vision-Language Pre-training Models via Surrogate-Specific Bias Correction
通过代理特定偏差校正提高视觉-语言预训练模型上的对抗迁移性
Lijia Yu, Jiuxin Cao, Yuchen Qiang, Changhao Chen, Yifei Huang, Bo Liu
机构
*
School of Cyber Science and Engineering, Southeast University(东南大学网络空间安全学院)
;
Purple Mountain Laboratories(紫金山实验室)
;
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs
从感知到决策:多模态大语言模型中听觉与视觉感知的信息流
Wish Suharitdamrong, Muhammad Awais, Xiatian Zhu, Sara Atito
机构
*
Surrey Institute for People-Centred AI (PAI)(萨里人本人工智能研究所)
;
University of Surrey(萨里大学)
;
Centre for Vision, Speech and Signal Processing (CVSSP)(视觉、语音和信号处理中心)
MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions
MeMo: 视觉受损条件下的实时视听目标说话人提取的注意力动量
Junjie Li, Wenxuan Wu, Shuai Wang, Zexu Pan, Kong Aik Lee, Helen Meng, Haizhou Li
机构
*
Department of Electrical and Electronic Engineering, Faculty of Engineering, The Hong Kong Polytechnic University(电子工程系,工程学院,香港理工大学)
;
Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong(系统工程与工程管理系,香港中文大学)
;
School of Artificial Intelligence (SAI), The Chinese University of Hong Kong, Shenzhen(人工智能学院(SAI),香港中文大学深圳校区)
;
School of Intelligence Science and Technology, Nanjing University(智能科学与技术学院,南京大学)
;
Tongyi Lab, Alibaba Group, Singapore(通义实验室,阿里巴巴集团,新加坡)
DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality Alignment
DeRA-MOS:通过解耦列表排序和模态对齐优化文本到音乐评估
Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen
机构
*
E.SUN Financial Holding Co., Ltd.(E.SUN财务控股公司)
;
United Link Co., Ltd.(联合链接有限公司)
;
Institute of Information Science, Academia Sinica(学术院信息科学研究所)
;
Department of Computer Science and Information Engineering, National Taiwan Normal University(台湾师范大学计算机科学与信息工程系)
Earth-OneVision: Extending Remote Sensing Multimodal Large Language Models to More Sensor Modalities and Tasks
Earth-OneVision:将遥感多模态大语言模型扩展到更多传感器模态和任务
Miaoxin Cai, Guanqun Wang, Wei Zhang, Guangyao Zhou, Yin Zhuang, Tong Zhang, Hao Wang, He Chen, Jun Li
机构
*
National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing (SBIIP), Beijing Institute of Technology(北京理工大学空间智能信息处理国家重点实验室)
;
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院空天信息创新研究院)
;
Key Laboratory of Technology in Geo-Spatial Information Processing and Application System, Chinese Academy of Sciences(中国科学院地理空间信息处理与应用系统技术重点实验室)
;
Advanced Research Institute of Multidisciplinary Sciences, Beijing Institute of Technology(北京理工大学前沿交叉科学研究院)
;
School of Mechatronical Engineering, Beijing Institute of Technology(北京理工大学机电学院)
;
School of Earth and Space Sciences, Peking University(北京大学地球与空间科学学院)
;
School of Electronics, Peking University(北京大学电子学院)
;
School of Computer Science and Hubei Key Laboratory of Intelligent Geo-Information Processing(华中科技大学计算机科学与技术学院&湖北省智能地理信息处理重点实验室)
GenEyePose: Patient-Free, Knowledge-Based Saccadic Eye Movement Modeling for Digital Neurophysiologic Biomarker Development
GenEyePose:用于数字神经生理学生物标志物开发的无患者、基于知识的扫视眼动建模
Tianyu Lin, Jooyoung Ryu, Puvada Sreevarsha, Rahul Srinivasaragavan, Riya Satavlekar, Susan Kim, Nidhi Soley, Yujie Yan, Ishan Vatsaraj, Carl Harris, Aimon Rahman, Vishal Patel, Joseph Greenstein, Casey Taylor, Kemar E. Green
机构
*
Whiting School of Engineering, Johns Hopkins University(约翰霍普金斯大学惠廷工程学院)
;
Department of Neurology, Johns Hopkins Medicine(约翰霍普金斯医学院神经内科)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
MSAM:多语义自适应挖掘用于跨模态无人机视频-文本检索
Jinghao Huang, Yaxiong Chen, Ganchao Liu
机构
*
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院)
;
School of Artificial Intelligence, Optics and Electronics (iOPEN), Northwestern Polytechnical University(西北工业大学人工智能、光学与电子学院(iOPEN))
机构
*
AnnLab(安实验室)
;
Institute of Semiconductors, Chinese Academy of Sciences(中国科学院半导体研究所)
;
Zhongguancun Academy(中关村学院)
;
State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室)
;
State Key Laboratory of High Performance Ceramics(高性能陶瓷国家重点实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
School of Electronic, Electrical and Communication Engineering(电子电气与通信工程学院)
;
University of ChineseAcademy of Sciences(中国科学院大学)
Soul Computing: A Theoretical Framework and Technical Architecture for Intelligent Agents with Independent Consciousness
灵魂计算:具有独立意识的智能体的理论框架与技术架构
Jinshan Zhang, Xishi Zhou, Qiu Peng, Jianwei Yin
机构
*
Innovation and Management Center, School of Software Technology, Zhejiang University (Ningbo)(浙江大学(宁波)软件学院创新与管理中心)
;
School of Software Technology, Zhejiang University, Ningbo(浙江大学软件学院(宁波))
CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution
CharTide: 数据中心的图表到代码生成通过三视角微调和查询驱动进化
Xiangxi Zheng, Kuang He, Jiayi Hu, Ping Yu, Rui Yan, Yuan Yao, Peng Hou, Anxiang Zeng, Alex Jinpeng Wang
机构
*
Nanjing University(南京大学)
;
LLM Team, Shopee Pte. Ltd.(Shopee 联邦学习团队)
;
East China Normal University(华东师范大学)
;
Nanjing University of Science and Technology(南京理工大学)
;
Central South University(中南大学)