Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-Tuning
通过字形驱动微调增强多模态大语言模型用于古汉字演变分析
Rui Song, Lida Shi, Ruihua Qi, Yingji Li, Hao Xu
机构
*
College of Computer Science and Technology, Jilin University, China(吉林大学计算机科学与技术学院)
;
Key Laboratory of Ancient Chinese Script, Culture Relics and Artificial Intelligence, Jilin University, China(吉林大学古文字、文物与人工智能重点实验室)
;
School of Artificial Intelligence, Jilin University, China(吉林大学人工智能学院)
;
School of Archaeology, Jilin University, China(吉林大学考古学院)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);分类 cs.AI
RCP: Representation Consistency Pruner for Mitigating Distribution Shift in Large Vision-Language Models
RCP:一种用于缓解大视觉-语言模型中分布偏移的表示一致性剪枝方法
Jianwei Zhang, Chaoning Zhang, Sihan Cao, Wang Liu, Pengcheng Zheng, Jiaxin Huang, Caiyan Qin, Yalan Ye, Wei Dong, Yang Yang
机构
*
School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)
;
Machine Learning Department, Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学机器学习系)
;
School of Robotics and Advanced Manufacture, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)机器人与先进制造学院)
机构
*
State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室)
;
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
长尾分布感知的混合专家路由器:用于大规模视觉-语言模型的混合专家架构
Chaoxiang Cai, Longrong Yang, Minghe Weng, Xuewei Li, Zequn Qin, Xi Li
机构
*
School of Software Technology, Zhejiang University(浙江大学软件学院)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
School of Electronic and Information, Shanghai Dianji University(上海电机学院电子信息学院)
PopResume: Causal Fairness Evaluation of LLM/VLM Resume Screeners with Population-Representative Dataset
PopResume:基于人口代表性数据集的LLM/VLM简历筛选器因果公平性评估
Sumin Yu, Juhyeon Park, Taesup Moon
机构
*
ECE, Seoul National University(首尔国立大学电子与计算机工程系)
;
IPAI, Seoul National University(首尔国立大学人工智能研究所)
;
ASRI / INMC / AIIS, Seoul National University(首尔国立大学人工智能研究所)
Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
基于强化学习的语言引导标记压缩在大视觉-语言模型中
Sihan Cao, Jianwei Zhang, Pengcheng Zheng, Jiaxin Yan, Caiyan Qin, Yalan Ye, Wei Dong, Peng Wang, Yang Yang, Chaoning Zhang
机构
*
School of Computer Science and Engineering(计算机科学与工程学院)
;
University of Electronic Science and Technology of China(电子科学与技术大学)
;
School of Robotics and Advanced Manufacture(机器人与先进制造学院)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
College of Information and Control Engineering(信息与控制工程学院)
机构
*
Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong, China(计算机科学与工程系,香港中文大学,香港,中国)
;
Institute of Medical Intelligence and XR, The Chinese University of Hong Kong, Hong Kong, China(医学智能与XR研究所,香港中文大学,香港,中国)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);分类 cs.CV
SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
SPEX:一种用于光谱遥感图像土地覆盖提取的视觉-语言模型
Dongchen Si, Di Wang, Erzhong Gao, Xiaolei Qin, Liu Zhao, Jing Zhang, Minqiang Xu, Jianbo Zhan, Jianshe Wang, Lin Liu, Bo Du, Liangpei Zhang
机构
*
College of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)
;
iFlytek Co., Ltd.(iFlytek公司)
;
National Engineering Research Center of Speech and Language Information Processing(语音与语言信息处理国家工程研究中心)
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Zhongguancun Academy(中关村学院)
;
National Engineering Research Center for Multimedia Software(多媒体软件国家工程研究中心)
;
Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University(湖北省多媒体与网络通信工程重点实验室,武汉大学)
;
State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University(测绘遥感信息工程国家重点实验室,武汉大学)