机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Tencent(腾讯)
;
Tsinghua University(清华大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Beijing Film Academy(北京电影学院)
;
Stanford University(斯坦福大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Singapore Institute of Technology(新加坡理工学院)
Scaling-Aware Adapter for Structure-Grounded LLM Reasoning
Scaling-Aware Adapter for Structure-Grounded LLM Reasoning
Zihao Jing, Qiuhao Zeng, Ruiyi Fang, Yan Yi Li, Yan Sun, Boyu Wang, Pingzhao Hu
机构
*
Department of Computer Science, Western University, London, Canada(加拿大伦敦西方大学计算机科学系)
;
Department of Biochemistry, Western University, London, Canada(加拿大伦敦西方大学生物化学系)
机构
*
Harbin Institute of Technology, Shenzhen, China.(哈尔滨工业大学(深圳))
;
Peng Cheng Laboratory, China.(鹏城实验室)
;
Huazhong University of Science and Technology, China(华中科技大学)
专题命中
VLM训练与架构
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.AI、cs.LG
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
探究跨模态技能注入:场景、方法与超参数
Zhiyu Xu, Lean Wang, Yuanxin Liu, Lei Li, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun
机构
*
State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学)
;
WeChat AI, Tencent Inc., China(腾讯公司,中国)
;
The University of Hong Kong(香港大学)
VLCE: A Knowledge-Enhanced Framework for Image Description in Disaster Assessment
VLCE:一种用于灾害评估图像描述的知识增强框架
Md. Mahfuzur Rahman, Kishor Datta Gupta, Marufa Kamal, Fahad Rahman, Sunzida Siddique, Ahmed Rafi Hasan, Mohd Ariful Haque, Roy George
机构
*
Clark Atlanta University(克拉克亚特兰大大学)
;
BRAC University(布拉克大学)
;
United International University(国际联合大学)
;
Daffodil International University(花王国际大学)
Comments10 pages, 3 figures. Expanded version of an extended abstract accepted at NeurIPS 2025 Workshop on VLM4RWD. Presents methodology and preliminary experimental results
Surveillance Video-Based Traffic Accident Detection Using Transformer Architecture
基于监控视频的交通事故检测使用变换器架构
Tanu Singh, Pranamesh Chakraborty, Long T. Truong
机构
*
Department of Civil Engineering, Indian Institute of Technology Kanpur(印度理工学院坎浦尔分校土木工程系)
;
School of Computing, Engineering and Mathematical Sciences, La Trobe University(拉特罗布大学计算科学与工程数学科学学院)
专题命中
VLM训练与架构
:vision language model(abstract);VLM(abstract);LLaVA(abstract);分类 cs.CV、cs.AI
机构
*
School of Engineering, Westlake University, Hangzhou, China(西湖大学工程学院)
;
School of Cyberspace Security, Nanjing University of Science and Technology, Nanjing, China(南京理工大学网络安全学院)
专题命中
VLM训练与架构
:LLaVA(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
机构
*
University of Maryland(马里兰大学)
;
Indian Institute of Technology Bombay(印度班加罗尔理工学院)
;
Princeton University(普林斯顿大学)
;
University of Colorado Boulder(科罗拉多大学博尔德分校)
;
Capital One(Capital One公司)
;
University of Central Florida(佛罗里达中央大学)
专题命中
VLM训练与架构
:LLaVA(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG
机构
*
National Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(国家多媒体信息处理重点实验室,计算机科学学院,北京大学)
;
ByteDance(字节跳动)
;
Intel Labs China(英特尔中国实验室)
;
CUHK MMLab(香港中文大学MMLab)