机构
*
Nanyang Technological University(南洋理工大学)
;
Centre for Frontier AI Research, A*STAR(A*STAR前沿人工智能研究中心)
;
Allen Institute for AI(艾伦人工智能研究所)
;
University of Washington(华盛顿大学)
机构
*
University of Maryland College Park(马里兰大学帕克分校)
;
Nanjing University of Posts and Telecommunications(南京邮电大学)
;
Stanford University(斯坦福大学)
;
Motional AD Inc.(Motional AD公司)
Comments14 pages, 4 figures, 13 tables. Code, evaluation harness, and the released Temporal LLLite adapter weights are at https://github.com/otanl/dreamlite-stream (also mirrored to Hugging Face and Zenodo)
机构
*
School of Information Science/National Key Laboratory of Deep Space Exploration University of Science and Technology of China(中国科学技术大学信息科学学院/深空探测国家实验室)
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
SpaceDrive: 在基于视觉语言模型的自动驾驶中引入空间感知
Peizheng Li, Zhenghao Zhang, David Holtz, Hang Yu, Yutong Yang, Yuzhi Lai, Rui Song, Andreas Geiger, Andreas Zell
机构
*
Mercedes-Benz AG(梅赛德斯-奔驰集团)
;
University of Tübingen(图宾根大学)
;
Tübingen AI Center(图宾根人工智能中心)
;
TU Munich(慕尼黑工业大学)
;
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
University of Stuttgart(斯图加特大学)
;
UCLA(加州大学洛杉矶分校)
专题命中
VLM训练与架构
:VLM(title,summary_cn);vision language model(abstract);分类 cs.CV
The Hidden Evolution of Disguised Visual Context inside the VLM
VLM内部伪装视觉上下文的隐藏演化
Wish Suharitdamrong, Tony Alex, Xiatian Zhu, Muhammad Awais, Sara Atito
机构
*
Surrey Institute for People-Centred AI, University of Surrey(萨里大学以人为本人工智能研究所)
;
Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey(萨里大学视觉、语音与信号处理中心)
机构
*
The University of Hong Kong(香港大学)
;
Nanjing University(南京大学)
;
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
;
Fudan University(复旦大学)
Progressive Multimodal Alignment for Continual Instruction Tuning
用于持续指令微调的渐进式多模态对齐
Duzhen Zhang, Yahan Yu, Qiaoyi Su, Jiahua Dong, Tielin Zhang
机构
*
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences(中国科学院脑科学与智能技术卓越创新中心)
;
Kyoto University(京都大学)
;
Migu Culture Technology Co.,Ltd.(咪咕文化科技有限公司)
;
State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology(脑认知与类脑智能技术国家重点实验室)
专题命中
VLM训练与架构
:MLLM(summary_cn,abstract);multimodal large language model(abstract,abstract_cn);分类 cs.CV、cs.AI
机构
*
McGill University(麦吉尔大学)
;
Mila Quebec AI Institute(魁北克人工智能研究所)
;
University of Michigan - Ann Arbor(密歇根大学安娜堡分校)
;
Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)
机构
*
Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)
;
Zhongguancun Academy(中关村学院)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Nanyang Technological University(南洋理工大学)
;
Department of Mechanical Engineering, Imperial College London(伦敦帝国理工学院机械工程系)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);LLaVA(abstract);InternVL(abstract);分类 cs.CV、cs.AI
CommentsWithdrawn by the author. On further review, the archived artifacts underlying this audit are too incomplete to support the reported statistics, and the paper's conclusions do not follow from the available evidence. The work is withdrawn in full; earlier versions should not be cited
Explaining, Verifying, and Aligning Semantic Hierarchies in Vision-Language Model Embeddings
解释、验证和对齐视觉语言模型嵌入中的语义层次结构
Gesina Schwalbe, Mert Keser, Moritz Bayerkuhnlein, Edgar Heinert, Annika Mütze, Marvin Keller, Sparsh Tiwari, Georgii Mikriukov, Diedrich Wolter, Jae Hee Lee, Matthias Rottmann
机构
*
University of Lübeck(吕贝克大学)
;
Technical University of Munich(慕尼黑工业大学)
;
AUMOVIO SE
;
Osnabrück University(奥斯纳布吕克大学)
;
Leibniz Institute for Agricultural Engineering and Bioeconomy(莱布尼茨农业工程与生物经济研究所)
;
University of Hamburg(汉堡大学)