CommentsInitial controlled diagnostic study on 23 natural drawing sets and three VLMs; broader model, building, repeated-inference, and human coverage is planned for a subsequent version
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
University of Freiburg(弗莱堡大学)
;
Hangzhou City University(杭州城市大学)
;
XGRIDS(XGRIDS公司)
;
Fudan University(复旦大学)
;
Hong Kong Baptist University(香港浸会大学)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving
UniDrive: 面向自动驾驶可解释风险理解的统一视觉-语言与定位框架
Xiaowei Gao, Pengxiang Li, Yitai Cheng, Ruihan Xu, James Haworth, Stephen Law, Yun Ye
机构
*
organization= Department of Earth Science \& Engineering, Imperial College London , city= London , postcode= SW7 2AZ , country= United Kingdom
;
organization= SpaceTimeLab, Department of Civil, Environmental
;
Geomatic Engineering, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom
;
organization= Department of Computing, The Hong Kong Polytechnic University , city= Hong Kong , country= China
;
organization= Trinity College, University of Oxford , city= Oxford , postcode= OX1 3BH , country= United Kingdom
;
organization= Department of Geography, University College London , city= London , postcode= WC1E 6BT , country= United Kingdom
;
organization= Centre for Global Infrastructure Resilience, The Bartlett School of Sustainable Construction, University College London , city= London , postcode= WC1E 7HB , country= United Kingdom
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation
RRS-10K:用于罕见遥感图像解释的多任务视觉语言模型基准测试
Yuqiao Lai, Jiancheng Qi, Fei Wang, Yuxin Liu, Kun Li, Ye Chen, Yan Gao, Yanyan Wei
机构
*
Laboratory of Intelligent Language Processing, National University of Defense Technology(国防科技大学智能语言处理实验室)
;
Hefei University of Technology(合肥工业大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
;
United Arab Emirates University(阿联酋大学)
机构
*
Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人工智能与机器人研究所)
;
Xingchen AGI Lab, China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd(星辰通用人工智能实验室,中国电信人工智能技术(北京)有限公司)
;
University of Science and Technology Beijing(北京科技大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Shanghai Jiao Tong University(上海交通大学)
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV
机构
*
Zhejiang University(浙江大学)
;
Ant Group(蚂蚁集团)
;
The State Key Laboratory of Blockchain and Data Security, Zhejiang University(浙江大学区块链与数据安全国家重点实验室)
;
Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术产业开发区(滨江)区块链与数据安全研究院)
专题命中
视觉定位与Grounding
:multimodal large language model(title,abstract);grounding(abstract);分类 cs.AI
机构
*
East China University of Science and Technology(华东理工大学)
;
Tsinghua University(清华大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
University of Science and Technology of China(中国科学技术大学)
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV
Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data
Prompting-MammAlps:用于相机陷阱数据的细粒度文本到视频检索
Valentin Gabeff, Baptiste Maquignaz, Jennifer Shan, Sepideh Mamooler, Gencer Sumbul, Blair Costelloe, Devis Tuia, Alexander Mathis
机构
*
Ecole Polytechnique Fédérale de Lausanne (EPFL)(洛桑联邦理工学院)
;
Max Planck Institute of Animal Behavior(马克斯·普朗克动物行为研究所)
;
University of Konstanz(康斯坦茨大学)