Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System
推进基于多模态大语言模型(MLLM)的无人机(UAV)图像理解与推理:基准测试及无训练多智能体系统
机构 * Fudan University(复旦大学) ; College of Future Information Technology(未来信息技术学院) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 多模态评测 :MLLM(title,title_cn);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 本研究构建UAVQA-Bench基准,识别MLLM用于无人机图像理解的三类失效模式,提出含DSPE、CAIR、DAAS的无训练多智能体系统UAV-MAS,其32B版本在基准上准确率超Gemini 3 Pro 4.0个百分点