GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery
GeoEyes: 面向超高清遥感影像证据驱动理解的按需视觉聚焦
Fengxiang Wang, Mingshuo Chen, Yueying Li, Yajie Yang, Yifan Zhang, Long Lan, Xue Yang, Hongda Sun, Yulin Wang, Di Wang, Jun Song, Jing Zhang, Bo Du
机构
*
National University of Defense Technology, China(国防科技大学)
;
Beijing University of Posts and Telecommunications, China(北京邮电大学)
;
University of the Chinese Academy of Sciences, China(中国科学院大学)
;
Shanghai Jiao Tong University, China(上海交通大学)
;
Wuhan University, China(武汉大学)
;
Chinese Academy of Science, China(中国科学院)
;
Tsinghua University, China(清华大学)
;
Renmin University of China, China(中国人民大学)
专题命中
视觉问答
:multimodal large language model(abstract);分类 cs.CV、cs.AI
LeafNet: A Large-Scale Dataset and Comprehensive Benchmark for Foundational Vision-Language Understanding of Plant Diseases
LeafNet:一个大规模数据集和全面基准,用于植物病害的基础视觉-语言理解
Khang Nguyen Quoc, Phuong D. Dao, Luyl-Da Quach
机构
*
School of Electrical Engineering, Korea University(韩国大学电气工程学院)
;
Department of Integrative Biology, The University of Texas at Austin(德克萨斯大学奥斯汀分校整合生物学系)
;
Department of Software Engineering, FPT University(FPT大学软件工程系)
Detecting Contextual Hallucinations in LLMs with Frequency-Aware Attention
基于频率感知注意力的LLM幻觉检测
Siya Qi, Yudong Chen, Runcong Zhao, Qinglin Zhu, Zhanghao Hu, Wei Liu, Yulan He, Zheng Yuan, Lin Gui
机构
*
Department of Informatics, King's College London, UK(伦敦国王学院信息学院)
;
Department of Statistics, University of Warwick, UK(沃里克大学统计系)
;
School of Computer Science, The University of Sheffield, UK(谢菲尔德大学计算机科学学院)
;
The Alan Turing Institute, UK(艾伦·图灵研究所)
Scaling Audio-Text Retrieval with Multimodal Large Language Models
通过多模态大语言模型扩展音频-文本检索
Jilan Xu, Carl Thomé, Danijela Horak, Weidi Xie, Andrew Zisserman
机构
*
Visual Geometry Group, University of Oxford(牛津大学视觉几何组)
;
Epidemic Sound
;
School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
专题命中
VLM训练与架构
:multimodal large language model(title,abstract);MLLM(abstract)
Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter
Mod-Adapter:通过调制适配器实现无需微调的多概念个性化
Weizhi Zhong, Huan Yang, Zheng Liu, Huiguo He, Zijian He, Xuesong Niu, Di Zhang, Guanbin Li
机构
*
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
Kolors Team, Kuaishou Technology(快手科技Kolors团队)
;
Shenzhen Loop Area Institute, Shenzhen, China(深圳循环区研究所)
;
Guangdong Key Laboratory of Big Data Analysis and Processing, Guangzhou, China(广东大数据分析与处理重点实验室)
;
South China University of Technology, Guangzhou, China(华南理工大学)
机构
*
Technical University of Darmstadt(德累斯顿技术大学)
;
Max-Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
;
University of Tuebingen(图宾根大学)
;
Tubingen AI Center(图宾根人工智能中心)
;
Max Planck Institute for Informatics(马克斯·普朗克信息研究所)
Simplifying Outcomes of Language Model Component Analyses with ELIA
用ELIA简化语言模型组件分析的结果
Aaron Louis Eidt, Nils Feldhus
机构
*
Technische Universität Berlin(柏林技术大学)
;
Fraunhofer Heinrich Hertz Institute(弗劳恩霍夫海因里希·赫茨研究所)
;
BIFOLD – Berlin Institute for the Foundations of Learning and Data(柏林学习与数据基础研究所)
VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration
VILLAIN 在 AVerImaTeC 上:通过多智能体协作验证图像-文本声明
Jaeyoon Jung, Yejun Yoon, Kunwoo Park
机构
*
School of AI Convergence, Soongsil University(人工智能融合学院,顺世大学)
;
MAUM AI Inc.(MAUM人工智能公司)
;
Department of Intelligent Semiconductors, Soongsil University(智能半导体系,顺世大学)