Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models
诊断大视觉语言模型中的密集同类属性绑定错误
Yuanzhi Xu, Qian Gao, Jun Fan, Guohui Ding, Zhenyu Yang, Yuteng Xiao, Sixue Lin
机构
*
Qilu University of Technology (Shandong Academy of Sciences)(齐鲁工业大学(山东省科学院))
;
China Telecom Digital Intelligence Technology Co., Ltd.(中国电信数字智能科技有限公司)
;
Shenyang Aerospace University(沈阳航空航天大学)
;
University of Nottingham Ningbo China(宁波诺丁汉大学)
VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding
VideoGAIA:面向通用人工智能助手的智能体视频理解基准
Fan Zhang, Guangming Yao, Jinyang Wu, Hao Wu, Zheng Lian, Xinyu Geng, Jingdong Chen, Yi Yuan, Pheng-Ann Heng
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Ant Group(蚂蚁集团)
;
Tsinghua University(清华大学)
;
Tongji University(同济大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
专题命中
视觉问答
:multimodal large language model(abstract);分类 cs.CV
机构
*
Guangdong Provincial Key Laboratory of Malignant Tumor Epigenetics and Gene Regulation, Guangdong-Hong Kong Joint Laboratory for RNA Medicine, Department of Medical Oncology, Breast Tumor Centre, Phase I Clinical Trial Centre(广东省恶性肿瘤表观遗传与基因调控重点实验室,粤港澳RNA医学联合实验室,医学肿瘤科,乳腺肿瘤中心,I期临床试验中心)
;
Guangdong Provincial Key Laboratory of Cancer Pathogenesis and Precision Diagnosis and Treatment, AI Big Data Laboratory, Department of Medical Oncology(广东省肿瘤发生与精准诊断与治疗重点实验室,人工智能大数据实验室,医学肿瘤科)
;
Chongqing Key Laboratory of Bio-perception(重庆生物感知重点实验室)
机构
*
School of Computing, University of Georgia(佐治亚大学计算机学院)
;
Department of Computer Science and Engineering, University of Texas, Arlington(德克萨斯大学阿灵顿分校计算机科学与工程系)
;
Massachusetts General Hospital, Harvard Medical School(麻省总医院哈佛医学院)
;
Department of Biomedical Engineering, New Jersey Institute of Technology(新泽西理工学院生物医学工程系)
Why Vision Fails as a Universal Bridge: Rectifying Modality Asynchrony in Multilingual MLLMs
为何视觉无法作为通用桥梁:修正多语言多模态大语言模型(MLLMs)中的模态异步性
Yihang Du, Juhao Liang, Zhengzhao Lai, Siyu Li, Yan Hu
机构
*
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Shenzhen Loop Area Institute(深圳河套学院)
;
National Health Data Institute (Shenzhen)(国家健康数据研究院(深圳))
专题命中
视觉推理
:visual reasoning(abstract);multimodal large language model(abstract)
CommentsThe authors are withdrawing this manuscript due to errors identified in the experimental evaluation and result aggregation, which affect several reported quantitative results and some conclusions. These issues require substantial re-evaluation of the experiments and analysis
AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning
AeroGround:用于空-地协同推理的综合基准
Shenghong Yi, Lin Zhang, Muzian Li, Jiakang Yuan, Haoyu Zhang, Peng Ye, Jiayuan Fan, Huafeng Qin, Tao Chen
机构
*
Shanghai Innovation Institute(上海创新研究院)
;
College of Future Information Technology, Fudan University(复旦大学未来信息技术学院)
;
College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)
;
Chongqing Technology and Business University(重庆工商大学)
;
The Chinese University of Hong Kong(香港中文大学)
FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making
FZ-VLM:用于肺结节特征表征与临床决策的两阶段Florence-Zephyr视觉语言模型框架
Pramit Dutta, Jenita Manokaran, Richa Mittal, Ryan Appleby, Eranga Ukwatta
机构
*
College of Engineering, University of Guelph(圭尔夫大学工程学院)
;
Holland Bloorview Kids Rehabilitation Hospital(霍兰德布鲁尔维尤儿童康复医院)
;
Guelph General Hospital(圭尔夫综合医院)
;
Ontario Veterinary College, University of Guelph(圭尔夫大学安大略兽医学院)
专题命中
视觉定位与Grounding
:VLM(title,title_cn);vision language model(title,abstract);vision-language model(abstract);分类 cs.CV、cs.AI