FRABench and UFEval: Unified Fine-grained Evaluation with Task and Aspect Generalization
Shibo Hong, Jiahao Ying, Haiyuan Liang, Mengdi Zhang, Jun Kuang, Jiazheng Zhang, Yixin Cao
机构
*
School of Computer Science, Fudan University(复旦大学计算机学院)
;
Meituan Group(美团集团)
;
School of Computer Science, Singapore Management University(新加坡管理学院)
专题命中
其他VLM
:multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG
Task-Aware KV Compression For Cost-Effective Long Video Understanding
Minghao Qin, Yan Shu, Peitian Zhang, Kun Lun, Huaying Yuan, Juenjie Zhou, Shitao Xiao, Bo Zhao, Zheng Liu
机构
*
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Shanghai Jiao Tong University(上海交通大学)
;
University of Trento(特伦特大学)
;
Renmin University of China(中国人民大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Hong Kong Polytechnic University(香港理工大学)
;
Institute of Automation CAS Beijing China(中国科学院自动化研究所)
专题命中
其他VLM
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
Rohit Saxena, Pasquale Minervini, Frank Keller
专题命中
其他VLM
:vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
CommentsThis paper includes a dataset of research posters with abstracts. We provide two cited examples ( arXiv:2211.11880 and arXiv:2210.07571 ) to illustrate reference summaries
机构
*
New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Kling Team, Kuaishou Technology(快手科技 Kling 团队)
;
Peking University(北京大学)
;
Nanjing University(南京大学)
专题命中
其他VLM
:multimodal large language model(title)
CommentsWe found problems in the code while rechecking our implementation. These issues led to noticeable numerical discrepancies, making some of the reported results and conclusions potentially unreliable. Therefore, we request to withdraw this submission
Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms
Tyler Loakman, Joseph James, Chenghua Lin
机构
*
Department of Computer Science, University of Sheffield, UK(谢菲尔德大学计算机科学系)
;
Department of Computer Science, University of Manchester, UK(曼彻斯特大学计算机科学系)
Self-Improvement in Multimodal Large Language Models: A Survey
Shijian Deng, Kai Wang, Tianyu Yang, Harsh Singh, Yapeng Tian
机构
*
The University of Texas at Dallas(德克萨斯大学达拉斯分校)
;
University of Toronto(多伦多大学)
;
University of Notre Dame(诺特丹大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
专题命中
其他VLM
:multimodal large language model(title)