Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards
多模态大语言模型的演进安全态势:新兴威胁与防护措施综述
Xi Li, Shu Zhao, Xiaohan Zou, Fei Zhao, Fuxiao Liu, Yusen Zhang, Cheng Han, Yushun Dong, Jiaqi Wang
机构
*
University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)
;
NVIDIA(英伟达公司)
;
Penn State University(宾夕法尼亚州立大学)
;
Columbia University(哥伦比亚大学)
;
University of Missouri-Kansas City(密苏里大学堪萨斯分校)
;
Florida State University(佛罗里达州立大学)
;
Auburn University(奥本大学)
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG
Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding
当下与过往时间:针对时间演进的文档理解的大语言模型基准测试
Mahbub E Sobhani, Md. Faiyaz Abdullah Sayeedi, Fahmid Hasan Chowdhury, Md Adnan Arefeen, Farig Sadeque, Md. Faizul Bari, Swakkhar Shatabda
机构
*
BRAC University(BRAC大学)
;
United International University(联合国际大学)
;
North South University(北南大学)
;
Spectrum Software & Consulting Ltd.(斯佩克特软件咨询有限公司)
专题命中
评测与基准
:large language model(title);language model(title);LLM(abstract,abstract_cn);分类 cs.AI
Performance of large language models in the optical diagnosis of colorectal polyps
大型语言模型在结直肠息肉光学诊断中的性能
Joshua C. Vences, William T. Tran, Nikko Gimpaya, Catharine M. Walsh, Rishad J. Khan, Robert Bechara, Asher C. Wiggins, Celine N. Rousan, Kaitlyn V. G. L. Morgado, Angie Ibrahim, Kevin H. M. Kuo, Daniel von Renteln, Alexander Hann, Dennis L. Shung, Michael A. Scaffidi, Charles Ménard, Joshua Landy, Samir C. Grover
专题命中
评测与基准
:large language model(title,abstract);language model(title,abstract);分类 cs.AI
AI总结
本研究评估Claude Opus 4等5款大型语言模型对结直肠息肉的光学诊断性能,发现其区分息肉亚型的准确率接近专家共识,但灵敏度与特异度未达ESGE标准,需进一步研究方可临床应用。
Comments\c{opyright} 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
SciVisAgentBench:用于评估科学数据分析和可视化代理的基准测试
Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Nathaniel Gorski, Jianxin Sun, Guoxi Liu, Helgi I. Ingolfsson, David Lenz, Hanqi Guo, Hongfeng Yu, Teja Leburu, Michael Molash, Bei Wang, Tom Peterka, Chaoli Wang, Shusen Liu
机构
*
University of Notre Dame(圣母大学)
;
Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)
;
University of Utah(犹他大学)
;
University of Nebraska–Lincoln(内布拉斯加大学林肯分校)
;
The Ohio State University(俄亥俄州立大学)
;
Argonne National Laboratory(阿贡国家实验室)
;
Anthropic PBC
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.AI
Comments15 pages. Pre-registered experimental program with a public, tiered claims ledger; includes powered negative results, a label-validity audit, a cross-judge shared-prior measurement (n_eff ~ 2 of 16 votes), and a first-party 15-expert human baseline. Pre-registrations, statistical harness, human responses, and the full experiment ledger are released