When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops
机构
*
Department of Computer Science and Engineering, The Chinese University of Hong Kong(香港中文大学计算机科学与工程系)
;
Institute of Medical Intelligence and XR, The Chinese University of Hong Kong(香港中文大学医学智能与扩展现实研究所)
机构
*
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Khalifa University(哈里发大学)
;
Indian Institute of Technology Delhi(印度理工学院德里分校)
机构
*
University of Chinese Academy of Sciences(中国科学院大学)
;
State Key Lab of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China(中国科学院人工智能安全国家重点实验室,计算技术研究所,北京,中国)
;
Harbin Institute of Technology (Weihai)(哈尔滨工业大学(威海))
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);分类 cs.CV
Physiology-Aware CNN and Zero-Shot Multimodal LLMs for ECG Image Classification: A Comparative Study
生理感知CNN与零样本多模态大语言模型在心电图图像分类中的比较研究
Khalil Ahammad, Derek Abbott, Mohsen Dorraki
机构
*
School of Computer Science and Information Technology, Adelaide University(阿德莱德大学计算机科学与信息学院)
;
Australian Institute for Machine Learning (AIML)(澳大利亚机器学习研究所)
;
Pi MedTech(Pi医疗科技)
;
School of Electrical and Electronic Engineering, Adelaide University(阿德莱德大学电子与电气工程学院)
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);分类 cs.LG
StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs
StylisticBias: 少数人类视觉线索驱动多模态大语言模型中的大部分社会偏见
Shaghayegh Kolli, Timo Cavelius, Nafiseh Nikeghbal, Samantha Dalal, Jana Diesner
机构
*
Technical University of Munich(慕尼黑工业大学)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
;
Princeton Center for Information and Technology Policy(普林斯顿信息与技术政策中心)
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);分类 cs.CV