Rationalizing Predictions by Adversarial Information Calibration
专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI
Comments arXiv admin note: substantial text overlap with arXiv:2012.08884
Journal ref Artificial Intelligence, Volume 315, February 2023
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI
Comments arXiv admin note: substantial text overlap with arXiv:2012.08884
Journal ref Artificial Intelligence, Volume 315, February 2023
专题命中 幻觉与事实性 :alignment(abstract);分类 cs.CL、cs.AI
Comments AAAI 2023 Main Track Long Paper
专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CL、cs.AI
Comments EMNLP 2022
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG
专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.LG
Comments Accepted by EMNLP 2022 (Findings)
专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI、cs.LG
Comments 28 pages,4 figures
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG
Comments Best Paper Award @ Advances in Neural Information Processing Systems - Machine Learning for Autonomous Driving Workshop (NeurIPS 2021 ML4AD)
专题命中 幻觉与事实性 :safety(abstract);分类 cs.CY、cs.LG
Journal ref IEEE Access, Vol. 10, pp. 58375-58418, 2022
专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.CY
Comments This paper got accepted at AAAI 2022, AI for Social Impact track
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG
Comments 12 pages, 9 figures, International Conference on Robotics and Automation (ICRA) 2022
专题命中 幻觉与事实性 :safety(abstract);分类 cs.CL、cs.AI
专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.CY、cs.LG
Comments AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society (AIES) 2021
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG
Comments Accepted at IJCNN 2021, to appear in IEEE proceedings. Equal contributions from US, RK and WZ
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG
Comments Accepted at CVPR 2019 Workshop "DThree19: Dependable Deep Detectors"
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG
Comments arXiv admin note: text overlap with arXiv:1901.02219
Journal ref Proceedings of the 12th International Conference on Agents and Artificial Intelligence - Volume 2: ICAART, 2020, ISBN 978-989-758-395-7, pages 522-529
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.CY
Comments To be presented at the AI for Social Good workshop at NeurIPS 2019
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG
Comments Accepted to IEEE Intelligent Transportation Systems Conference - ITSC 2019
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI、cs.LG
机构 * Georgia Institute of Technology(佐治亚理工学院)
专题命中 幻觉与事实性 :alignment(abstract,comments);分类 cs.AI
Comments Keywords: Explainable Computer Vision, Large Vision-Language Models, AI Interpretability, Explainable AI, Visual Saliency, Attribution Maps, Cross-Modal Attribution, Human Attention Alignment, AI Transparency
专题命中 幻觉与事实性 :safety(abstract,journal_ref);分类 cs.AI
Journal ref In: Gallina B., Skavhaug A., Schoitsch E., Bitsch F. (eds) Computer Safety, Reliability, and Security. SAFECOMP 2018. Lecture Notes in Computer Science, vol 11094. Springer, Cham
代理验证的大语言用户体验微模拟:一种用于早期决策支持的工件优先协议
专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI
AI总结 该研究提出代理验证的LLM驱动UX微模拟流程,通过代理语料库验证模拟结果,结合多指标对比基线、消融实验分析智能体策略,提供可复现的早期UX决策支持方案。
Comments 27 pages, 5 figures, 15 tables
多智能体语言模型中的道德风险
专题命中 幻觉与事实性 :safety(abstract);分类 cs.AI
AI总结 研究多智能体语言模型中的道德风险,引入对话道德风险游戏,评估七个模型并分解行为,用多种优化机制更新,发现效果异质,强调应报告机制级行为而非仅团队成功。
Comments This revision substantially expands the empirical evaluation to eleven open-weight and three frontier models, adding matched query-cost, team-reward, group-size, and private-share incentive analyses. It also extends the weight-level and GEPA results, frozen-prompt information-structure interventions, statistical uncertainty analyses, and qualitative prompt/reasoning-trace studies
基于不确定性的贝叶斯解释框架用于电力质量问题分类
机构 * School of Engineering, Deakin University(德肯大学工程学院) ; ARC Training Centre in Energy Technologies for Future Grids, School of Engineering, University of Wollongong(未来电网能源技术培训中心,沃尔灵宗大学工程学院)
专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG
AI总结 本文提出一种贝叶斯解释框架,通过生成相关性归因分布来建模解释不确定性,提升电力质量问题分类器的透明度和可靠性。
CLAIM:基于不确定性度量的大语言模型开放域主动澄清领先方案
机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
专题命中 幻觉与事实性 :alignment(abstract);分类 cs.AI
AI总结 本研究提出基于不确定性度量的CLAIM框架,无需人工标注,结合监督学习与强化学习训练统一澄清决策模型,实现开放域人机交互中低成本鲁棒的主动澄清。
Comments 11 pages, 4 figures, and 3 tables
自我进化临床系统之路:将医疗智能体从辅助扩展到自主
机构 * Hunan University(湖南大学) ; ByteDance(字节跳动) ; Duke University(杜克大学) ; Westlake University(西湖大学) ; The University of Hong Kong(香港大学) ; Nanyang Technological University(南洋理工大学) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; University of Macau(澳门大学) ; The Ohio State University(俄亥俄州立大学)
专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI
AI总结 研究探讨大语言模型等对医疗智能体的重塑,从临床部署出发,将其形式化为决策系统并给出自主性分类。沿统一框架扩展,强调临床环境扩展为关键方向,定位临床自我进化为前沿,还研究了多领域应用及挑战,提供医学成像系统路线图。
Comments Project page: https://github.com/zhcz328/Awesome-Medical-Agents
CARE:用于可靠医学视觉问答的置信感知推理
机构 * Ant Group(蚂蚁集团) ; University of Michigan(密歇根大学) ; City University of Hong Kong(香港城市大学)
专题命中 幻觉与事实性 :trustworthy(abstract);分类 cs.AI
AI总结 该研究针对医学多模态大语言模型的置信度校准问题,提出CARE框架,通过双阶段流程优化准确率与校准度,在三个医学VQA基准上取得最优性能,为临床决策支持提供可信基础。
Comments Accepted by MICCAI 2026
VIDS-Seg:面向儿科心脏超声分割的可靠不确定性量化
机构 * University of Basel(巴塞尔大学)
专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG
AI总结 该研究提出VIDS-Seg方法,在成人超声数据集训练后零样本应用于儿科心脏超声分割,可可靠检测模型在儿科亚群的隐性失效,兼具分割精度与不确定性匹配优势。
视觉-语言表示学习中的动态分布感知不确定性跟踪
机构 * State Key Laboratory of Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室) ; ByteDance Inc(字节跳动公司)
专题命中 幻觉与事实性 :safety(abstract);分类 cs.LG
AI总结 针对视觉-语言模型事后不确定性量化方法忽略测试分布动态性的问题,提出DDA-UQ框架,通过高斯混合模型建模嵌入空间,实现动态不确定性估计,性能优于现有最优方法。