Never trust, always verify : a roadmap for Trustworthy AI?
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG
Comments Accepted in CompleNet-2021 (oral presentation)
Journal ref In: Teixeira, A.S., Pacheco, D., Oliveira, M., Barbosa, H., Gonçalves, B., Menezes, R. (eds) Complex Networks XII. CompleNet-Live 2021. Springer Proceedings in Complexity. Springer, Cham
专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY
Comments 57 pages, 7 figures, 3 tables. Public project deliverable, fortiss whitepaper
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG
Comments arXiv admin note: text overlap with arXiv:2105.04615, arXiv:2104.07060
专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI
Comments To appear in ACL 2022, 5 pages, 2 figures
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY
专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.LG
Comments 20 pages,8 figures, 15 tables
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY
Journal ref Science (2021) Vol 374, Issue 6573, pp. 1327-1329
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.CY
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG
Comments 27 pages, 13 figures
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY
Comments Accepted for publication in the Proc. of the 2nd Workshop on Ethics in Software Engineering Research and Practice
专题命中 安全评测 :safety(title,abstract);分类 cs.CY、cs.LG
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG
Journal ref Journal of Biomedical Informatics, 113 (2021), 103655
专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG
Comments 6 pages, 2 figures, 1 table. Oral at the Beyond Backpropagation Workshop, NeurIPS 2020
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CY、cs.LG
Comments ACM CODS-COMAD 2021 Tutorial
专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.CY
Comments 1st International Workshop on New Foundations for Human-Centered AI @ ECAI 2020
专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG
专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG
Comments ICML 2020
用于可靠乳腺超声诊断的空间基础概念瓶颈模型
专题命中 安全评测 :trustworthy(title,comments);alignment(abstract);分类 cs.AI
AI总结 研究乳腺超声诊断中概念瓶颈模型可信度受监督限制问题,提出空间基础概念瓶颈模型(SG-CBM),利用病变轮廓弱监督,通过导出特定区域训练概念图,经交叉验证等提升诊断指标与概念证据空间对齐,强调数据质量监督设计及可信度验证的必要。
Comments Accepted to the Workshop on Data Quality Aware, High-Performance, and Trustworthy AI Systems for Healthcare at IEEE/ACM CHASE 2026
谁的对齐?比较不同组织决策情境下的大语言模型过程对齐
机构 * University of Cambridge(剑桥大学)
专题命中 安全评测 :alignment(title,abstract);分类 cs.AI
AI总结 本文提出一种决策策略捕获方法测量过程对齐,发现LLM在ECHR第6条决策中过程对齐与输出准确性高度相关,但在德国消费信贷决策中关系消失,揭示了多元对齐挑战。
Comments Accepted to Pluralistic Alignment Workshop @ ICML 2026, Seoul, South Korea
通过基于人设的对抗性链式思考视觉语言模型验证实现被动施工现场安全监控
机构 * Department of Computer Science, University of Maryland, College Park, MD, USA(大学马里兰学院计算机科学系,马里兰州科利尔帕克,MD,美国)
专题命中 安全评测 :safety(title,abstract);分类 cs.AI
AI总结 本文提出了一种被动的施工现场安全监控方法,通过三阶段架构处理视频数据,结合细调的YOLO11、SAM 3和Qwen3-VL-8B-Instruct模型,利用基于人设的对抗性链式思考协议提高合规性验证和幻觉控制,主要贡献是第三阶段提示设计,提升了12%的精度。
Comments 10 pages, 4 figures. First place, Ironsite.ai Spatial Intelligence Hackathon, University of Maryland, February 2026. Code available at https://github.com/ananthsriram1/ironsite-hackathon-project-safety_assistant
迈向基于法律和安全原则的神经符号因果规则合成、验证与评估
机构 * Hasso Plattner Institute \ of Potsdam Prof.-Dr.-Helmert Str. 2-3, D-14482 Potsdam, Germany ; Hasso Plattner Institute \ of Potsdam
专题命中 安全评测 :safety(title,abstract);分类 cs.AI;trustworthy(journal_ref)
AI总结 本文提出一种神经符号因果框架,结合一阶逻辑抽象树、结构因果模型和深度强化学习,通过Meta层缓解目标误指定问题,实现可扩展的规则维护。
Journal ref Neurosymbolic eXplainable Trustworthy Systems @ AAMAS 2026
四轴决策对齐用于长周期企业AI代理
机构 * Vasundra Srinivasan
专题命中 安全评测 :alignment(title,abstract);分类 cs.AI
AI总结 本文提出四轴对齐框架,用于评估长周期企业AI代理的决策行为,涵盖事实精度、推理连贯性、合规重建和校准回避,通过实验揭示了决策对齐的重要性。
Comments 21 pages, 5 figures, 8 tables. PDFLaTeX. Code and artifacts: https://github.com/vasundras/decision-alignment-long-horizon-agents
三元循环:一种协商AI共播直播中对齐的框架
机构 * University College London(伦敦大学学院)
专题命中 安全评测 :alignment(title,abstract);分类 cs.AI
AI总结 本文提出三元循环框架,用于协商AI共播直播中的对齐问题,通过三者之间的双向适应过程,解决多用户社交环境中的动态反馈循环问题。
Comments 6 pages, 1 figure, Proceedings the Human-AI Interaction Alignment Workshop at CHI 2026 (CHI26 BiAlign Workshop)
可接受性对齐
专题命中 安全评测 :alignment(title,abstract);分类 cs.AI
AI总结 本文提出MAP-AI架构,通过蒙特卡洛方法评估决策策略的可接受性,实现AI对齐的动态决策理论属性。
Comments 24 pages, 2 figures, 2 tables.. Decision-theoretic alignment under uncertainty