arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 1732 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 幻觉与事实性 1732 篇

2207.07228 2026-04-23 eess.SP 78%

Multi-FEAT: Multi-Feature Edge Alignment for Targetless Camera-LiDAR Calibration

多特征边缘对齐:无目标相机-激光雷达校准

Bichi Zhang, Holger Caesar, Raj Thilak Rajan

专题命中 幻觉与事实性 :alignment(title,abstract)

AI总结 本文提出Multi-FEAT方法,通过将3D激光雷达点云投影到2D全景图,并利用多种激光雷达特征信息补充稀疏边界,结合相机边缘提取和特征匹配函数优化校准参数,实验表明其在KITTI数据集上优于现有无目标校准方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05210 2026-04-08 cs.CV 78%

Integration of Object Detection and Small VLMs for Construction Safety Hazard Identification

对象检测与小型视觉语言模型的整合用于建筑安全危险识别

Muhammad Adil, Mehmood Ahmed, Muhammad Aqib, Vicente A. Gonzalez, Gaang Lee, Qipei Mei

机构 * Infrastructure Human Tech (IHT) Lab, Department of Civil and Environmental Engineering, University of Alberta, Edmonton, Alberta, Canada(基础设施人类技术(IHT)实验室,土木与环境工程系,阿尔伯塔大学,埃德蒙顿,阿尔伯塔,加拿大)

专题命中 幻觉与事实性 :safety(title,abstract)

AI总结 本文提出了一种结合对象检测与多模态推理的轻量级框架,以提高建筑工地安全危险识别的准确性和效率,实验表明该方法在多个小型视觉语言模型上均提升了检测性能和解释质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01981 2026-03-03 stat.AP 78%

Quantifying Uncertainty in Void Swelling Prediction: A Conformal Prediction Framework for Reactor Safety Margins

量化空洞膨胀预测中的不确定性:一种用于反应堆安全余量的符合预测框架

Minhee Kim, Yong Yang

专题命中 幻觉与事实性 :safety(title,abstract)

AI总结 本文提出一种结合集成机器学习与符合预测框架,用于量化空洞膨胀预测中的不确定性,以提升反应堆安全余量评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17405 2026-01-27 cs.CV 78%

HAAF: Hierarchical Adaptation and Alignment of Foundation Models for Few-Shot Pathology Anomaly Detection

HAAF: 用于少样本病理异常检测的层次适应与对齐框架

Chunze Yang, Wenjie Zhao, Yue Tang, Junbo Lu, Jiusong Ge, Qidong Liu, Zeyu Gao, Chen Li

机构 * Sch Comp Sci \& Technol, Xi'an Jiaotong University Xi'an China University of Cambridge Cambridge United Kingdom Sch Comp Sci \& Technol, Xi'an Jiaotong University University of Cambridge

专题命中 幻觉与事实性 :alignment(title,abstract)

AI总结 HAAF通过层次适应与对齐框架,解决少样本病理异常检测中的粒度不匹配问题,提升模型在低资源场景下的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07845 2026-01-14 cs.CV 78%

Edge-AI Perception Node for Cooperative Road-Safety Enforcement and Connected-Vehicle Integration

边缘AI感知节点用于协作道路安全执法和连接车辆整合

Shree Charran R, Rahul Kumar Dubey

机构 * Indian Institute of Science (IISc)(印度科学研究院) Bosch Global Software Technologies Private Limited (BGSW)(博世全球软件技术私人有限公司)

专题命中 幻觉与事实性 :safety(title,abstract)

AI总结 本文提出一种边缘AI感知节点,用于多类交通违规分析和安全事件传播,通过高精度检测和OCR处理提升执法效率,同时支持V2X协议与连接车辆协同提升道路安全。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.06433 2025-12-25 cs.CV 78%

Diagnose Like A REAL Pathologist: An Uncertainty-Focused Approach for Trustworthy Multi-Resolution Multiple Instance Learning

像真正的病理科医生一样诊断:一种以不确定性为核心的多分辨率多实例学习方法

Sungrae Hong, Sol Lee, Jisu Shin, Jiwon Jeong, Mun Yong Yi

机构 * Korea Advanced Institute of Science and Technology(韩国科学技术院)

专题命中 幻觉与事实性 :trustworthy(title,abstract)

AI总结 本文提出UFC-MIL方法,通过多分辨率图像和不确定性校准,提升多实例学习的诊断可靠性。

Comments Accepted by IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14008 2025-10-22 cs.MA 78%

Stop Reducing Responsibility in LLM-Powered Multi-Agent Systems to Local Alignment

Jinwei Hu, Yi Dong, Shuang Ao, Zhuoyun Li, Boxuan Wang, Lokesh Singh, Guangliang Cheng, Sarvapali D. Ramchurn, Xiaowei Huang

专题命中 幻觉与事实性 :alignment(title,abstract)

Comments Updated manuscript of our previous version (arXiv:2502.01714). Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.19252 2025-10-08 cs.CV 78%

Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search

Yuta Oshima, Masahiro Suzuki, Yutaka Matsuo, Hiroki Furuta

机构 * The University of Tokyo(东京大学) Google DeepMind(谷歌DeepMind)

专题命中 幻觉与事实性 :alignment(title,abstract)

Comments Accepted to NeurIPS2025. Website: https://sites.google.com/view/t2v-dlbs and Code: https://github.com/shim0114/T2V-Diffusion-Search

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13919 2025-09-18 cs.CV 78%

Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration

Yuanchen Wu, Ke Yan, Shouhong Ding, Ziyin Zhou, Xiaoqiang Li

机构 * School of Computer Engineering(计算机工程学院) Key Laboratory of Multimedia Trusted Perception(多媒体可信感知关键实验室) Efficient Computing, Xiamen University(高效计算,厦门大学) Tencent Youtu Lab(腾讯优图实验室)

专题命中 幻觉与事实性 :alignment(title,abstract)

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15298 2025-09-08 cs.CV 78%

TPA: Temporal Prompt Alignment for Fetal Congenital Heart Defect Classification

Darya Taratynova, Alya Almsouti, Beknur Kalmakhanbet, Numan Saeed, Mohammad Yaqub

机构 * Department of Machine Learning Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) Abu Dhabi, UAE(机器学习系,Mohamed bin Zayed人工智能大学(MBZUAI),阿布扎比,阿联酋) Department of Computer Vision Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) Abu Dhabi, UAE(计算机视觉系,Mohamed bin Zayed人工智能大学(MBZUAI),阿布扎比,阿联酋)

专题命中 幻觉与事实性 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.13198 2025-06-18 cs.RO 78%

LBAP: Improved Uncertainty Alignment of LLM Planners using Bayesian Inference

James F. Mullen, Dinesh Manocha

机构 * University of Maryland, College Park(马里兰大学学院公园分校)

专题命中 幻觉与事实性 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03202 2025-03-06 cs.CV 78%

Variance-Aware Loss Scheduling for Multimodal Alignment in Low-Data Settings

Sneh Pillai

专题命中 幻觉与事实性 :alignment(title,abstract)

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18654 2025-03-03 cs.CV 78%

Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment

Pritam Sarkar, Sayna Ebrahimi, Ali Etemad, Ahmad Beirami, Sercan Ö. Arık, Tomas Pfister

专题命中 幻觉与事实性 :alignment(title,abstract)

Comments Published in ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17995 2024-11-28 cs.CV 78%

Revisiting Misalignment in Multispectral Pedestrian Detection: A Language-Driven Approach for Cross-modal Alignment Fusion

Taeheon Kim, Sangyun Chung, Youngjoon Yu, Yong Man Ro

专题命中 幻觉与事实性 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10461 2024-10-31 cs.HC 78%

Exploring Parent-Child Perceptions on Safety in Generative AI: Concerns, Mitigation Strategies, and Design Implications

Yaman Yu, Tanusree Sharma, Melinda Hu, Justin Wang, Yang Wang

专题命中 幻觉与事实性 :safety(title,abstract)

Comments 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03655 2024-03-08 cs.RO cs.MA 78%

PUMA: Fully Decentralized Uncertainty-aware Multiagent Trajectory Planner with Real-time Image Segmentation-based Frame Alignment

Kota Kondo, Claudius T. Tewari, Mason B. Peterson, Annika Thomas, Jouko Kinnari, Andrea Tagliabue, Jonathan P. How

专题命中 幻觉与事实性 :alignment(title,abstract)

Comments 7 pages, 13 figures, conference paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.05690 2023-12-13 cs.HC 78%

Is Ignorance Bliss? The Role of Post Hoc Explanation Faithfulness and Alignment in Model Trust in Laypeople and Domain Experts

Tessa Han, Yasha Ektefaie, Maha Farhat, Marinka Zitnik, Himabindu Lakkaraju

专题命中 幻觉与事实性 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07495 2023-11-14 cs.SE 78%

The Last Decade in Review: Tracing the Evolution of Safety Assurance Cases through a Comprehensive Bibliometric Analysis

Mithila Sivakumar, Alvine Boaye Belle, Jinjun Shan, Opeyemi Adesina, Song Wang, Marsha Chechik, Marios Fokaefs, Kimya Khakzad Shahandashti, Oluwafemi Odu

专题命中 幻觉与事实性 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11439 2023-05-22 cs.CV 78%

Few-Shot Learning with Visual Distribution Calibration and Cross-Modal Distribution Alignment

Runqi Wang, Hao Zheng, Xiaoyue Duan, Jianzhuang Liu, Yuning Lu, Tian Wang, Songcen Xu, Baochang Zhang

专题命中 幻觉与事实性 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.05617 2023-04-13 cs.SE cs.SY eess.SY 78%

AutoRepair: Automated Repair for AI-Enabled Cyber-Physical Systems under Safety-Critical Conditions

Deyun Lyu, Jiayang Song, Zhenya Zhang, Zhijie Wang, Tianyi Zhang, Lei Ma, Jianjun Zhao

专题命中 幻觉与事实性 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.16720 2022-12-01 eess.SY cs.SY 78%

Quadratic Programming for Continuous Control of Safety-Critical Multi-Agent Systems Under Uncertainty

Si Wu, Tengfei Liu, Magnus Egerstedt, Zhong-Ping Jiang

专题命中 幻觉与事实性 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.06410 2020-12-14 cs.RO cs.SY eess.SY 78%

Learning How to Trade-Off Safety with Agility Using Deep Covariance Estimation for Perception Driven UAV Motion Planning

Onur Akgun, Kamil Canberk Atik, Mustafa Erdem, Mehmetcan Kaymaz, Bugrahan Yamak, N. Kemal Ure

专题命中 幻觉与事实性 :safety(title,abstract)

Comments A paper on intelligent motion planning for agile drones. It is currently being reviewed for ICRA 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.01168 2020-10-07 cs.AI cs.CL cs.LG 78%

Evaluating the Calibration of Knowledge Graph Embeddings for Trustworthy Link Prediction

Tara Safavi, Danai Koutra, Edgar Meij

专题命中 幻觉与事实性 :trustworthy(title);分类 cs.CL、cs.AI、cs.LG

Comments EMNLP 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.08198 2020-03-03 cs.RO 78%

Safety Considerations in Deep Control Policies with Safety Barrier Certificates Under Uncertainty

Tom Hirshberg, Sai Vemprala, Ashish Kapoor

专题命中 幻觉与事实性 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1810.10170 2019-04-09 cond-mat.mtrl-sci 78%

Band gap and band alignment prediction of nitride based semiconductors using machine learning

Yang Huang, Changyou Yu, Weiguang Chen, Yuhuai Liu, Chong Li, Chunyao Niu, Fei Wang, Yu Jia

专题命中 幻觉与事实性 :alignment(title,abstract)

Comments Contact: Yang Huang, yah048@ucsd.edu

Journal ref the Journal of Materials Chemistry C, 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.22959 2026-07-28 cs.CV cs.AI 新提交 77%

HALLELUAI: A Hallucination-Aware AI System for Ultra-Realistic Image-to-Video Generation at Scale

HALLELUAI:一个用于大规模超逼真图像到视频生成的幻觉感知人工智能系统

Aniket Sakpal, Yang Jiang, Rouzbeh Davoudi, Shayan Hassantabar, Mani Najmabadi

机构 * Expedia Group(亿客行集团)

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.AI

AI总结 针对AI生成视频时质量控制难的问题,HALLELUAI系统集成视频调节与智能再生模块,能评估风险并迭代修复。经人工参与评估,该系统可输出超逼真视频,推动了可信AI视频内容发展,实现大规模图像到视频生成。

Comments 10 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15481 2026-04-14 cs.LG 77%

LLM-as-Judge on a Budget

在预算限制下的人工智能语言模型作为裁判

Aadirupa Saha, Aniket Wagde, Branislav Kveton

机构 * UIC(伊利诺伊大学芝加哥分校) Adobe

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.LG

AI总结 本文提出一种基于多臂老虎机理论的变异性自适应方法,用于在有限预算下优化对提示-响应对的查询分配,以最小化评估误差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03524 2026-04-07 cs.AI 77%

Structural Rigidity and the 57-Token Predictive Window: A Physical Framework for Inference-Layer Governability in Large Language Models

结构刚性和57个标记的预测窗口:一种用于大语言模型推理层可控性的物理框架

Gregory M. Ruddell

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文提出一种物理框架,通过分析大语言模型的推理行为,揭示了预提交信号的存在条件,并展示了结构刚性在不同几何区域中的统一度量。

Comments Extends arXiv:2603.21415. 30 pages. Also available on Zenodo (10.5281/zenodo.19393882)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09870 2026-03-04 cs.CL 77%

Steer2Edit: From Activation Steering to Component-Level Editing

Steer2Edit:从激活引导到组件级编辑

Chung-En Sun, Ge Yan, Zimo Wang, Tsui-Wei Weng

机构 * Department of Computer Science(计算机科学系) Halıcıoğlu Data Science Institute, UC San Diego(哈利奇奥卢数据科学研究所,加州大学圣迭戈分校)

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.CL

AI总结 Steer2Edit通过将引导信号转化为可解释的参数更新,实现组件级权重编辑,提升模型的安全性、真实性和推理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14933 2025-05-22 cs.LG 77%

Foundations of Unknown-aware Machine Learning

Xuefeng Du

专题命中 幻觉与事实性 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.LG

Comments PhD Dissertation

详情

展开后加载摘要…

URL PDF HTML 收藏