arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-04-03 至 2026-04-03 共收录 42 信号源:cs.CL, cs.AI, cs.LG

1. 评测与基准 42 篇

2604.01366 2026-04-03 cs.AI 89%

CogBias: Measuring and Mitigating Cognitive Bias in Large Language Models

CogBias: 测量和缓解大型语言模型中的认知偏差

Fan Huang, Songheng Zhang, Haewoon Kwak, Jisun An

机构 * Indiana University Bloomington(印第安纳大学伯明顿分校) Singapore Management University(新加坡管理大学)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文研究大型语言模型中的认知偏差,通过定义四种偏差类型,评估三种模型并发现偏差在不同家族中表现不同,通过激活引导技术减少偏差并保持模型能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07645 2026-04-03 cs.SE cs.AI cs.CR 89%

A Self-Improving Architecture for Dynamic Safety in Large Language Models

为大语言模型动态安全设计的自改进架构

Tyler Slater

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 评测与基准 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI

AI总结 本文提出自改进安全框架SISF,通过反馈循环实现LLM系统在运行时自动检测安全故障并生成防御策略,实验显示其能有效降低攻击成功率,提升系统鲁棒性。

Comments Under review at the journal Information and Software Technology (Special Issue on Software Architecture for AI-Driven Systems)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21884 2026-04-03 cs.CL cs.AI 86%

Support-Contra Asymmetry in LLM Explanations

支持-矛盾不对称性在大语言模型解释中的体现

Avinash Patil

机构 * Hewlett Packard Enterprise(慧与科技公司)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了大语言模型生成的解释与外部模型提取的预测词汇证据之间的关系,发现正确预测的解释更倾向于引用支持性词汇,而错误预测的解释则更多引用矛盾性词汇,揭示了解释与证据之间的不对称性。

Comments 17 Pages, 12 Figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02135 2026-04-03 cs.CL 85%

GaelEval: Benchmarking LLM Performance for Scottish Gaelic

GaelEval: 用于苏格兰盖尔语的大型语言模型性能评估

Peter Devine, William Lamb, Beatrice Alex, Ignatius Ezeani, Dawn Knight, Mícheál J. Ó Meachair, Paul Rayson, Martin Wynne

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文提出GaelEval,首个针对盖尔语的多维评估基准,包含语法任务、翻译基准和文化知识问答,评估19个LLM发现专有模型表现优于开源模型,盖尔语提示法有小幅优势。

Comments 13 pages, to be published in Proceedings of LLMs4SSH (workshop co-located with LREC 2026; Mallorca, Spain; May 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00478 2026-04-03 cs.AI 85%

The Silicon Mirror: Dynamic Behavioral Gating for Anti-Sycophancy in LLM Agents

硅镜:用于LLM代理反趋炎附势的动态行为门控

Harshee Jignesh Shah

机构 * Independent Researcher(独立研究者)

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);RLHF(abstract)

AI总结 本文提出硅镜框架,通过动态检测用户说服策略并调整AI行为以维持事实准确性,实验显示其显著降低LLM的趋炎附势倾向。

Comments 7 pages, 8 figures, 5 tables. Code and evaluation data available at https://github.com/Helephants/langgraph-layered-context

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01615 2026-04-03 cs.AI cs.SE 85%

Analysis of LLM Performance on AWS Bedrock: Receipt-item Categorisation Case Study

对AWS Bedrock上LLM性能的分析:收据项目分类案例研究

Gabby Sanchez, Sneha Oommen, Cassandra T. Britto, Di Wang, Jung-De Chiou, Maria Spichkova

机构 * RMIT University(皇家墨尔本理工大学)

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);prompting(abstract)

AI总结 本文系统评估了AWS Bedrock上四种指令调优模型在收据项目分类中的性能,分析了准确率、响应稳定性及成本效率,并探讨了不同提示方法对准确性和成本的影响。

Comments Preprint. Accepted to the 19th International Conference on Evaluation of Novel Approaches to Software Engineering (ENASE 2026). Final version to be published by SCITEPRESS, http://www.scitepress.org

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22862 2026-04-03 cs.SE cs.CL 85%

The Evolution of Tool Use in LLM Agents: From Single-Tool Call to Multi-Tool Orchestration

大语言模型代理中工具使用的演变:从单工具调用到多工具编排

Haoyuan Xu, Chang Li, Xinyan Ma, Xianhao Ou, Zihan Zhang, Tao He, Xiangyu Liu, Zixiang Wang, Jiafeng Liang, Zheng Chu, Runxuan Liu, Rongchuan Mu, Dandan Tu, Ming Liu, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学) Harvard University(哈佛大学) Huawei Technologies Co., Ltd.(华为技术有限公司)

专题命中 评测与基准 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文探讨了大语言模型代理中工具使用从单次调用到长期编排的演变,分析了多工具代理的最新进展,涵盖任务规划、安全控制、效率优化及实际应用等领域。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02276 2026-04-03 cs.AI cs.CL cs.LG 85%

De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules

De Jure:基于迭代LLM自优化的结构化监管规则提取

Keerat Guliani, Deepkamal Gill, David Landsman, Nima Eshraghi, Krishna Kumar, Lovedeep Gondara

机构 * The Vanguard Group, Inc.(先锋集团)

专题命中 评测与基准 :LLM(title,abstract);prompting(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 De Jure通过四个阶段自动提取结构化监管规则,无需人工标注或领域知识,提升监管文本处理的效率和可追溯性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00979 2026-04-03 cs.CL cs.AI 84%

Dual Optimal: Make Your LLM Peer-like with Dignity

双最优:使你的大语言模型具有尊严的同行

Xiangqi Wang, Yue Huang, Haomin Zhuang, Kehan Guo, Xiangliang Zhang

机构 * University of Notre Dame(圣母大学)

专题命中 评测与基准 :LLM(title,abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出Dignified Peer框架,通过反阿谀和可信性对抗逃避服务模式,利用PersonaKnob数据集和容忍约束Lagrangian DPO算法,构建具有尊严和同行能力的语言模型代理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01987 2026-04-03 cs.CV cs.LG 83%

Curia-2: Scaling Self-Supervised Learning for Radiology Foundation Models

Curia-2:用于放射学基础模型的自监督学习扩展

Antoine Saporta, Baptiste Callard, Corentin Dancette, Julien Khlaut, Charles Corbière, Leo Butsanets, Amaury Prat, Pierre Manceron

机构 * Raidium Department of Vascular and Oncological Interventional Radiology, Hôpital Européen Georges Pompidou, AP-HP, Paris, France(巴黎欧洲乔治·蓬皮杜医院血管与肿瘤介入放射科,AP-HP) Faculté de Santé, Université Paris-Cité, Paris, France(巴黎西岱大学健康学院)

专题命中 评测与基准 :foundation model(title,abstract);language model(abstract);分类 cs.LG

AI总结 Curia-2通过改进预训练策略和表征质量,提升了放射学数据的捕捉能力,首次实现了多模态CT和MRI基础模型的亿参数视觉变换器架构,并在两个新评估轨道中验证了其性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01588 2026-04-03 cs.AI 83%

NED-Tree: Bridging the Semantic Gap with Nonlinear Element Decomposition Tree for LLM Nonlinear Optimization Modeling

NED-Tree: 通过非线性元素分解树弥合语义鸿沟以实现LLM非线性优化建模

Zhijing Hu, Yufan Deng, Haoyang Liu, Changjun Fan

机构 * College of System Engineering, National University of Defense Technology(国防科技大学系统工程学院) University of Science and Technology of China(中国科学技术大学)

专题命中 评测与基准 :LLM(title);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出NED-Tree框架,通过句级提取策略和递归树结构解决LLM在非线性优化建模中的语义鸿沟问题,实验显示其在10个基准测试中达到72.51%的平均准确率。

Comments 17 pages, 7 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02207 2026-04-03 cs.AI cs.CL 81%

Blinded Radiologist and LLM-Based Evaluation of LLM-Generated Japanese Translations of Chest CT Reports: Comparative Study

盲评放射科医生和基于LLM的评估:LLM生成的日语胸CT报告翻译的比较研究

Yosuke Yamagishi, Atsushi Takamatsu, Yasunori Hamaguchi, Tomohiro Kikuchi, Shouhei Hanaoka, Takeharu Yoshikawa, Osamu Abe

机构 * Division of Radiology and Biomedical Engineering, Graduate School of Medicine, The University of Tokyo(东京大学医学研究生院放射学与生物医学工程系) Department of Radiology, The University of Tokyo Hospital(东京大学医院放射科) Department of Radiology, Kanazawa University Graduate School of Medical Sciences(金泽大学医学研究生院放射科) Department of Computational Diagnostic Radiology and Preventive Medicine, The University of Tokyo Hospital(东京大学医院计算诊断放射学与预防医学系) Department of Radiology, School of Medicine, Jichi Medical University(自治医科大学医学院放射科)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.CL、cs.AI

AI总结 研究比较了放射科医生和LLM评估对LLM生成的日语胸CT报告翻译的教育适用性,发现LLM生成的翻译在自然流畅性上被普遍认可,但放射科医生之间存在显著差异,LLM评估偏好LLM输出,与放射科医生无显著一致。

Comments 25 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01977 2026-04-03 cs.CR cs.AI cs.CL cs.LG cs.SE 80%

RuleForge: Automated Generation and Validation for Web Vulnerability Detection at Scale

RuleForge: 用于大规模Web漏洞检测的自动化生成与验证

Ayush Garg, Sophia Hager, Jacob Montiel, Aditya Tiwari, Michael Gentile, Zach Reavis, David Magnotti, Wayne Fullen

机构 * Johns Hopkins University(约翰霍普金斯大学) Amazon Web Services(亚马逊云服务)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 RuleForge通过自动化生成基于JSON的检测规则,结合LLM作为判断系统和反馈机制,提升漏洞检测的准确性和效率,减少误报。

Comments 11 pages, 10 figures. To be submitted to CAMLIS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12864 2026-04-03 cs.CL cs.AI cs.LG 80%

LEXam: Benchmarking Legal Reasoning on 340 Law Exams

LEXam:基于340法律考试的法律推理基准测试

Yu Fan, Jingwei Ni, Jakob Merane, Yang Tian, Yoan Hermstrüwer, Yinya Huang, Mubashara Akhtar, Etienne Salimbeni, Florian Geering, Oliver Dreyer, Daniel Brunner, Markus Leippold, Mrinmaya Sachan, Alexander Stremitzer, Christoph Engel, Elliott Ash, Joel Niklaus

机构 * ETH Zurich(苏黎世联邦理工学院) University of Zurich(苏黎世大学) University of Lausanne(洛桑大学) Max Planck Institute for Research on Collective Goods(马克斯·普朗克集体物品研究所) Omnilex University of St. Gallen(圣加仑大学) Swiss Federal Supreme Court(瑞士联邦最高法院)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 LEXam基准测试通过340法律考试数据,包含7537道英语和德语法律考试题目,涵盖多种法律课程和学位级别,旨在评估大语言模型在长文本法律推理中的挑战与性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01637 2026-04-03 cs.CR cs.AI 79%

Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection

SecLens:针对安全漏洞检测LLM的角色特定评估

Subho Halder, Siddharth Saxena, Kashinath Kadaba Shrish, Thiyagarajan M

机构 * Mattersec Labs(Mattersec实验室) Kalmantic Labs(Kalmantic实验室)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.AI

AI总结 SecLens引入多利益相关者评估框架,通过角色特定加权配置评估LLM在安全漏洞检测中的性能差异,揭示多目标问题的本质。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00528 2026-04-03 cs.CV cs.AI 79%

Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding

思考、行动、构建:一种基于视觉语言模型的代理框架用于零样本3D视觉定位

Haibo Wang, Zihao Lin, Zhiyang Xu, Lifu Huang

机构 * University of California, Davis(加州大学戴维斯分校) Virginia Tech(弗吉尼亚理工大学)

专题命中 评测与基准 :language model(title,abstract);分类 cs.AI

AI总结 本文提出TAB框架,通过2D到3D重建范式直接处理原始RGB-D流,利用2D VLMs解析复杂空间语义并结合多视图几何构建3D结构,克服传统方法依赖预处理点云的局限,实验证明其在ScanRefer和Nr3D上优于零样本方法和全监督基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08462 2026-04-03 cs.AI 79%

M3-BENCH: Process-Aware Evaluation of LLM Agents' Social Behaviors in Mixed-Motive Games

M3-BENCH:混合动机游戏中LLM代理社会行为过程感知评估

Sixiong Xie, Zhuofan Shi, Haiyang Shen, Yun Ma, Xiang Jing

机构 * Peking University(北京大学) National Key Laboratory of Data Space Technology and System(数据空间技术与系统全国重点实验室)

专题命中 评测与基准 :LLM(title,abstract);分类 cs.AI

AI总结 M3-BENCH通过24个混合动机游戏,从行为轨迹、推理过程和沟通内容三个视角评估LLM代理的社会能力,揭示了仅关注结果的评估方法无法捕捉到的社交能力差异,尤其是推理模型在内部 deliberation 与有效沟通之间的矛盾。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01047 2026-04-03 cs.SE cs.AI 78%

HAFixAgent: History-Aware Program Repair Agent

HAFixAgent:基于历史的程序修复代理

Yu Shi, Hao Li, Bram Adams, Ahmed E. Hassan

机构 * Queen's University(女王大学)

专题命中 评测与基准 :LLM(abstract,comments);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出HAFixAgent,通过注入历史衍生的仓库启发式方法,提升大规模多补丁bug修复的效率与鲁棒性,实验证明其在Defects4J和BugsInPy上均优于现有方法。

Comments support both Defects4J and BugsInPy; use the same LLM for all baseline comparisons; add sensitivity analysis for imperfect fault localization

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01259 2026-04-03 cs.RO 78%

Bench2Drive-VL: Benchmarks for Closed-Loop Autonomous Driving with Vision-Language Models

Bench2Drive-VL: 用于基于视觉语言模型的闭环自动驾驶的基准测试

Xiaosong Jia, Yuqian Shao, Zhenjie Yang, Qifeng Li, Zhiyuan Zhang, Junchi Yan

机构 * Shanghai Key Laboratory of Multimodal Embodied AI(上海市多模态具身智能重点实验室)

专题命中 评测与基准 :language model(title,abstract)

AI总结 本文提出Bench2Drive-VL,通过闭环评估方法提升视觉语言模型在自动驾驶中的性能,包含DriveCommenter生成多样化问题对、统一协议接口、灵活推理框架及完整开发生态,开源代码和标注数据。

Comments All codes and annotated datasets are available at \url{https://github.com/Thinklab-SJTU/Bench2Drive-VL} and \url{https://huggingface.co/datasets/Telkwevr/Bench2Drive-VL-base}

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29399 2026-04-03 cs.AI cs.DB 77%

ELT-Bench-Verified: Benchmark Quality Issues Underestimate AI Agent Capabilities

ELT-Bench-Verified: 评估质量问题低估了AI代理的能力

Christopher Zanoli, Andrea Giovannini, Tengjun Jin, Ana Klimovic, Yotam Perlitz

机构 * IBM Research(IBM研究院) ETH Zurich(苏黎世联邦理工学院) University of Illinois (UIUC)(伊利诺伊大学厄巴纳-香槟分校)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文通过重新评估ELT-Bench发现代理能力被低估,提出Auditor-Corrector方法修正评估质量,构建ELT-Bench-Verified提升评估准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00039 2026-04-03 cs.DB cs.AI cs.IR 77%

AutoPK: Leveraging LLMs and a Hybrid Similarity Metric for Advanced Retrieval of Pharmacokinetic Data from Complex Tables and Documents

AutoPK:利用大语言模型和混合相似性度量进行药代动力学数据的高级检索

Hossein Sholehrasa, Amirhossein Ghanaatian, Doina Caragea, Lisa A. Tell, Jim E. Riviere, Majid Jaberi-Douraki

机构 * Kansas State University(堪萨斯州立大学) University of California-Davis(加利福尼亚大学戴维斯分校)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 AutoPK通过大语言模型和混合相似性度量,实现了从复杂表格和文档中高效提取药代动力学数据,显著提升了精度和召回率,适用于兽医药理学和公共卫生决策。

Comments Published in IEEE ICTAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01586 2026-04-03 cs.CV cs.AI 77%

SHOE: Semantic HOI Open-Vocabulary Evaluation Metric

SHOE:语义HOI开放词汇评估度量

Maja Noack, Qinqian Lei, Taipeng Tian, Bihan Dong, Robby T. Tan, Yixin Chen, John Young, Saijun Zhang, Bo Wang

机构 * University of Mississippi(密西西比大学) National University of Singapore(新加坡国立大学) Independent Researcher(独立研究员) ASUS Intelligent Cloud Services (AICS)(华硕智能云服务(AICS))

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出SHOE评估框架,通过结合预测与真实HOI标签的语义相似性,改进开放词汇HOI检测的评估方法,实验表明其与人类判断一致率达85.73%。

Comments Accepted to GRAIL-V Workshop at CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02989 2026-04-03 cs.CL 77%

A Comparative Study of Competency Question Elicitation Methods from Ontology Requirements

基于本体需求的竞争力问题 elicitation 方法比较研究

Reham Alharbi, Valentina Tamma, Terry R. Payne, Jacopo de Berardinis

机构 * College of Computer Science and Engineering, Taibah University(塔伊巴大学计算机科学与工程学院) School of Computer Science and Informatics, University of Liverpool(利物浦大学计算机科学与信息学院)

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文比较了三种竞争力问题生成方法:人工编制、模式实例化和LLM生成,发现不同方法有不同特点,LLM生成的问题需进一步优化才能用于需求建模。

Comments Revised version (v2) accepted for the 23rd European Semantic Web Conference (ESWC-2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01572 2026-04-03 cs.CR 75%

AI-Assisted Hardware Security Verification: A Survey and AI Accelerator Case Study

AI辅助的硬件安全验证:综述与AI加速器案例研究

Khan Thamid Hasan, Md Ajoad Hasan, Nashmin Alam, Md. Touhidul Islam, Upoma Das, Farimah Farahmandi

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 本文综述了AI辅助硬件安全验证的最新进展,通过NVDLA案例研究展示AI在自动化验证流程中的应用,强调AI输出需基于仿真证据和形式验证确保可信安全。

Comments This paper will be presented at IEEE VLSI Test Symposium (VTS) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29376 2026-04-03 cs.CV 75%

Assessing Multimodal Chronic Wound Embeddings with Expert Triplet Agreement

基于专家三元组一致性的多模态慢性伤口嵌入评估

Fabian Kabus, Julia Hindel, Jelena Bratulić, Meropi Karakioulaki, Ayush Gupta, Cristina Has, Thomas Brox, Abhinav Valada, Harald Binder

机构 * Institute of Medical Biometry and Statistics (IMBI), Medical Faculty and Medical Center, University of Freiburg(弗莱堡大学医学院与医学中心医学统计与生物统计研究所) Department of Computer Science, Faculty of Engineering, University of Freiburg(弗莱堡大学工程学院计算机科学系) Department of Dermatology, Medical Faculty and Medical Center, University of Freiburg(弗莱堡大学医学院与医学中心皮肤科)

专题命中 评测与基准 :large language model(abstract);language model(abstract);foundation model(abstract)

AI总结 本文提出通过专家三元组比较评估嵌入空间,引入TriDerm框架整合图像、边界掩码和专家报告,融合视觉与文本模态提升专家一致性至73.5%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06408 2026-04-03 cs.HC 75%

CommentScope: A Comment-Embedded Assisted Reading System for a Long Text

CommentScope:一种嵌入评论的辅助阅读系统用于长文本

Shuai Chen, Lei Han, Haoran Zhang, Kaihao Liu, Zhaoman Zhong

专题命中 评测与基准 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 CommentScope通过嵌入评论提升长文本阅读效率,采用LLM分类和可视化模块,实现评论分类与展示,提升信息获取与阅读流畅度。

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06763 2026-04-03 cs.CV cs.AI 70%

SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding

SafePLUG: 通过像素级洞察和时间锚定赋能多模态大语言模型以理解交通事故

Zihao Sheng, Zilin Huang, Yansong Qu, Jiancong Chen, Yuhao Luo, Yen-Jung Chen, Yue Leng, Sikai Chen

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Purdue University(普渡大学) Google(谷歌)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出SafePLUG框架,通过像素级理解和时间锚定提升多模态大语言模型在交通事故分析中的能力,引入新数据集并实现区域问答、像素分割、时间事件定位等任务的高性能表现。

Comments The code, dataset, and model checkpoints will be made publicly available at: https://zihaosheng.github.io/SafePLUG

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02034 2026-04-03 cs.AI 70%

AI in Insurance: Adaptive Questionnaires for Improved Risk Profiling

保险中的AI:适应性问卷以提升风险评估

Diogo Silva, João Teixeira, Bruno Lima

机构 * Deloitte(德勤) Faculty of Engineering, University of Porto(波尔图大学工程学院) LIACC, Faculty of Engineering, University of Porto(波尔图大学工程学院人工智能与计算机科学实验室)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出ARQuest框架,利用大语言模型和替代数据源创建个性化适应性问卷,通过社交媒体图像分析等技术提升风险评估的准确性与用户满意度。

Journal ref International Workshop on Agentic Engineering (AGENT 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01941 2026-04-03 cs.CV cs.AI 70%

Captioning Daily Activity Images in Early Childhood Education: Benchmark and Algorithm

在早期教育中的日常活动图像描述:基准和算法

Sixing Li, Zhibin Gu, Ziqi Zhang, Weiguo Pan, Bing Li, Ying Wang, Hongzhe Liu

机构 * School of Robotics, Beijing Union University(北京联合大学机器人学院) College of Computer and Cyber Security, Hebei Normal University(河北师范大学计算机与网络安全学院) Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所) National Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所模式识别国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(中国科学院大学人工智能学院) PeopleAI, Inc.(人民人工智能公司) People Youhe Education Technology Co., Ltd.(人民优和教育科技有限公司)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出ECAC基准和RSRS算法,通过专家标注和细粒度标签提升早期教育图像描述的专业性,实验显示KinderMM-Cap-3B在TTS上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16219 2026-04-03 cs.CR cs.AI 70%

SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detection

SentinelNet:通过基于信用的动态威胁检测保障多智能体协作

Yang Feng, Xudong Pan

机构 * The University of Edinburgh(爱丁堡大学) Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 评测与基准 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 SentinelNet通过基于信用的动态威胁检测机制,主动识别并缓解多智能体协作中的恶意行为,实现高检测率和系统恢复率。

Comments Accepted at The ACM Web Conference 2026 (WWW 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏