arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-03-25 至 2026-03-25 共收录 246 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 19 篇

2603.22612 2026-03-25 cs.CR cs.HC 75%

BioShield: A Context-Aware Firewall for Securing Bio-LLMs

BioShield: 一种面向生物大语言模型的上下文感知防火墙

Protiva Das, Sovon Chakraborty, Sidhant Narula, Lucas Potter, Xavier-Lewis Palmer, Pratip Rana, Daniel Takabi, Mohammad Ghasemigol

专题命中 领域大模型 :LLM(abstract);large language model(abstract);language model(abstract)

AI总结 BioShield通过上下文感知的提示扫描和响应验证,为生物领域大语言模型提供多层次安全防护,有效防止双重用途攻击。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23322 2026-03-25 stat.AP cs.AI cs.CY physics.geo-ph 70%

Leveraging LLMs and Social Media to Understand User Perception of Smartphone-Based Earthquake Early Warnings

利用大型语言模型和社会媒体理解智能手机地震预警的用户感知

Hanjing Wang, S. Mostafa Mousavi, Patrick Robertson, Richard M. Allen, Alexie Barski, Robert Bosch, Nivetha Thiruverahan, Youngmin Cho, Tajinder Gadh, Steve Malkos, Boone Spooner, Greg Wimpey, Marc Stogaitis

机构 * Department of Earth and Planetary Sciences, Harvard University(哈佛大学地球与行星科学系) Google LLC(谷歌公司) Seismological Laboratory, University of California, Berkeley(加州大学伯克利分校地震实验室) Google Germany GmbH(谷歌德国公司)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文通过分析社交媒体数据,探讨用户对智能手机地震预警系统的感知,发现用户信任与预警及时性密切相关,揭示了系统准确性的用户定义差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23501 2026-03-25 cs.CV cs.AI cs.CL 62%

MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage

MedObvious:通过临床分诊暴露视觉语言模型中的医学莫拉维奇悖论

Ufaq Khan, Umair Nawaz, L D M S S Teja, Numaan Saeed, Muhammad Bilal, Yutong Xie, Mohammad Yaqub, Muhammad Haris Khan

机构 * Mohamed bin Zayed University of Artificial Intelligence, UAE(马尔代夫布扎伊德人工智能大学,阿联酋) National Institute of Technology, Silchar(西尔CHAR国家理工学院) Birmingham City University, UK(伯明翰城市大学,英国)

专题命中 领域大模型 :language model(abstract);分类 cs.CL、cs.AI

AI总结 研究提出MedObvious基准测试,通过临床分诊任务验证视觉语言模型的输入验证能力,发现现有模型在处理不一致或无效输入时仍存在可靠性问题。

Comments 11 Pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19843 2026-03-25 eess.SY cs.LG cs.SY 57%

Artificial intelligence for partial differential equations in computational mechanics: A review

计算力学中的偏微分方程人工智能:综述

Yizheng Wang, Jinshuai Bai, Zhongya Lin, Qimin Wang, Cosmin Anitescu, Jia Sun, Mohammad Sadegh Eshaghi, Yuantong Gu, Xi-Qiao Feng, Xiaoying Zhuang, Timon Rabczuk, Yinghua Liu

机构 * Department of Engineering Mechanics, Tsinghua University, Beijing 100084, China(清华大学工程力学系) Simulation Technology, Institute of Photonics, Leibniz University Hannover, Hannover 30167, Germany(汉诺威莱布尼茨大学光子研究所仿真技术部) Drilling Mechanical Department, CNPC Engineering Technology RD Company Limited, Beijing 102206, China(中石油工程科技研发有限公司钻探机械部) School of Mechanical, Medical and Process Engineering, Queensland University of Technology, Brisbane, QLD 4000, Australia(昆士兰科技大学机械、医疗与加工工程学院) ARC Industrial Transformation Training Centre—Joint Biomechanics, Queensland University of Technology, Brisbane, QLD 4000, Australia(昆士兰科技大学 ARC 工业转型培训中心—联合生物力学)

专题命中 领域大模型 :foundation model(abstract);分类 cs.LG

AI总结 本文综述了人工智能求解偏微分方程在计算力学中的应用,包括固体力学、流体力学和生物力学,总结了基于物理信息神经网络、深度能量方法等算法及理论贡献。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23116 2026-03-25 cs.CV 50%

Automatic Segmentation of 3D CT scans with SAM2 using a zero-shot approach

使用SAM2的零样本方法自动分割3DCT扫描

Miquel Lopez Escoriza, Pau Amargant Alvarez

机构 * Department of Computer Science, EPFL, Switzerland(瑞士联邦理工学院计算机科学系)

专题命中 领域大模型 :foundation model(abstract)

AI总结 本文研究了在不进行微调或领域特定训练的情况下,利用Segment Anything Model 2对体积CT数据进行自动分割的零样本方法,通过改进SAM2的视频记忆机制以适应3D数据,展示了零样本方法在医学图像分割中的可行性。

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22618 2026-03-25 cs.HC 50%

Emotional Support with Conversational AI: Talking to Machines About Life

与人工智能进行情感支持:与机器谈论生活

Olivia Yan Huang, Monika Stodolska, Sharifa Sultana

专题命中 领域大模型 :prompting(abstract)

AI总结 本文探讨了人工智能伴侣聊天机器人在情感支持中的交互过程,分析了用户与AI的互动如何被社区解读,并提出情感支持是社会技术协商的过程。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 知识编辑与模型理解 16 篇

2603.22332 2026-03-25 cs.LG cs.AI 90%

Large Language Models for Missing Data Imputation: Understanding Behavior, Hallucination Effects, and Control Mechanisms

大型语言模型在缺失数据填补中的应用:理解行为、幻觉效应与控制机制

Arthur Dantas Mangussi, Ricardo Cardoso Pereira, Ana Carolina Lorena, Pedro Henriques Abreu

机构 * Computer Science Division, Aeronautics Institute of Technology(航空技术研究所计算机科学系) Technology Institute, Federal University of São Paulo(圣保罗联邦大学技术学院) University of Coimbra, CISUC/LASI – Centre for Informatics and Systems of the University of Coimbra, Department of Informatics Engineering(科英布拉大学)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI、cs.LG

AI总结 本文研究了大型语言模型在表格数据集缺失数据填补中的鲁棒性,通过零样本提示工程方法,对比了五种常用LLM与六种先进填补基线方法,发现LLM在真实数据集表现优异,但计算成本高,而传统方法在合成数据集更优。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22303 2026-03-25 cs.LG cs.AI 90%

Sample Transform Cost-Based Training-Free Hallucination Detector for Large Language Models

基于样本转换成本的无训练hallucination检测器用于大型语言模型

Zeyang Ding, Xinglin Hu, Jicong Fan

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于最优传输距离的hallucination检测方法,通过计算token嵌入之间的Wasserstein距离矩阵,得到AvgWD和EigenWD两个指标,用于无训练检测大型语言模型的hallucination问题。

Comments 24 pages, 5 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23268 2026-03-25 cs.LG cs.AI 86%

SafeSeek: Universal Attribution of Safety Circuits in Language Models

SafeSeek: 语言模型中安全回路的通用归因

Miao Yu, Siyuan Fu, Moayad Aloqaily, Zhenhong Zhou, Safa Otoum, Xing fan, Kun Wang, Yufei Guo, Qingsong Wen

机构 * University of Science United Arab Emirates University Nanyang Technological University Zayed University Intelligent Science \& Technology Academy of CASIC

专题命中 知识编辑与模型理解 :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出SafeSeek框架,通过优化识别语言模型中的完整安全回路,验证了在后门攻击和安全对齐场景中的有效性,展示了安全回路的识别与利用方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17108 2026-03-25 cs.CV 86%

LLM-Powered Flood Depth Estimation from Social Media Imagery: A Vision-Language Model Framework with Mechanistic Interpretability for Transportation Resilience

基于大语言模型的社交媒体影像洪水深度估计:一种具有机制可解释性的视觉-语言模型框架用于交通韧性

Nafis Fuad, Xiaodong Qian

专题命中 知识编辑与模型理解 :language model(title,abstract);LLM(title)

AI总结 本文提出FloodLlama,一种基于视觉-语言模型的洪水深度估计框架,通过合成数据集和多模态传感管道实现厘米级实时洪水深度估计,提升交通网络韧性。

Comments There is a update in result, which is needed to be addressed

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22567 2026-03-25 cs.CE cs.MA 85%

TrustTrade: Human-Inspired Selective Consensus Reduces Decision Uncertainty in LLM Trading Agents

TrustTrade: 人类启发的选 择性共识降低LLM交易代理决策不确定性

Minghan Li, Rachel Gonsalves, Weiyue Li, Sunghoon Yoon, Mengyu Wang

专题命中 知识编辑与模型理解 :LLM(title,abstract);large language model(abstract);language model(abstract)

AI总结 TrustTrade通过引入多代理选择性共识框架,减少LLM交易代理在高噪声市场中的决策不确定性,提升风险意识和稳定性。

Comments 24 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23037 2026-03-25 cs.CV cs.AI cs.CL cs.LG cs.RO 82%

YOLOv10 with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection and trustworthy multimodal AI in computer vision perception

YOLOv10结合Kolmogorov-Arnold网络和视觉-语言基础模型用于可解释的目标检测和可信的多模态AI在计算机视觉感知

Marios Impraimakis, Daniel Vazquez, Feiyu Zhou

机构 * University of Bath(巴斯大学) Zhejiang University(浙江大学)

专题命中 知识编辑与模型理解 :foundation model(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出一种基于Kolmogorov-Arnold网络和视觉-语言基础模型的YOLOv10框架,通过七种几何和语义特征提升目标检测的可解释性,并在模糊、遮挡或低纹理场景中实现可信的置信度估计。

Comments 14 pages, 23 Figures, 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08379 2026-03-25 cs.AI cs.LG 81%

SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models

SOM方向优于单一方向:语言模型中的多方向拒绝抑制

Giorgio Piras, Raffaele Mura, Fabio Brau, Luca Oneto, Fabio Roli, Battista Biggio

专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出利用自组织映射(SOM)提取多方向拒绝特征,通过分析有害提示表示与无害提示表示的差异,验证了多方向抑制方法在提升模型安全性和拒绝能力上的有效性。

Comments Accepted at AAAI 2026

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01929 2026-03-25 cs.GR cs.AI cs.CV cs.LG 79%

Image Generation from Contextually-Contradictory Prompts

从语境矛盾提示生成图像

Saar Huberman, Or Patashnik, Omer Dahary, Ron Mokady, Daniel Cohen-Or

机构 * Tel Aviv University(特拉维夫大学) BRIA AI(BRIA人工智能)

专题命中 知识编辑与模型理解 :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种阶段感知提示分解框架,通过代理提示引导去噪过程,解决提示中概念矛盾导致的语义不准确问题,提升图像生成的准确性。

Comments Project page: https://tdpc2025.github.io/SAP/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22299 2026-03-25 cs.LG cs.AI 73%

Between the Layers Lies the Truth: Uncertainty Estimation in LLMs Using Intra-Layer Local Information Scores

层间之中蕴藏真相:利用层内局部信息评分在LLMs中进行不确定性估计

Zvi N. Badash, Yonatan Belinkov, Moti Freiman

机构 * Faculty of Data and Decision Sciences, Technion --- Israel Institute of Technology(数据与决策科学学院,技术离子研究所) Faculty of Computer Science, Technion --- Israel Institute of Technology(计算机科学学院,技术离子研究所) Faculty of Biomedical Engineering, Technion --- Israel Institute of Technology(生物医学工程学院,技术离子研究所)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种轻量级方法,通过单次前向传递评估内部表示中的跨层一致模式,实现LLMs的不确定性估计,在跨数据集迁移中优于传统探针方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22295 2026-03-25 cs.CL cs.AI 73%

Whether, Not Which: Mechanistic Interpretability Reveals Dissociable Affect Reception and Emotion Categorization in LLMs

是否而非哪一个:机制可解释性揭示了LLMs中可分离的情感接收与情绪分类

Michael Keeman

机构 * Keido Labs(Keido实验室)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文通过机制可解释性方法揭示了LLMs中情感处理的两种可分离机制,证明了情绪电路并非依赖关键词,为AI安全评估提供了新标准。

Comments 38 pages, 11 figures, 16 tables. Code and data: https://github.com/keidolabs/affect-reception

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23301 2026-03-25 cs.CL 70%

Steering LLMs for Culturally Localized Generation

引导大语言模型实现文化本地化生成

Simran Khanuja, Hongbin Liu, Shujian Zhang, John Lambert, Mingqing Chen, Rajiv Mathews, Lun Wang

机构 * Google DeepMind(谷歌DeepMind) Carnegie Mellon University(卡内基梅隆大学)

专题命中 知识编辑与模型理解 :post-training(abstract);prompting(abstract);分类 cs.CL

AI总结 本文通过机制可解释性揭示LLM中的文化表征,提出文化嵌入(CuE)方法,提升文化忠实度并引导生成长尾文化概念,为文化本地化提供可控方法。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22812 2026-03-25 cs.CL 70%

Efficient Hallucination Detection: Adaptive Bayesian Estimation of Semantic Entropy with Guided Semantic Exploration

高效幻觉检测:基于适应性贝叶斯估计的语义熵与引导语义探索

Qiyao Sun, Xingming Li, Xixiang He, Ao Cheng, Xuanyu Ji, Hailun Lu, Runke Huang, Qingyong Hu

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出一种适应性贝叶斯估计框架,通过动态调整采样预算提升幻觉检测效率,在低预算场景下减少50%样本使用并提升AUROC 12.6%。

Comments Accepted to a AAAI 2026 (Oral Presentation, <5% acceptance rate), Project page: https://qingyonghu.github.io/Efficient-Hallucination-Detection/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22648 2026-03-25 cs.HC cs.AI 70%

AwesomeLit: Towards Hypothesis Generation with Agent-Supported Literature Research

AwesomeLit: 向基于代理支持的文献研究进行假设生成

Zefei Xie, Yuhan Guo, Kai Xu

机构 * School of Computer Science, University of Nottingham, UK(诺丁汉大学计算机科学学院) School of Intelligence Science and Technology, Peking University, China(北京大学智能科学与技术学院)

专题命中 知识编辑与模型理解 :large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文提出AwesomeLit系统,通过可视化工具帮助用户探索未知领域,识别研究方向并提升信心。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22655 2026-03-25 cs.LG cs.AI 62%

Generalizing Dynamics Modeling More Easily from Representation Perspective

从表示角度更易于泛化动态建模

Yiming Wang, Zhengnan Zhang, Genghe Zhang, Jiawen Dan, Changchun Li, Chenlong Hu, Chris Nugent, Jun Liu, Ximing Li, Bo Yang

机构 * College of Software and Key Laboratory of Symbolic Computation and Knowledge Engineering of Ministry of Education, Jilin University, China(软件学院,教育部符号计算与知识工程重点实验室,吉林大学,中国) School of Computing, Belfast, Northern Ireland, UK(computing 学院,贝尔法斯特,北爱尔兰,英国)

专题命中 知识编辑与模型理解 :language model(abstract);分类 cs.AI、cs.LG

AI总结 本文提出PDEDER方法,通过预训练语言模型和Lyapunov指数目标,实现更稳定的潜在空间动态建模,提升跨系统泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15253 2026-03-25 cs.CV 50%

HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning

HalDec-Bench:图像描述中幻觉检测的基准测试

Kuniaki Saito, Risa Shinoda, Shohei Tanaka, Tosho Hirasawa, Fumio Okura, Yoshitaka Ushiku

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文提出HalDec-Bench,用于评估图像描述中幻觉检测的性能,通过多样化模型生成的描述和人工标注揭示了模型在不同难度任务中的表现差异,并发现检测器倾向于认为响应开头的句子正确。

Comments This work was intended as a replacement of arXiv:2511.20515 and any subsequent updates will appear there

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20515 2026-03-25 cs.CV 50%

HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning

HalDec-Bench:图像描述中幻觉检测的基准测试

Kuniaki Saito, Risa Shinoda, Shohei Tanaka, Tosho Hirasawa, Fumio Okura, Yoshitaka Ushiku

机构 * OMRON SINICX The University of Tokyo(东京大学) The University of Osaka(大阪大学)

专题命中 知识编辑与模型理解 :language model(abstract)

AI总结 本文提出HalDec-Bench,用于评估图像描述中幻觉检测的性能,通过多样化VLM生成的描述和人类标注揭示模型差异,发现检测器倾向于认可响应开头的句子,并通过强VLM过滤减少数据噪声。

Comments Previously this version appeared as arXiv:2603.15253 which was submitted as a new work by accident

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他LLM 14 篇

2512.11336 2026-03-25 cs.CV 89%

UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models

UFVideo:迈向统一的细粒度视频协作理解的大型语言模型

Hewen Pan, Cong Wei, Dashuang Liang, Zepeng Huang, Pengfei Gao, Ziqi Zhou, Lulu Xue, Pengfei Yan, Xiaoming Wei, Minghui Li, Shengshan Hu

机构 * Huazhong University of Science and Technology(华中科技大学) Meituan(美团)

专题命中 其他LLM :large language model(title,abstract);language model(title,abstract);LLM(abstract)

AI总结 UFVideo通过统一多粒度协作理解能力,实现视频理解的全面覆盖,展示了其在多粒度视频任务中的灵活性和优势。

Comments CVPR 2026 Camera Ready, Github Code: https://github.com/Heven-Pan/UFVideo

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22341 2026-03-25 cs.CR cs.AI cs.CL 86%

T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search

T-MAP: 通过轨迹感知的进化搜索实现红队攻击LLM代理

Hyomin Lee, Sangwoo Park, Yumin Choi, Sohyun An, Seanie Lee, Sung Ju Hwang

机构 * KAIST(韩国科学技术院) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他LLM :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出T-MAP方法,通过轨迹感知的进化搜索发现对抗性提示,有效突破安全防护并实现有害目标,实验显示其在攻击实现率上优于基线方法,适用于多步骤工具执行的LLM代理安全测试。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22882 2026-03-25 cs.LG cs.CV 85%

TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy Exploration

TreeTeaming:通过分层策略探索实现视觉语言模型的自主红队测试

Chunxiao Li, Lijun Li, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 其他LLM :language model(title,abstract);LLM(abstract);large language model(abstract);分类 cs.LG

AI总结 TreeTeaming通过动态演化策略探索方法,实现对视觉语言模型的自主红队测试,显著提升攻击成功率和策略多样性,同时降低攻击毒性。

Comments CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22387 2026-03-25 cs.SE cs.AI cs.MA 85%

AI-Generated Code Is Not Reproducible (Yet): An Empirical Study of Dependency Gaps in LLM-Based Coding Agents

AI生成的代码不可重复(尚且):对基于LLM的编码代理依赖差距的实证研究

Bhanu Prakash Vangala, Ali Adibifar, Ashish Gehani, Tanu Malik

专题命中 其他LLM :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文实证研究了基于LLM的编码代理生成代码的可重复性,发现仅68.3%的项目能在干净环境中成功执行,且不同语言存在显著差异,同时发现依赖项存在显著隐藏依赖。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23474 2026-03-25 cs.CY 82%

Evidence of political bias in search engines and language models before major elections

选举前大型搜索引擎和语言模型中政治偏见的证据

Íris Damião, Paulo Almeida, João Franco, Nuno Santos, Pedro C. Magalhães, Joana Gonçalves-Sá

专题命中 其他LLM :language model(title,abstract);large language model(abstract)

AI总结 研究通过审计四个搜索引擎和两个语言模型,发现其在选举前对政治实体和议题存在偏见,尤其在欧洲偏向右翼,在美国偏向共和党议题,语言模型则相对平衡但仍有右翼偏见。

Comments 20 pages, 4 figures; Supplementary Information : Page 22 - 74

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22709 2026-03-25 cs.CL eess.AS 79%

Who Spoke What When? Evaluating Spoken Language Models for Conversational ASR with Semantic and Overlap-Aware Metrics

谁说了什么?通过语义和重叠感知指标评估对话语音模型的对话ASR

Naohiro Tawara, Samuele Cornell, Alexander Polok, Marc Delcroix, Lukáš Burget, Shinji Watanabe

机构 * NTT, Inc.(日本NTT公司) CMU(卡内基梅隆大学) BUT(捷克但尼奇技术大学)

专题命中 其他LLM :language model(title);LLM(abstract);分类 cs.CL

AI总结 本文通过语义和重叠感知指标评估对话ASR模型,发现基于LLM的系统在双人对话中表现良好,但随着说话人数和重叠增加而下降,而模块化流水线方法更稳健。

Comments Submitted to INTERSPEECH 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17908 2026-03-25 cs.CL cs.AI cs.IR 79%

Principled Context Engineering for RAG: Statistical Guarantees via Conformal Prediction

基于RAG的原理化上下文工程:通过符合预测实现统计保证

Debashish Chakraborty, Eugene Yang, Daniel Khashabi, Dawn Lawrie, Kevin Duh

机构 * HLTCOE, Johns Hopkins University(HLTCOE,约翰霍普金斯大学)

专题命中 其他LLM :LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出通过符合预测框架实现RAG的上下文工程,通过统计可控的过滤方法减少冗余上下文,提升事实准确性。

Comments Accepted at ECIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22799 2026-03-25 cs.CL 77%

Span Modeling for Idiomaticity and Figurative Language Detection with Span Contrastive Loss

基于跨度对比损失的隐喻性和修辞语言检测模型

Blake Matheny, Phuong Minh Nguyen, Minh Le Nguyen

机构 * Japan Advanced Institute of Science and Technology(日本科学技术先进研究院)

专题命中 其他LLM :large language model(abstract);language model(abstract);prompting(abstract);分类 cs.CL

AI总结 本文提出结合槽损失和跨度对比损失的BERT和RoBERTa模型,提升隐喻性检测性能,在现有数据集上取得最佳序列准确率。通过消融研究验证了SCL的有效性及泛化能力,并提出几何均值F1和序列准确率的综合评估方法。

详情

展开后加载摘要…

URL PDF HTML 收藏