arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-05-27 至 2026-05-27 共收录 437 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 38 篇

2605.26445 2026-05-27 cs.CL 86%

Curation and Extraction of Drug-Related Entities from Reddit Platform

从Reddit平台策划和提取药物相关实体

Zewei Wang, Zihan Xu, Yishu Wei, Michael Chary, Yifan Peng

机构 * Population Health Sciences, Weill Cornell Medicine, New York City, USA(威立·科恩医学中心人口健康科学系,纽约市,美国) School of Computing and Information Systems, University of Melbourne, Melbourne, Australia(墨尔本大学计算机与信息系统学院,墨尔本,澳大利亚) Emergency Medicine, Weill Cornell Medicine, New York City, USA(急诊医学,威立·科恩医学中心,纽约市,美国)

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract);分类 cs.CL

AI总结 为解决医生对非法药物真实使用情况了解有限的问题,本文构建了ReDose数据集(6435条Reddit帖子),并采用BERT、LLM和RAG模型进行药物、剂量和效果实体提取,其中BiomedBERT在药物实体上F1达0.843,Llama-3 70B优于GPT-4,但效果提取仍具挑战。

Comments Accepted by IEEE International Conference on Healthcare Informatics (ICHI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08318 2026-05-27 quant-ph 86%

A Model Context Protocol Server for Quantum Execution in Hybrid Quantum-HPC Environments

用于混合量子-HPC环境中量子执行的模型上下文协议服务器

Masaki Shiraishi, Ikko Hamamura, Tatsuya Ishigaki, Tadashi Kadowaki

专题命中 领域大模型 :LLM(summary_cn,abstract);large language model(abstract);language model(abstract)

AI总结 提出基于MCP服务器的AI驱动框架,使LLM代理能通过自然语言自动执行量子计算工作流,包括采样和期望值计算等原语。

Comments Accepted to QC4C3 workshop at QCNC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27066 2026-05-27 cs.CL cs.IR 84%

Large Language Model-Powered Query-Driven Event Timeline Summarization in Industrial Search

工业搜索中基于大语言模型的查询驱动事件时间线摘要

Mingyue Wang, Xingyu Xie, Hang Yang, Li Gao, Lixin Su, Ge Chen, Dawei Yin, Daiting Shi

机构 * Baidu Inc.(百度公司)

专题命中 领域大模型 :large language model(title);language model(title);分类 cs.CL

AI总结 提出QDET系统,通过多任务微调和强化学习实现查询驱动的事件时间线摘要,在百度搜索中显著提升用户参与度。

Comments Accepted at KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26926 2026-05-27 cs.AI 84%

From Norms to Indicators (N2I-RAG): An Agentic Retrieval-Augmented Generation Framework for Legal Indicator Computation

从规范到指标 (N2I-RAG): 一种用于法律指标计算的智能检索增强生成框架

Youssef Al Mouatamid, Marie Bonnin, Jihad Zahir

机构 * LISI Laboratory(LISI实验室) Cadi Ayyad University(卡迪·阿亚德大学) Univ Brest(布列塔尼大学) IRD, Univ Brest, CNRS, Ifremer, LEMAR(IRD、布列塔尼大学、CNRS、Ifremer、LEMAR)

专题命中 领域大模型 :LLM(summary_cn,abstract);language model(abstract);分类 cs.AI

AI总结 提出N2I-RAG框架,通过自适应检索、基于LLM的智能体和验证机制,实现从法律文本到指标的透明、可追溯的自动计算,在法国海洋环境法语料库上优于基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27249 2026-05-27 cs.AI cs.CL 82%

Gumbel Machine: Counterfactual Student Writing Generation via Gumbel Noise Steering

Gumbel机器:通过Gumbel噪声引导生成反事实学生写作

Hunter McNichols, Alexander Scarlatos, Mihai Dascalu, Danielle McNamara, Andrew Lan

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) University Politehnica of Bucharest(布加勒斯特理工大学) Arizona State University(亚利桑那州立大学)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出Gumbel机器,一种利用β-Hindsight控制解码算法生成既符合评分标准又与学生原文相似的反事实文本的模块化方法。

Comments preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07028 2026-05-27 cs.MA cs.AI cs.CL 82%

Strategic Persuasion with Trait-Conditioned Multi-Agent Systems for Iterative Legal Argumentation

基于特质条件的多智能体系统在迭代法律论证中的战略说服

Philipp D. Siedler

机构 * Aleph Alpha Research(Aleph Alpha研究)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 提出战略法庭框架,通过特质条件化的大语言模型智能体模拟多轮法律辩论,发现异质团队表现更优,并引入强化学习特质编排器动态优化辩护策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10450 2026-05-27 cs.LG cs.AI math.OC 82%

Constructing Industrial-Scale Optimization Modeling Benchmark

构建工业规模优化建模基准

Zhong Li, Hongliang Lu, Tao Wei, Yuxuan Chen, Wenyu Liu, Yuan Lan, Fan Zhang, Zaiwen Wen

机构 * Great Bay University(大湾大学) Peking University(北京大学) Huawei Technologies Co., Ltd(华为技术有限公司)

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 提出MIPLIB-NL基准,通过结构感知逆向构建方法从真实混合整数线性规划中生成自然语言规范与求解器代码,以评估大语言模型在工业规模优化建模中的性能。

Comments This paper was accepted by ICML'26 for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09544 2026-05-27 cs.CL 81%

MetaGraph: A Large-Scale Meta-Analysis of GenAI in Financial NLP (2022-2025)

MetaGraph:金融NLP中GenAI的大规模元分析(2022-2025)

Paolo Pedinotti, Peter Baumann, Nathan Jessurun, Leslie Barrett, Enrico Santus

机构 * Bloomberg(贝莱德)

专题命中 领域大模型 :LLM(summary_cn,abstract);分类 cs.CL

AI总结 提出MetaGraph方法,利用本体引导的LLM从科学语料中提取类型化知识图谱,对681篇GenAI在金融领域的论文进行结构化趋势分析,揭示了三个阶段:早期LLM驱动的任务和数据集扩展、对局限性和风险的日益关注、以及向模块化系统导向方法的转变。

Comments 8 pages, appendices, GEM, ACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03771 2026-05-27 cs.CY cs.AI 81%

Trustworthiness of Legal Considerations for the Use of LLMs in Education

LLM在教育中使用的法律考量可信度

Sara Alaswad, Tatiana Kalganova, Wasan Awad

机构 * College of Information Technology(信息科技学院) Brunel University of London(伦敦布鲁内尔大学) Ahlia University(阿利亚大学)

专题命中 领域大模型 :LLM(title_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 本文通过比较全球主要地区(欧盟、英国、美国、中国、海湾合作委员会国家)的AI监管框架,提出针对海湾合作委员会国家的合规中心AI治理框架,以促进教育中AI系统的合法、伦理和文化适应性部署。

Comments 11 pages, 3 figures, 6 tables

Journal ref Proc. IEEE DASA 2025, Manama, Bahrain, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26589 2026-05-27 cs.LG cs.AI stat.ML 81%

Few-shot Cross-country Generalization of Tabular Machine Learning and Foundation Models for Childhood Anemia Prediction under Distribution Shift

分布漂移下儿童贫血预测的表格机器学习与基础模型的少样本跨国家泛化

Yusuf Brima, Marcellin Atemkeng, Lansana Hassim Kallon, David Niyukuri, Antoine Vacavant, Samuel Saidu, Ding-Geng Chen

机构 * Department of Mathematics, Rhodes University, South Africa(数学系,罗德斯大学,南非) National Institute for Theoretical and computational Sciences (NITheCS), Stellenbosch, 7600, South Africa(理论与计算科学国家研究所(NITheCS),斯泰伦博斯,7600,南非) Interdisciplinary Research Program in Public Health, University of Burundi, Burundi(公共卫生跨学科研究计划,布恩迪大学,布恩迪) Universite Clermont Auvergne, Clermont Auvergne INP, CNRS, Institut Pascal, Clermont–Ferrand, France(克莱蒙特-奥弗涅大学,克莱蒙特-奥弗涅INP,CNRS,帕西尔研究所,克莱蒙特-费尔南,法国) Department of International Public Health, Liverpool School of Tropical Medicine, Liverpool, UK(国际公共卫生系,利物浦热带医学学校,利物浦,英国) College of Health Solutions, Arizona State University, Phoenix, USA(健康解决方案学院,亚利桑那州立大学,凤凰城,美国) Department of Statistics, University of Pretoria, Pretoria, South Africa(统计系,普里特oria大学,普里特oria,南非)

专题命中 领域大模型 :foundation model(title,abstract);分类 cs.AI、cs.LG

AI总结 本研究评估了基于Transformer的表格基础模型TabPFN在跨国家、数据稀缺环境下预测儿童贫血的性能,发现其优于经典监督方法,尤其在低数据场景下表现出更好的区分度和校准能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26533 2026-05-27 cs.CV cs.AI cs.CL cs.LG 80%

A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection

一种用于工业检测中自动缺陷推理与报告生成的混合视觉-语言架构

Malikussaid, Imad Gohar

机构 * School of Computing, Telkom University(Telkom大学计算机学院) Faculty of Engineering and Technology, School of Computing and Artificial Intelligence(工程与技术学院,计算与人工智能学院)

专题命中 领域大模型 :LLM(abstract,abstract_cn);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出一种解耦的边缘可部署管道,结合YOLO26-x-obb检测器、确定性编码模块和QLoRA微调的Qwen-2.5-1.5B模型,实现风电叶片缺陷定位与结构化报告生成,在BLEU-4、幻觉率和专家评分上显著优于零样本VLM基线。

Comments 23 pages, 6 figures, 9 equations, and 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00944 2026-05-27 cs.CY 80%

Auditing the Reliability of Multimodal Generative Search

审计多模态生成式搜索的可靠性

Erfan Samieyan Sahneh, Luca Maria Aiello

专题命中 领域大模型 :LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 通过大规模审计Gemini 2.5 Pro系统,发现3.7%至18.7%的基于视频的生成式搜索声明未被引用来源支持,主要失败模式为不可验证的特异性与夸大声明。

Comments 14 pages, LaTeX, typos corrected, adding Data Availability section + examples in appendix, New plot(overview)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09886 2026-05-27 cs.CL 79%

Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze Surprisal

缩小差距:探究为何语言模型惊奇度优于完形填空惊奇度

Sathvik Nair, Byung-Doh Oh

机构 * University of Maryland(马里兰大学) Nanyang Technological University(南洋理工大学)

专题命中 领域大模型 :language model(title,abstract);分类 cs.CL

AI总结 本研究通过三个假设(低分辨率、语义相似词区分、低频词概率准确性)解释了语言模型概率在预测处理努力上优于完形填空数据的原因。

Comments 18 pages, 10 figures, accepted to ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18919 2026-05-27 cs.CV 78%

Advancing Metallic Surface Defect Detection via Anomaly-Guided Pretraining on a Large Industrial Dataset

通过在大规模工业数据集上的异常引导预训练推进金属表面缺陷检测

Chuni Liu, Hongjie Li, Jiaqi Du, Yangyang Hou, Qian Sun, Lei Jin, Ke Xu

机构 * Collaborative Innovation Center of Steel Technology, University of Science and Technology Beijing(钢铁技术协同创新中心,北京科技大学)

专题命中 领域大模型 :pretraining(title,abstract)

AI总结 提出异常引导自监督预训练(AGSSP)方法,通过两阶段框架利用异常先验引导表示学习,在金属表面缺陷检测中显著提升性能,mAP@0.5提升高达10%。

Comments Accepted for publication in Pattern Recognition

Journal ref Pattern Recognition, Volume 179, Part C, 2026, 113788

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27115 2026-05-27 cs.AI 77%

Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation

基于对抗感知的多教师同策略蒸馏以实现领域保留下的通用能力恢复

Tianlei Chen, Jiao Ou, Ziyuan Liu, Ruiming Tang, Jian Liang, Han Li

机构 * Kuaishou Technology, Beijing, China(快手科技,北京,中国)

专题命中 领域大模型 :LLM(abstract,abstract_cn);post-training(abstract);分类 cs.AI

AI总结 针对多教师同策略蒸馏在提示覆盖不完全时出现的恢复-保留对抗和弱信号平坦化问题,提出CaMOPD方法,通过解耦交替训练和基于差距的样本选择,在保持领域性能的同时有效恢复通用能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24219 2026-05-27 cs.AI 77%

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

超越最终答案:多智能体工业工作流中的轨迹级幻觉审计

Harshada Badave, Santosh Borse, Andrea Gomez, Harshitha Narahari, Sara Carter, Vishwa Bhatt, Aishani Rachakonda, Shuxin Lin, Dhaval Patel

机构 * IBM Columbia University(哥伦比亚大学)

专题命中 领域大模型 :LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出Trajel数据集和评估框架,通过五类幻觉分类法审计多智能体工业工作流中的轨迹级幻觉,发现现有基准忽略的常见失败模式,并证明轨迹感知检测优于标准事后验证。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01489 2026-05-27 cs.AI cs.CL 73%

SciResearcher: Scaling Deep Research Agents for Frontier Scientific Reasoning

SciResearcher: 面向前沿科学推理的深度研究智能体规模化

Tianshi Zheng, Rui Wang, Xiyun Li, Kelvin Kiu Wai Tam, Newt Nguyen Kim Hue Nam, Wei Fan, Yangqiu Song, Tianqing Fang

机构 * HKUST(香港科技大学) CUHK(香港大学) Tencent AI Lab(腾讯AI实验室)

专题命中 领域大模型 :foundation model(abstract);post-training(abstract);分类 cs.CL、cs.AI

AI总结 提出SciResearcher框架,通过合成基于学术证据的概念与计算任务并训练智能体,在HLE-Bio/Chem-Gold等基准上达到最优性能。

Comments 23 pages, 6 figures, 15 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27045 2026-05-27 cs.CL 70%

ExTax: Explainable Disinformation Detection via Persuasion, Emotion, and Narrative Role Taxonomies

ExTax:基于说服、情感和叙事角色分类学的可解释虚假信息检测

Shang Luo, Yingguang Yang, Zhenchen Sun, Yang Liu, Bin Chong, Jingru Chen, Yancheng Chen, Jiayu Liang, Kefu Xu, Hao Peng, Philip S. Yu

机构 * Peking University(北京大学) University of Science and Technology of China(中国科学技术大学) North China University of Science and Technology(华北理工大学) Tsinghua University(清华大学) Nanjing University of Aeronautics and Astronautics(南京航空航天大学) University of Chinese Academy of Sciences(中国科学院大学) Soochow University(苏州大学) Beihang University(北航) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

专题命中 领域大模型 :LLM(abstract,abstract_cn);分类 cs.CL

AI总结 提出ExTax框架,统一说服修辞、情感操纵和叙事角色为17维分类空间,通过熵驱动动态标签平滑和多头注意力融合分类与上下文特征,实现可解释的虚假信息检测,在跨域基准上达到0.8456 Macro F1。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20291 2026-05-27 cs.LG 70%

Weasel: Out-of-Domain Generalization for Web Agents via Importance-Diversity Data Selection

Weasel: 通过重要性-多样性数据选择实现Web智能体的域外泛化

Fatemeh Pesaran Zadeh, Seyeon Choi, Xing Han Lù, Siva Reddy, Gunhee Kim

机构 * Seoul National University(首尔国立大学) McGill University(麦吉尔大学) Mila -- Quebec AI Institute(蒙特利尔AI研究所) Canada CIFAR AI Chair(加拿大CIFAR人工智能主席)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 提出Weasel方法,通过优化平衡单步重要性与状态、网站、交互模式成对多样性的目标,选择固定预算的轨迹子集,结合目标中心AXTree剪枝和风格一致理由替换,提升Web智能体离线训练的域外泛化性能并降低训练成本。

Comments ICML 2026. Code is released at https://github.com/fatemehpesaran310/weasel

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08267 2026-05-27 cs.CL 70%

Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning

Med-CoReasoner: 通过语言感知的协同推理减少医学推理中的语言差异

Fan Gao, Sherry T. Tong, Jiwoong Sohn, Jiahao Huang, Junfeng Jiang, Ding Xia, Piyalitt Ittichaiwong, Kanyakorn Veerakanjana, Hyunjae Kim, Qingyu Chen, Edison Marrese Taylor, Kazuma Kobayashi, Akiko Aizawa, Irene Li

机构 * The University of Tokyo(东京大学) ETH Zürich(苏黎世联邦理工学院) National Institute of Informatics(日本信息处理学会) Siriraj Informatics and Data Innovation Center(Siriraj信息与数据创新中心) Yale University(耶鲁大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 提出Med-CoReasoner框架,通过并行英语和本地语言推理、结构化概念抽象及概念级对齐与检索,将本地临床知识整合到英语逻辑框架中,以缩小医学推理中的多语言差距,在MultiMed-X基准上平均提升5%的多语言推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08600 2026-05-27 cs.CL 70%

NSF-SciFy: Mining the NSF Awards Database for Scientific Claims

NSF-SciFy:从NSF资助数据库中挖掘科学主张

Delip Rao, Weiqiu You, Eric Wong, Chris Callison-Burch

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 领域大模型 :language model(abstract);prompting(abstract);分类 cs.CL

AI总结 提出NSF-SciFy数据集,从40万篇NSF摘要中提取280万条科学主张,通过零样本提示联合提取主张和研究提案,并在三个下游任务中微调语言模型取得显著提升。

Comments ACL 2026. 19 pages, 7 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27332 2026-05-27 cs.SE cs.AI cs.CV 57%

EdgeFlow: Edge-Map Augmented VLM-Based Flowchart Processing for Industrial Requirements Engineering

EdgeFlow: 基于边缘图增强的VLM流程图处理用于工业需求工程

Zhifei Dou, Shabnam Hassani, Ou Wei

机构 * Huawei Research Canada(华为加拿大研究)

专题命中 领域大模型 :language model(abstract);分类 cs.AI

AI总结 提出EdgeFlow方法,通过向视觉语言模型(VLM)输入添加Canny边缘图作为结构先验,无需训练数据或微调即可提升流程图到Mermaid代码的转换精度,在工业数据集上节点F1提升17.39%,边F1提升16.94%。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27131 2026-05-27 cs.ET cs.AI cs.DB 57%

Beyond the Data Mesh Illusion: Designing Modern AI-augmented Lakehouses to Bridge the Gap Between Theory and Practice

超越数据网格幻象:设计现代AI增强型湖仓以弥合理论与实践差距

Oliver Angélil, Jan Migon

机构 * ishango.ai Zurich, Switzerland(ishango.ai 瑞士苏黎世) Independent Researcher(独立研究者)

专题命中 领域大模型 :LLM(abstract_cn);分类 cs.AI

AI总结 针对企业数据平台中领域自服务与整体治理之间的张力,提出一种基于现代湖仓架构的AI增强型中心辐射模型,通过中心卓越中心提供共享服务与AI治理,领域团队逐步承担更多责任,以平衡灵活性与控制,并通过数据产品采纳率、查找时间和洞察时间三个指标评估架构效果。

Comments 11 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21602 2026-05-27 cs.LG cs.CV 57%

An Empirical Study of Machine Learning Robustness and Scalability for Imbalanced Tabular Clinical Data in Emergency and Critical Care

机器学习在急诊和重症监护中不平衡表格临床数据的鲁棒性与可扩展性实证研究

Yusuf Brima, Marcellin Atemkeng

机构 * Computer Vision Group, Institute of Cognitive Science, Osnabrück University(计算机视觉组,认知科学研究所,奥斯纳布吕克大学) Department of Mathematics, Rhodes University(数学系,罗德斯大学) National Institute for Theoretical and Computational Sciences (NITheCS)(国家理论与计算科学研究所(NITheCS))

专题命中 领域大模型 :foundation model(abstract);分类 cs.LG

AI总结 本研究在MIMIC-IV-ED和eICU数据集上评估六类模型在不平衡临床表格数据上的性能,发现树模型在可扩展性上最优,而表格基础模型在性能与效率间提供新的权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15891 2026-05-27 cs.CV 50%

RadJEPA: Radiology Encoder for Chest X-Rays via Joint Embedding Predictive Architecture

RadJEPA:基于联合嵌入预测架构的胸部X光放射学编码器

Anas Anwarul Haq Khan, Mariam Husain, Pratik Jalan, Kshitij Jadhav

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Bombay(印度理工学院孟买分校计算机科学与工程系) Department of Biomedical Engineering, Johns Hopkins University(约翰霍普金斯大学生物医学工程系) Koita Centre for Digital Health, Indian Institute of Technology Bombay(印度理工学院孟买分校Koita数字健康中心)

专题命中 领域大模型 :pretraining(abstract)

AI总结 提出RadJEPA,一种无需语言监督的自监督框架,通过联合嵌入预测架构在约84万张无标签胸部X光图像上预训练,学习预测掩码区域的潜在表示,在放射学报告生成等任务中达到或超越现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19602 2026-05-27 cs.CV 50%

No Data? No Problem: Robust Vision-Tabular Learning with Missing Values

无数据?没问题:面向缺失值的鲁棒视觉-表格学习

Marta Hasny, Laura Daza, Keno Bressem, Maxime Di Folco, Julia Schnabel

机构 * School of Computation, Information and Technology, Technical University of Munich, Germany(计算、信息与技术学院,慕尼黑技术大学,德国) Institute of Machine Learning in Biomedical Imaging, Helmholtz Munich, Germany(生物医学成像中的机器学习研究所,海德堡慕尼黑,德国) School of Biomedical Engineering and Imaging Sciences, King’s College London, UK(生物医学工程与成像科学学院,伦敦国王学院,英国) Department of Diagnostic and Interventional Radiology, TUM University Hospital, Technical University of Munich, Germany(诊断与介入放射科,慕尼黑技术大学医院,德国) Munich Center for Machine Learning, Germany(慕尼黑机器学习中心,德国)

专题命中 领域大模型 :pretraining(abstract)

AI总结 提出RoVTL框架,通过对比预训练中的表格属性缺失增强和下游任务中的Tabular More vs. Fewer损失,实现从0%到100%表格数据可用性下的鲁棒多模态学习。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 知识编辑与模型理解 25 篇

2605.27288 2026-05-27 cs.CL cs.AI cs.LG 92%

It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty

并非总是谄媚:基于认知不确定性测量LLM的从众行为

Kevin H. Guo, Chao Yan, Avinash Baidya, Katherine Brown, Xiang Gao, Juming Xiong, Zhijun Yin, Bradley A. Malin

机构 * Vanderbilt University(范德比尔特大学) Vanderbilt University Medical Center(范德比尔特大学医学中心) Intuit AI Research(Intuit AI研究院)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出MUSE框架,通过区分谄媚从众和不确定性驱动的从众,揭示LLM在用户反驳时改变立场的行为机制,并发现两种从众均随用户感知专业性和建议合理性增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27016 2026-05-27 cs.CL cs.AI cs.LG stat.ML 92%

Evaluating the Relevance of Uncertainty Estimators for LLM Hallucination

评估不确定性估计器与LLM幻觉的相关性

Yedidia Agnimo, Anna Korba, Annabelle Blangero, Nicolas Chesneau, Karteek Alahari

机构 * CREST, ENSAE Institut Polytechnique de Paris(CREST,巴黎高等理工学院) Ekimetrics France(法国Ekimetrics) Centre Inria de l’Université Grenoble Alpes(格勒诺布尔阿尔卑斯大学信息研究院)

专题命中 知识编辑与模型理解 :LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 通过系统实证研究,评估信息论、基于采样和反思性等不确定性估计器与LLM幻觉之间的关联,发现关联性高度可变且通常较弱,挑战了将不确定性作为幻觉直接信号的做法。

Comments 35 pages, 7 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04931 2026-05-27 cs.LG cs.AI 91%

Emergent Causal-Geometric Dynamics Across Depth in Large Language Models

大型语言模型中跨深度的涌现因果几何动力学

Shahar Haim, Daniel C McNamee

机构 * Champalimaud Centre for the Unknown(查普拉米乌德未知中心)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 通过结合几何分析与因果干预,揭示了解码器-only大型语言模型中从上下文处理到预测形成的跨层转变,并发现后期层中角度结构参数化下一词分布相似性并实现选择性因果控制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26332 2026-05-27 cs.CV cs.AI 89%

Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models

被擦除但可被利用:针对已遗忘文本到图像扩散模型的黑盒嵌入感知提示攻击

Arian Komaei Koma, Seyed Amir Kasaei, AmirMahdi Sadeghzadeh, Mohammad Hossein Rohban

机构 * Department of Computer Engineering(计算机工程系)

专题命中 知识编辑与模型理解 :prompting(title,abstract);LLM(abstract,abstract_cn);large language model(abstract);language model(abstract)

AI总结 提出一种黑盒嵌入感知对抗提示攻击BEAP,利用大语言模型迭代生成有效对抗提示,以恢复被遗忘概念,并在攻击成功率上提升超过60%。

详情

展开后加载摘要…

URL PDF HTML 收藏