arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

语言大模型 / LLM

大语言模型、预训练、指令微调、后训练和语言模型应用。

2026-07-07 至 2026-07-07 共收录 710 信号源:cs.CL, cs.AI, cs.LG

1. 领域大模型 57 篇

2604.11730 2026-07-07 cs.CV cs.HC cs.LG 版本更新 70%

Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions

视频中矛盾/犹豫识别用于个性化数字健康干预

Manuela González-González, Soufiane Belharbi, Muhammad Osama Zeeshan, Masoumeh Sharafi, Muhammad Haseeb Aslam, Lorenzo Sia, Nicolas Richet, Marco Pedersoli, Alessandro Lameiras Koerich, Simon L Bacon, Eric Granger

机构 * LIVIA, Dept. of Systems Engineering, ETS Montreal, Canada(ETS蒙特利尔大学系统工程系LIVIA实验室) LIVIA, Dept. of Software and IT Engineering, ETS Montreal, Canada(ETS蒙特利尔大学软件与信息工程系LIVIA实验室) Dept. of Health, Kinesiology, & Applied Physiology, Concordia University, Montreal, Canada(康科迪亚大学健康、运动科学与应用生理学系) Montreal Behavioural Medicine Centre, CIUSSS Nord-de-l’Ile-de-Montréal, Canada(蒙特利尔行为医学中心,蒙特利尔北岛卫生与社会服务局)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 本文研究了通过深度学习模型在视频中进行矛盾/犹豫识别,以提升数字健康干预的个性化和成本效益,实验基于新的BAH视频数据集,发现需改进多模态模型以准确识别矛盾/犹豫。

Comments 11 pages, 4 figures, ACII 2026. arXiv admin note: substantial text overlap with arXiv:2505.19328

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07285 2026-07-07 cs.CL cs.CY 版本更新 70%

Why teaching resists automation in an AI-inundated era: Human judgment, non-modular work, and the limits of delegation

为何在人工智能泛滥的时代教学抵抗自动化:人类判断、非模块化工作与委托的局限

Songhee Han

机构 * Florida State University(佛罗里达州立大学)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文探讨了在人工智能普及背景下,教学工作难以自动化的原因,指出教学本质上具有解释性、关联性和专业判断,无法被完全自动化或委托给技术。

Comments Revised version; accepted for publication in TechTrends

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08029 2026-07-07 cs.LG cs.CV 版本更新 70%

CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space

CLARITY:用于通过在潜在空间中建模情境感知疾病轨迹来指导治疗决策的医学世界模型

Tianxingjian Ding, Yuanhao Zou, Chen Chen, Mubarak Shah, Yu Tian

机构 * Institute of Artificial Intelligence, University of Central Florida(中佛罗里达大学人工智能研究所)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.LG

AI总结 CLARITY通过在潜在空间中建模时间间隔和患者特定数据,预测疾病演变轨迹,生成个性化治疗方案,并在医学领域实现最先进的治疗规划性能。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22869 2026-07-07 cs.CL 70%

JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge

JBE-QA:日本司法考试问答数据集用于评估法律领域知识

Zhihan Cao, Fumihito Nishino, Hiroaki Yamada, Nguyen Ha Thanh, Yusuke Miyao, Ken Satoh

机构 * Institute of Science Tokyo, Japan(东京科学研究所,日本) ROIS-DS, Center for Juris-informatics, Japan(司法信息中心,日本) National Institute of Informatics, Japan(日本信息机构) University of Tokyo, Japan(东京大学,日本)

专题命中 领域大模型 :large language model(abstract);language model(abstract);分类 cs.CL

AI总结 本文提出JBE-QA数据集,用于评估大语言模型的法律知识,涵盖民法、刑法和宪法,通过分解问题为独立判断并添加上下文字段,共包含3464个平衡标注问题,评估26种LLM表现,显示具备推理能力的专有模型表现最佳。

Comments Three tables and one figure

Journal ref In Proc. of the 15th Language Resources and Evaluation Conference (LREC 2026) (pp. 5317-5327)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08377 2026-07-07 cs.CV 版本更新 67%

UniVideo: Unified Understanding, Generation, and Editing for Videos

UniVideo:视频的统一理解、生成与编辑

Cong Wei, Quande Liu, Zixuan Ye, Qiulin Wang, Xintao Wang, Pengfei Wan, Kun Gai, Wenhu Chen

机构 * University of Waterloo(多伦多大学) Kling Team, Kuaishou Technology(快手技术团队)

专题命中 领域大模型 :large language model(abstract);language model(abstract)

AI总结 提出UniVideo框架,采用双流设计,结合多模态大语言模型与多模态DiT进行视频生成等任务。统一多种视频任务于单范式,实验表明其性能出色,还支持任务组合与能力迁移,且发布了模型和代码。

Comments Project Website https://congwei1230.github.io/UniVideo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18176 2026-07-07 cs.CV 版本更新 67%

Atlas is Your Perfect Context: One-Shot Customization for Generalizable Foundational Medical Image Segmentation

图谱是理想上下文:用于可泛化基础医学图像分割的一次性定制

Ziyu Zhang, Yi Yu, Simeng Zhu, Ahmed Aly, Yunhe Gao, Ning Gu, Yuan Xue

机构 * Nanjing University(南京大学) The Ohio State University(俄亥俄州立大学) Stanford University(斯坦福大学)

专题命中 领域大模型 :foundation model(abstract);pretraining(abstract)

AI总结 研究医学图像中解剖结构分割问题,提出图谱引导框架AtlasSegFM,通过图谱查询配准生成提示,用冻结基础模型细化分割,经自适应融合模块结合图谱先验与模型输入及预测,实现一次性定制。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10448 2026-07-07 cs.IR cond-mat.mtrl-sci 版本更新 67%

MatSKRAFT: A framework for large-scale materials knowledge extraction from scientific tables

MatSKRAFT:一个用于从科学表格中大规模提取材料知识的框架

Kausik Hira, Mohd Zaki, Mausam, N. M. Anoop Krishnan

专题命中 领域大模型 :large language model(abstract);language model(abstract)

AI总结 介绍MatSKRAFT框架,能从表格数据中自动提取整合材料科学知识。核心方法是转化表格为基于图的表示,用约束驱动GNN处理。贡献是性能超当代前沿模型,构建含大量条目的综合数据库,揭示新关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20052 2026-07-07 cs.CL cs.AI 62%

PromptRad: Knowledge-Enhanced Multi-Label Prompt-Tuning for Low-Resource Radiology Report Labeling

PromptRad: 基于知识的多标签提示微调用于低资源放射报告标注

Ying-Jia Lin, Tzu-Chin Lo, Ping-Chien Li, Chi-Tung Cheng, Chien-Hung Liao, Hung-Yu Kao

机构 * Department of Artificial Intelligence and AI Research Center, Chang Gung University(人工智能系及AI研究中心,长庚大学) Department of Radiology, Sijhih Cathay General Hospital(放射科,西吉医院) Department of Medical Imaging and Intervention, Chang Gung Memorial Hospital(医学影像与介入科,长庚纪念医院) Department of Trauma and Emergency Surgery, Chang Gung Memorial Hospital(创伤与急诊外科,长庚纪念医院) Department of Computer Science, National Tsing Hua University(计算机科学系,国立清华大学)

专题命中 领域大模型 :language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出PromptRad,一种基于知识的多标签提示微调方法,用于在低资源环境下进行放射报告标注,通过引入UMLS元词典中的同义词增强类别表示,以更少的标注数据实现优于传统方法的性能。

Comments BioNLP 2026 @ ACL (camera-ready version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00597 2026-07-07 cs.CL cs.IR 新提交 57%

Multi-Turn Agentic Scientific Literature Search via Workflow Induction

通过工作流归纳实现多轮智能科学文献搜索

Jisen Li, Bingxuan Li, Nanyi Jiang, Xuying Ning, Xiyao Wang, Yifan Shen, Heng Wang, Yuqing Jian, Xiaoxia Wu, Ben Athiwaratkun, Pan Lu, Jiaxuan You, Bingxin Zhao

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Together AI University of Pennsylvania(宾夕法尼亚大学) Stanford University(斯坦福大学)

专题命中 领域大模型 :preference optimization(abstract);分类 cs.CL

AI总结 提出PaperPilot,通过构建可执行DAG工作流进行多轮文献搜索,利用用户反馈优化查询和工作流,显著提升搜索性能并消除执行错误。

Comments 17 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12502 2026-07-07 physics.soc-ph cs.AI 新提交 57%

A Mathematical Theory of Value: a synthesis on goal-directed agency under resource constraints

价值的数学理论:资源约束下目标导向行为的综合

Cheng Qian

机构 * Cheng Qian(陈倩)

专题命中 领域大模型 :language model(abstract);分类 cs.AI

AI总结 本文提出价值是目标导向主体在资源约束下转化资源为目标进度的速率,通过尺度不变性公理导出对数度量,并推导出价值编码定理,实现价值与信息论的统一。

Comments Also available at https://doi.org/10.5281/zenodo.20487041 (v6). v2: the pre-registered continuation-gate experiment has been run -- prediction (i) confirmed on real frontier agents; prediction (ii) retired to mathematical scope. Code, pre-registrations, and raw run records: https://github.com/macrokit/value

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22858 2026-07-07 cs.CL cs.IR 57%

RAG System for Supporting Japanese Litigation Procedures: Faithful Response Generation Complying with Legal Norms

基于RAG的日本诉讼程序支持系统:符合法律规范的忠实响应生成

Yuya Ishihara, Atsushi Keyaki, Hiroaki Yamada, Ryutaro Ohara, Mihoko Sumida

机构 * Hitotsubashi University, Japan(早稻田大学,日本) Institute of Science Tokyo, Japan(东京科学研究所,日本) Nakamura, Tsunoda & Matsumoto, Japan(日本纳卡拉、藤田与松本事务所)

专题命中 领域大模型 :LLM(abstract);分类 cs.CL

AI总结 本文设计了一种基于RAG的LLM系统,用于支持日本医疗诉讼程序,确保响应符合法律规范,通过检索模块检索相关外部知识并保持响应的忠实性。

Comments This is a preprint version of a paper reviewed and accepted at BREV-RAG 2025: Beyond Relevance-based EValuation of RAG Systems, a SIGIR-AP 2025 workshop

Journal ref Proceedings of the International Workshop on Beyond Relevance-based EValuation of RAG Systems (BREV-RAG) 2025 (pp. 59-65)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05222 2026-07-07 cs.CV 新提交 50%

A Multimodal Reasoning Typology for Grounding Chart-Image Coherence in Science Communication

面向科学传播中图表-图像连贯性锚定的多模态推理类型学

Avina Nakarmi, Sohom Sen, Xun Song, Sreyashi Samaddar, Aritra Dasgupta

机构 * New Jersey Institute of Technology(新泽西理工学院) Brooklyn College(布鲁克林学院)

专题命中 领域大模型 :language model(abstract)

AI总结 本研究提出R1-R5多模态推理类型学,刻画科学文献中图表、图像与文本的协同推理缺口,可系统识别连贯性锚定状态,缩小不同人群解读科学结论的认知差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04694 2026-07-07 cs.CV 新提交 50%

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

解决缺失的第一步:视觉语言模型能否标准化原始异构医学数据?

Xin Chen, Dongliang Xu, Cunhao Zhu, Xudong Luo, Haoyang Lyu, Xiaoxiao Sun, Serena Yeung-Levy, Yue Yao

机构 * Shandong University(山东大学) Stanford University(斯坦福大学)

专题命中 领域大模型 :language model(abstract)

AI总结 研究视觉语言模型应用于医学AI时原始医学数据标准化问题,通过构建基准测试,让模型处理原始数据集文件夹,评估其多种能力,发现即使最佳模型端到端成功率也低,凸显该环节是医学AI诊断关键瓶颈。

Comments 16 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04249 2026-07-07 cs.CV 新提交 50%

Beyond Random Sampling: Distribution-Aware Alignment for Semi-Supervised Medical Image Segmentation

超越随机采样:用于半监督医学图像分割的分布感知对齐

Weihao Yan, Yeqiang Qian, Yi Dong, Ming Yang

机构 * School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院) Department of Ultrasound, Xinhua Hospital Affiliated to Shanghai Jiao Tong University School of Medicine(上海交通大学医学院附属新华医院超声科)

专题命中 领域大模型 :foundation model(abstract)

AI总结 研究针对半监督医学图像分割中随机采样策略在低数据量时表征偏差的问题,提出基于分布对齐的高效框架,用分布感知样本选择策略和记忆引导复制粘贴模块,提升分割性能。

Comments 19 pages, 5 figures, accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02571 2026-07-07 cs.CV 新提交 50%

Dual-Adaptive SAM3: Hierarchical Routing over Low-Rank Expert Layers for Parameter-Efficient Medical Image Segmentation

双自适应SAM3:基于低秩专家层的分层路由用于参数高效的医学图像分割

Ying Chen, Jinyue Li, Kun Wang, Qiankun Li, Yang Liu

机构 * Shenzhen Research Institute, The Chinese University of Hong Kong(香港中文大学深圳研究院) University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学) Imperial Global Singapore (IGS), Imperial College London(伦敦帝国理工学院新加坡帝国全球(IGS))

专题命中 领域大模型 :language model(abstract)

AI总结 提出双自适应SAM3框架,通过任务感知的动态专家路由器和参数感知的分解参数化专家设计,兼顾分割精度与参数效率,在医学图像分割上有高表现。

Comments Accepted by MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02617 2026-07-07 cs.SE 版本更新 50%

His2Trans: A Knowledge-Guided Agentic Framework for Project-Level C-to-Rust Migration

基于构建意识的增量C到Rust迁移:通过骨架优先翻译和历史知识重用

Shengbo Wang, Mingwei Liu, Guangsheng Ou, Yuwen Chen, Zike Li, Yanlin Wang

专题命中 领域大模型 :LLM(abstract)

AI总结 本文提出His2Trans框架,通过构建意识的骨架构建和历史知识重用,实现C到Rust的增量迁移,提升构建可行性并减少警告数量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15330 2026-07-07 cs.CV 版本更新 50%

CoDoL: Conditional Domain Prompt Learning for Out-of-Distribution Generalization

CoDoL:用于分布外泛化的条件域提示学习

Min Zhang, Yuyin Wang, Zhongxiang Dai, Zhikang Chen, Jie Zhou, Miao Liu, Sen Cui

机构 * East China Normal University(东华师范大学) Xidian University(西安电子科技大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) The University of Oxford(牛津大学) Tsinghua University(清华大学)

专题命中 领域大模型 :language model(abstract)

AI总结 针对基于提示的CLIP方法存在的文本描述不准确、视觉语言嵌入对齐有限问题,提出CoDoL方法,利用域信息形成提示,还提出DMN生成输入条件令牌,实验验证其在分布外泛化的有效性。

Comments Accepted to TMLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 知识编辑与模型理解 43 篇

2506.07406 2026-07-07 cs.LG cs.AI 版本更新 91%

InverseScope: Scalable Activation Inversion for Interpreting Large Language Models

InverseScope:用于解释大语言模型的可扩展激活反演

Yifan Luo, Zhennan Zhou, Bin Dong

机构 * School of Mathematical Science, Peking University(北京大学数学科学学院) School of Science, Westlake University(西湖大学理学院) Beijing International Center for Mathematical Research(北京国际数学研究中心) New Cornerstone Science Laboratory, Peking University(北京大学新基石科学实验室)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(abstract,abstract_cn);分类 cs.AI、cs.LG

AI总结 研究大语言模型内部表征解释难题,提出InverseScope框架,通过输入反演解释神经激活,用新颖架构提高采样效率,揭示模型表征空间结构,可扩展到14B参数模型并推广到分布外输入。

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02923 2026-07-07 cs.CL cs.AI 版本更新 90%

Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias

理事会模式:一种异质多智能体共识框架,用于减少大语言模型的幻觉和偏见

Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, Ran Wang

机构 * Lead Researcher(研究员) Research Assistant(研究员) Academic Advisor(学术顾问) Research Consultant(研究顾问)

专题命中 知识编辑与模型理解 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI

AI总结 本文提出理事会模式,通过多智能体共识框架减少LLM的幻觉和偏见,实验显示在HaluEval和TruthfulQA上 hallucination率降低35.9%,质量得分提升10.2个百分点,且在偏见方面表现更优。

Comments 24 pages, 8 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03350 2026-07-07 cs.CR cs.AI cs.SE 新提交 90%

LLM-Enhanced Hierarchical Heterogeneous Graph Representation Learning for Malicious Python Package Detection

用于恶意Python包检测的基于大语言模型增强的分层异构图表示学习

Hang Gao, Xiaoyu Chen, Baoquan Cui, Zhen Tang, Peng Qiao, Fengge Wu, Jian Zhang

机构 * Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 知识编辑与模型理解 :LLM(title,summary_cn);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 提出用于恶意Python包检测的LLM增强分层异构图表示学习框架,构建分层异构代码图,利用LLMs推理功能语义角色,开发分层异构图神经网络并结合功能级归因机制,实验表明该框架性能优越。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18895 2026-07-07 q-fin.RM cs.LG 89%

Could Large Language Models work as Post-hoc Explainability Tools in Credit Risk Models?

大语言模型能否在信用风险模型中作为事后可解释性工具?

Wenxi Geng, Dingyuan Liu, Liya Li, Yiqing Wang

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Georgia Institute of Technology(佐治亚理工学院) Southern Methodist University(南方 Methodist 大学)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);LLM(abstract);分类 cs.LG

AI总结 本文研究大语言模型是否能作为信用风险模型的事后可解释性接口,评估其在保持特征重要性排名和生成自主解释方面的能力。

Comments 39 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04430 2026-07-07 cs.CL 新提交 88%

Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees

具有可证明对齐保证的大语言模型中的不确定性感知弃权

Sijin Dong, Hiroyuki Shinnou

机构 * Ibaraki University(茨城大学)

专题命中 知识编辑与模型理解 :large language model(title,abstract);language model(title,abstract);分类 cs.CL

AI总结 研究大语言模型在问答系统中无可靠置信估计的问题,提出基于置信区间的校准框架CIC,将任意不确定性分数转化为风险可控的选择性回答规则,保证所选阈值控制错误率,兼具风险控制和回答效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02686 2026-07-07 cs.AI cs.LG 新提交 88%

ASK in the Dark: Uncertainty-Gated LLM Assistance under Partial Observability

黑暗中的询问:部分可观测性下的不确定性门控语言模型辅助

Juarez Monteiro, Nathan Gavenski, Guilherme Lima, Francisco Galuppo, Odinaldo Rodrigues, Adriano Veloso

机构 * Kunumi Institute(库努米研究所) King’s College London(伦敦国王学院)

专题命中 知识编辑与模型理解 :LLM(title);SLM(abstract,abstract_cn);language model(abstract);small language model(abstract)

AI总结 研究部分可观测性下强化学习智能体,指出普通不确定性门控方法失败原因是上下文问题。提出ASK+,为语言模型提供轨迹感知上下文和结构化推理,证明预测熵信号可行,在多环境中取得更好效果,强调提示设计重要性。

Comments Accepted at the IJCAI-ECAI Joint Workshop on Planning for Complex Real-World Applications and Bridging the Gap Between AI Planning and (Reinforcement) Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15334 2026-07-07 cs.CL cs.AI 版本更新 87%

No Reliable Evidence of Self-Reported Sentience in Small Large Language Models

小语言模型中自我报告的感知能力缺乏可靠证据

Caspar Kaiser, Sean Enderby

机构 * University of Warwick, Warwick Business School(沃里克大学,沃里克商学院)

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(title);分类 cs.CL、cs.AI

AI总结 通过询问多个开源语言模型关于其自身意识的问题,并利用基于内部激活训练的分类器验证回答,对语言模型是否认为自己有感知能力进行测试,发现模型否认有感知,分类器也无反例,且Qwen模型中越大越确定。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28823 2026-07-07 cs.CL 版本更新 86%

What are They Thinking? Delineation, Probing, and Tracking of Concepts in LLMs

他们在想什么?LLM中概念的界定、探测与追踪

Mohamed Abdelwahab, Michelle Yu Collins, Sihan Chen, Yi Cheng Zhao, Zafarullah Mahmood, Jiading Zhu, Soliman Ali, Jonathan Rose

机构 * The Edward S. Rogers Sr. Department of Electrical and Computer Engineering(埃德华·S·罗杰斯 Sr. 部门电子与计算机工程系)

专题命中 知识编辑与模型理解 :LLM(title_cn,summary_cn);分类 cs.CL

AI总结 本文提出通过线性探针低成本地检测LLM嵌入中的概念,并展示了概念界定、探针训练与跨上下文追踪的方法,为大规模模型监控奠定基础。

Comments Accepted to the 6th Workshop on Trustworthy Natural Language Processing (TrustNLP 2026)

Journal ref Proceedings of the 6th Workshop on Trustworthy NLP (TrustNLP 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15903 2026-07-07 cs.LG cs.AI 版本更新 86%

Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor

迈向减轻基于大语言模型的智能体幻觉:渐进泛化边界探索与看门狗监测

Siyuan Liu, Wenjing Liu, Zhiwei Xu, Xin Wang, Bo Chen, Tao Li

机构 * College of Computer Science, Nankai University(南开大学计算机科学学院) Haihe Lab of ITAI College of intelligent Science and Technology, Inner Mongolia University of Technology(内蒙古工业大学智能科学与技术学院海河实验室) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Department of Electrical and Computer Engineering, Stony Brook University(石溪大学电气与计算机工程系) Department of Computer Science, Michigan Technological University(密歇根技术大学计算机科学系)

专题命中 知识编辑与模型理解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG

AI总结 研究针对大语言模型赋能智能体的幻觉问题,提出HalMit黑箱框架,用概率分形采样技术并行生成查询触发异常响应,以识别泛化边界来检测幻觉,提升智能体可靠性。

Journal ref Proceedings of the 28th European Conference on Artificial Intelligence (ECAI 2025), IOS Press, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04963 2026-07-07 cs.AI 新提交 85%

STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training

STAPO:用于大语言模型智能体训练的选择性轨迹感知策略优化

Qiuyi Qi, Tian Liang, Mutian Bao, Jinjian Zhang, Dongnan Liu, Wei Zhou, Linjian Mo, Ming Kong, Jie Liu, Feng Zhang, Qiang Zhu

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团) City University of Hong Kong(香港城市大学) College of Artificial Intelligence, Shanghai Institute for Advanced Study, Zhejiang University(浙江大学上海高等研究院人工智能学院) School of Earth Sciences, Zhejiang University(浙江大学地球科学学院) Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系)

专题命中 知识编辑与模型理解 :LLM(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI

AI总结 针对强化学习训练大语言模型智能体时轨迹忽视问题,提出归一化熵,引入STAPO框架,利用其定位异常步骤,通过联合机制优化,提升轨迹感知并保持训练稳定性,实验验证其有效性和鲁棒性。

Comments ACL 2026 MainConference

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06269 2026-07-07 cs.LG cs.AI 版本更新 85%

Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models

利用机制可解释性对大语言模型进行对抗攻击

Thomas Winninger, Boussad Addad, Katarzyna Kapusta

机构 * Thales(泰勒斯)

专题命中 知识编辑与模型理解 :large language model(title);language model(title);分类 cs.AI、cs.LG

AI总结 研究利用机制可解释性对大语言模型进行对抗攻击,通过识别接受子空间,用梯度优化使嵌入重新路由,降低计算成本,在多种模型上快速实现高成功率攻击,为攻防研究开辟新方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.04525 2026-07-07 cs.CL cs.AI 新提交 84%

Language Models Represent and Transform Concepts with Shared Geometry

语言模型通过共享几何表示和转换概念

Zhimin Hu, Lanhao Niu, Sashank Varma

专题命中 知识编辑与模型理解 :language model(title,abstract);large language model(abstract);分类 cs.CL、cs.AI

AI总结 研究神经网络中概念表示问题,将概念表示形式化为点云流形、上下文转换为向量场,在大语言模型中实例化该框架,发现模型共享概念表示及上下文转换的通用几何结构。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.03657 2026-07-07 cs.CV cs.AI 新提交 83%

ViPo-MLLM: Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation

ViPo-MLLM:用于无注释手语翻译的视觉-姿势多模态大语言模型

Ahmed Abul Hasanaath, Bicheng Xu, Mir Rayat Imtiaz Hossain, Leonid Sigal, Hamzah Luqman

机构 * King Fahd University of Petroleum and Minerals(国王法赫德石油矿物大学) University of British Columbia(不列颠哥伦比亚大学)

专题命中 知识编辑与模型理解 :LLM(title,abstract);language model(abstract);分类 cs.AI

AI总结 研究尝试无注释手语翻译,提出ViPo-MLLM框架结合时空RGB和人体姿势特征,用专用编码器与交叉模态注意力,经结构化提示和训练后由大语言模型处理。该模型在数据集取得新成果,验证了机制有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏