arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

共收录 149
2601.22495 2026-06-17 cs.LG 版本更新

Gradual Fine-Tuning for Flow Matching Models

流匹配模型的渐进微调

Gudrun Thorkelsdottir, Arindam Banerjee

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 提出渐进微调(GFT)框架,通过退火策略在目标分布样本下微调流生成模型,理论保证逼近真实目标,实验表明稳定性、效率与多样性优于现有方法。

Comments Preprint. Added methodology and experimental sections

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10137 2026-06-17 cs.LG math.DG stat.CO stat.ML 版本更新

Variational autoencoders with latent high-dimensional steady geometric flows for dynamics

具有潜在高维稳态几何流的变分自编码器用于动力学

Andrew Gracyk

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 提出VAE-DLM方法,在潜在空间中引入稳态几何流,通过物理信息方法求解高维流,增强潜在表示的表达能力,在PDE型数据上降低OOD误差15%-35%。

Comments Edits and improved tables

Journal ref 23rd International Conference of Numerical Analysis and Applied Mathematics (ICNAAM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15210 2026-06-17 cs.SD cs.AI cs.LG 版本更新

Explicit Context-Driven Neural Acoustic Modeling for High-Fidelity RIR Generation

显式上下文驱动的神经声学建模用于高保真RIR生成

Chen Si, Qianyi Wu, Chaitanya Amballa, Romit Roy Choudhury

机构 * University of California San Diego(加州大学圣地亚哥分校) Monash University(墨尔本大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 提出MiNAF模型,通过查询房间网格并提取距离分布作为显式局部几何特征,引导神经隐式模型生成更准确的房间脉冲响应(RIR),在多项指标上达到竞争性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05693 2026-06-16 cs.LG cs.IR 版本更新

MolE-RAG: Molecular Structure-Enhanced Retrieval-Augmented Generation for Chemistry

MolE-RAG:面向化学的分子结构增强检索增强生成

Joey Chan, Wonbin Kweon, Ashley Shin, Niharika Bhattacharjee, Pengcheng Jiang, Yue Guo, Jiawei Han

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of California, San Diego(加州大学圣地亚哥分校)

AI总结 提出无需训练的分子中心检索增强生成框架MolE-RAG,通过整合检索文献、分子特定信息和结构相似分子三种上下文,显著提升LLM在分子性质预测任务中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30837 2026-06-16 cs.CR cs.LG 版本更新

Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense

先派侦察兵:提示注入防御中自适应检测器分配的预推理方法

Shuhao Zhang, Jiarui Li, Qi Cao, Ruiyi Zhang, Pengtao Xie

机构 * UC San Diego(加州大学圣迭戈分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 针对提示注入检测器异构且不可靠的问题,提出SCOUT框架,通过预测每个检测器对每个样本的可靠性和延迟,动态分配检测器,实现安全性与效率的权衡。

Comments We propose SCOUT, a detector allocation framework that predicts each detector's accuracy and latency on a given input before running it, letting operators control the safety-utility trade-off with a single threshold and route to an LLM judge only when needed

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09163 2026-06-16 cs.AI 版本更新

FORTIS: Benchmarking Over-Privilege in Agent Skills

FORTIS:评估代理技能中的过度特权

Shawn Li, Chenxiao Yu, Han Wang, Wei Yang, Ryan Rossi, Franck Dernoncourt, Xiyang Hu, Philip Yu, Chaowei Xiao, Huan Zhang, Yue Zhao

机构 * University of Southern California(南加州大学) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Adobe Research(Adobe研究) Arizona State University(亚利桑那州立大学) University of Illinois Chicago(伊利诺伊大学芝加哥分校) Johns Hopkins University(约翰霍普金斯大学)

AI总结 研究发现,当前代理技能层普遍存在过度特权问题,模型在选择和执行技能时常超出任务需求,导致性能不佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17990 2026-06-16 cs.AI 版本更新

WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics

WorkflowPerturb:用于评估多智能体工作流度量的校准压力测试

Madhav Kanda, Sharad Agarwal, Rodrigo Fonseca, Alok Gautam Kumbhare, Pedro Las-Casas

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Microsoft(微软公司)

AI总结 提出WorkflowPerturb基准,通过对黄金工作流施加分级扰动来评估多智能体工作流度量,揭示度量分数校准不良问题,支持变更管理中的严重性感知解释。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13197 2026-06-16 cs.RO cs.CV cs.LG 版本更新

Imitating What Works: Simulation-Filtered Modular Policy Learning from Human Videos

模仿有效的方法:基于仿真过滤的人类视频模块化策略学习

Albert J. Zhai, Kuo-Hao Zeng, Jiasen Lu, Ali Farhadi, Shenlong Wang, Wei-Chiu Ma

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Allen Institute for AI(Allen人工智能研究所) University of Washington(华盛顿大学) Cornell University(康奈尔大学)

AI总结 提出Perceive-Simulate-Imitate框架,通过仿真过滤人类视频中的抓取-轨迹对,学习任务导向的抓取与后抓取运动策略,无需机器人数据即可实现鲁棒操作。

Comments Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20686 2026-06-16 q-bio.BM cs.DC cs.LG cs.PF 版本更新

MegaFold: Efficient Training of Next-Generation 3D Attention Protein Models on Cross-Platform GPUs

MegaFold: 跨平台GPU上高效训练下一代3D注意力蛋白质模型

Hoa La, Ahan Gupta, Alex Morehead, Jianlin Cheng, Minjia Zhang

机构 * UIUC SSAIL Lab(UIUC SSAIL实验室)

AI总结 针对AlphaFold3类模型因3D注意力机制导致训练效率低的问题,提出MegaFold系统,通过高效内核、分片策略、算子融合和流水线优化,在NVIDIA和AMD GPU上实现更长序列训练和加速。

Comments 13 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05163 2026-06-16 cs.CL cs.LG 版本更新

Enhancing LLM Safety Through a Theoretical Minimax Game Lens

通过理论极小极大博弈视角增强LLM安全性

Yihe Deng, Yu Yang, Junkai Zhang, Wei Wang, Bo Li

机构 * University of California, Los Angeles(加州大学洛杉矶分校) VirtueAI University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 提出极小极大强化学习框架,通过数据生成器与分类器协同进化生成高质量多语言安全数据,理论证明收敛到纳什均衡,使小模型在英文基准上超越SOTA近10%且推理速度提升4.5倍。

Comments 24 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00107 2026-06-16 cs.LG cs.AI eess.SP 版本更新

Virtual Sensing to Enable Real-Time Monitoring of Inaccessible Locations & Unmeasurable Parameters

虚拟传感实现不可达位置与不可测参数的实时监测

Kazuma Kobayashi, Farid Ahmed, Jaewan Park, Subhankar Sarkar, Souvik Chakraborty, Syed Bahauddin Alam

机构 * Plasma & Radiological Engineering Department, Grainger College of Engineering, Nuclear, University of Illinois Urbana-Champaign(等离子体与辐射工程系,格拉inger工程学院,核能,伊利诺伊大学厄巴纳-香槟分校) Mechanical Science and Engineering Department, Grainger College of Engineering, University of Illinois Urbana-Champaign(机械科学与工程系,格拉inger工程学院,伊利诺伊大学厄巴纳-香槟分校) National Center for Supercomputing Applications, Urbana, IL, USA(国家超级计算应用中心,伊利诺伊州厄巴纳,美国) Department of Applied Mechanics, Indian Institute of Technology Delhi, New Delhi, India(应用力学系,印度理工学院德里,新德里,印度) Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi(Yardi人工智能学院,印度理工学院德里)

AI总结 针对能量系统中物理传感器无法部署的实时监测问题,提出基于神经算子的虚拟传感框架MIMONet,将稀疏边界测量映射到内部场,在多种热流体系统中实现亚毫秒级高精度推理。

Comments New analysis and results are added

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03063 2026-06-15 math.ST cs.IT cs.LG math.IT stat.ME stat.ML stat.TH 版本更新

Stability of a Generalized Debiased Lasso with Applications to Resampling-Based Variable Selection

广义去偏Lasso的稳定性及其在基于重抽样的变量选择中的应用

Jingbo Liu

机构 * Department of Statistics, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校统计系) Department of Electrical and Computer Engineering, the Grainger College of Engineering(格拉inger工程学院电子与计算机工程系)

AI总结 提出基于稳定性原理的广义去偏Lasso估计量,通过设计矩阵单列扰动下的简单更新公式,在比例增长机制下实现渐近精确近似,显著降低重抽样变量选择的计算成本。

Comments to appear in Bernoulli

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06196 2026-06-15 cs.CL cs.HC 版本更新

EiCAP: Beyond Fluency, Probing and Improving Emotional Intelligence in LLMs via Psychologically Grounded Multi-Turn Dialogue

EiCAP:超越流畅性,通过心理学基础的多轮对话探究和提升大语言模型的情感智能

Nizi Nazar, Pardis Sadat Zahraei, Dilek Hakkani-Tür, Natasa Milic-Frayling, Ehsaneddin Asgari

机构 * Qatar Computing Research Institute (QCRI), Hamad Bin Khalifa University(卡塔尔计算研究所(QCRI),哈马德·本·卡伊夫大学) University of Illinois Urbana-Champaign (UIUC)(伊利诺伊大学厄巴纳-香槟分校)

AI总结 提出基于心理学六层情感智能分类法的EiCAP框架,包含评估基准EiCAP-Bench和微调语料EiCAP-SFT,发现通用对话微调不提升情感智能,而基于情感智能的LoRA微调显著提升模型在所有24个子类别上的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11004 2026-06-12 cs.CL 版本更新

NOVA: NOise-aware Verbal Confidence CAlibration for Robust Large Language Models in RAG Systems

NOVA: 面向RAG系统中鲁棒大语言模型的噪声感知言语置信度校准

Jiayu Liu, Rui Wang, Qing Zong, Yumeng Wang, Cheng Qian, Qingcheng Zeng, Tianshi Zheng, Haochen Shi, Dadi Guo, Baixuan Xu, Chunyang Li, Yangqiu Song

机构 * HKUST(香港科技大学) UIUC(伊利诺伊大学香槟分校) Northwestern University(西北大学)

AI总结 提出NOVA框架,通过规则引导的监督微调,解决检索增强生成中噪声上下文导致的过度自信问题,在域内和域外分别提升ECE 10.9%和8.0%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03660 2026-06-12 cs.LG 版本更新

Single vs. Multiple Branches in DeepONet and S-DeepONet: Network Architecture Follows Coupling in Multiphysics Systems

DeepONet和S-DeepONet中的单分支与多分支:网络架构遵循多物理系统中的耦合

Jaewan Park, Kazuma Kobayashi, Qibang Liu, Seid Koric, Diab Abueidda, Syed Bahauddin Alam

机构 * National Center for Supercomputing Applications, University of Illinois at Urbana-Champaign(国家超级计算应用中心,伊利诺伊大学厄巴纳-香槟分校) The Grainger College of Engineering, Mechanical Science and Engineering, University of Illinois at Urbana-Champaign(格拉inger工程学院,机械科学与工程系,伊利诺伊大学厄巴纳-香槟分校) The Grainger College of Engineering, Nuclear, Plasma & Radiological Engineering, University of Illinois at Urbana-Champaign(格拉inger工程学院,核物理与辐射工程系,伊利诺伊大学厄巴纳-香槟分校) Department of Industrial and Manufacturing Systems Engineering, Kansas State University(工业与制造系统工程系,堪萨斯州立大学) Civil and Urban Engineering Department, New York University Abu Dhabi, UAE(土木与城市工程系,纽约大学阿布扎比分校,阿联酋)

AI总结 研究比较单分支与多分支神经算子架构在强耦合多物理系统中的表现,发现单分支网络在紧耦合场景下通过共享潜在表示优于多分支,而多分支适用于解耦或单物理任务,代理模型加速高达1.8×10^4倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01274 2026-06-12 cs.CV cs.AI 版本更新

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding

ReFoCUS: 用于上下文理解的强化引导帧优化

Hosu Lee, Junho Kim, Hyunjun Kim, Yong Man Ro

机构 * Korea Advanced Institute of Science & Technology(韩国科学技术院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 提出ReFoCUS框架,首次将在线策略梯度强化学习集成到视频大语言模型的帧级优化中,通过自回归和查询条件选择架构学习帧选择策略,无需显式帧级监督,提升视频问答推理准确性。

Comments Project page: https://interlive-team.github.io/ReFoCUS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25820 2026-06-11 cs.LG 版本更新

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models

基于扩散的多模态大语言模型的视觉冗余控制并行解码

Yulin Yuan, Hongshuo Zhao, Xiangming Meng

机构 * Zhejiang University(浙江大学) ZJUI-UIUC Institute(ZJUI-UIUC研究院)

AI总结 针对扩散型多模态大语言模型并行解码中视觉冗余问题,提出视觉冗余指数(VRI)和无需训练的视觉冗余控制解码(VRCD)方法,通过令牌到图像的注意力优先选择视觉互补位置,在多个基准上提升准确率。

Comments 18 pages, 5 figures, preprint. Code is available at https://github.com/infiniteYuanyl/VRCD

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22962 2026-06-11 cs.LG 版本更新

Scaling Laws of Global Weather Models

全球天气模型的缩放定律

Yuejiang Yu, Langwen Huang, Alexandru Calotoiu, Torsten Hoefler

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 本文分析数据驱动天气模型中模型大小、数据集大小和计算预算与验证损失之间的缩放定律,发现Aurora数据缩放最强,GraphCast参数效率高但硬件利用率低,计算最优分析表明增加训练数据比增大模型更有效,且模型形状上宽度优于深度。

Comments Accepted at ICML 2026. 21 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20024 2026-06-10 cs.LG 版本更新

Replicable Bandits with UCB based Exploration

基于UCB探索的可复现Bandits

Rohan Deb, Udaya Ghai, Karan Singh, Arindam Banerjee

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Amazon(亚马逊) Carnegie Mellon University(卡内基梅隆大学)

AI总结 研究随机多臂老虎机和线性老虎机中基于UCB探索的可复现算法,提出RepUCB和RepLinUCB,分别实现最优遗憾界,显著降低可复现性代价。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11502 2026-06-09 cs.CL 版本更新

Robust Biomedical Publication Type and Study Design Classification with Knowledge-Guided Perturbations

基于知识引导扰动的鲁棒生物医学出版物类型与研究设计分类

Shufan Ming, Joe D. Menke, Neil R. Smalheiser, Halil Kilicoglu

机构 * School of Information Sciences(信息科学学院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Department of Psychiatry(精神病学系) University of Illinois Chicago(伊利诺伊大学芝加哥分校)

AI总结 本文提出基于受控语义扰动的评估框架,通过实体遮蔽和领域对抗训练提升生物医学出版物类型分类的鲁棒性,发现通过抑制非任务定义特征可缓解鲁棒性与领域准确性之间的权衡。

Comments Accepted by IEEE ICHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17911 2026-06-09 cs.CL cs.AI 版本更新

Condition-Gated Reasoning for Context-Dependent Biomedical Question Answering

基于条件的推理用于依赖上下文的生物医学问答

Jash Rajesh Parekh, Wonbin Kweon, Joey Chan, Rezarta Islamaj, Robert Leaman, Pengcheng Jiang, Chih-Hsuan Wei, Zhizheng Wang, Zhiyong Lu, Jiawei Han

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) National Institutes of Health(美国国立卫生研究院)

AI总结 本文提出CondMedQA基准和Condition-Gated Reasoning框架,通过构建条件感知知识图谱,提升生物医学问答中条件依赖的推理能力。

URL PDF HTML 收藏
2510.16028 2026-06-09 cs.CR cs.AI cs.LG cs.SY eess.SY 版本更新

TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks

TAO:面向浮点神经网络的容忍感知乐观验证

Jianzhu Yao, Hongxu Su, Taobo Liao, Zerui Cheng, Huan Zhang, Xuechao Wang, Pramod Viswanath

机构 * Princeton University(普林斯顿大学) HKUST (GZ)(香港科技大学(广州)) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 提出TAO协议,通过算子级容忍区域和Merkle锚定的争议游戏,在不依赖可信硬件或确定性内核的情况下验证浮点神经网络输出,开销仅0.3%。

Comments 18 pages, 8 figures

Journal ref Proceedings of the 21st European Conference on Computer Systems, (2026) 1515-1532

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19829 2026-06-09 cs.AI 版本更新

Knowing How to Edit: Reliable Evaluation Signals for Diagnosing and Optimizing Prompts at Query Level

一种统一评估指导的查询相关提示优化框架

Ke Chen, Yifeng Wang, Hassan Almosapeeh, Haohan Wang

机构 * School of Information Sciences, University of Illinois Urbana-Champaign(信息科学学院,伊利诺伊大学厄巴纳-香槟分校) College of Engineering, Carnegie Mellon University(工程学院,卡内基梅隆大学)

AI总结 提出一个基于性能导向的提示评估框架,并开发一个无需执行的评估器来预测多维质量分数,进而指导一个度量感知优化器以可解释的查询相关方式重写提示,在多个数据集和骨干模型上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10239 2026-06-09 cs.HC cs.CL 版本更新

Breaking the Curse of Knowledge: Designing Personalized Jargon Support for Real-Time Online Meetings

打破知识的诅咒:为实时在线会议设计个性化术语支持

Yifan Song, Yijun Liu, Wing Yee Au, Hon Yung Wong, Brian P. Bailey, Tal August

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Fujitsu Research of America(富士通美国研究)

AI总结 提出ParseJargon系统,利用用户画像和会话内反馈实现个性化术语识别,提升在线会议中跨学科听众的理解和参与度。

Comments Portions of this work appeared in CHI '26 Extended Abstracts ("Breaking the Curse of Knowledge: Toward Personalized Jargon Support in Online Meetings") and ACL '26 System Demonstrations ("ParseJargon: Personalized Real-time Jargon Support in Online Meetings")

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23292 2026-06-08 cs.AI cs.LG 版本更新

Agentic Physical AI toward a Domain-Specific Foundation Model for Energy Systems: A Case Study on Nuclear Reactor Control

面向能源系统的领域特定基础模型的具身物理人工智能:以核反应堆控制为例

Yoon Pyo Lee, Samrendra Roy, Kazuma Kobayashi, Sajedul Talukder, Diab Abueidda, Seid Koric, Souvik Chakraborty, Syed Bahauddin Alam

机构 * The Grainger College of Engineering, Nuclear, Plasma & Radiological Engineering, University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校格雷格学院工程学院、核等工程学院) Department of Nuclear Engineering, Hanyang University(汉阳大学核工程系) University of Texas - El Paso(德克萨斯大学埃尔帕索分校) National Center for Supercomputing Applications(国家超级计算应用中心) Department of Applied Mechanics, Indian Institute of Technology Delhi(印度德里理工学院应用力学系) Yardi School of Artificial Intelligence, Indian Institute of Technology Delhi(印度德里理工学院亚里人工智能学院)

AI总结 本研究提出通过紧凑语言模型作为具身物理人工智能,利用基于物理模拟器验证的策略优化替代感知推理,在核反应堆控制任务中实现领域特定基础模型,并展示了规模扩展带来的可靠性提升和策略集中化行为。

Comments Accepted for publication in npj Artificial Intelligence

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05360 2026-06-08 cs.HC cs.AI 版本更新

OGA-AID: Clinician-in-the-loop AI Report Drafting Assistant for Multimodal Observational Gait Analysis in Post-Stroke Rehabilitation

OGA-AID:用于中风后康复多模态观察性步态分析的临床医生在环AI报告起草助手

Khoi T. N. Nguyen, Nghia D. Nguyen, Hui Yu Koh, Patrick W. H. Kwong, Karen Sui Geok Chua, Ananda Sidarta, Baosheng Yu

机构 * Rehabilitation Research Institute of Singapore, Nanyang Technological University, Singapore(新加坡康复研究中心,南洋理工大学,新加坡) Lee Kong Chian School of Medicine, Nanyang Technological University, Singapore(李光前医学院,南洋理工大学,新加坡) The Grainger College of Engineering, University of Illinois Urbana-Champaign, United States(伊利诺伊大学厄巴纳-香槟分校格雷格学院,美国) Department of Rehabilitation Sciences, The Hong Kong Polytechnic University, Hong Kong(香港理工大学康复科学系,香港) VinUni-Illinois Smart Health Center, VinUniversity, Vietnam(Vin大学Vin-伊利诺伊智能健康中心,越南) Institute of Rehabilitation Excellence, Tan Tock Seng Hospital, NHG Health, Singapore(卓越康复研究所,坦托克桑格医院,NHG健康,新加坡)

AI总结 提出OGA-AID,一种临床医生在环的多智能体大语言模型系统,通过协调三个专业智能体合成患者运动记录、运动学轨迹和临床资料,生成结构化步态评估报告,在真实患者数据上优于单次多模态基线,并展示了AI辅助分析与人类临床判断的互补关系。

Comments 2026 CV4Clinic CVPR Workshop Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04123 2026-06-08 cs.CY cs.AI cs.LG cs.SE 版本更新

Measuring Agents in Production

生产环境中的智能体测量

Melissa Z. Pan, Negar Arabzadeh, Riccardo Cogo, Yuxuan Zhu, Alexander Xiong, Lakshya A Agrawal, Huanzhi Mao, Emma Shen, Sid Pallerla, Liana Patel, Shu Liu, Tianneng Shi, Xiaoyuan Liu, Jared Quincy Davis, Emmanuele Lacavalla, Alessandro Basile, Shuyi Yang, Paul Castro, Daniel Kang, Koushik Sen, Dawn Song, Joseph E. Gonzalez, Ion Stoica, Matei Zaharia, Marquita Ellis

机构 * University of California at Berkeley(加州大学伯克利分校) IBM Research(IBM研究院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Stanford University(斯坦福大学)

AI总结 通过对86个已部署系统的调查和20个案例研究,发现生产环境中的LLM智能体主要采用简单可控的方法,可靠性是首要挑战,并依赖系统级设计和人工评估。

Comments Accepted to the 43rd International Conference on Machine Learning (ICML 2026) as Oral Presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.23204 2026-06-08 cs.AI 版本更新

TSAQA: Time Series Analysis Question And Answering Benchmark

TSAQA:时间序列分析问答基准

Baoyu Jing, Sanhorn Chen, Lecheng Zheng, Boyu Liu, Zihao Li, Jiaru Zou, Tianxin Wei, Zhining Liu, Zhichen Zeng, Ruizhong Qiu, Xiao Lin, Yuchen Yan, Dongqi Fu, Jingchao Ni, Jingrui He, Hanghang Tong

机构 * University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Virginia Polytechnic Institute and State University(弗吉尼亚理工学院和州立大学) Amazon(亚马逊) Meta AI University of Houston(休斯顿大学)

AI总结 提出TSAQA基准,涵盖6种时间序列分析任务(含新型PZ格式),评估LLM在13领域21万样本上的表现,最佳模型仅65.08分。

Comments Comments: 35 pages, 7 figures. Accepted to the GEM Workshop at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10896 2026-06-08 cs.CL 版本更新

DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference

DialDefer: 检测和缓解LLM对话性遵从的框架

Parisa Rabbani, Priyam Sahoo, Ruben Mathew, Aishee Mondal, Harshita Ketharaman, Nimet Beyza Bozdag, Dilek Hakkani-Tür

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 提出DialDefer框架,通过对话性遵从分数检测和缓解LLM在对话评估中因提问框架导致的判断偏移,发现框架效应显著但准确率稳定,且模型对人类与AI的不同归因产生最大偏移。

Comments 10 pages main content, 7 figures, 35 pages total with appendix

Journal ref ACL 2026 - Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

详情

展开后加载摘要…

URL PDF HTML 收藏