arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Washington(华盛顿大学)

2026-06-30 至 2026-06-30 共收录 11
2606.29689 2026-06-30 cs.CL

Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

MLLM 能否像人类一样进行批评?评估多模态大语言模型中的开放式审美推理

Sajjad Ghiasvand, Maryam Amirizaniani, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani

机构 * UCSB(加州大学圣塔芭芭拉分校) University of Washington(华盛顿大学) UCLA(加州大学洛杉矶分校)

AI总结 本文评估多模态大语言模型在开放式审美批评中的表现,发现基于参考的相似度指标存在误导,模型在选择性、特异性和多样性方面与人类批评存在系统性差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29648 2026-06-30 cs.CL cs.AI cs.LG cs.MA

Hybrid Retriever Evolution for Multimodal Document Reasoning Agents

混合检索器演化:面向多模态文档推理代理

Bohan Yao, Shruthan Radhakrishna, Vikas Yadav

机构 * ServiceNow University of Washington(华盛顿大学)

AI总结 提出失败驱动的演化框架,让元代理自动学习任务代理如何协调多种检索器进行多步文档问答,实现自适应检索路由,在MMLongBench-Doc和DocBench上提升高达19.6分。

Comments 17 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29047 2026-06-30 cs.CE cs.LG physics.flu-dyn

Weak Dominant Balance for Robust Identification of Dynamically Consistent Fluid Flow Structure

弱主导平衡:用于鲁棒识别动态一致流体流动结构

Samuel Ahnert, Esther Lagemann, H. Jane Bae, Kunihiko Taira, Ricardo Vinuesa, Christian Lagemann, Steven L. Brunton

机构 * Department of Mechanical Engineering, University of Washington, Seattle, WA, USA(华盛顿大学机械工程系) AI Institute in Dynamic Systems, University of Washington, Seattle, WA, USA(动态系统人工智能研究所) Lynn Booth and Kent Kresa Department of Aerospace, California Institute of Technology, Pasadena, CA, USA(加州理工学院航空航天系) Department of Mechanical and Aerospace Engineering, University of California, Los Angeles, CA, USA(加州大学洛杉矶分校机械与航空航天工程系) Department of Aerospace Engineering, University of Michigan, Ann Arbor, MI, USA(密歇根大学航空航天工程系)

AI总结 提出弱主导平衡框架,通过弱形式投影避免数值微分噪声,实现从高噪声数据中识别物理机制,并成功应用于湍流管道流动的三阶偏微分方程分解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28733 2026-06-30 cs.AI

Agentic Abstention: Do Agents Know When to Stop Instead of Act?

Agentic Abstention: 智能体知道何时停止而非行动吗?

Han Luo, Bingbing Wen, Lucy Lu Wang

机构 * University of Leeds(利兹大学) Southwest Jiaotong University(西南交通大学) University of Washington(华盛顿大学) Allen Institute for AI(人工智能研究院)

AI总结 研究LLM智能体在不确定环境下何时应停止行动的问题,提出Agentic Abstention概念,通过实验分析不同模型和框架的弃权行为,并引入CONVOLVE方法提升及时弃权率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28628 2026-06-30 eess.IV cs.CV cs.LG

Envisage: Diffusion-Based Rhinoplasty Goal Visualization with Mask-Decomposed Evaluation

Envisage:基于扩散的鼻整形目标可视化与掩码分解评估

Mudit Agarwal, Amit D. Bhrany

机构 * University of Washington(华盛顿大学) University of Washington School of Medicine(华盛顿大学医学院) Department of Otolaryngology–Head and Neck Surgery(耳鼻喉科–头颈外科部门)

AI总结 提出Envisage扩散修复管线,从单张正面照片可视化鼻整形目标,并引入SurgicalScore掩码分解协议评估局部编辑质量,克服全脸身份指标在硬合成编辑下的混淆。

Comments 29 pages, 4 figures, 22 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27273 2026-06-30 cs.SD 版本更新

Few-Shot Synthetic Accented Speech for ASR Fine-Tuning: What Helps and When?

少样本合成方言语音用于ASR微调:什么有助于什么?

Yurii Halychanskyi, Nimet Beyza Bozdag, Mark Hasegawa-Johnson, Dilek Hakkani-Tür, Volodymyr Kindratenko

机构 * University of Washington(华盛顿大学)

AI总结 研究比较了合成方言语音在ASR微调中的有效性,发现随机音素扰动比目标方言音素编辑更有效,且真实语音与合成语音混合可稳定低资源微调。

Comments Accepted as a contributed talk and poster at the ICML 2026 Workshop on Machine Learning for Audio

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10109 2026-06-30 cs.AI cs.HC cs.LG

LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals

基于自我报告的LLM代理能够实现通用个体模拟

Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, Michael S. Bernstein

机构 * Computer Science Department, Stanford University(斯坦福大学计算机科学系) Department of Communication Studies, Northwestern University(西北大学传播学系) Department of Communication, University of Washington(华盛顿大学传播学系) Google DeepMind(谷歌DeepMind) Department of Sociology, Stanford University(斯坦福大学社会学系) Sciences Po(巴黎政治学院)

AI总结 本文研究了基于自我报告数据的LLM代理在模拟个体行为方面的有效性,通过不同数据源构建代理并验证其在多种任务中的准确性和跨群体公平性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17753 2026-06-30 cs.CY cs.AI

The 2025 AI Agent Index: Documenting Technical and Safety Features of Deployed Agentic AI Systems

2025人工智能代理指数:记录已部署代理式人工智能系统的技术和安全特性

Leon Staufer, Kevin Feng, Kevin Wei, Luke Bailey, Yawen Duan, Mick Yang, A. Pinar Ozisik, Stephen Casper, Noam Kolt

机构 * University of Cambridge(剑桥大学) University of Washington(华盛顿大学) Harvard Law School(哈佛法学院) Stanford University(斯坦福大学) Concordia AI(康科迪亚AI) University of Pennsylvania(宾夕法尼亚大学) Massachusetts Institute of Technology(麻省理工学院) Hebrew University of Jerusalem(耶路撒冷希伯来大学)

AI总结 本文提出2025人工智能代理指数,记录30种先进代理式AI系统的起源、设计、能力、生态系统及安全特性,揭示代理发展中的趋势和开发者透明度问题。

Comments To be publishesd at ACM FAccT 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07338 2026-06-30 stat.ML cs.DM cs.LG stat.ME

Towards Complete Causal Explanation with Expert Knowledge

迈向完整因果解释的专家知识

Aparajithan Venkateswaran, Emilija Perković

机构 * Microsoft(微软) University of Washington(华盛顿大学)

AI总结 本文研究如何通过专家知识限制最大祖先图的马尔可夫等价类,提出新的图定向规则和算法,扩展了Meek(1995)以处理潜在混杂因素。

Comments 86 pages (main paper 26 pages, supplementary material 60 pages), 21 figures, 7 algorithms, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.08775 2026-06-30 cs.HC cs.AI

What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM Use

提示中应工程什么?基于需求驱动的LLM使用训练

Qianou Ma, Weirui Peng, Chenyang Yang, Hua Shen, Kenneth Koedinger, Tongshuang Wu

机构 * Carnegie Mellon University(卡内基梅隆大学) Columbia University(哥伦比亚大学) University of Washington(华盛顿大学)

AI总结 本文提出ROPE框架,通过评估和训练套件帮助用户生成清晰需求,提升LLM应用构建效果。

Comments 15 pages; TOCHI 2025

Journal ref TOCHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02831 2026-06-30 cs.CV cs.AI

FLAME 3 Dataset: Unleashing the Power of Radiometric Thermal UAV Imagery for Wildfire Management

FLAME 3数据集:释放无人机辐射热成像用于野火管理的潜力

Bryce Hopkins, Leo ONeill, Michael Marinaccio, Mobin Habibpour, Eric Rowell, Russell Parsons, Sarah Flanary, Irtija Nazim, Carl Seielstad, Fatemeh Afghah

机构 * Holcombe Department of Electrical and Computer Engineering, Clemson University(克莱姆森大学霍尔科姆电气与计算机工程系) Pacific Southwest Research Station, U.S. Forest Services(美国林业局太平洋西南研究站) School of Environmental and Forest Sciences, University of Washington(华盛顿大学环境与森林科学学院) US Forest Service, Rocky Mountain Research Station, Fire Sciences Laboratory(美国林业局落基山研究站火灾科学实验室) Department of Mechanical Engineering, Clemson University(克莱姆森大学机械工程系) Department of Forest Management, University of Montana(蒙大拿大学森林管理系)

AI总结 本文提出FLAME 3数据集,通过无人机收集同步可见光与辐射热影像,推动基于辐射热成像的野火管理AI应用,提供新的数据类型和采集方法。

Comments 15 pages, 8 Figures, 9 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏