arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Texas at Austin(得克萨斯大学奥斯汀分校)

共收录 1212
2607.04037 2026-07-07 cs.LG cs.AI 新提交

Reward-Gated On-Policy Distillation

奖励门控策略蒸馏

Mohammad Sadegh Akhondzadeh, Vijay Lingam, Atula Tejaswi, Chanakya Ekbote, Sujay Sanghavi, Aleksandar Bojchevski

机构 * University of Cologne(科隆大学) AWS AI(亚马逊云科技人工智能) UT Austin(德克萨斯大学奥斯汀分校) MIT(麻省理工学院)

AI总结 研究如何进行策略蒸馏,核心方法是引入奖励门控策略蒸馏(RG-OPD)利用验证器反馈决定是否信任教师逻辑,主要贡献是在推理和编码基准测试中表现出色,优于其他蒸馏方法。

Comments 8 pages, 2 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02959 2026-07-07 cs.CV cs.LG 新提交

Incentivizing Vision Language Models to Search for Long Video Question Answering

激励视觉语言模型进行长视频问答搜索

Harsh Goel, S P Sharan, Sahil Shah, Minkyu Choi, Joungbin An, Kristen Grauman, Sandeep P. Chinchali

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 研究将长视频问答从被动单流程感知任务转变为多轮检索过程,核心方法是自然语言驱动搜索及强化学习后训练,贡献是提升长视频理解基准测试分数。

Comments To appear at the European Conference on Computer Vision (ECCV 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02668 2026-07-07 cs.CL 新提交

Improving LLMs via Validator-to-Generator Alignment

通过验证器与生成器对齐改进大语言模型

Juan Diego Rodriguez, Jocelyn Zhang, Katrin Erk, Greg Durrett

机构 * Department of Computer Science, The University of Texas at Austin(德克萨斯大学奥斯汀分校计算机科学系) Departments of Linguistics and Computer Science, University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校语言学与计算机科学系) Department of Computer Science & Center for Data Science, New York University(纽约大学计算机科学系及数据科学中心)

AI总结 研究大语言模型不一致问题,提出基于话语频率的生成器-验证器一致性新公式,介绍\FCPA方法,实验表明该方法能提升生成器性能及两者一致性,且保持验证器质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23053 2026-07-07 cs.LG math.OC 版本更新

ML-Guided Primal Heuristics for Mixed Binary Quadratic Programs

基于机器学习的混合二元二次规划的原始启发式方法

Weimin Huang, Natalie M. Isenberg, Ján Drgoňa, Draguna L Vrabie, Bistra Dilkina

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出基于机器学习的混合二元二次规划求解启发式方法,通过改进神经网络架构和损失函数,提升求解效率和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.09445 2026-07-07 cs.CV 版本更新

AsymLoc: Towards Asymmetric Feature Matching for Efficient Visual Localization

AsymLoc:迈向高效视觉定位的非对称特征匹配

Mohammad Omama, Gabriele Berton, Eric Foxlin, Yelin Kim

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Amazon(亚马逊)

AI总结 本文提出AsymLoc框架,通过几何驱动匹配目标和联合检测器-描述符蒸馏目标,实现轻量学生模型与大教师模型的非对称特征匹配,提升视觉定位效率与精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00392 2026-07-07 cs.SE cs.AI 版本更新

Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents

EvolveTool-Bench:评估LLM生成的工具库作为软件 artifact 的质量

Alibek Kaliyev, Artem Maryanskyy

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) Uber Technologies(优步科技公司)

AI总结 本文提出EvolveTool-Bench,通过评估LLM生成的工具库在软件工程流程中的质量,揭示任务完成率之外的软件质量风险,强调需将工具库视为首要软件 artifact。

Comments 11 pages, 4 figures; accepted at KDD 2026 Workshop on Agentic AI Evaluation and Trustworthiness

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22770 2026-07-07 cs.LG cs.AI 版本更新

From Arithmetic to Logic: The Resilience of Logic and Lookup-Based Neural Networks Under Parameter Bit-Flips

从算术到逻辑:逻辑与基于查找的神经网络在参数位翻转下的韧性

Alan T. L. Bacellar, Sathvik Chemudupati, Shashank Nag, Allison Seigler, Priscila M. V. Lima, Felipe M. G. França, Lizy K. John

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Federal University of Rio de Janeiro(里约热内卢联邦大学) Instituto de Telecomunicações, Porto, Portugal(葡萄牙波尔图电信研究所) Google LLC(谷歌公司)

AI总结 研究探讨了神经网络在硬件位翻转误差下的韧性,发现低精度、高稀疏性、受限激活和浅层深度等特性在理论模型下更优,逻辑和查找型网络实现了这些设计趋势的联合极限,实验验证了其在故障容忍中的稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01709 2026-07-07 q-fin.PR cs.LG

Reinforcement Learning for Option Hedging: Static Implied-Volatility Fit versus Shortfall-Aware Performance

期权对冲的强化学习:静态隐含波动率拟合与短失意识表现

Ziheng Chen, Minxuan Hu, Jiayu Yi, Wenxi Sun

机构 * Department of Mathematics, University of Texas at Austin(德克萨斯大学奥斯汀分校数学系) Cornell Ann S. Bowers College of Computing and Information Science, Cornell University(康奈尔大学安·S·鲍威尔计算机与信息科学学院) School of Social Sciences, Nanyang Technological University(南洋理工大学社会科学学院) Krieger School of Arts and Sciences, Johns Hopkins University(约翰霍普金斯大学克里格尔艺术与科学学院)

AI总结 本文提出RLOP方法,通过结合风险厌恶和交易成本改进QLBS框架,评估期权定价模型在静态和动态维度的表现,强调实际对冲效果的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22851 2026-07-07 cs.LG cs.AI cs.CL 版本更新

Adaptive Margin RLHF via Preference over Preferences

通过偏好之上的偏好进行自适应边际基于人类反馈的强化学习

Yaswanth Chittepu, Prasann Singhal, Greg Durrett, Scott Niekum

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 研究在基于人类反馈的强化学习中奖励模型学习的边际优化问题,提出利用偏好之上的偏好推断自适应边际,介绍具体实例DPO-PoP,提升判别和生成性能,并给出采样策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12853 2026-07-07 eess.IV cs.CV 版本更新

BrainNormalizer: Anatomy-Informed Pseudo-Healthy Brain Reconstruction from Tumor MRI via Edge-Guided ControlNet

BrainNormalizer:通过边缘引导控制网络从肿瘤MRI进行解剖学信息伪健康脑重建

Min Gu Kwak, Yeonju Lee, Hairong Wang, Kristin R. Swanson, Jing Li

机构 * University of Pittsburgh(匹兹堡大学) Georgia Institute of Technology(佐治亚理工学院) University of Texas at Austin(德克萨斯大学奥斯汀分校) Cedars-Sinai Medical Center(西德萨斯医疗中心) Mayo Clinic Arizona(梅奥诊所亚利桑那分部)

AI总结 针对脑肿瘤致结构变形难区分肿瘤与解剖变异问题,提出BrainNormalizer框架,用两阶段训练学习解剖先验等,通过特定策略实现伪健康脑重建,实验表明其有优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09629 2026-07-07 cs.CV 版本更新

Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors

利用学习到的纹理先验增强单目3D手部重建

Giorgos Karvounas, Nikolaos Kyriazis, Iason Oikonomidis, Georgios Pavlakos, Antonis A. Argyros

机构 * ICS-FORTH(希腊研究所) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Crete(克里特大学)

AI总结 重新审视纹理在单目3D手部重建中的作用,提出轻量级纹理模块,嵌入像素观察到UV纹理空间,实现预测与观察手部外观的密集对齐损失,增强HaMeR系统,提高精度和真实感。

Comments Accepted at WACV 2026. Project page: https://gkarv.github.io/hand-texture-module/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20666 2026-07-07 cs.HC cs.CL 版本更新

TAMA: A Human-AI Collaborative Thematic Analysis Framework Using Multi-Agent LLMs for Clinical Interviews

TAMA:一种使用多智能体大语言模型进行临床访谈的人机协作主题分析框架

Huimin Xu, Seungjun Yi, Terence Lim, Jiawei Xu, Andrew Well, Carlos Mery, Aidong Zhang, Yuji Zhang, Heng Ji, Keshav Pingali, Yan Leng, Ying Ding

机构 * School of Information, University of Texas at Austin(信息学院,德克萨斯大学奥斯汀分校) Department of Biomedical Engineering, University of Texas at Austin(生物医学工程系,德克萨斯大学奥斯汀分校) College of Natural Sciences, University of Texas at Austin(自然科学院,德克萨斯大学奥斯汀分校) Graphen, Inc.(Graphen公司) Department of Cardiac Surgery, Division of Pediatric Cardiac Surgery, Vanderbilt University School of Medicine(心脏外科系,范德比尔特大学医学中心) Pediatric Heart Institute, Monroe Carell Jr. Children’s Hospital at Vanderbilt(儿童心脏研究所,范德比尔特儿童医院) Department of Computer Science, University of Virginia(计算机科学系,弗吉尼亚大学) Department of Computer Science, University of Illinois at Urbana-Champaign(计算机科学系,伊利诺伊大学厄巴纳-香槟分校) Department of Computer Science, University of Texas at Austin(计算机科学系,德克萨斯大学奥斯汀分校) McCombs School of Business, University of Texas at Austin(麦克阿瑟商学院,德克萨斯大学奥斯汀分校) Dell Medical School, University of Texas at Austin(德克萨斯大学奥斯汀分校德莱尔医学院)

AI总结 提出人机协作主题分析框架TAMA,利用多智能体系统结构化对话及心脏专家专业知识,用于临床访谈分析,在罕见病访谈转录本分析中性能优于单智能体方法。

Comments Manuscript accepted to ACM Transactions on Computing for Healthcare

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.02490 2026-07-03 cs.CL cs.CV 新提交

Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning

基于强化学习的视觉语言模型视觉接地自反思

Liyan Tang, Fangcong Yin, Greg Durrett

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) New York University(纽约大学)

AI总结 提出VRRL框架,通过随机掩码轨迹前缀和缓冲回滚机制,增强视觉语言模型在分布外场景下的视觉接地自反思能力,显著提升准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01538 2026-07-03 cs.CL 新提交

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale

语言模型真的能进行上下文检索吗?在百万Token规模的文档中淹没

Siddharth Gollapudi, Nilesh Gupta, Prasann Singhal, Sewon Min

机构 * UC Berkeley(加州大学伯克利分校) UT Austin(得克萨斯大学奥斯汀分校)

AI总结 本文首次系统研究百万Token语料库上的上下文检索,提出0.6B参数的BlockSearch模型,通过注意力稀释分析引入长度感知调整,在MS MARCO和NQ上匹配稠密检索,在LIMIT上得分高出3倍。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00673 2026-07-02 cs.RO 新提交

Path Planning in Physically Viable World Models

物理可行世界模型中的路径规划

Su Ann Low, Cheng-Hsi Hsiao, Xingjian Li, Adam J. Thorpe, Ufuk Topcu, Krishna Kumar

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 提出物理可行世界模型,通过物理仿真生成修改后的场景,使机器人能在部署前评估未来地形变化对路径可行性的影响。

Comments 18 pages, 7 figures, submitted to CORL

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.00492 2026-07-02 cs.CV 新提交

GenSP: Consistent Spherical Parameterization via Learning Shape Generative Models

GenSP:通过学习形状生成模型实现一致的球面参数化

Sai Karthikey Pentapati, Shashank Gupta, Rajesh Sureddi, Yuezhi Yang, Alan C. Bovik, Qixing Huang

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Colorado Boulder(科罗拉多大学博尔德分校)

AI总结 提出GenSP框架,通过学习神经生成模型预测从单位球面到形状的连续映射,并通过逆映射获得一致的球面参数化,显著降低几何畸变并提升跨形状一致性。

Comments Accepted at ECCV 2026. Sai Karthikey Pentapati and Shashank Gupta contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01208 2026-07-02 cs.CL cs.AI cs.LG 新提交

Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation

蒸馏检测:通过弹药筒蒸馏揭露大语言模型中的隐蔽偏见

Shayan Talaei, Abhinav Chinta, Devvrit Khatri, Amin Karbasi, Azalia Mirhoseini, Amin Saberi

机构 * Stanford University(斯坦福大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) Foundation AI–Cisco Systems Inc.(Foundation AI–思科系统公司)

AI总结 提出Distill to Detect (D2D)方法,通过蒸馏模型与基座之间的分布偏移到KV缓存前缀适配器中,放大隐蔽偏见信号至可检测程度,并基于Fisher加权投影理论解释其有效性。

Comments Accepted to the ICML 2026 Workshops on TAIGR, AI4GOOD, Mechanistic Interpretability, and CoLoRAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25169 2026-07-02 math.ST cs.LG stat.TH 新提交

Laplace-Fisher Gate Identities for Optimal Matrix-Gated Blended Score Estimation

Laplace--Fisher 门恒等式用于最优矩阵门控混合得分估计

Alois Duston, Tan Bui-Thanh

机构 * The Oden Institute for Computational Engineering and Sciences(奥登计算工程与科学研究院) The University of Texas at Austin(德克萨斯大学奥斯汀分校) The Department of Aerospace Engineering & Engineering Mechanics(航空航天工程与机械系)

AI总结 针对非归一化目标采样中的得分估计问题,提出矩阵值门控混合得分估计方法,推导出方差最优门公式(Laplace-Fisher 门恒等式),并证明有限参考一致性,应用于贝叶斯逆问题的归一化密度评估。

Comments Provisional report

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12165 2026-07-02 cs.CV 版本更新

Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video

利用被动场景声音和真实世界视频进行音频-视觉相机姿态估计

Daniel Adebi, Sagnik Majumder, Kristen Grauman

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出一种结合方向到达谱和双耳嵌入的音频-视觉框架,用于真实世界视频中的相机姿态估计,在视觉信息受损时表现出鲁棒性。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18765 2026-07-02 cs.CV cs.AI 版本更新

NI-Tex: Non-isometric Image-based Garment Texture Generation

NI-Tex: 基于非等距图像的服装纹理生成

Hui Shan, Ming Li, Haitao Yang, Kai Zheng, Sizhe Zheng, Yanwei Fu, Xiangru Huang

机构 * Zhejiang University(浙江大学) Shanghai Innovation Institute(上海创新研究院) Westlake University(西湖大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) Fudan University(复旦大学)

AI总结 针对非等距图像与3D服装网格间的纹理生成问题,提出结合物理模拟数据集、非等距图像编辑和迭代烘焙的框架,生成高质量PBR纹理。

Comments Accepted to CVPR 2026 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.15427 2026-07-02 cs.LG cs.DS 版本更新

Semi-Bandit Learning for Monotone Stochastic Optimization

单调随机优化的半强盗学习

Arpit Agarwal, Rohan Ghuge, Viswanath Nagarajan, Zhengjia Zhuo

机构 * Indian Institute of Technology Bombay(印度理工学院孟买分校) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Michigan(密歇根大学)

AI总结 针对未知分布下的单调随机优化问题,提出一种半强盗在线学习算法,在仅观测被探测变量样本的情况下,实现相对于最优近似算法的√(T log T)遗憾界,并扩展到删失和二元反馈场景。

Comments Full version (and extension) of FOCS 2024 paper. Fixes some missing assumptions in our results for continuous distributions. Also adds extensions to censored and binary feedback settings (along with applications) Revision: We improved the $k$ dependence

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15692 2026-07-02 cs.HC cs.CL cs.CV 交叉投稿

Surfacing Variations to Calibrate Perceived Reliability of MLLM-generated Image Descriptions

表面化变异以校准MLLM生成图像描述的可信度感知

Meng Chen, Akhil Iyer, Amy Pavel

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of California, Berkeley(加州大学伯克利分校)

AI总结 通过系统性地展示多个MLLM响应间的变异,帮助盲人或低视力用户无需视觉检查即可检测不可靠信息,实验表明该方法将识别不可靠声明的能力提升4.9倍。

Comments 18 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29793 2026-07-01 cs.CL q-fin.GN 版本更新

Fund2Persona: A Framework for Building and Refining Financial Advisor Personas from Fund Disclosure Data

Fund2Persona:基于基金披露数据构建与精炼金融顾问角色的框架

Suhwan Park, Hoyoung Lee, Zhangyang Wang, Alejandro Lopez-Lira, Young Cha, Chanyeol Choi, Jaewon Choi, Yongjae Lee

机构 * UNIST(蔚山科学技术院) LinqAlpha University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Florida(佛罗里达大学) Blackstone(黑石集团) Hanwha Life(韩华生命)

AI总结 提出Fund2Persona框架,利用基金披露、持仓变动、市场背景和管理人评论构建金融顾问角色,并通过智能体循环精炼,在持仓重建和管理人评论对齐任务上优于通用基线,生成更具体有用的建议。

Comments 17 pages, 5 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31993 2026-07-01 cs.RO 新提交

OopsieVerse: A Safety Benchmark with Damage-Aware Simulation for Robot Manipulation

OopsieVerse: 一个具有损伤感知仿真的机器人操作安全基准

Arnav Balaji, Arpit Bahety, Sriniket Ambatipudi, Daniel Lam, Junhong Xu, Roberto Martín-Martín

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 提出OopsieVerse框架,通过DamageSim检测和量化接触力、温度等损伤信号,在OmniGibson和RoboCasa仿真器中实现损伤感知,用于安全策略学习与评估。

Comments Project website: https://robin-lab.cs.utexas.edu/oopsieverse/. The first two authors contributed equally; order decided by dice roll. Accepted to Robotics: Science and Systems (RSS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31909 2026-07-01 cs.RO 新提交

CoDex: Learning Compositional Dexterous Functional Manipulation without Demonstrations

CoDex: 无需演示学习组合式灵巧功能性操作

Bowen Jiang, William Painter Reger, Roberto Martin-Martin

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 提出零演示框架CoDex,结合视觉语言模型与强化学习,自主发现并执行涉及物体内部机制操控的灵巧操作策略,在多种工具任务中验证了其有效性。

Comments IEEE International Conference on Robotics and Automation (ICRA) 2026. Project page: https://robin-lab.cs.utexas.edu/CoDex/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31813 2026-07-01 cs.LG cs.AI 新提交

Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in RLVR

保持几何的正交初始化用于RLVR中的低秩适应

Ruijia Zhang, Jiacheng Zhu, Hanqing Zhu, Laixi Shi

机构 * Johns Hopkins University(约翰霍普金斯大学) Meta Superintelligence Labs(Meta超级智能实验室) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 针对RLVR中LoRA初始化问题,提出保持几何的正交初始化方法(RLPO和RLMO),理论证明其最小化与全微调差距,实验表明稳定训练且优于标准LoRA。

Comments 30 pages, accepted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.06790 2026-07-01 cs.RO cs.LG cs.SY eess.SY 新提交

Learning All-Terrain Locomotion for a Planetary Rover with Actively Articulated Suspension

学习具有主动铰接悬挂的行星探测车的全地形运动

Arthur Bouton, Tristan D. Hasseler, Michael Paton, Travis Brown, Jacob Levy, William Reid, Joshua Martin, Hari Nayar

机构 * Jet Propulsion Laboratory, California Institute of Technology(喷气推进实验室,加州理工学院) Center for Autonomy, University of Texas at Austin(自主性中心,德克萨斯大学奥斯汀分校) Space Systems Laboratory, University of Maryland(空间系统实验室,马里兰大学)

AI总结 提出一种带有主动万向悬挂的四轮行星探测车概念,利用强化学习训练单一神经网络控制器,实现自主障碍协商和全地形运动,通过策略整合和零样本迁移在物理车上验证。

Comments 21 pages, 26 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16757 2026-07-01 eess.AS cs.AI 版本更新

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation

重新审视音频-语言预训练以学习通用音频表示

Wei-Cheng Tseng, Xuanru Zhou, Mingyue Huo, Yiwen Shao, Hao Zhang, Dong Yu

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) Zhejiang University(浙江大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Tencent AI Lab Seattle(腾讯AI实验室西雅图)

AI总结 本文通过构建大规模数据集CaptionStew并系统对比对比学习和字幕生成目标,揭示了音频-语言预训练在数据效率和可扩展性上的权衡,为通用音频表示学习提供了实证指导。

Comments ACL 2026 Main. Code available at https://github.com/AudenAI/Auden

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18521 2026-07-01 cs.CV 版本更新

Improved Immiscible Diffusion: Accelerate Diffusion Training by Reducing Its Miscibility

改进的不可混扩散:通过降低可混性加速扩散训练

Yiheng Li, Feng Liang, Dan Kondratyuk, Masayoshi Tomizuka, Kurt Keutzer, Chenfeng Xu

机构 * UC Berkeley(加州大学伯克利分校) UT Austin(德克萨斯大学奥斯汀分校)

AI总结 提出通过降低任意层级的可混性来加速扩散模型训练,实现多种实现方式(如KNN噪声选择和图像缩放),在多种任务上获得高达4倍以上的训练加速。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30491 2026-06-30 cs.CL cs.AI

SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation

SIMAX: 一种可扩展且可解释的多保真度标注医患对话模拟框架

Zhuhan Bao, Rui Yang, Bohao Yang, Zhiyi Liu, Sicheng Shu, Ruio Heerschap, Le Li, Doris Yang, Elisabeth Bond, Haoyuan Wang, Nicoleta Economou-Zavlanos, Joshua M. Biro, Matthew McDermott, Nan Liu, Anand Chowdhury, Kai Sun, Kathryn Pollak, Ed Hammond, Chuan Hong

机构 * Department of Biostatistics and Bioinformatics, Duke University School of Medicine(杜克大学医学学院生物统计学与生物信息学系) Duke-NUS AI + Medical Sciences Initiative, Duke-NUS Medical School(杜克-新加坡国立大学医学科学院AI+医学科学计划) Centre for Biomedical Data Science, Duke-NUS Medical School(杜克-新加坡国立大学医学学院生物医学数据科学中心) Department of Statistical Science, Duke University(杜克大学统计科学系) Leiden University Medical Centre(莱顿大学医学中心) Department of Mathematics, University of Texas at Austin(德克萨斯大学奥斯汀分校数学系) Department of Internal Medicine, Yale School of Medicine(耶鲁医学院内科学系) Department of Biostatistics, Epidemiology and Informatics, Perelman School of Medicine, University of Pennsylvania(宾夕法尼亚大学佩尔曼医学院生物统计学、流行病学与信息学系) The Graduate Group in Applied Mathematics and Computational Science, School of Arts and Sciences, University of Pennsylvania(宾夕法尼亚大学艺术与科学学院应用数学与计算科学联合组) Medstar Health National Center for Human Factors in Healthcare, Washington, DC, USA(Medstar健康国家人因工程中心,华盛顿特区,美国) Department of Biomedical Informatics, Columbia University(哥伦比亚大学生物医学信息学系) Cancer Prevention and Control, Duke Cancer Institute, Durham, NC, USA(杜克癌症研究所癌症预防与控制部,达勒姆,北卡罗来纳州,美国) Department of Population Health Sciences, Duke University School of Medicine(杜克大学医学学院流行病学与公共卫生系) Division of Rheumatology and Immunology, Duke University School of Medicine(杜克大学医学学院风湿病学与免疫学系) Pre-hospital and Emergency Research Centre, Health Services Research and Population Health, Duke-NUS Medical School(杜克-新加坡国立大学医学学院院前急救与应急研究中心,健康服务研究与人口健康) NUS Artificial Intelligence Institute, National University of Singapore(新加坡国立大学人工智能研究所) Division of Pulmonary, Allergy and Critical Care Medicine, Duke University School of Medicine(杜克大学医学学院呼吸科、过敏科与危重医学系) Duke Center for Health Informatics, Duke University(杜克大学健康信息学中心) Duke Clinical Research Institute, Durham, NC, USA(杜克临床研究中心,达勒姆,北卡罗来纳州,美国)

AI总结 提出SIMAX框架,通过预定义场景、角色和沟通行为生成可控医患对话,自动评估显示语音自然度和转录保真度良好,可用于开发和验证沟通编码系统。

详情

展开后加载摘要…

URL PDF HTML 收藏