arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Carnegie Mellon University(卡内基梅隆大学)

共收录 2342
2602.08336 2026-04-08 cs.CL cs.CV

From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models

从推理到像素:统一多模态模型中对齐差距的基准测试

Cheng Yang, Chufan Shi, Bo Shui, Yaokang Wu, Muzi Tao, Huijuan Wang, Ivan Yee Lee, Yong Liu, Xuezhe Ma, Taylor Berg-Kirkpatrick

机构 * University of California San Diego(加利福尼亚大学圣迭戈分校) University of Southern California(南加利福尼亚大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文通过UReason基准测试,探讨统一多模态模型中模态对齐问题,发现去上下文生成在图像生成任务中表现更优,揭示了文本推理与生成图像之间存在对齐差距。

Comments Project page: https://ureason.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09365 2026-04-08 cs.CL cs.AI

Frame of Reference: Addressing the Challenges of Common Ground Representation in Situational Dialogs

参考框架:解决情境对话中共同地面表示的挑战

Biswesh Mohapatra, Théo Charlot, Giovanni Duca, Mayank Palan, Laurent Romary, Justine Cassell

机构 * Inria(法国国家信息与自动化研究所) Nantes Université(南特大学) University of Trento(特伦托大学) VJTI Mumbai(孟买维杰扬特理工学院) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文研究情境对话中共同地面表示的挑战,通过动态共享环境中的关系参考建立共同地面,并提出基于强化学习改进表示方法的策略。

Comments Work accepted at ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02949 2026-04-08 cs.CL cs.CV

ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly

ProMQA-Assembly:多模态装配任务问答数据集

Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Susan Holm, Yuran Wang, Vincent Zhou, Ken Fukuda, Teruko Mitamura

机构 * Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所) National Institute of Advanced Industrial Science and Technology (AIST)(国立研究开发法人产业技术综合研究所(AIST))

AI总结 本文提出ProMQA-Assembly数据集,包含646个多模态问答对,用于评估装配任务中的人机交互系统,通过半自动化标注方法生成问题并结合细粒度动作标签提升多样性,验证了推理模型在复杂多模态任务中的表现。

Comments LREC 2026. Code and data: https://github.com/kimihiroh/promqa-assembly

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12863 2026-04-08 cs.SD cs.AI cs.CV eess.AS

Unified Cross-modal Translation of Score Images, Symbolic Music, and Performance Audio

统一的跨模态评分图像、符号音乐和表演音频翻译

Jongmin Jung, Dongmin Kim, Sihun Lee, Seola Cho, Hyungjoon Soh, Irmak Bukey, Chris Donahue, Dasaem Jeong

机构 * Department of Artificial Intelligence, Sogang University(西江大学人工智能系) Sogang Future Lab, Sogang University(西江大学未来实验室) Department of Physics Education, Seoul National University(首尔大学物理教育系) Computer Science Department, Carnegie Mellon University(卡内基梅隆大学计算机科学系) Department of Art & Technology, Sogang University(西江大学艺术与技术系)

AI总结 本文提出统一模型,通过大规模数据集和模态分词实现多模态翻译,提升光学音乐识别的符号错误率至13.67%并实现评分图像条件音频生成。

Comments Submitted to IEEE Transactions on Audio, Speech and Language Processing (TASLPRO)

Journal ref IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pp. 1876-1891, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08528 2026-04-08 cs.CL cs.SD eess.AS

On The Landscape of Spoken Language Models: A Comprehensive Survey

关于语音语言模型的景观:全面综述

Siddhant Arora, Kai-Wei Chang, Chung-Ming Chien, Yifan Peng, Haibin Wu, Yossi Adi, Emmanuel Dupoux, Hung-Yi Lee, Karen Livescu, Shinji Watanabe

机构 * Carnegie Mellon University(卡内基梅隆大学) National Taiwan University(国立台湾大学) Toyota Technological Institute at Chicago(丰田芝加哥技术研究所) Hebrew University of Jerusalem(耶路撒冷希伯来大学) ENS - PSL, EHESS, CNRS(巴黎高等师范学院 - 巴黎文理研究大学、社会科学高等研究院、法国国家科学研究中心)

AI总结 本文综述了语音语言模型的发展,分析了其架构、训练和评估方法,探讨了关键挑战与未来方向。

Comments Published in Transactions on Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04898 2026-04-07 cs.AI cs.CL cs.LG

QED-Nano: Teaching a Tiny Model to Prove Hard Theorems

QED-Nano:用小型模型证明难题定理

LM-Provers, Yuxiao Qu, Amrith Setlur, Jasper Dekoninck, Edward Beeching, Jia Li, Ian Wu, Lewis Tunstall, Aviral Kumar

机构 * CMU(卡内基梅隆大学) Hugging Face ETH Zurich(苏黎世联邦理工学院) Project Numina

AI总结 本文提出QED-Nano模型,通过三个阶段训练在数学竞赛证明任务中超越大模型,以较低成本实现高性能证明生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04808 2026-04-07 cs.LG cs.AI

Selecting Decision-Relevant Concepts in Reinforcement Learning

强化学习中决策相关概念的选择

Naveen Raman, Stephanie Milani, Fei Fang

机构 * Carnegie Mellon University(卡内基梅隆大学) New York University(纽约大学)

AI总结 本文提出决策相关选择算法,通过状态抽象视角自动选择决策相关的概念,提升强化学习策略的可解释性和性能。

Comments 16 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04704 2026-04-07 cs.CL

IDIOLEX: Unified and Continuous Representations for Idiolectal and Stylistic Variation

IDIOLEX:统一且连续的语用与风格变异表示

Anjali Kantharuban, Aarohi Srivastava, Fahim Faisal, Orevaoghene Ahia, Antonios Anastasopoulos, David Chiang, Yulia Tsvetkov, Graham Neubig

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Notre Dame(圣母大学) George Mason University(乔治梅森大学) University of Washington(华盛顿大学) Archimedes Research Unit(阿基米德研究单元)

AI总结 本文提出IDIOLEX框架,通过结合句子来源信息与语言特征,学习连续的风格和方言表示,用于分析和分类,并用于对齐语言模型的风格。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04648 2026-04-07 cs.LG

From Curiosity to Caution: Mitigating Reward Hacking for Best-of-N with Pessimism

从好奇心到谨慎:通过悲观主义缓解最佳-N中的奖励黑客

Zhuohao Yu, Zhiwei Steven Wu, Adam Block

机构 * Carnegie Mellon University(卡内基梅隆大学) Columbia University(哥伦比亚大学)

AI总结 本文提出一种基于悲观主义的缓解方法,通过降低不典型响应的奖励估计来避免奖励黑客,证明其在最佳-N采样中有效。

Comments 29 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01487 2026-04-07 cs.AI cs.SI

AgentSocialBench: Evaluating Privacy Risks in Human-Centered Agentic Social Networks

AgentSocialBench: 评估人类中心代理社交网络中的隐私风险

Prince Zizhuang Wang, Shuli Jiang

机构 * Carnegie Mellon University(卡内基梅隆大学)

AI总结 研究探讨了人类中心代理社交网络中的隐私挑战,提出AgentSocialBench基准,揭示代理协作中隐私保护的复杂性及现有LLM在隐私保护上的不足。

Comments 43 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22560 2026-04-07 cs.RO

Allometric Scaling Laws for Bipedal Robots

双足机器人的等比定律

Naomi Oke, Aja M. Carter, Ben Gu, Steven Man, Cordelia Pride, Sarah Bergbreiter, Aaron M. Johnson

机构 * Department of Mechanical Engineering, Carnegie Mellon University(卡内基梅隆大学机械工程系) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)

AI总结 本文研究双足机器人腿部长度与质量、速度等参数的等比关系,发现机器人质量与腿部长度平方成正比,与生物等比定律不同。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08406 2026-04-07 cs.HC cs.CL

Sandpiper: Orchestrated AI-Annotation for Educational Discourse at Scale

Sandpiper:大规模教育对话中的协同AI标注

Daryl Hedley, Doug Pietrzak, Jorge Dias, Ian Burden, Bakhtawar Ahtisham, Zhuqian Zhou, Kirk Vanacore, Josh Marland, Rachel Slama, Justin Reich, Kenneth Koedinger, René Kizilcec

机构 * National Tutoring Observatory(国家辅导观察站) FreshCognate Cornell University(康奈尔大学) Massachusetts Institute of Technology(麻省理工学院) Carnegie Mellon University(卡内基梅隆大学)

AI总结 Sandpiper通过结合交互式研究仪表板与大语言模型,实现高体积对话数据与人类定性专家的协同分析,提升研究效率和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07943 2026-04-07 cs.AI

IV Co-Scientist: Multi-Agent LLM Framework for Causal Instrumental Variable Discovery

IV共科学家:多智能体LLM框架用于因果工具变量发现

Ivaxi Sheth, Zhijing Jin, Bryan Wilder, Dominik Janzing, Mario Fritz

机构 * CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心) Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) Jinesis AI Lab(Jinesis AI实验室) University of Toronto(多伦多大学) Vector Institute(向量研究所) EuroSafeAI Carnegie Mellon University(卡内基梅隆大学) Amazon(亚马逊)

AI总结 本文探讨LLM在发现因果工具变量中的作用,提出IV共科学家多智能体系统,通过测试和评估LLM恢复已知工具及避免无效工具的能力,展示LLM在大规模观测数据中发现有效工具的潜力。

Comments Paper accepted at CleaR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11109 2026-04-07 cs.CV cs.AI cs.GR

Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning

通过交错多模态推理的视觉-反图形代理

Shaofeng Yin, Jiaxin Ge, Zora Zhiruo Wang, Chenyang Wang, Xiuyu Li, Michael J. Black, Trevor Darrell, Angjoo Kanazawa, Haiwen Feng

机构 * University of California, Berkeley(加州大学伯克利分校) Carnegie Mellon University(卡内基梅隆大学) Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) Impossible, Inc.(Impossible公司)

AI总结 本文提出VIGA框架,通过符号逻辑与视觉感知的交叉验证,解决视觉语言模型在单次设置中重建图像的挑战,支持2D文档生成、3D重建等任务。

Comments Project page: https://fugtemypt123.github.io/VIGA-website/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18052 2026-04-07 cs.CL cs.CY

The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies

PIMMUR 原则:确保 LLM 社会集体行为的有效性

Jiaxu Zhou, Jen-tse Huang, Xuhui Zhou, Man Ho Lam, Xintao Wang, Hao Zhu, Wenxuan Wang, Maarten Sap

机构 * Chinese University of Hong Kong(香港中文大学) Johns Hopkins University(约翰斯·霍普金斯大学) Carnegie Mellon University(卡内基梅隆大学) Fudan University(复旦大学) Stanford University(斯坦福大学) Renmin University of China(中国人民大学)

AI总结 本文通过审计39项研究,识别出六个普遍缺陷,指出89.7%的研究违反至少一个原则,导致模拟有效性受损。通过复现实验发现,当遵循PIMMUR原则时,报告的集体现象往往消失或反转,表明许多'涌现'行为可能是方法学伪影而非真实社会动态。

Comments 13 pages, 9 figures, 3 tables; add more papers in our systematic audit (39 in total)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13130 2026-04-07 cs.CV cs.AI cs.CL

ZINA: Multimodal Fine-grained Hallucination Detection and Editing

ZINA:多模态细粒度幻觉检测与编辑

Yuiga Wada, Kazuki Matsuda, Komei Sugiura, Graham Neubig

机构 * Keio AI Research Center(庆应义塾大学人工智能研究中心) Keio University(庆应义塾大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出ZINA方法,用于多模态大语言模型的细粒度幻觉检测与编辑,通过构建VisionHall数据集验证了其在检测和编辑任务中的优越性。

Comments CVPR 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.03083 2026-04-07 stat.ME cs.LG stat.ML

Causal K-Means Clustering

因果k均值聚类

Kwangho Kim, Jisu Kim, Edward H. Kennedy

机构 * Department of Statistics, Korea University(高丽大学统计系) Department of Statistics, Seoul National University(首尔大学统计系) Department of Statistics and Data Science, Carnegie Mellon University(卡内基梅隆大学统计与数据科学系)

AI总结 本文提出因果k均值聚类方法,用于识别未知的子群体结构,通过非参数效率理论和双机器学习开发偏倚校正估计器,实现快速根n收敛和渐近正态性,适用于多治疗水平的现代结果研究。

Journal ref J. R. Stat. Soc. Ser. B, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04036 2026-04-07 cs.IR cs.CL

MisEdu-RAG: A Misconception-Aware Dual-Hypergraph RAG for Novice Math Teachers

MisEdu-RAG:一种面向初学者数学教师的误解意识双超图RAG

Zhihan Guo, Rundong Xue, Yuting Lu, Jionghao Lin

机构 * The University of Hong Kong(香港大学) Xi'an Jiaotong University(西安交通大学) Carnegie Mellon University(卡内基梅隆大学) Monash University(莫纳什大学)

AI总结 针对初学者数学教师难以诊断和纠正学生错误的问题,提出基于双超图的RAG框架,通过组织教学知识和学生错误案例,提升响应质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03964 2026-04-07 cs.AI

SKILLFOUNDRY: Building Self-Evolving Agent Skill Libraries from Heterogeneous Scientific Resources

SKILLFOUNDRY:从异构科学资源构建自演化智能体技能库

Shuaike Shen, Wenduo Cheng, Mingqian Ma, Alistair Turcan, Martin Jinye Zhang, Jian Ma

机构 * Ray and Stephanie Lane Computational Biology Department, School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院Ray and Stephanie Lane计算生物学系) Machine Learning Department, School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院机器学习系)

AI总结 SKILLFOUNDRY通过自演化框架将异构资源转化为验证过的智能体技能包,提升科学代理在基准和领域任务中的性能,扩展覆盖范围,为更强大的科学代理提供基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.03716 2026-04-07 cs.CV cs.GR

CGHair: Compact Gaussian Hair Reconstruction with Card Clustering

CGHair: 基于卡聚类的紧凑高斯发丝重建

Haimin Luo, Srinjay Sarkar, Albert Mosella-Montoro, Francisco Vicente Carrasco, Fernando De la Torre

机构 * Carnegie Mellon University(卡内基梅隆大学) ShanghaiTech University(上海科技大学)

AI总结 本文提出一种紧凑的高保真发丝重建方法,通过卡聚类减少存储和渲染成本,同时保持视觉质量。

Comments Accepted to CVPR 2026. This arXiv version is not the final published version

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06356 2026-04-07 cs.LG cs.SI

Mitigating Structural Overfitting: A Distribution-Aware Rectification Framework for Missing Feature Imputation

缓解结构过拟合:一种分布感知的缺失特征填补框架

Yifan Song, Fenglin Yu, Yihong Luo, Xingjian Tao, Siya Qiu, Kai Han, Jing Tang

机构 * Carnegie Mellon University(卡内基梅隆大学) Shanghai University of Finance and Economics(上海财经大学)

AI总结 本文提出DART框架,通过全局结构增强和语义校正机制,解决图学习中缺失特征填补的结构过拟合问题,并在六个公开数据集和Sailing数据集上验证了其有效性。

Comments Accepted by SIGIR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12301 2026-04-07 cs.CL cs.LG cs.SD eess.AS

WhisperRT -- Turning Whisper into a Causal Streaming Model

WhisperRT -- 将Whisper转化为因果流式模型

Tomer Krichli, Bhiksha Raj, Joseph Keshet

机构 * Andrew and Erna Viterbi Faculty of Electrical and Computer Engineering, Technion–Israel Institute of Technology(以色列理工学院安德鲁与厄娜·维特比电气与计算机工程学院) Language Technologies Institute, School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院语言技术研究所)

AI总结 本文提出将Whisper模型转化为低延迟流式模型的方法,通过使编码器因果化并生成与可用时间上下文对齐的token,实现更高效的流式语音识别。

Comments 14 pages, 7 Figures, This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06649 2026-04-06 cs.LG cs.AI

Local Reinforcement Learning with Action-Conditioned Root Mean Squared Q-Functions

局部强化学习与基于动作的均方Q函数

Frank Wu, Mengye Ren

机构 * Carnegie Mellon University(卡内基梅隆大学) New York University(纽约大学)

AI总结 本文提出基于动作的均方Q函数,通过时间差分学习实现局部强化学习,优于现有无反向传播方法,在MinAtar和DeepMind控制套件中表现优异。

Comments 18 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22264 2026-04-06 cs.CV cs.AI

SmartCLIP: Modular Vision-language Alignment with Identification Guarantees

SmartCLIP: 基于识别保证的模块化视觉-语言对齐

Shaoan Xie, Lingjing Kong, Yujia Zheng, Yu Yao, Zeyu Tang, Eric P. Xing, Guangyi Chen, Kun Zhang

机构 * Carnegie Mellon University(卡内基梅隆大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) The University of Sydney(悉尼大学)

AI总结 本文提出SmartCLIP,通过理论条件实现文本与视觉表征的灵活对齐,确保语义信息完整保留并解耦视觉表征,提升多任务性能。

Comments CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13995 2026-04-06 cs.CL cs.AI cs.CY

ELEPHANT: Measuring and understanding social sycophancy in LLMs

ELEPHANT:测量和理解LLMs中的社交阿谀

Myra Cheng, Sunny Yu, Cinoo Lee, Pranav Khadpe, Lujain Ibrahim, Dan Jurafsky

机构 * Stanford University(斯坦福大学) Carnegie Mellon University(卡内基梅隆大学) University of Oxford(牛津大学)

AI总结 研究通过ELEPHANT基准测量LLMs中的社交阿谀行为,发现模型在保持用户形象方面比人类更极端,且在道德冲突中倾向于支持用户立场。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10720 2026-04-03 cs.LG

Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality

超越黑箱:通过因果最小性在生成模型中实现可识别的解释与控制

Lingjing Kong, Shaoan Xie, Guangyi Chen, Yuewen Sun, Xiangchen Song, Eric P. Xing, Kun Zhang

机构 * Carnegie Mellon University(卡内基梅隆大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

AI总结 本文通过因果最小性原则,在生成模型中建立可解释性基础,提出层级选择模型框架,实现对生成模型的可控性与解释性提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04665 2026-04-03 cs.LG cs.SY eess.SY

A Simultaneous Approach for Training Neural Differential-Algebraic Systems of Equations

神经微分代数方程的联合训练方法

Laurens R. Lueg, Victor Alves, Daniel Schicksnus, John R. Kitchin, Carl D. Laird, Lorenz T. Biegler

机构 * Carnegie Mellon University(卡内基梅隆大学) RWTH Aachen University(亚琛工业大学)

AI总结 本文提出一种联合训练方法用于神经微分代数方程,通过离散化非线性优化问题同时求解神经网络参数和DAE解,考虑混合模型并改进求解器性能,实现高精度、泛化性和计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.03432 2026-04-03 math.OC cs.GT cs.LG

A Polynomial-Time Algorithm for Variational Inequalities under the Minty Condition

变分不等式在 Minty 条件下的多项式时间算法

Ioannis Anagnostides, Gabriele Farina, Tuomas Sandholm, Brian Hu Zhang

机构 * Carnegie Mellon University(卡内基梅隆大学) Massachusetts Institute of Technology(麻省理工学院) Strategy Robot, Inc.(Strategy Robot公司) Strategic Machine, Inc.(Strategic Machine公司) Optimized Markets, Inc.(Optimized Markets公司)

AI总结 本文提出在 Minty 条件下求解 ε-变分不等式的新算法,具有多项式时间复杂度,解决了传统方法中依赖于 1/ε 的指数问题,并展示了其在博弈论中的应用。

Comments V3 polishes the writing and makes a correction to Theorem 5.2

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15961 2026-04-03 cs.RO

IA-TIGRIS: An Incremental and Adaptive Sampling-Based Planner for Online Informative Path Planning

IA-TIGRIS: 一种用于在线信息路径规划的增量和自适应采样规划器

Brady Moon, Nayana Suvarna, Andrew Jong, Satrajit Chatterjee, Junbin Yuan, Muqing Cao, Sebastian Scherer

机构 * Robotics Institute, School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院机器人研究所) Department of Mechanical Engineering, Brigham Young University(杨百翰大学机械工程系)

AI总结 本文提出IA-TIGRIS,一种用于实时任务的增量和自适应采样路径规划方法,通过增量优化和动态更新信念地图,提升信息获取效率,实验证明其在无人机应用中信息增益提升达38%。

Comments Published in IEEE Transactions on Robotics, 19 pages, 19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00489 2026-04-02 cs.CL

Adapting Text LLMs to Speech via Multimodal Depth Up-Scaling

通过多模态深度扩展适应文本LLM到语音

Kazuki Yano, Jun Suzuki, Shinji Watanabe

机构 * Tohoku University(东北大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出多模态深度扩展方法,通过在冻结的文本LLM中插入新层并仅训练这些层来适应语音数据,实验表明其在ASR性能上接近全微调,且文本能力退化更少。

详情

展开后加载摘要…

URL PDF HTML 收藏