arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

Carnegie Mellon University(卡内基梅隆大学)

2026-04-08 至 2026-04-08 共收录 13
2604.06138 2026-04-08 cs.SD cs.AI

Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization

生成用于长文本音频摘要的合成医生-患者对话

Yanis Labrak, David Grünert, Séverin Baroudi, Jiyun Chun, Pawel Cyrta, Sergio Burdisso, Ahmed Hassoon, David Liu, Adam Rothschild, Reed Van Deusen, Petr Motlicek, Andrew Perrault, Ricard Marxer, Thomas Schaaf

机构 * Idiap Research Institute(Idiap 研究所) University of Zurich(苏黎世大学) The Ohio State University(俄亥俄州立大学) Université de Toulon, Aix Marseille Univ, LIS, CNRS(土伦大学,艾克斯-马赛大学,计算机与系统实验室,法国国家科学研究中心) Stenograf Johns Hopkins University Bloomberg School of Public Health(约翰·霍普金斯大学布隆伯格公共卫生学院) Colorado School of Mines(科罗拉多矿业学院) Allegheny Health Network(阿勒格尼健康网络) University of Pittsburgh Medical Center(匹兹堡大学医学中心) ILLS, CNRS(ILLS,法国国家科学研究中心) Solventum Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出生成合成医生-患者对话数据的方法,用于长文本音频摘要任务,通过多阶段流程生成包含音频和参考笔记的数据集,评估显示级联方法优于端到端模型。

Comments Submitted for review at Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06126 2026-04-08 cs.LG cs.AI

Gym-Anything: Turn any Software into an Agent Environment

Gym-Anything: 将任意软件转化为智能体环境

Pranjal Aggarwal, Graham Neubig, Sean Welleck

机构 * CMU(卡内基梅隆大学)

AI总结 本文提出Gym-Anything框架,通过多智能体任务将任意软件转化为交互式计算机使用环境,生成涵盖医疗、天文、工程等领域的10K+长周期任务,提升现实场景下的智能体研究效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00503 2026-04-08 cs.CV

Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models

Diff4Splat:基于潜在动态重建模型的可控4D场景生成

Panwang Pan, Chenguo Lin, Jingjing Zhao, Chenxin Li, Yuchen Lin, Haopeng Li, Honglei Yan, Kairun Wen, Yunlong Lin, Yixuan Yuan, Yadong Mu

机构 * Peking University(北京大学) Xiamen University(厦门大学) CUHK(香港中文大学) Carnegie Mellon University(卡内基梅隆大学)

AI总结 Diff4Splat通过单张图像和相机轨迹生成可控的4D场景,结合视频扩散模型的生成先验与大规模4D数据集学习的几何和运动约束,在单次前向传递中直接预测可变形3D高斯场,无需测试时优化。

Comments Accepted to CVPR 2026. Project page: https://paulpanwang.github.io/Diff4Splat

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05084 2026-04-08 cs.LG stat.ML

Distribution-dependent Generalization Bounds for Tuning Linear Regression Across Tasks

跨任务线性回归中调节正则化超参数的分布依赖泛化界限

Maria-Florina Balcan, Saumya Goyal, Dravyansh Sharma

机构 * School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院) Machine Learning Department, Carnegie Mellon University(卡内基梅隆大学机器学习系) Toyota Technological Institute at Chicago(芝加哥丰田技术学院) Northwestern University(西北大学)

AI总结 本文研究了跨多个相关任务调节线性回归正则化超参数的方法,提出了分布依赖的泛化误差界,改进了传统方法在高维数据下的表现。

Comments 55 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13040 2026-04-08 cs.CV

MAMMA: Markerless & Automatic Multi-Person Motion Action Capture

MAMMA:无标记且自动的多人物动作捕捉

Hanz Cuevas-Velasquez, Anastasios Yiannakidis, Soyong Shin, Giorgio Becherini, Markus Höschle, Joachim Tesch, Taylor Obersat, Tsvetelina Alexiadis, Eni Halilaj, Michael J. Black

机构 * Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文提出MAMMA,一种无标记动作捕捉管道,通过多视角视频准确恢复SMPL-X参数。该方法通过预测密集的2D接触感知表面地标,实现即使在重叠情况下也能估计人物对应关系,且在无需标记的情况下实现高精度动作捕捉。

Comments Main paper and supplementary material

Journal ref CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05519 2026-04-08 eess.AS cs.HC cs.LG cs.SD eess.SP

Active noise cancellation on open-ear smart glasses

开放式智能眼镜上的主动降噪

Kuang Yuan, Freddy Yifei Liu, Tong Xiao, Yiwen Song, Chengyi Shen, Saksham Bhutani, Justin Chan, Swarun Kumar

机构 * Department of Electrical and Computer Engineering, Carnegie Mellon University(卡内基梅隆大学电气与计算机工程系) Department of Medical Physics and Acoustics, Carl von Ossietzky Universität Oldenburg(奥尔登堡大学医学物理与声学系) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

AI总结 本文提出首个实时主动降噪系统,用于开放式智能眼镜,通过微型开放式扬声器和麦克风阵列实现环境噪声抑制,实验显示在100-1000Hz频段内可实现9.6dB至11.2dB的降噪效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05484 2026-04-08 cs.RO cs.CV

CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment

CoEnv:通过组合环境驱动具身体验多智能体协作

Li Kang, Yutao Fan, Rui Li, Heng Zhou, Yiran Qin, Zhemeng Zhang, Songtao Huang, Xiufeng Song, Zaibin Zhang, Bruno N. Y. Chen, Zhenfei Yin, Dongzhan Zhou, Wangmeng Zuo, Lei Bai

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Harbin Institute of Technology(哈尔滨工业大学) University of Science and Technology of China(中国科学技术大学) CUHK-Shenzhen(香港中文大学(深圳)) Fudan University(复旦大学) Dalian University of Technology(大连理工大学) Carnegie Mellon University(卡内基梅隆大学) University of Oxford(牛津大学)

AI总结 本文提出CoEnv框架,通过模拟与现实结合的组合环境,实现多智能体协作的高效执行与安全部署,提升任务成功率和效率。

Comments 31 pages, 8 figures, including supplementary material. Project page: https://faceong.github.io/CoEnv/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18911 2026-04-08 cs.LG

From Human-Level AI Tales to AI Leveling Human Scales

从人类水平AI故事到AI人类尺度

Peter Romero, Fernando Martínez-Plumed, Zachary R. Tidler, Matthieu Téhénan, Sipeng Chen, Álvaro David Gómez Antón, Luning Sun, Manuel Cebrian, Lexin Zhou, Yael Moros Daval, Daniel Romero-Alvarado, Félix Martí Pérez, Kevin Wei, José Hernández-Orallo

机构 * Valencian Research Institute of Artificial Intelligence, Universitat Politècnica de València(瓦伦西亚人工智能研究所,瓦伦西亚理工大学) Leverhulme Centre for the Future of Intelligence, University of Cambridge(勒弗休姆未来智能中心,剑桥大学) The Psychometrics Centre, University of Cambridge(心理测量中心,剑桥大学) Georgia Institute of Technology(佐治亚理工学院) Harvard University(哈佛大学) Center for Automation and Robotics, Spanish National Research Council(自动化与机器人中心,西班牙国家研究委员会) Department of Computer Science, Princeton University(普林斯顿大学计算机科学系) Carnegie Mellon University(卡内基梅隆大学) University of Cambridge, Department of Computer Sciences and Technology(剑桥大学计算机科学与技术系)

AI总结 本文提出一种框架,通过校准项目以'世界人口'为基准,建立人类锚定的共同尺度,以更准确评估AI能力。

Comments 23 pages, 10 figures. submitted to ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08336 2026-04-08 cs.CL cs.CV

From Reasoning to Pixels: Benchmarking the Alignment Gap in Unified Multimodal Models

从推理到像素:统一多模态模型中对齐差距的基准测试

Cheng Yang, Chufan Shi, Bo Shui, Yaokang Wu, Muzi Tao, Huijuan Wang, Ivan Yee Lee, Yong Liu, Xuezhe Ma, Taylor Berg-Kirkpatrick

机构 * University of California San Diego(加利福尼亚大学圣迭戈分校) University of Southern California(南加利福尼亚大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文通过UReason基准测试,探讨统一多模态模型中模态对齐问题,发现去上下文生成在图像生成任务中表现更优,揭示了文本推理与生成图像之间存在对齐差距。

Comments Project page: https://ureason.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09365 2026-04-08 cs.CL cs.AI

Frame of Reference: Addressing the Challenges of Common Ground Representation in Situational Dialogs

参考框架:解决情境对话中共同地面表示的挑战

Biswesh Mohapatra, Théo Charlot, Giovanni Duca, Mayank Palan, Laurent Romary, Justine Cassell

机构 * Inria(法国国家信息与自动化研究所) Nantes Université(南特大学) University of Trento(特伦托大学) VJTI Mumbai(孟买维杰扬特理工学院) Carnegie Mellon University(卡内基梅隆大学)

AI总结 本文研究情境对话中共同地面表示的挑战,通过动态共享环境中的关系参考建立共同地面,并提出基于强化学习改进表示方法的策略。

Comments Work accepted at ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.02949 2026-04-08 cs.CL cs.CV

ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly

ProMQA-Assembly:多模态装配任务问答数据集

Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Susan Holm, Yuran Wang, Vincent Zhou, Ken Fukuda, Teruko Mitamura

机构 * Language Technologies Institute, Carnegie Mellon University(卡内基梅隆大学语言技术研究所) National Institute of Advanced Industrial Science and Technology (AIST)(国立研究开发法人产业技术综合研究所(AIST))

AI总结 本文提出ProMQA-Assembly数据集,包含646个多模态问答对,用于评估装配任务中的人机交互系统,通过半自动化标注方法生成问题并结合细粒度动作标签提升多样性,验证了推理模型在复杂多模态任务中的表现。

Comments LREC 2026. Code and data: https://github.com/kimihiroh/promqa-assembly

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12863 2026-04-08 cs.SD cs.AI cs.CV eess.AS

Unified Cross-modal Translation of Score Images, Symbolic Music, and Performance Audio

统一的跨模态评分图像、符号音乐和表演音频翻译

Jongmin Jung, Dongmin Kim, Sihun Lee, Seola Cho, Hyungjoon Soh, Irmak Bukey, Chris Donahue, Dasaem Jeong

机构 * Department of Artificial Intelligence, Sogang University(西江大学人工智能系) Sogang Future Lab, Sogang University(西江大学未来实验室) Department of Physics Education, Seoul National University(首尔大学物理教育系) Computer Science Department, Carnegie Mellon University(卡内基梅隆大学计算机科学系) Department of Art & Technology, Sogang University(西江大学艺术与技术系)

AI总结 本文提出统一模型,通过大规模数据集和模态分词实现多模态翻译,提升光学音乐识别的符号错误率至13.67%并实现评分图像条件音频生成。

Comments Submitted to IEEE Transactions on Audio, Speech and Language Processing (TASLPRO)

Journal ref IEEE Transactions on Audio, Speech and Language Processing, vol. 34, pp. 1876-1891, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08528 2026-04-08 cs.CL cs.SD eess.AS

On The Landscape of Spoken Language Models: A Comprehensive Survey

关于语音语言模型的景观:全面综述

Siddhant Arora, Kai-Wei Chang, Chung-Ming Chien, Yifan Peng, Haibin Wu, Yossi Adi, Emmanuel Dupoux, Hung-Yi Lee, Karen Livescu, Shinji Watanabe

机构 * Carnegie Mellon University(卡内基梅隆大学) National Taiwan University(国立台湾大学) Toyota Technological Institute at Chicago(丰田芝加哥技术研究所) Hebrew University of Jerusalem(耶路撒冷希伯来大学) ENS - PSL, EHESS, CNRS(巴黎高等师范学院 - 巴黎文理研究大学、社会科学高等研究院、法国国家科学研究中心)

AI总结 本文综述了语音语言模型的发展,分析了其架构、训练和评估方法,探讨了关键挑战与未来方向。

Comments Published in Transactions on Machine Learning Research

详情

展开后加载摘要…

URL PDF HTML 收藏