arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Southern California(南加州大学)

2026-05-28 至 2026-05-28 共收录 8
2605.28805 2026-05-28 cs.CL cs.AI cs.CV cs.LG

OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration

OmniVerifier-M1: 具有显式结构化重校准的多模态元验证器

Xinchen Zhang, Bowei Liu, Jiale Liu, Chufan Shi, Yizhen Zhang, Junhong Liu, Youliang Zhang, Zhiheng Li, Yujiu Yang, Ling Yang

机构 * Tsinghua University(清华大学) Pennsylvania State University(宾夕法尼亚州立大学) University of Southern California(南加州大学) Microcyto Princeton University(普林斯顿大学)

AI总结 提出OmniVerifier-M1,通过符号化元验证(如边界框)和解耦强化学习,实现多模态大模型的可靠细粒度验证与动态区域级自校正。

Comments ICML 2026. Project: https://github.com/Cominclip/OmniVerifier

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27772 2026-05-28 cs.SD cs.LG

Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox

音频大语言模型是听还是读?使用VoxParadox分析和缓解副语言失败

Jiacheng Pang, Ashutosh Chaubey, Mohammad Soleymani

机构 * Institute for Creative Technologies, University of Southern California, Los Angeles, USA(创意技术研究所,南加州大学,洛杉矶,美国)

AI总结 针对音频大语言模型在副语言理解上的不足,提出对抗性基准VoxParadox和Prompt-Conditioned Layer Mixer方法,显著提升模型对副语言线索的利用能力。

Comments Accepted as a conference paper at ICML 2026. Project page: https://voxparadox.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27759 2026-05-28 cs.RO

Colosseum V2: Benchmarking Generalization for Vision Language Action Models

Colosseum V2:视觉语言动作模型的泛化能力基准测试

Jeremy Morgan, Prajwal Vijay, Hyeonho Oh, Jincen Song, Ashvin Arora, Alina Du, Gaurav Sukhatme, Jesse Thomason, Ishika Singh

机构 * Department of Computer Science, University of Southern California(南加州大学计算机科学系) Department of Electrical Engineering, Indian Institute of Technology Madras(印度理工学院Madras分校电子工程系) Fu Foundation School of Engineering and Applied Science, Columbia University(哥伦比亚大学工程与应用科学学院)

AI总结 提出Colosseum V2大规模仿真基准,通过28个任务和两种机器人形态,系统评估VLA模型在分布偏移下的泛化能力,揭示其在高层次理解与鲁棒行为之间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27533 2026-05-28 cs.RO

Inducing Calmness With Pocket-Sized Robotics: Reducing Movement and Heart Rate in Children through Hand-Held Tactile Interactions

用口袋大小的机器人诱导平静:通过手持触觉交互降低儿童的心率和运动

Morten Roed Frederiksen, Kasper Støy, Maja Matarić

机构 * Data Systems and Robotics, IT-University of Copenhagen(数据系统与机器人,丹麦IT大学) Interaction Lab, University of Southern California(交互实验室,南加州大学)

AI总结 本研究通过手持触觉设备上的节奏振动匹配游戏,发现触觉交互能显著降低儿童的生理唤醒(心率下降3.56 bpm)和身体躁动(整体运动减少38%),从而促进平静和专注状态。

Comments 34 pages, 2 tables, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27388 2026-05-28 cs.CL cs.AI cs.SI

Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities

通过反应语气建模社区态度:评估LLM与在线社区语言行为对齐的人机协作框架

Nuan Wen, Xuezhe Ma

机构 * Information Sciences Institute University of Southern California(南加州大学信息科学研究所)

AI总结 提出CARE框架,通过细粒度言语气势分析,评估LLM模拟社区对真实新闻的反应,揭示其存在“现实主义差距”,表明当前对齐策略不足以捕捉在线群体的社会语言动态。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.15864 2026-05-28 cs.CV cs.CL

Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination

VLMs 是在看还是只是在说?揭示视觉重新检查的幻觉

Chufan Shi, Cheng Yang, Yaokang Wu, Linghao Jin, Bo Shui, Taylor Berg-Kirkpatrick, Xuezhe Ma

机构 * University of Southern California(南加州大学) University of California San Diego(加州大学圣地亚哥分校) Carnegie Mellon University(卡内基梅隆大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 通过图像交换探测框架 VisualSwap 和 800 对图像基准 VS-Bench,发现视觉语言模型在推理时声称的“重新检查图像”多为文本模式,而非真正的视觉重新检查,且思考模型更易受影响,用户指令可恢复视觉基础但自我反思无效。

Comments ICML 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24938 2026-05-28 cs.LG cs.AI cs.CL

Rethinking Layer Redundancy: Calibration Matters More Than Search in LLM Depth Pruning

重新思考层冗余:校准比搜索在LLM深度剪枝中更重要

Minkyu Kim, Vincent-Daniel Yun, Youngrae Kim, Suin Cho, Woosang Lim, Sunwoo Lee

机构 * Neural Superintelligence Lab, MODULABS(神经超智能实验室,MODULABS) University of Southern California(南加州大学) Boston University(波士顿大学) Seoul National University(首尔国立大学) Inha University(inha大学)

AI总结 本文通过实验发现,在大型语言模型深度剪枝中,校准配置对剪枝模式和性能的影响远大于搜索算法的选择。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00349 2026-05-28 cs.AI cs.MA

COOP$^2$: Defining, Observing, and Repairing Cooperation in LLM Multi-Agent Systems

COOP$^2$: 定义、观察和修复LLM多智能体系统中的合作

Hanqing Yang, Narjes Nourzad, Shiyu Chen, Marie Siew, Jingdi Chen, Carlee Joe-Wong

机构 * Carnegie Mellon University(卡内基梅隆大学) University of Southern California(南加州大学) Singapore University of Technology and Design(新加坡科技设计大学) University of Arizona(亚利桑那大学)

AI总结 提出COOP$^2$框架,通过将高层合作动态与任务进度关联,定义可验证的合作任务,并开发COOP$^2$-Repair方法预测约束失败并引导修复,提升LLM多智能体系统的任务成功率和约束满足度。

详情

展开后加载摘要…

URL PDF HTML 收藏