arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Southern California(南加州大学)

共收录 1288
2208.13701 2026-03-16 stat.ME cs.LG math.OC stat.ML

Data-Driven Influence Functions for Optimization-Based Causal Inference

数据驱动的影响函数用于基于优化的因果推断

Michael I. Jordan, Yixin Wang, Angela Zhou

机构 * Department of EECS and Statistics(电气工程与计算机科学系和统计学系) University of California, Berkeley(加州大学伯克利分校) Department of Statistics(统计学系) University of Michigan(密歇根大学) Department of Data Sciences and Operations(数据科学与运营系) University of Southern California, Marshall School of Business(南加州大学马歇尔商学院)

AI总结 本文研究了一种构造性算法,通过有限差分近似统计函数的Gâteaux导数,重点在于因果推断中出现的函数。研究概率分布未知但需从数据估计的情况,并探讨经验、数值和分析Gâteaux导数之间的关系。

Comments Revision

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12625 2026-03-16 cs.IR cs.AI cs.CV

VLM4Rec: Multimodal Semantic Representation for Recommendation with Large Vision-Language Models

VLM4Rec:基于大视觉-语言模型的多模态语义表示推荐系统

Ty Valencia, Burak Barlas, Varun Singhal, Ruchir Bhatia, Wei Yang

机构 * University of Southern California(南加州大学)

AI总结 VLM4Rec通过语义对齐而非直接特征融合,提升多模态推荐性能,实验显示其优于原始视觉特征和融合方法。

Comments 13 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12260 2026-03-16 cs.RO

HumDex: Humanoid Dexterous Manipulation Made Easy

HumDex: 人形机器人灵活操控的便捷实现

Liang Heng, Yihe Tang, Jiajun Xu, Henghui Bao, Di Huang, Yue Wang

机构 * USC Physical Superintelligence (PSI) Lab(USC物理超智能实验室) WorldEngine AI

AI总结 本文提出HumDex系统,通过IMU运动追踪和学习导向方法,实现人形机器人灵活操控的高效数据采集与精准执行。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19531 2026-03-16 cs.LG cs.AI

A Statistical Approach for Modeling Irregular Multivariate Time Series with Missing Observations

一种用于具有缺失观测的不规则多变量时间序列的统计方法

Dingyi Nie, Yixing Wu, C. -C. Jay Kuo

机构 * University of Southern California(南加州大学)

AI总结 本文提出通过提取时间无关的汇总统计量来处理不规则多变量时间序列中的缺失值,通过计算四个关键特征并结合标准分类器,在四个生物医学数据集上取得最佳性能,同时降低计算复杂度。

Comments Accepted for publication in APSIPA Transactions on Signal and Information Processing

Journal ref APSIPA Transactions on Signal and Information Processing, 15(1): 61-75, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15342 2026-03-16 cs.CV

LowDiff: Efficient Diffusion Sampling with Low-Resolution Condition

LowDiff: 基于低分辨率条件的高效扩散采样

Jiuyi Xu, Qing Jin, Meida Chen, Andrew Feng, Yang Sui, Yangming Shi

机构 * Colorado School of Mines(科罗拉多矿学院) Independent Researcher(独立研究者) University of Southern California Institute for Creative Technologies(南加州大学创意技术研究所) Rice University(里士满大学)

AI总结 LowDiff通过级联方法生成逐步提升的高分辨率图像,利用统一模型逐步细化图像,实现更高效的扩散采样,提升生成效率和质量。

Comments 16 pages, 7 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10888 2026-03-12 cs.SD

VoxCare: Studying Natural Communication Behaviors of Hospital Caregivers through Wearable Sensing of Egocentric Audio

VoxCare: 通过可穿戴的自体音频传感研究医院护理人员的自然沟通行为

Tiantian Feng, Kleanthis Avramidis, Anfeng Xu, Deqi Wang, Brandon M Booth, Shrikanth Narayanan

机构 * University of Southern California(南加州大学) The University of Memphis(孟菲斯大学)

AI总结 VoxCare通过可穿戴音频传感研究医院护理人员的自然沟通行为,利用实时声学特征提取和教师-学生框架识别语音活动,揭示沟通频率、持续时间和语音唤醒度,为理解医疗专业人员行为提供数据驱动的方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10178 2026-03-12 cs.CV cs.CL

Video-Based Reward Modeling for Computer-Use Agents

基于视频的计算机使用代理奖励建模

Linxin Song, Jieyu Zhang, Huanxin Sheng, Taiwei Shi, Gupta Rahul, Yang Liu, Ranjay Krishna, Jian Kang, Jieyu Zhao

机构 * University of Southern California(南加州大学) University of Washington(华盛顿大学) MBZUAI(机器智能研究院) Amazon AGI(亚马逊人工通用智能)

AI总结 本研究提出基于视频的计算机使用代理奖励建模方法,通过ExeVR-53k数据集和对抗性指令翻译技术,实现高精度的任务成功预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09206 2026-03-11 cs.CV cs.LG

MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data

MM-Zero: 从零数据自我进化多模型视觉语言模型

Zongxia Li, Hongyang Du, Chengsong Huang, Xiyang Wu, Lantao Yu, Yicheng He, Jing Xie, Xiaomin Wu, Zhichao Liu, Jiarui Zhang, Fuxiao Liu

机构 * University of Maryland(马里兰大学) Brown University(布朗大学) Washington University in St. Louis(圣路易斯华盛顿大学) Adobe University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Southern California(南加州大学) NVIDIA

AI总结 MM-Zero通过多角色框架实现VLM零数据自我进化,提升多模态推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11790 2026-03-11 cs.LG cs.CR

JULI: Jailbreak Large Language Models by Self-Introspection

通过自我反思 jailbreak 大型语言模型:JULI

Jesson Wang, Zhanhao Hu, David Wagner

机构 * University of Southern California(南加州大学) University of California, Berkeley(加州大学伯克利分校)

AI总结 JULI 通过操纵令牌日志概率,利用微小插件块 BiasNet 实现对 API 调用 LLMs 的 jailbreak,无需模型权重或生成过程权限,且在黑盒环境下有效。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09034 2026-03-11 eess.AS cs.SD

Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition

神经音频编解码器中容量与鲁棒性之间的权衡

Jordan Prescott, Thanathai Lertpetchpun, Shrikanth Narayanan

机构 * Signal Analysis and Interpretation Laboratory, University of Southern California, USA(南加州大学信号分析与解释实验室)

AI总结 本文研究了神经音频编解码器中容量与鲁棒性之间的权衡,发现中间深度量化能平衡对抗扰动与语音内容质量,从而最小化转录错误。

Comments Submitted to Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09011 2026-03-11 cs.RO cs.AI cs.HC

Improving through Interaction: Searching Behavioral Representation Spaces with CMA-ES-IG

通过交互改进:利用CMA-ES-IG搜索行为表示空间

Nathaniel Dennler, Zhonghao Shi, Yiran Tao, Andreea Bobu, Stefanos Nikolaidis, Maja Matarić

机构 * CSAIL/AeroAstro, Massachusetts Institute of Technology(CSAIL/AeroAstro,麻省理工学院) Department of Computer Science, University of Southern California(计算机科学系,南加州大学)

AI总结 本文提出CMA-ES-IG算法,通过整合用户体验考虑,提升机器人行为偏好学习的效率和鲁棒性。

Comments Under submission to IJRR

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08936 2026-03-11 cs.SD cs.AI cs.CL cs.MM eess.AS

VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs

VoxEmo:基于语音大语言模型的语音情感识别基准测试

Hezhao Zhang, Huang-Cheng Chou, Shrikanth Narayanan, Thomas Hain

机构 * Department of Computer Science, University of Sheffield, United Kingdom(英国谢菲尔德大学计算机科学系) Signal Analysis and Interpretation Laboratory (SAIL), Ming Hsieh Department of Electrical and Computer Engineering, University of Southern California, Los Angeles, CA 90089, USA(美国南加州大学电气与计算机工程系信号分析与解释实验室(SAIL))

AI总结 VoxEmo是一个基于语音大语言模型的语音情感识别基准测试,通过提供多样化的提示复杂度和分布感知的软标签协议,评估模型在真实世界中的情感识别能力。

Comments submitted to Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07972 2026-03-10 cs.AI

Adaptive Collaboration with Humans: Metacognitive Policy Optimization for Multi-Agent LLMs with Continual Learning

适应性协作与人类:多智能体大语言模型的元认知策略优化与持续学习

Wei Yang, Defu Cao, Jiacheng Pang, Muyan Weng, Yan Liu

机构 * University of Southern California(南加州大学)

AI总结 HILA框架通过元认知策略优化与持续学习,实现多智能体与人类的协同协作,提升复杂任务处理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07800 2026-03-10 cs.RO

Preference-Conditioned Reinforcement Learning for Space-Time Efficient Online 3D Bin Packing

基于偏好条件的强化学习用于空间时间高效的在线3D装箱

Nikita Sarawgi, Omey M. Manyar, Fan Wang, Thinh H. Nguyen, Daniel Seita, Satyandra K. Gupta

机构 * Viterbi School of Engineering, University of Southern California(美国南加州大学维特比工程学院) Amazon Robotics(亚马逊机器人)

AI总结 STEP方法通过偏好条件强化学习,在保持装箱密度的同时将操作时间减少44%。

Comments 8 pages, 5 figures. Accepted to IEEE International Conference on Robotics and Automation 2026. Project Website: https://step-packing.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07796 2026-03-10 cs.RO

Inverse Resistive Force Theory (I-RFT): Learning granular properties through robot-terrain physical interactions

逆电阻力理论(I-RFT):通过机器人与地形的物理互动学习颗粒特性

Shipeng Liu, Feng Xue, Yifeng Zhang, Tarunika Ponnusamy, Feifei Qian

机构 * University of Southern California, Los Angeles, CA 90089, USA(美国南加州大学)

AI总结 I-RFT通过机器人与地形的物理互动学习颗粒特性,结合颗粒电阻力理论与高斯过程,实现对地形属性的准确估计与不确定性量化,提升自主探索效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07550 2026-03-10 cs.CL cs.AI

Learning-free L2-Accented Speech Generation using Phonological Rules

无需学习的L2口音语音生成:基于语音学规则

Thanathai Lertpetchpun, Yoonjeong Lee, Jihwan Lee, Tiantian Feng, Dani Byrd, Shrikanth Narayanan

机构 * Signal Analysis and Interpretation Lab, University of Southern California, USA(信号分析与解读实验室,美国南加州大学) Department of Linguistics, University of Southern California(语言学系,美国南加州大学)

AI总结 本文提出无需学习的L2口音语音生成方法,通过语音学规则与多语言TTS模型结合,在无需带口音数据的情况下实现音素级口音操控。

Comments Submitted to Interspeech2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07534 2026-03-10 cs.CL

Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data

Accent Vector: 多语言TTS中无需带 accents数据的可控 accent操控

Thanathai Lertpetchpun, Thanapat Trachu, Jihwan Lee, Tiantian Feng, Dani Byrd, Shrikanth Narayanan

机构 * Signal Analysis and Interpretation Lab, University of Southern California, USA(信号分析与解释实验室,南加州大学) Thomas Lord Department of Computer Science, University of Southern California, USA(托马斯·劳德计算机科学系,南加州大学) Department of Linguistics, University of Southern California(语言学系,南加州大学)

AI总结 本文提出Accent Vector,一种无需带口音数据的多语言TTS可控口音操控方法,通过微调和向量操作实现精细口音控制。

Comments Submitted to Interspeech2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07459 2026-03-10 cs.HC cs.AI

"Better Ask for Forgiveness than Permission": Practices and Policies of AI Disclosure in Freelance Work

更好的请求宽恕而非许可:自由职业工作中AI披露的实践与政策

Angel Hsing-Chi Hwang, Senya Wong, Baixiao Chen, Jessica He, Hyo Jin Do

机构 * University of Southern California(南加州大学) Emory University(埃默里大学) IBM Research(IBM研究院)

AI总结 本文探讨了自由职业工作中AI披露的实践与政策,揭示了工人与客户在披露期望上的差异及政策不明确带来的问题,提出需要更清晰的指导以促进信任和责任。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10110 2026-03-10 cs.RO cs.AI cs.LG

IMPACT: Intelligent Motion Planning with Acceptable Contact Trajectories via Vision-Language Models

IMPACT: 通过视觉-语言模型实现可接受接触轨迹的智能运动规划

Yiyang Ling, Karan Owalekar, Oluwatobiloba Adesanya, Erdem Bıyık, Daniel Seita

机构 * Thomas Lord Department of Computer Science, Viterbi School of Engineering, University of Southern California(托马斯·劳德计算机科学系,维特里比工程学院,南加州大学)

AI总结 IMPACT通过视觉-语言模型实现高效的接触丰富运动规划,在杂乱环境中优于其他方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06947 2026-03-10 cs.RO

Feasibility Restoration under Conflicting STL Specifications with Pareto-Optimal Refinement

在冲突的STL规范下进行可行性恢复的帕累托最优细化

Tianhao Wu, Yiwei Lyu

机构 * Department of Computer Science, University of Southern California(南加州大学计算机科学系) Department of Computer Science and Engineering, Texas A&M University(德克萨斯农工大学计算机科学与工程系)

AI总结 本文提出了一种两阶段框架,在冲突的STL规范下通过最小松弛恢复可行性,并通过多目标优化实现可解释的决策。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06816 2026-03-10 cs.CL cs.AI q-bio.NC

"Dark Triad" Model Organisms of Misalignment: Narrow Fine-Tuning Mirrors Human Antisocial Behavior

黑暗三联征模型生物:对齐偏差:狭窄微调映射人类反社会行为

Roshni Lulla, Fiona Collins, Sanaya Parekh, Thilo Hagendorff, Jonas Kaplan

机构 * Brain & Creativity Institute, University of Southern California(大脑与创造力研究所,南加州大学) Department of Psychology, University of Southern California(心理学系,南加州大学) Interchange Forum for Reflecting on Intelligent Systems, University of Stuttgart(智能系统反思交流论坛,斯图加特大学)

AI总结 本文通过黑暗三联征框架,研究LLM中的对齐偏差问题,通过微调诱导反社会行为,揭示LLM中潜在的人格结构。

Comments 38 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06600 2026-03-10 cs.LG cs.AI

FuzzingRL: Reinforcement Fuzz-Testing for Revealing VLM Failures

FuzzingRL: 用于揭示视觉语言模型故障的强化模糊测试

Jiajun Xu, Jiageng Mao, Ang Qi, Weiduo Yuan, Alexander Romanus, Helen Xia, Vitor Campagnolo Guizilini, Yue Wang

机构 * University of Southern California(南加州大学) Toyota Research Institute(丰田研究机构)

AI总结 FuzzingRL通过强化模糊测试生成挑战性问题,揭示视觉语言模型的故障点并降低其准确性。

Comments 18 pages, 4 figures. † These authors jointly supervised this work: Jiageng Mao and Yue Wang

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18570 2026-03-10 cs.LG

VISTA: Vision-Language Inference for Training-Free Stock Time-Series Analysis

VISTA:面向无训练股票时间序列分析的视觉-语言推理

Tina Khezresmaeilzadeh, Parsa Razmara, Seyedarmin Azizi, Mohammad Erfan Sadeghi, Erfan Baghaei Potraghloo

机构 * University of Southern California(南加州大学)

AI总结 VISTA通过结合文本和视觉信息,利用无训练的视觉-语言模型实现股票时间序列预测,实验结果显示其在预测精度上显著优于传统方法。

Comments Accepted to the CVPR 2025 Workshop on Transformers for Vision (T4V): accepted-papers" target="_blank" rel="noopener">https://sites.google.com/view/t4v-cvpr25/accepted-papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05887 2026-03-09 eess.AS cs.AI

Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec

重建!不要编码:面向高可懂性和低延迟流式神经音频编解码器的自监督表示重建损失

Junhyeok Lee, Xiluo He, Jihwan Lee, Helin Wang, Shrikanth Narayanan, Thomas Thebaud, Laureano Moro-Velazquez, Jesús Villalba, Najim Dehak

机构 * Center for Language and Speech Processing, Johns Hopkins University, USA(语言与语音处理中心,约翰霍普金斯大学,美国) Signal Analysis and Interpretation Laboratory, University of Southern California, USA(信号分析与解释实验室,南加州大学,美国)

AI总结 本文提出自监督表示重建损失,用于提升流式神经音频编解码器的可懂性和低延迟性能。

Comments Submitted to Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05449 2026-03-06 cs.CV cs.AI cs.GR

RealWonder: Real-Time Physical Action-Conditioned Video Generation

RealWonder: 基于实时物理动作的视频生成

Wei Liu, Ziyu Chen, Zizhang Li, Yue Wang, Hong-Xing Yu, Jiajun Wu

机构 * Stanford University(斯坦福大学) University of Southern California(南加州大学)

AI总结 RealWonder通过物理模拟实现实时动作条件视频生成,利用3D重建、物理模拟和简化视频生成器,在单张图像基础上生成高质量视频,适用于多种物理场景。

Comments The first two authors contributed equally. The last two authors advised equally. Project website: https://liuwei283.github.io/RealWonder/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04861 2026-03-06 cs.AI cs.LG cs.RO

Causally Robust Reward Learning from Reason-Augmented Preference Feedback

基于理由增强的偏好反馈的因果鲁棒奖励学习

Minjune Hwang, Yigit Korkmaz, Daniel Seita, Erdem Bıyık

机构 * Thomas Lord Department of Computer Science, University of Southern California(美国南加州大学计算机科学系)

AI总结 ReCouPLe通过自然语言理由提供因果信号,提升奖励学习在分布偏移和新任务中的性能。

Comments Published in International Conference on Learning Representations (ICLR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04746 2026-03-06 cs.AI cs.HC econ.GN q-fin.EC

Visioning Human-Agentic AI Teaming: Continuity, Tension, and Future Research

展望人机协同AI:连续性、张力与未来研究

Bowen Lou, Tian Lu, T. S. Raghu, Yingjie Zhang

机构 * University of Southern California(南加州大学) Arizona State University(亚利桑那州立大学) Peking University(北京大学)

AI总结 本文提出Team SA理论作为人机协同AI转型的整合锚点,探讨在适应性自主下持续对齐的挑战与未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00810 2026-03-05 cs.LG cs.NE

Soft Quality-Diversity Optimization

软质量-多样性优化

Saeed Hedayatian, Stefanos Nikolaidis

机构 * University of Southern California(南加州大学) Archimedes AI(阿基米德人工智能)

AI总结 本文提出软QD框架,通过避免离散化解决QD优化问题,推出SQUAD算法在高维问题中具有更好的可扩展性。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21675 2026-03-05 cs.LG math.OC

Scalable Second-order Riemannian Optimization for $K$-means Clustering

可扩展的二次黎曼优化用于K均值聚类

Peng Xu, Chun-Ying Hou, Xiaohui Chen, Richard Y. Zhang

机构 * Department of Statistics, University of Illinois Urbana-Champaign(统计系,伊利诺伊大学厄巴纳-香槟分校) Department of Electrical and Computer Engineering, University of Illinois Urbana-Champaign(电气与计算机工程系,伊利诺伊大学厄巴纳-香槟分校) Department of Mathematics, University of Southern California(数学系,南加州大学)

AI总结 本文提出了一种基于二次黎曼优化的K均值聚类方法,通过分解流形结构实现线性时间复杂度,显著提升收敛速度并保持统计准确性。

Journal ref The Fourteenth International Conference on Learning Representations (ICLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01222 2026-03-05 cs.CL cs.AI

WebDS: An End-to-End Benchmark for Web-based Data Science

WebDS: 一种端到端的基于网络的数据科学基准

Ethan Hsu, Hong Meng Yam, Ines Bouissou, Aaron Murali John, Raj Thota, Josh Koe, Vivek Sarath Putta, G K Dharesan, Alexander Spangher, Shikhar Murty, Tenghao Huang, Christopher D. Manning

机构 * Stanford University(斯坦福大学) Pinetree Research(Pinetree研究公司) University of California, Berkeley(加州大学伯克利分校) Singapore University of Technology and Design(新加坡科技设计大学) University of Southern California(南加州大学)

AI总结 WebDS提出了一种端到端的基于网络的数据科学基准,旨在评估代理在复杂多步骤任务中的表现,揭示当前LLM在实际数据科学任务中的性能差距。

Comments 14 pages, ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏