arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Texas at Austin(得克萨斯大学奥斯汀分校)

共收录 1213
2512.17008 2026-01-27 cs.LG

Turn-PPO: Turn-Level Advantage Estimation with PPO for Improved Multi-Turn RL in Agentic LLMs

Turn-PPO:基于PPO的回合级优势估计以提升代理语言模型中的多回合强化学习

Junbo Li, Peng Zhou, Rui Meng, Meet P. Vadera, Lihong Li, Yang Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Amazon(亚马逊)

AI总结 本研究提出turn-PPO,通过回合级MDP公式化提升多回合强化学习在代理语言模型中的表现,验证其在WebShop和Sokoban数据集上的有效性。

Journal ref EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.13968 2026-01-27 cs.CV cs.AI cs.CL

RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation

RotBench: 对多模态大语言模型识别图像旋转能力的评估

Tianyi Niu, Jaemin Cho, Elias Stengel-Eskin, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学夏洛特分校) Allen Institute for Artificial Intelligence(人工智能研究院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 RotBench评估了多模态大语言模型在识别图像旋转角度方面的性能,发现大多数模型难以区分90°和270°旋转,但能识别0°和180°图像,揭示了模型空间推理能力与人类的差距。

Comments EACL 2026 Camera-Ready. Code and data: https://github.com/tianyiniu/RotBench

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21288 2026-01-27 cs.GR cs.AI

SpringTime: Learning Simulatable Models of Cloth with Spatially-varying Constitutive Properties

SpringTime: 学习具有空间变化本构性质的布料模拟模型

Guanxiong Chen, Shashwat Suri, Yuhao Wu, Yixian Cheng, Ganidhu Abeysirigoonawardena, Etienne Vouga, David I. W. Levin, Dinesh K. Pai

机构 * University of British Columbia(不列颠哥伦比亚大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Toronto(多伦多大学)

AI总结 SpringTime通过学习运动观测数据,高效建模具有空间变化本构性质的布料模拟,克服膜锁定问题,提升训练速度和重建精度。

Comments Submitted to Graphics Interface '26 (In review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16965 2026-01-26 cs.AI

Spatial-Agent: Agentic Geo-spatial Reasoning with Scientific Core Concepts

空间智能体:基于科学核心概念的地理空间推理

Riyang Bao, Cheng Yang, Dazhou Yu, Zhexiang Tang, Gengchen Mai, Liang Zhao

机构 * Emory University(埃默里大学) Rutgers University(罗格斯大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 Spatial-Agent通过基于空间信息科学的理论框架,实现了可解释的地理空间推理,优于现有方法并产生可执行的工作流。

Comments 15pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16906 2026-01-26 cs.LG cs.HC

The Trajectory Alignment Coefficient in Two Acts: From Reward Tuning to Reward Learning

两幕中的轨迹对齐系数:从奖励调优到奖励学习

Calarina Muslimani, Yunshu Du, Kenta Kawamoto, Kaushik Subramanian, Peter Stone, Peter Wurman

机构 * Sony AI(索尼人工智能) University of Alberta(阿尔伯塔大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出轨迹对齐系数TAC,用于指导奖励调优和奖励学习,通过实验验证其在复杂领域中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18261 2026-01-26 cs.IR cs.AI

LLM Reasoning for Cold-Start Item Recommendation

基于大语言模型的冷启动物品推荐

Shijun Li, Yu Wang, Jin Wang, Ying Li, Joydeep Ghosh, Anne Cocos

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出利用大语言模型的推理能力,改进冷启动物品推荐,通过多种微调方法提升推荐性能,在Netflix数据上取得8%的提升。

Comments Published on Proceedings of the ACM on Web Conference 2026 (WWW 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16490 2026-01-26 cs.LG cs.AI cs.CC cs.CV

On Computational Limits of FlowAR Models: Expressivity and Efficiency

流AR模型的计算极限:表达力与效率

Yang Cao, Chengyue Gong, Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song

机构 * Wyoming Seminary(怀俄明私立学校) UT-Austin(得克萨斯大学奥斯汀分校) The University of Hong Kong(香港大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Simons Institute, UC Berkeley(伯克利大学模拟研究所)

AI总结 本研究分析了FlowAR模型的电路复杂性,证明其可被TC⁰电路模拟,并探讨了其计算效率与表达能力的限制。

Comments AISTATS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12150 2026-01-26 cs.RO cs.AI cs.LG

HEIGHT: Heterogeneous Interaction Graph Transformer for Robot Navigation in Crowded and Constrained Environments

HEIGHT:用于拥挤和受限环境中机器人导航的异构交互图变压器

Shuijing Liu, Haochen Xia, Fatemeh Cheraghi Pouria, Kaiwen Hong, Neeloy Chakraborty, Zichao Hu, Joydeep Biswas, Katherine Driggs-Campbell

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 HEIGHT通过异构交互图变压器提升机器人在拥挤和受限环境中的导航性能,有效捕捉时空异构交互,提高路径安全性和效率。

Comments Accepted to IEEE Transactions of Automation Science and Engineering (T-ASE)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15473 2026-01-23 cs.LG cs.AI

Panther: Faster and Cheaper Computations with Randomized Numerical Linear Algebra

Panther:利用随机数值线性代数实现更快更便宜的计算

Fahd Seddik, Abdulrahman Elbedewy, Gaser Sami, Mohamed Abdelmoniem, Yahia Zakaria

机构 * University of British Columbia(不列颠哥伦比亚大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) Cairo University(开罗大学)

AI总结 Panther通过整合RandNLA算法,为深度学习提供高效的内存节省方案,实现更快更便宜的计算。

Comments 5 pages, 3 figures, 2 listings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15417 2026-01-23 cs.LG cs.AI

Ambient Dataloops: Generative Models for Dataset Refinement

环境数据循环:生成模型的数据集精修

Adrián Rodríguez-Muñoz, William Daspit, Adam Klivans, Antonio Torralba, Constantinos Daskalakis, Giannis Daras

机构 * Massachusetts Institute of Technology(麻省理工学院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 Ambient Dataloops通过数据集与模型的共进化过程,提升数据质量并优化生成模型性能,实现高质量图像生成和蛋白质设计。

Comments 27 pages, 9 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06193 2026-01-23 cs.CL cs.AI

Do You Feel Comfortable? Detecting Hidden Conversational Escalation in AI Chatbots

你感觉舒适吗?检测AI聊天机器人中隐藏的对话升级

Jihyung Park, Saleh Afroogh, David Atkinson, Junfeng Jiao

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 GAUGE通过实时检测AI聊天机器人中隐藏的情感升级,提升对话中隐性伤害的识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15239 2026-01-22 stat.ML cs.LG math.ST stat.TH

Multi-context principal component analysis

多情境主成分分析

Kexin Wang, Salil Bhate, João M. Pereira, Joe Kileel, Matylda Figlerowicz, Anna Seigal

机构 * Harvard University(哈佛大学) Broad Institute of MIT and Harvard(哈佛-麻省理工Broad研究所) University of Georgia(佐治亚大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 多情境主成分分析(MCPCA)是一种理论和算法框架,用于识别跨不同情境子集共享的变异因素,应用于基因表达和语言模型数据,揭示隐藏的变异轴。

Comments 47 pages, 8 figures. Supplementary tables are provided as downloadable file

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14037 2026-01-22 cs.CV

Human detectors are surprisingly powerful reward models

人类检测器出人意料地强大

Kumar Ashutosh, XuDong Wang, Xi Yin, Kristen Grauman, Adam Polyak, Ishan Misra, Rohit Girdhar

机构 * Meta University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出 HuDA 模型,通过简单奖励函数提升视频生成中的人体运动质量,胜率高达 73%。

Comments Technical report

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18803 2026-01-22 cs.LG math.OC

Deceptive Sequential Decision-Making via Regularized Policy Optimization

通过正则化策略优化实现欺骗性序列决策

Yerin Kim, Alexander Benvenuti, Bo Chen, Mustafa Karabag, Abhishek Kulkarni, Nathaniel D. Bastian, Ufuk Topcu, Matthew Hale

机构 * Georgia Institute of Technology(佐治亚理工学院) University of Texas at Austin(德克萨斯大学奥斯汀分校) United States Military Academy(美国军事学院)

AI总结 本文提出三种正则化策略,用于在策略优化中主动欺骗对手关于系统奖励的误解,从而在保持高累积奖励的同时误导对手。

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24574 2026-01-21 cs.CL cs.AI cs.LG

Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time

在测试时理解并引导推理模型的认知行为

Zhenyu Zhang, Xiaoxia Wu, Zhongzhu Zhou, Qingyang Wu, Yineng Zhang, Pragaash Ponnusamy, Harikaran Subbaraj, Jue Wang, Shuaiwen Leon Song, Ben Athiwaratkun

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) Together AI University of Sydney(悉尼大学)

AI总结 CREST通过引导推理模型的认知行为,提高推理准确性和效率,减少计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.10597 2026-01-21 cs.CV cs.AI

Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models

Text2Seg: 通过文本引导的视觉基础模型实现遥感图像语义分割

Jielu Zhang, Zhongliang Zhou, Gengchen Mai, Mengxuan Hu, Zihan Guan, Sheng Li, Lan Mu

机构 * University of Georgia(佐治亚大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Virginia(弗吉尼亚大学)

AI总结 Text2Seg通过文本引导的视觉基础模型实现遥感图像语义分割,显著提升零样本预测性能。

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12307 2026-01-21 cs.MA cs.CL cs.LG

Rethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline

重新思考多智能体工作流的价值:一个强大的单智能体基线

Jiawei Xu, Arief Koesdwiady, Sisong Bei, Yan Han, Baixiang Huang, Dakuo Wang, Yutong Chen, Zheshen Wang, Peihao Wang, Pan Li, Ying Ding

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Amazon(亚马逊) Emory University(埃默里大学) Northeastern University(东北大学) Georgia Institute of Technology(佐治亚理工学院)

AI总结 本研究通过单个代理的多轮对话模拟多智能体工作流,提出OneFlow算法,实现高效且准确的多代理流程,为多智能体系统研究提供强基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.09660 2026-01-21 math.OC cs.LG

Fast Two-Time-Scale Stochastic Gradient Method with Applications in Reinforcement Learning

快速双时间尺度随机梯度方法及其在强化学习中的应用

Sihan Zeng, Thinh T. Doan

机构 * J.P. Morgan AI Research(摩根大通AI研究) UT Austin, Department of Aerospace Engineering & Engineering Mechanics(得克萨斯大学奥斯汀分校航空航天工程与工程力学系)

AI总结 本文提出了一种快速双时间尺度随机梯度方法,通过引入平均步骤提升收敛速度,并在强化学习中实现了优于现有方法的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11905 2026-01-21 cs.AI cs.LG math.ST stat.TH

LIBRA: Language Model Informed Bandit Recourse Algorithm for Personalized Treatment Planning

LIBRA:基于语言模型的带状 recourse 算法用于个性化治疗计划

Junyu Cao, Ruijiang Gao, Esmaeil Keyvanshokooh, Jianhao Ma

机构 * McCombs School of Business, University of Texas at Austin(德克萨斯大学奥斯汀分校麦克斯韦商学院) Naveen Jindal School of Management, University of Texas at Dallas(德克萨斯大学达拉斯分校奈文·金达管理学院) Mays Business School, Texas A&M University(德克萨斯农工大学梅斯商学院) Wharton School, University of Pennsylvania(宾夕法尼亚大学沃顿商学院)

AI总结 LIBRA 是一种结合大语言模型和带状学习的算法,用于在个性化治疗中实现更高效的决策和鲁棒性。

Comments 50 pages. Previous version with human-AI collaboration: arXiv:2410.14640

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11581 2026-01-21 cs.CL cs.AI

Enhancing the QA Model through a Multi-domain Debiasing Framework

通过多领域去偏框架增强问答模型

Yuefeng Wang, ChangJae Lee

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出多领域去偏框架,通过知识蒸馏和领域扩展技术,有效缓解问答模型中的偏见问题,提升其在对抗性环境下的性能和可靠性。

Comments 5 pages, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09692 2026-01-15 cs.CL cs.AI cs.LG

Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection

基于生成数据的路由:无标注LLM技能估计与专家选择

Tianyi Niu, Justin Chih-Yao Chen, Genta Indra Winata, Shi-Xiong Zhang, Supriyo Chakraborty, Sambit Sahu, Yue Zhang, Elias Stengel-Eskin, Mohit Bansal

机构 * UNC Chapel Hill(北卡罗来纳大学教堂山分校) Capital One(Capital One公司) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出CASCAL,一种通过共识投票和层次聚类提升LLM路由性能的查询-only路由器,在弱生成器数据下表现更优。

Comments Code: https://github.com/tianyiniu/RoutingGenData

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09385 2026-01-15 cs.SD cs.CL cs.MM

SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing

SLAM-LLM: 一种模块化、开源的多模态大语言模型框架及语音、语言、音频和音乐处理的最佳实践

Ziyang Ma, Guanrou Yang, Wenxi Chen, Zhifu Gao, Yexing Du, Xiquan Li, Zhisheng Zheng, Haina Zhu, Jianheng Zhuo, Zheshu Song, Ruiyang Xu, Tiranrui Wang, Yifan Yang, Yanqiao Zhu, Zhikang Niu, Liumeng Xue, Yinghao Ma, Ruibin Yuan, Shiliang Zhang, Kai Yu, Eng Siong Chng, Xie Chen

机构 * X-LANCE Lab, School of Computer Science, MoE Key Lab of Artificial Intelligence Shanghai Jiao Tong University(X-LANCE实验室,计算机科学学院,人工智能教育部重点实验室,上海交通大学) Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团) Peng Cheng Laboratory(鹏城实验室) University of Texas at Austin(德克萨斯大学奥斯汀分校) Tianjin University(天津大学) Hong Kong University of Science and Technology(香港科学大学) Queen Mary University of London(伦敦玛丽女王大学) Nanyang Technological University(南洋理工大学) Shanghai Innovation Institute(上海创新研究院)

AI总结 SLAM-LLM是一种开源多模态大语言模型框架,专注于语音、语言、音频和音乐处理,提供模块化配置和高性能检查点以加速研究开发。

Comments Published in IEEE Journal of Selected Topics in Signal Processing (JSTSP)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20612 2026-01-15 cs.LG

Policy Compatible Skill Incremental Learning via Lazy Learning Interface

通过懒惰学习接口实现策略兼容的技能增量学习

Daehee Lee, Dongsu Lee, TaeYoon Kwack, Wonje Choi, Honguk Woo

机构 * Sungkyunkwan University(成均馆大学) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出SIL-C框架,通过懒惰学习接口实现技能与策略的兼容性,提升下游任务性能无需重新训练策略。

Comments NeurIPS 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19565 2026-01-14 cs.LO cs.AI

Deductive Systems for Logic Programs with Counting

带有计数聚合的逻辑程序的演绎系统

Jorge Fandinno, Vladimir Lifschitz

机构 * University of Nebraska Omaha(内布拉斯加大学奥马哈分校) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出了一种扩展的演绎系统,用于证明包含计数聚合的逻辑程序之间的强等价性。

Comments Under consideration in Theory and Practice of Logic Programming (TPLP)

Journal ref Theory and Practice of Logic Programming 25 (2025) 924-964

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02371 2026-01-14 cs.CY cs.AI cs.MA cs.NI

Permission Manifests for Web Agents

基于Web代理的权限声明

Samuele Marro, Alan Chan, Xinxing Ren, Lewis Hammond, Jesse Wright, Gurjyot Wanga, Tiziano Piccardi, Nuno Campos, Tobin South, Jialin Yu, Sunando Sengupta, Eric Sommerlade, Alex Pentland, Philip Torr, Jiaxin Pei

机构 * University of Oxford(牛津大学) Institute for Decentralized AI(去中心化人工智能研究所) Centre for the Governance of AI(人工智能治理中心) Coral Protocol(珊瑚协议) Cooperative AI Foundation(协作人工智能基金会) Webair Johns Hopkins University(约翰霍普金斯大学) Witan Labs(Witan实验室) Stanford University(斯坦福大学) Microsoft(微软) UT Austin(得克萨斯大学奥斯汀分校)

AI总结 本文提出 agent-permissions.json,一种轻量级声明,用于规范 Web 代理的交互权限,以提升自动化应用与网站所有者的协调性。

Comments Authored by the Lightweight Agent Standards Working Group https://las-wg.org/

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12992 2026-01-14 cs.RO cs.CL cs.CV cs.MA

UNCAP: Uncertainty-Guided Neurosymbolic Planning Using Natural Language Communication for Cooperative Autonomous Vehicles

UNCAP:基于自然语言通信的不确定性引导神经符号规划

Neel P. Bhatt, Po-han Li, Kushagra Gupta, Rohan Siva, Daniel Milan, Alexander T. Hogue, Sandeep P. Chinchali, David Fridovich-Keil, Zhangyang Wang, Ufuk Topcu

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 UNCAP通过自然语言通信和不确定性引导方法,提升多车辆协同自动驾驶系统的可扩展性和安全性。

Journal ref AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07309 2026-01-14 cs.CL cs.LG

PIE: Performance Interval Estimation for Free-Form Generation Tasks

PIE:用于自由形式生成任务的性能区间估计

Chi-Yang Hsu, Alexander Braylan, Yiheng Su, Matthew Lease, Omar Alonso

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Amazon(亚马逊)

AI总结 PIE提出了一种用于自由形式生成任务的性能区间估计方法,通过回归实现更准确的指标评分和校准的不确定性区间。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10326 2026-01-14 cs.AI cs.GT cs.LG cs.MA

VGC-Bench: Towards Mastering Diverse Team Strategies in Competitive Pokémon

VGC-Bench: 向掌握多样化团队策略的竞技宝可梦迈进

Cameron Angliss, Jiaxun Cui, Jiaheng Hu, Arrasy Rahman, Peter Stone

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 VGC-Bench通过提供标准化评估和人类对战数据集,研究如何让AI代理在多样化团队策略的竞技宝可梦中实现稳健适应与泛化。

Comments AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07245 2026-01-13 cs.AI cs.CL cs.LG

Learning to Trust the Crowd: A Multi-Model Consensus Reasoning Engine for Large Language Models

学习信任群体:一种多模型共识推理引擎用于大语言模型

Pranav Kallem

机构 * Department of Computer Science(计算机科学系) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本研究提出多模型共识推理引擎,通过整合多个大语言模型的输出,提升回答的准确性和可靠性,实验表明其在多个基准测试中显著优于单一模型和多数投票。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07334 2026-01-13 cs.LG

Graph-KV: Breaking Sequence via Injecting Structural Biases into Large Language Models

Graph-KV: 通过向大语言模型注入结构偏差打破序列

Haoyu Wang, Peihao Wang, Mufei Li, Shikun Liu, Siqi Miao, Zhangyang Wang, Pan Li

机构 * Georgia Institute of Technology(佐治亚理工学院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 Graph-KV通过引入结构偏差打破序列限制,提升大语言模型在图结构任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏