arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Texas at Austin(得克萨斯大学奥斯汀分校)

共收录 1213
2412.07010 2026-05-04 cs.LG physics.comp-ph

TAEN: A Model-Constrained Tikhonov Autoencoder Network for Forward and Inverse Problems

TAEN:一种用于正反问题的模型约束Tikhonov自编码器网络

Hai V. Nguyen, Tan Bui-Thanh, Clint Dawson

机构 * Department of Aerospace Engineering and Engineering Mechanics, the University of Texas at Austin(德克萨斯大学航空航天工程与工程力学系) The Oden Institute for Computational Engineering and Sciences, the University of Texas at Austin(德克萨斯大学奥登计算工程与科学研究院)

AI总结 本文提出TAEN模型,通过单个观测样本学习正反问题的替代模型,结合数据随机化策略和理论基础,实现高效准确的求解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02277 2026-05-01 cs.LG cs.AI

Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs

垃圾DNA假说:修剪小的预训练权重不可逆且单调地损害LLM中的“困难”下游任务

Lu Yin, Ajay Jaiswal, Shiwei Liu, Souvik Kundu, Zhangyang Wang

机构 * University of Surrey Eindhoven University of Technology University of Texas at Austin Intel Labs University of Oxford

AI总结 该研究提出垃圾DNA假说,指出LLM预训练权重中存在关键知识,修剪小权重会单调损害困难下游任务性能,且即使允许持续训练也无法弥补损失。

Comments Published at ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06526 2026-05-01 cond-mat.mtrl-sci cs.LG

Predicting Atomistic Transitions with Transformers

用变压器预测原子级转变

Henry Tischler, Wenting Li, Qi Tang, Danny Perez, Thomas Vogel

机构 * Computing and Artificial Intelligence Division, Los Alamos National Laboratory(计算与人工智能部门,洛斯阿拉莫斯国家实验室) School of Engineering and Computer Science, University of Denver(工程与计算机科学学院,丹佛大学) Department of Physics and Astronomy, University of Denver(物理与天文学系,丹佛大学) Department of Electrical and Computer Engineering, University of Texas at Austin(电气与计算机工程系,德克萨斯大学奥斯汀分校) School of Computational Science and Engineering, Georgia Institute of Technology(计算科学与工程学院,佐治亚理工学院) Theoretical Division, Los Alamos National Laboratory(理论部门,洛斯阿拉莫斯国家实验室) X Computational Physics Division, Los Alamos National Laboratory(X计算物理部门,洛斯阿拉莫斯国家实验室)

AI总结 本文利用变压器模型高效预测纳米簇中的原子级转变,通过评估物理有效性并生成多种微态,降低计算成本。

Comments Presented at the 2025 Conference on Data Analysis (CoDA), February 25-28, Santa Fe, New Mexico

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18796 2026-05-01 cs.LG q-bio.QM

Graph-Based Biomarker Discovery and Interpretation for Alzheimer's Disease

基于图的阿尔茨海默病生物标志物发现与解释

Maryam Khalid, Fadeel Sher Khan, John Broussard, Arko Barman

机构 * Rice University(里士大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Texas Health Science Center at Houston(德克萨斯大学健康科学中心休斯顿分校)

AI总结 本文提出BRAIN框架,通过图表示方法联合优化诊断准确性和生物标志物发现,揭示阿尔茨海默病诊断相关的生物标志物子网络。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16552 2026-04-30 cs.CV cs.AI

Co-generation of Layout and Shape from Text via Autoregressive 3D Diffusion

通过自回归3D扩散模型生成布局和形状

Zhenggang Tang, Yuehao Wang, Yuchen Fan, Jun-Kun Chen, Yu-Ying Yeh, Kihyuk Sohn, Zhangyang Wang, Qixing Huang, Alexander Schwing, Rakesh Ranjan, Dilin Wang, Zhicheng Yan

机构 * Meta Reality Labs(Meta现实实验室) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出一种新的文本到场景生成方法,通过自回归3D扩散模型生成布局和形状,解决文本描述与生成场景不一致的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05811 2026-04-30 cs.CV

Video Compression Meets Video Generation: Latent Inter-Frame Pruning with Attention Recovery

视频压缩与视频生成的交汇:具有注意力恢复的潜在帧剪枝

Dennis Menn, Yuedong Yang, Bokun Wang, Xiwen Wei, Mustafa Munir, Feng Liang, Radu Marculescu, Chenfeng Xu, Diana Marculescu

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) Meta

AI总结 本文提出LIPAR框架,通过利用视频潜在补丁中的时间冗余性,减少计算延迟,提升视频编辑效率,同时保持生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20876 2026-04-30 cs.CL

Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine

少做决定,多进行沟通:关于医学领域端到端事实核查构建效度的探讨

Sebastian Joseph, Lily Chen, Barry Wei, Michael Mackert, Iain J. Marshall, Paul Pu Liang, Ramez Kouzy, Byron C. Wallace, Junyi Jessy Li

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Stanford University(斯坦福大学) Indiana University School of Medicine(印第安纳大学医学院) King’s College London(伦敦国王学院) Massachusetts Institute of Technology(麻省理工学院) The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心) Northeastern University(东北大学)

AI总结 本文探讨医学领域端到端事实核查系统的构建效度,指出其在连接现实声明与科学证据、处理模糊声明及主观真实性标签方面的挑战,主张将其视为互动沟通问题。

Comments ACL 2026 Findings camera-ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03136 2026-04-29 cs.CL cs.AI cs.RO

Limited Linguistic Diversity in Embodied AI Datasets

具身AI数据集中的语言多样性有限

Selma Wanna, Agnes Luhtaru, Jonathan Salfity, Ryan Barron, Juston Moore, Cynthia Matuszek, Mitch Pryor

机构 * Los Alamos National Laboratory(洛斯阿拉莫斯国家实验室) Institute of Computer Science, University of Tartu(塔尔图大学计算机科学学院) Department of Mechanical Engineering, The University of Texas at Austin(德克萨斯大学奥斯汀分校机械工程系) University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

AI总结 本文分析了多个广泛使用的视觉-语言-动作数据集的语言特性,发现其指令存在高度重复和结构单一的问题,旨在推动更详细的 dataset 报告和语言覆盖的扩展策略。

Comments Accepted to ACL 2026 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25506 2026-04-29 cs.NI cs.AI

Assistants, Not Architects: The Role of LLMs in Networked Systems Design

助手,而非架构师:大语言模型在联网系统设计中的作用

Pratyush Sahu, Rahul Bothra, Venkat Arun, Brighten Godfrey, Akshay Narayan, Ahmed Saeed

机构 * Georgia Tech(佐治亚理工学院) UIUC(伊利诺伊大学香槟分校) UT Austin(得克萨斯大学奥斯汀分校) Brown University(布朗大学)

AI总结 本文探讨了大语言模型在联网系统架构设计中的局限性,并提出Kepler框架,结合专家驱动的规范与SMT优化,实现可行的设计方案并支持可解释的设计探索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25096 2026-04-29 cs.CL cs.HC

The Dynamics of Delusion: Modeling Bidirectional False Belief Amplification in Human-Chatbot Dialogue

妄想的动力学:人类-聊天机器人对话中双向虚假信念放大建模

Ashish Mehta, Jared Moore, Jacy Reese Anthis, William Agnew, Eric Lin, Peggy Yin, Desmond C. Ong, Nick Haber, Carol Dweck

机构 * Stanford University(斯坦福大学) Carnegie Mellon University(卡内基梅隆大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 研究通过分析聊天记录数据,发现人类与聊天机器人之间存在双向影响,人类短期影响聊天机器人,而聊天机器人长期影响人类,并自我强化妄想,为AI安全提供新见解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24833 2026-04-29 cs.RO cs.AI cs.GR cs.LG

MotionBricks: Scalable Real-Time Motions with Modular Latent Generative Model and Smart Primitives

MotionBricks: 可扩展的实时动作生成与模块化潜在生成模型及智能原语

Tingwu Wang, Olivier Dionne, Michael De Ruyter, David Minor, Davis Rempe, Kaifeng Zhao, Mathis Petrovich, Ye Yuan, Chenran Li, Zhengyi Luo, Brian Robison, Xavier Blackwell, Bernardo Antoniazzi, Xue Bin Peng, Yuke Zhu, Simon Yuen

机构 * NVIDIA Simon Fraser University(西蒙·弗雷泽大学) The University of Texas at Austin(得克萨斯大学奥斯汀分校)

AI总结 本文提出MotionBricks,通过模块化潜在生成模型和智能原语实现实时动作生成,解决工业应用中实时可扩展性和多模态控制的挑战,展示了高质量和实时性能。

Comments ACM Transactions on Graphics; SIGGRAPH 2026. Project page: https://nvlabs.github.io/motionbricks/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08567 2026-04-29 cs.CL cs.MA

Multi-User Large Language Model Agents

多用户大语言模型代理

Shu Yang, Shenzhe Zhu, Hao Zhu, José Ramón Enríquez, Di Wang, Alex Pentland, Michiel A. Bakker, Jiaxin Pei

机构 * Stanford University(斯坦福大学) KAUST UT Austin(得克萨斯大学奥斯汀分校) MIT(麻省理工学院)

AI总结 本文首次系统研究多用户LLM代理,提出统一交互协议并设计压力测试场景,揭示前沿模型在优先级维护、隐私保护和协调效率方面的系统性缺陷。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23858 2026-04-28 cs.CV

Latent Inter-Frame Pruning: A Training-Free Method Bridging Traditional Video Compression and Modern Diffusion Transformers for Efficient Generation

潜在帧剪枝:一种无训练方法,连接传统视频压缩与现代扩散变换器以实现高效生成

Dennis Menn, Chih-Hsien Chou

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Futurewei Technologies, Inc.(未来科技公司)

AI总结 本文提出一种无训练的潜在帧剪枝方法,通过剪枝重复的潜在块来减少计算负担并提高生成速度,同时引入注意力恢复机制以解决训练与推理间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23069 2026-04-28 cs.CL

ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents

ContextWeaver:面向LLM代理的选型与依赖结构化记忆构建

Yating Wu, Yuhao Zhang, Sayan Ghosh, Sourya Basu, Anoop Deoras, Jun Huan, Gaurav Gupta

机构 * UT Austin(得克萨斯大学奥斯汀分校) AWS AI Labs(AWS人工智能实验室)

AI总结 本文提出ContextWeaver,通过构建推理步骤图谱,实现LLM代理的记忆选择与依赖结构化,提升长上下文交互性能,减少推理步骤与token消耗。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14541 2026-04-27 cs.LG cs.AI cs.AR

Report for NSF Workshop on AI for Electronic Design Automation

NSF关于人工智能在电子设计自动化领域的研讨会报告

Deming Chen, Vijay Ganesh, Weikai Li, Yingyan Celine Lin, Yong Liu, Subhasish Mitra, David Z. Pan, Ruchir Puri, Jason Cong, Yizhou Sun

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Georgia Institute of Technology(佐治亚理工学院) University of California at Los Angeles(加州大学洛杉矶分校) Cadence Design Systems, Inc.(Cadence设计系统公司) Stanford University(斯坦福大学) University of Texas at Austin(得克萨斯大学奥斯汀分校) IBM

AI总结 报告总结了NSF人工智能在电子设计自动化领域研讨会的讨论和建议,探讨了AI技术如何加速EDA设计流程,提出加强AI与EDA合作、投资基础AI研究等核心贡献。

Comments Accepted by IEEE Circuits and Systems Magazine (2026). This is the accepted version. The published version is available at https://ieeexplore.ieee.org/document/11466406

Journal ref IEEE Circuits and Systems Magazine, vol. 26, no. 1, First Quarter 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21548 2026-04-27 cs.CL cs.IT cs.LG math.IT

MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression

MultiTok: 可变长度分词用于高效大语言模型,源自LZW压缩

Noel Elias, Homa Esfahanizadeh, Kaan Kale, Sriram Vishwanath, Muriel Medard

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) Nokia Bell Labs(诺基亚贝尔实验室) Boğaziçi University(博雅大学) Georgia Institute of Technology(佐治亚理工学院) Massachusetts Institute of Technology (MIT)(麻省理工学院(MIT))

AI总结 本文提出基于LZW压缩的MultiTok分词方法,通过压缩重复短语提升LLM训练效率,实现更高效训练与相似准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21119 2026-04-24 cs.CV cs.AI cs.SD

Materialistic RIR: Material Conditioned Realistic RIR Generation

物质导向的RIR:基于材料的现实RIR生成

Mahnoor Fatima Saad, Sagnik Majumder, Kristen Grauman, Ziad Al-Halah

机构 * University of Utah(犹他大学) UT Austin(得克萨斯大学奥斯汀分校)

AI总结 本文提出一种基于材料的RIR生成方法,通过分离空间和材料影响,提升生成声学的真实感和材料敏感性,实验结果显示在声学和材料指标上均有显著提升。

Comments Accepted to CVPR 2026 Findings. Project page: https://mahnoor-fatima-saad.github.io/MatRIR.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20401 2026-04-23 cs.CR cs.AI

Onyx: Cost-Efficient Disk-Oblivious ANN Search

Onyx: 低成本的磁盘无关ANN搜索

Deevashwer Rathee, Jean-Luc Watson, Zirui Neil Zhao, G. Edward Suh, Raluca Ada Popa

机构 * UC Berkeley, NVIDIA(伯克利大学,NVIDIA) NVIDIA UT Austin, NVIDIA(奥斯汀大学,NVIDIA)

AI总结 Onyx通过优化ANN和ORAM层的资源利用,降低ANN搜索成本和延迟,采用紧凑中间表示和局部意识浅树设计提高效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07474 2026-04-23 cs.CL cs.AI

Cross-Modal Taxonomic Generalization in (Vision-) Language Models

跨模态分类泛化在(视觉-)语言模型中

Tianyang Xu, Marcelo Sandoval-Castaneda, Karen Livescu, Greg Shakhnarovich, Kanishka Misra

机构 * Toyota Technological Institute at Chicago(芝加哥丰田技术研究所) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 研究探讨了语言模型在无显性证据情况下通过跨模态泛化恢复超范畴知识的能力,发现语言模型能基于语言线索实现泛化。

Comments ACL 2026 (main conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13928 2026-04-23 cs.CL cs.AI

LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X

Shuo Xing, Junyuan Hong, Yifan Wang, Runjin Chen, Zhenyu Zhang, Ananth Grama, Zhengzhong Tu, Zhangyang Wang

机构 * Texas A&M University(德克萨斯农工大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) Purdue University(普渡大学)

Comments Updated experiments with corrected data

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07963 2026-04-23 cs.CL cs.AI

Caught in the Web of Words: Do LLMs Fall for Spin in Medical Literature?

陷入词语的网络:大型语言模型是否在医学文献中受偏见影响?

Hye Sun Yun, Karen Y. C. Zhang, Ramez Kouzy, Iain J. Marshall, Junyi Jessy Li, Byron C. Wallace

机构 * Northeastern University, Boston, MA, USA(东北大学,波士顿,马萨诸塞州,美国) The University of Texas MD Anderson Cancer Center, Houston, Texas, USA(德克萨斯大学MD安德森癌症中心,休斯顿,德克萨斯州,美国) King’s College London, London, UK(伦敦国王学院,伦敦,英国) The University of Texas at Austin, Austin, Texas, USA(德克萨斯大学奥斯汀分校,奥斯汀,德克萨斯州,美国)

AI总结 研究探讨了大型语言模型(LLM)在处理医学文献时是否受作者偏见影响,发现LLM比人类更容易受偏见影响,但可通过提示减少其影响。

Comments 26 pages, 17 figures, 4 tables, Conference on Health, Inference, and Learning (CHIL) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20688 2026-04-23 cs.LG cs.AI

Storm Surge Modeling, Bias Correction, Graph Neural Networks, Graph Convolution Networks

风暴潮建模、偏差校正、图神经网络、图卷积网络

Noujoud Nader, Stefanos Giaremis, Clint Dawson, Carola Kaiser, Karame Mohammadiporshokooh, Hartmut Kaiser

机构 * Center for Computation and Technology(计算技术中心) Louisiana State University(路易斯安那州立大学) Department of Physics(物理系) Aristotle University of Thessaloniki(塞萨洛尼基阿利克塞奥斯大学) Oden Institute for Computational Engineering and Sciences(计算工程与科学研究院) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出StormNet,一种用于风暴潮预测偏差校正的时空图神经网络,通过整合GCN、GAT和LSTM,有效降低预测误差,提升极端天气事件中的预测准确性与可靠性。

Comments 51 pages, 9 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.20551 2026-04-23 stat.ML cs.LG

On Bayesian Softmax-Gated Mixture-of-Experts Models

基于贝叶斯softmax门控专家混合模型

Nicola Bariletto, Huy Nguyen, Nhat Ho, Alessandro Rinaldo

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文研究了贝叶斯专家混合模型的softmax门控机制,分析了后验分布的渐近行为,推导了密度估计、参数估计和模型选择的收敛性保证,并提出了专家数量选择策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.01703 2026-04-23 cs.CY cs.AI cs.LG eess.AS

Community-Informed AI Models for Police Accountability

面向警察问责的社区导向AI模型

Benjamin A. T. Graham, Lauren Brown, Georgios Chochlakis, Morteza Dehghani, Raquel Delerme, Brittany Friedman, Ellie Graeden, Preni Golazizian, Rajat Hebbar, Parsa Hejabi, Aditya Kommineni, Mayagüez Salinas, Michael Sierra-Arévalo, Jackson Trager, Nicholas Weller, Shrikanth Narayanan

机构 * Department of Political Science and International Relations, University of Southern California(美国南加州大学政治学与国际关系系) School of Public Policy, University of Southern California(美国南加州大学公共政策学院) Signal Analysis and Interpretation Laboratory (SAIL), University of Southern California(美国南加州大学信号分析与解释实验室) Department of Computer Science, University of Southern California(美国南加州大学计算机科学系) Brain and Creativity Institute, University of Southern California(美国南加州大学脑与创造力研究所) Department of Sociology, University of Southern California(美国南加州大学社会学系) Center for Global Health Science and Security, Georgetown University(乔治城大学全球健康科学与安全中心) Department of Electrical and Computer Engineering, University of Southern California(美国南加州大学电气与计算机工程系) The Lewis Registry(李氏登记处) Department of Sociology, The University of Texas at Austin(德克萨斯大学奥斯汀分校社会学系) Department of Psychology, University of Southern California(美国南加州大学心理学系) Department of Political Science, University of California Riverside(加州大学河滨分校政治学系) Harvard Law School, Harvard University(哈佛大学法学院)

AI总结 本文提出一种社区导向的多视角AI工具开发方法,通过整合多方观点提升政府问责透明度,以洛杉矶警察局交通停靠记录分析为例展示其应用。

Comments 33 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19069 2026-04-22 cs.CL cs.AI

Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference

专家产品训练减少自然语言推理中的数据集痕迹

Aby Mammen Mathew

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文提出PoE训练方法,通过降低偏见模型过度自信的例子权重,减少数据集痕迹,提升推理准确性与减少偏见依赖。

Comments 10 pages, 3 figures, 4 tables. Single-author paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12516 2026-04-22 cs.RO

Zero to Autonomy in Real-Time: Online Adaptation of Dynamics in Unstructured Environments

从零到自主:在无结构环境中动态的在线适应

William Ward, Sarah Etter, Jesse Quattrociocchi, Christian Ellis, Adam J. Thorpe, Ufuk Topcu

机构 * Oden Institute for Computational Engineering & Science, University of Texas at Austin(德纳学院计算工程与科学研究所,德克萨斯大学奥斯汀分校) Department of Computer Science, University of Texas at Austin(计算机科学系,德克萨斯大学奥斯汀分校) DEVCOM Army Research Laboratory(陆军研究实验室)

AI总结 本文提出了一种结合函数编码器和递推最小二乘法的在线适应方法,通过流式里程计更新潜变量,实现实时动态适应,提升在无结构环境中的安全性和规划效率。

Comments Initial submission to RA-L

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00439 2026-04-22 cs.CL

Improving the Distributional Alignment of LLMs using Supervision

通过监督改进大型语言模型的分布对齐

Gauri Kambhatla, Sanjana Gautam, Angela Zhang, Alex Liu, Ravi Srinivasan, Junyi Jessy Li, Matthew Lease

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

AI总结 本文通过简单监督提升LLM在不同群体上的分布对齐,分析了三个数据集的对齐效果,并为未来研究提供基准。

Comments ACL Main 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20279 2026-04-22 cs.CV cs.CL

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

VLM-3R:融合指令对齐3D重建的视觉-语言模型

Zhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang, Runjin Chen, Hezhen Hu, Kevin Wang, Huaizhi Qu, Shijie Zhou, Dilin Wang, Zhicheng Yan, Hongyu Xu, Justin Theiss, Tianlong Chen, Jiachen Li, Zhengzhong Tu, Zhangyang Wang, Rakesh Ranjan

机构 * UT Austin(德克萨斯大学奥斯汀分校) XMU(厦门大学) TAMU(德克萨斯大学阿灵顿分校) UCR(加州大学河滨分校) UNC(北卡罗来纳大学教堂山分校) Meta UCLA(加州大学洛杉矶分校)

AI总结 本文提出VLM-3R,通过融合3D重建指令微调,实现单目视频的3D空间辅助与具身推理,提升了视觉-空间推理和时间3D上下文理解的准确性和可扩展性。

Comments Project Page: https://vlm-3r.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16683 2026-04-22 cs.CV cs.AI

GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations

GAIR:具有地理对齐隐式表示的定位感知自监督对比预训练

Zeping Liu, Ni Lao, Zhangyu Wang, Junfeng Jiao, Gengchen Mai

机构 * SEAI Lab, Department of Geography and the Environment, The University of Texas at Austin(地理与环境系SEAI实验室,德克萨斯大学奥斯汀分校) Google LLC, Mountain View, CA, USA(谷歌公司,山景城,加利福尼亚州,美国) SIT Lab, School of Computing and Information Science, The University of Maine(计算与信息科学系SIT实验室,缅因大学) Urban Information Lab, School of Architecture, The University of Texas at Austin(城市信息实验室,建筑系,德克萨斯大学奥斯汀分校)

AI总结 GAIR通过引入隐式神经表示模块,解决地理空间任务中多模态数据的局部化表示问题,实现跨模态的地理对齐,提升空间关系建模能力。

Comments Accepted by ISPRS Journal of Photogrammetry and Remote Sensing

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.18203 2026-04-21 cs.CL

Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs

多模态大语言模型中的乘法:文本、图像和音频输入的计算

Samuel G. Balter, Ethan Jerzak, Connor T. Jerzak

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校) National University of Singapore (NUS)(新加坡国立大学)

AI总结 研究多模态大语言模型在处理不同模态下的多数字乘法时的局限性,提出一个受控的多模态乘法基准,通过因子变化数字长度、稀疏性、表示形式和模态,定义算术负载作为总和和非零数字计数的乘积,揭示模型性能与算术负载的关系。

Comments To appear in ACL Findings (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏