arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

共收录 830
2601.10173 2026-01-16 cs.CR cs.AI cs.CL

ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack

ReasAlign: 基于推理增强的安全对齐以抵御提示注入攻击

Hao Li, Yankai Yang, G. Edward Suh, Ning Zhang, Chaowei Xiao

机构 * Washington University in St. Louis(华盛顿大学圣路易斯分校) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) NVIDIA(NVIDIA公司) Johns Hopkins University(约翰霍普金斯大学)

AI总结 ReasAlign通过结构化推理和测试时缩放机制,有效防御提示注入攻击,实现安全与效用的最佳平衡。

Comments 15 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10141 2026-01-16 cs.LG cs.AI

Understanding and Preserving Safety in Fine-Tuned LLMs

理解并保持在微调大语言模型中的安全性

Jiawen Zhang, Yangfan Hu, Kejia Chen, Lipeng He, Jiachen Ma, Jian Lou, Dan Li, Jian Liu, Xiaohu Yang, Ruoxi Jia

机构 * Zhejiang University(浙江大学) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) University of Waterloo(滑铁卢大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Sun Yat-sen University(中山大学)

AI总结 本文提出安全保持微调(SPF)方法,通过分析安全性与实用性梯度的几何关系,有效解决微调过程中安全性与实用性之间的矛盾,保持模型性能并恢复预训练的安全性对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09756 2026-01-16 cs.CR cs.AI cs.CL

Synthetic Data for Veterinary EHR De-identification: Benefits, Limits, and Safety Trade-offs Under Fixed Compute

兽医电子健康记录的合成数据:在固定计算下的好处、限制和安全权衡

David Brundage

机构 * University of Wisconsin-Madison, School of Veterinary Medicine(威斯康星大学麦迪逊分校兽医学院)

AI总结 本研究探讨了合成数据在兽医电子健康记录去标识化中的应用,发现合成数据在固定预算下无法替代真实数据,但适度混合可提升性能,但需注意训练暴露增加而非数据质量提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21364 2026-01-15 cs.LG cs.AI

Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders

无需牺牲的可解释性:混合解码器的忠实密集层分解

James Oldfield, Shawn Im, Sharon Li, Mihalis A. Nicolaou, Ioannis Patras, Grigorios G Chrysos

机构 * UW-Madison(威斯康星大学麦迪逊分校)

AI总结 本文提出混合解码器(MxDs)通过层级稀疏性实现密集层分解,保持原始解码器的表达能力,显著提升稀疏性-准确性前沿性能。

Comments NeurIPS 2025 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07975 2026-01-14 cs.CV

An Efficient Additive Kolmogorov-Arnold Transformer for Point-Level Maize Localization in Unmanned Aerial Vehicle Imagery

一种高效的加法柯尔莫哥洛夫-安德罗夫变换器用于无人机图像中点级玉米定位

Fei Li, Lang Qiao, Jiahao Fan, Yijia Xu, Shawn M. Kaeppler, Zhou Zhang

机构 * Department of Biological Systems Engineering, University of Wisconsin-Madison, Madison, WI 53706, USA(生物系统工程系,威斯康星大学麦迪逊分校) Department of Agronomy, University of Wisconsin-Madison, Madison, WI 53706, USA(农学系,威斯康星大学麦迪逊分校)

AI总结 本文提出AKT变换器,通过改进的注意力机制和高效模型结构,实现无人机图像中高精度点级玉米定位,提升农业遥感应用效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07854 2026-01-14 q-bio.OT cs.AI

Immunological Density Shapes Recovery Trajectories in Long COVID

免疫密度塑造长期新冠恢复轨迹

Jing Wang, Tong Zhang, Xing Niu, Jie Shen, Yiming Luo, Qiaomin Xie, Amar Sra, Zorina Galis, Jeremy Weiss

机构 * National Library of Medicine(国家医学图书馆) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Stevens Institute of Technology(史蒂文斯理工学院) Columbia University(哥伦比亚大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) George Washington University(乔治华盛顿大学) National Heart, Lung, and Blood Institute(国家心肺血液研究所)

AI总结 研究发现长期新冠症状严重程度由免疫密度决定,恢复主要通过重复接种疫苗实现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03475 2026-01-14 cs.LG stat.AP

Joint Progression Modeling (JPM): A Probabilistic Framework for Mixed-Pathology Progression

联合进展建模(JPM):一种用于混合病理进展的概率框架

Hongtao Hao, Joseph L. Austerweil

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Chiba Institute of Technology(千叶技术大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

AI总结 JPM通过概率框架处理混合病理进展,提升排序准确性21%,并与文献结果一致。

Comments 49 pages; Machine Learning for Health (ML4H) Symposium 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15785 2026-01-13 cs.LG cs.AI

Investigating a Model-Agnostic and Imputation-Free Approach for Irregularly-Sampled Multivariate Time-Series Modeling

探究一种模型无关且无需插值的不规则采样多变量时间序列建模方法

Abhilash Neog, Arka Daw, Sepideh Fatemi Khorasgani, Medha Sawhney, Aanish Pradhan, Mary E. Lofton, Bennett J. McAfee, Adrienne Breef-Pilz, Heather L. Wander, Dexter W Howard, Cayelan C. Carey, Paul Hanson, Anuj Karpatne

机构 * Virginia Tech(弗吉尼亚理工大学) Oak Ridge National Lab(橡树岭国家实验室) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

AI总结 本文提出一种无需插值的模型无关方法MissTSM,用于不规则采样多变量时间序列建模,在高缺失率和无周期结构条件下表现优异。

Comments Accepted at TMLR

Journal ref Transactions on Machine Learning Research (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06309 2026-01-13 cs.CV cs.AI

VideoWeave: A Data-Centric Approach for Efficient Video Understanding

VideoWeave:一种以数据为中心的高效视频理解方法

Zane Durante, Silky Singh, Arpandeep Khatua, Shobhit Agarwal, Reuben Tan, Yong Jae Lee, Jianfeng Gao, Ehsan Adeli, Li Fei-Fei

机构 * Stanford University(斯坦福大学) Microsoft Research(微软研究院) University of Wisconsin - Madison(威斯康星大学麦迪逊分校)

AI总结 VideoWeave通过重新组织训练数据提升视频语言模型的数据效率,无需修改模型架构,实现更高准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02716 2026-01-13 cs.LG cs.CL

A Unified Understanding and Evaluation of Steering Methods

对引导方法的统一理解和评估

Shawn Im, Sharon Li

机构 * Department of Computer Sciences University of Wisconsin-Madison(计算机科学系 威斯康星大学麦迪逊分校)

AI总结 本文提出统一框架,分析评估引导方法,揭示其有效性并展示方法优势,为LLM中的设计优化提供指导

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11741 2026-01-12 cs.AI cs.CL cs.LG

Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?

攀登推理的阶梯:在SFT之后,大语言模型能解决什么问题?

Yiyou Sun, Georgia Zhou, Haoyue Bai, Hao Wang, Dacheng Li, Nouha Dziri, Dawn Song

机构 * University of California, Berkeley(加州大学伯克利分校) University of Wisconsin, Madison(威斯康星大学麦迪逊分校) Allen Institute for AI(人工智能研究院)

AI总结 本文通过分析AIME24数据集,揭示了大语言模型在数学推理任务中不同难度层级的进阶要求,发现SFT对提升推理能力有限,而扩大数据规模更有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04277 2026-01-09 cs.LG

Unlocking the Pre-Trained Model as a Dual-Alignment Calibrator for Post-Trained LLMs

解封预训练模型作为后训练LLMs的双对齐校准器

Beier Luo, Cheng Wang, Hongxin Wei, Sharon Li, Xuefeng Du

机构 * Department of Statistics and Data Science, Southern University of Science and Technology(统计与数据科学系,南方科技大学) School of Computing, National University of Singapore(计算学院,新加坡国立大学) Department of Computer Sciences, University of Wisconsin-Madison(计算机科学系,威斯康星大学麦迪逊分校) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学)

AI总结 本文提出Dual-Align方法,通过双对齐策略校正后训练LLMs的置信度漂移和过程漂移,提升校准性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22009 2026-01-09 cs.CV

StreamFlow: Theory, Algorithm, and Implementation for High-Efficiency Rectified Flow Generation

StreamFlow:高效率校正流生成的理论、算法与实现

Sen Fang, Hongbin Zhong, Yalin Feng, Yanxin Zhang, Dimitris N. Metaxas

机构 * Rutgers University, New Jersey, USA(罗杰斯大学) Georgia Institute of Technology, Atlanta, Georgia, USA(佐治亚理工学院) Nanyang Technological University, Singapore(南洋理工大学) University of Wisconsin-Madison, Wisconsin, USA(威斯康星大学麦迪逊分校)

AI总结 本文提出StreamFlow,通过理论、算法和实现的综合优化,显著提升了基于流模型的图像生成效率,达到611%的加速效果。

Comments Improved the quality. Project Page at https://world-snapshot.github.io/StreamFlow/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02315 2026-01-06 cs.CV

Prithvi-Complimentary Adaptive Fusion Encoder (CAFE): unlocking full-potential for flood inundation mapping

普里提维-互补自适应融合编码器(CAFE):解锁洪水淹没制图的全部潜力

Saurabh Kaushik, Lalit Maurya, Beth Tellman

机构 * Center for Sustainability and the Global Environment (SAGE), University of Wisconsin–Madison(可持续性与全球环境中心(SAGE),威斯康星大学麦迪逊分校) Portsmouth AI and Data Science Centre (PAIDS), School of Computing, University of Portsmouth(波特兰人工智能与数据科学中心(PAIDS),计算学院,波特兰大学)

AI总结 普里提维-互补自适应融合编码器(CAFE)通过融合多通道多模态数据提升洪水制图的分割性能。

Comments Accepted at CV4EO Workshop @ WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01969 2026-01-06 cs.RO

What you reward is what you learn: Comparing rewards for online speech policy optimization in public HRI

你奖励什么,你就能学会什么:比较在线语音政策优化在公共人机交互中的奖励

Sichao Song, Yuki Okafuji, Kaito Ariu, Amy Koike

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) The University of Osaka(大阪大学)

AI总结 本文研究了在线语音政策优化在公共人机交互中的奖励机制比较,通过多臂老虎机问题分析不同奖励对政策适应性的影响,并提出实用设计经验。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09665 2026-01-06 cs.SI cs.CL

Tales of the 2025 Los Angeles Fire: Hotwash for Public Health Concerns in Reddit via LLM-Enhanced Topic Modeling

2025年洛杉矶火灾故事:通过LLM增强的专题建模在Reddit上进行公共健康担忧的热洗

Sulong Zhou, Qunying Huang, Shaoheng Zhou, Yun Hang, Xinyue Ye, Aodong Mei, Kathryn Phung, Yuning Ye, Uma Govindswamy, Zehan Li

机构 * Department of Landscape Architecture and Urban Planning & Urban Artificial Intelligence Lab, Texas A&M University(景观建筑与城市规划系及城市人工智能实验室,德克萨斯农工大学) Geography, University of Wisconsin-Madison(地理系,威斯康星大学麦迪逊分校) Google(谷歌) Department of Environmental and Occupational Health Sciences, School of Public Health, University of Texas Health Science Center at Houston(环境与职业健康科学系,公共卫生学院,德克萨斯健康科学中心休斯顿分部) McWilliams School of Biomedical Informatics, University of Texas Health Science Center at Houston(麦克威廉斯生物医学信息学学院,德克萨斯健康科学中心休斯顿分部)

AI总结 本研究通过LLM增强的专题建模分析2025年洛杉矶野火期间Reddit上的公众言论,识别出公共健康担忧和心理健康风险,构建了分层框架用于危机话语分析。

Comments Fix typos in Method Section. Add data/code availability

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01589 2026-01-06 quant-ph cs.LG

Learning Relationship between Quantum Walks and Underdamped Langevin Dynamics

学习量子行走与非阻尼朗之万动力学之间的关系

Yazhen Wang

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

AI总结 本文研究了量子行走与非阻尼朗之万动力学在学习任务中的关系,揭示了两者在计算和推断属性上的等同与非等同特性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09378 2026-01-06 cs.CY cs.CL

From Bench to Bedside: A Review of Clinical Trials in Drug Discovery and Development

从实验室到临床:药物发现与开发中临床试验的综述

Tianyang Wang, Ming Liu, Benji Peng, Xinyuan Song, Charles Zhang, Xintian Sun, Qian Niu, Junyu Liu, Silin Chen, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Yunze Wang, Yichao Zhang, Cheng Fei, Lawrence KQ Yan, Ziyuan Qin, Riyang Bao, Zekun Jiang

机构 * University of Liverpool, UK(利物浦大学) Purdue University, USA(普渡大学) Georgia Institute of Technology, USA(佐治亚理工学院) Emory University, USA(埃默里大学) Simon Fraser University, Canada(Simon Fraser大学) Kyoto University, Japan(京都大学) Zhejiang University, China(浙江大学) National Taiwan Normal University, Taiwan(台湾师范大学) Indiana University, USA(印第安纳大学) University of Edinburgh, UK(爱丁堡大学) The University of Texas at Dallas, USA(德克萨斯大学达拉斯分校) University of Wisconsin-Madison, USA(威斯康星大学麦迪逊分校) The Hong Kong University of Science(香港科学大学)

AI总结 本文综述了临床试验在药物发现与开发中的关键作用,探讨了各阶段特点、挑战及未来发展方向。

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00218 2026-01-05 cs.LG

Unknown Aware AI-Generated Content Attribution

未知意识的人工智能生成内容归因

Ellie Thieu, Jifan Zhang, Haoyue Bai

机构 * UW–Madison(威斯康星大学麦迪逊分校)

AI总结 本文提出了一种利用未标记野生数据提升人工智能生成内容归因性能的方法,通过约束优化增强目标生成器识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02152 2026-01-05 cs.LG cs.CL

Tabby: A Language Model Architecture for Tabular and Structured Data Synthesis

Tabby:一种用于表格和结构化数据合成的语言模型架构

Sonia Cromp, Satya Sai Srinath Namburi GNVV, Mohammed Alkhudhayri, Catherine Cao, Samuel Guo, Nicholas Roberts, Frederic Sala

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

AI总结 Tabby是一种用于表格和结构化数据合成的语言模型架构,通过门控专家混合机制提升数据生成质量,结合新型训练技术实现44%的质量提升。

Comments 21 pages, 8 figures. Appearing in TMLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17221 2026-01-01 cs.CV

DAVE: A VLM Vision Encoder for Document Understanding and Web Agents

DAVE: 一种用于文档理解与网络代理的视觉编码器

Brandon Huang, Hang Hua, Zhuoran Yu, Trevor Darrell, Rogerio Feris, Roei Herzig

机构 * MIT-IBM Watson AI Lab(MIT-IBM Watson AI实验室) UC Berkeley(加州大学伯克利分校) University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

AI总结 DAVE是一种专为文档理解和网络代理设计的视觉编码器,通过自监督和监督预训练结合模型融合策略,提升对文档和网络任务的适应性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24063 2026-01-01 cs.LG

How and Why LLMs Generalize: A Fine-Grained Analysis of LLM Reasoning from Cognitive Behaviors to Low-Level Patterns

LLMs如何泛化:从认知行为到低级模式的细粒度分析

Haoyue Bai, Yiyou Sun, Wenjie Hu, Shi Qiu, Maggie Ziyu Huan, Peiyang Song, Robert Nowak, Dawn Song

机构 * University of Wisconsin, Madison(威斯康星大学麦迪逊分校) University of California, Berkeley(加州大学伯克利分校) University of Pennsylvania(宾夕法尼亚大学) California Institute of Technology(加州理工学院)

AI总结 本文通过细粒度分析揭示LLMs在SFT和RL微调下的泛化差异,提出基于核心技能的基准测试,揭示RL模型在推理稳定性上的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00742 2026-01-01 cs.CV cs.AI eess.IV

Zoomer: Adaptive Image Focus Optimization for Black-box MLLM

Zoomer: 为黑盒大语言模型实现自适应图像聚焦优化

Jiaxu Qian, Chendong Wang, Yifan Yang, Chaoyun Zhang, Huiqiang Jiang, Xufang Luo, Yu Kang, Qingwei Lin, Anlan Zhang, Shiqi Jiang, Ting Cao, Tianjun Mao, Suman Banerjee, Guyue Liu, Saravan Rajmohan, Dongmei Zhang, Yuqing Yang, Qi Zhang, Lili Qiu

机构 * Microsoft(微软公司) Peking University(北京大学) University of Wisconsin Madison(威斯康星大学麦迪逊分校) University of Southern California(南加州大学)

AI总结 Zoomer通过自适应图像聚焦优化提升黑盒MLLM的多模态理解能力,显著提升准确性并减少token使用

Comments TMLR accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.08367 2026-01-01 cs.LG stat.ML

Active Learning with Neural Networks: Insights from Nonparametric Statistics

基于神经网络的主动学习:非参数统计学的见解

Yinglun Zhu, Robert Nowak

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

AI总结 本文提出基于非参数统计学的深度主动学习方法,首次提供了近最优的标签复杂性保证,无需低噪声假设,并扩展至Radon BV²空间。

Comments Correct typos and make minor structural revisions

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00043 2026-01-01 stat.ML cs.LG

Efficient Active Learning with Abstention

高效主动学习与回避

Yinglun Zhu, Robert Nowak

机构 * University of Wisconsin–Madison(威斯康星大学麦迪逊分校)

AI总结 本文提出了一种高效的主动学习算法,允许在困难样本上回避预测,从而显著降低标签复杂性,同时避免噪声寻求行为。

Comments Correct typos and make minor technical and structural revisions

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23214 2025-12-30 cs.CL cs.LG cs.PL cs.SE

Anka: A Domain-Specific Language for Reliable LLM Code Generation

Anka:一种用于可靠LLM代码生成的领域特定语言

Saif Khalfan Saif Al Mazrouei

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

AI总结 Anka是一种专为复杂代码生成设计的领域特定语言,通过约束语法显著减少错误,使LLM在多步骤任务中表现优于Python。

Comments 11 pages, 1 figure, 4 tables. Code and benchmarks available at https://github.com/BleBlo/Anka

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18564 2025-12-29 cs.AI

Vox Deorum: A Hybrid LLM Architecture for 4X / Grand Strategy Game AI -- Lessons from Civilization V

Vox Deorum:一种用于4X/大战略游戏AI的混合LLM架构——来自《文明V》的启示

John Chen, Sihan Cheng, Can Gurkan, Ryan Lay, Moez Salahuddin

机构 * University of Arizona(亚利桑那大学) Northwestern University(西北大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Independent Researcher(独立研究者)

AI总结 Vox Deorum提出了一种混合LLM架构用于4X游戏AI,通过宏观战略推理与子系统协同,展示了LLM在复杂游戏中的应用潜力。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19913 2025-12-24 stat.ML cs.LG hep-ex

Quasiprobabilistic Density Ratio Estimation with a Reverse Engineered Classification Loss Function

准概率密度比估计与逆向工程分类损失函数

Matthew Drnevich, Stephen Jiggins, Kyle Cranmer

机构 * Physics Department, New York University(纽约大学物理系) Physics Department, University of Wisconsin--Madison(威斯康星大学麦迪逊分校物理系) Data Science Institute, University of Wisconsin--Madison(威斯康星大学麦迪逊分校数据科学研究院)

AI总结 本文提出了一种适用于准概率密度比估计的凸损失函数,并通过粒子物理中的双希格斯生产实验实现了最先进的结果。

Comments 25 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12511 2025-12-24 cs.SE cs.AI cs.PL

SACTOR: LLM-Driven Correct and Idiomatic C to Rust Translation with Static Analysis and FFI-Based Verification

SACTOR:基于大语言模型的C到Rust翻译工具,结合静态分析和FFI验证

Tianyang Zhou, Ziyi Zhang, Haowen Lin, Somesh Jha, Mihai Christodorescu, Kirill Levchenko, Varun Chandrasekaran

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Wisconsin–Madison(威斯康星大学麦迪逊分校) Google(谷歌)

AI总结 SACTOR通过结合静态分析和FFI验证,利用大语言模型实现C到Rust的正确且习惯性翻译,提升了代码安全性和效率

Comments 35 pages, 15 figures Previously named as "LLM-Driven Multi-step Translation from C to Rust using Static Analysis"

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11006 2025-12-24 cs.CV cs.AI

Fine-Grained Instruction-Guided Graph Reasoning for Vision-and-Language Navigation

细粒度指令引导的图推理用于视觉-语言导航

Yaohua Liu, Xinyuan Song, Yunfu Deng, Yifan Xie, Binkai Ou, Yan Zhong

机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) Department of Computer Science, Emory University(埃默里大学计算机科学系) Department of Computer Science, University of Wisconsin-Madison(威斯康星大学麦迪逊分校计算机科学系) Tsinghua University(清华大学) Innovation and Research and Development Department, BoardWare Information System Company(BoardWare信息系统公司创新与研发部) School of Mathematics, Peking University(北京大学数学学院)

AI总结 本文提出细粒度指令引导的图推理框架OIKG,通过解耦角度与视觉提示并增强空间表示,提升视觉-语言导航中指令理解和跨模态对齐能力。

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏