arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Michigan(密歇根大学安娜堡分校)

共收录 1281
2603.23835 2026-03-26 stat.ML cs.LG math.ST stat.TH

Beyond Consistency: Inference for the Relative risk functional in Deep Nonparametric Cox Models

超越一致性:深度非参数Cox模型中相对风险函数的推断

Sattwik Ghosal, Xuran Meng, Yi Li

机构 * Department of Biostatistics, University of Michigan(密歇根大学生物统计学系)

AI总结 本文研究深度非参数Cox模型中相对风险函数的推断问题,提出新的渐近分布理论,解决梯度优化误差传播、点估计偏差控制及不确定性量化等挑战。

Comments 24 pages, 5 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23408 2026-03-25 cs.CV

GeoSANE: Learning Geospatial Representations from Models, Not Data

GeoSANE:从模型而非数据中学习地理空间表示

Joelle Hanna, Damian Falk, Stella X. Yu, Damian Borth

机构 * University of St.Gallen(圣加伦大学) University of Michigan and UC Berkeley(密歇根大学和伯克利大学) ESA Φ \Phi -Lab(欧洲航天局Φ实验室)

AI总结 GeoSANE通过整合多个模型的权重生成统一的地理空间表示,提升跨任务和模态的性能,优于从头训练或知识蒸馏的模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22903 2026-03-25 cs.RO

Task-Aware Positioning for Improvisational Tasks in Mobile Construction Robots via an AI Agent with Multi-LMM Modules

面向即兴任务的移动建筑机器人AI代理的定位任务感知

Seongju Jang, Francis Baek, SangHyun Lee

机构 * University of Michigan(密歇根大学) Georgia Institute of Technology(佐治亚理工学院)

AI总结 本文提出一种AI代理,通过多LMM模块实现对自然语言描述的即兴任务的理解与定位,使移动建筑机器人能自主完成非预定义任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22781 2026-03-25 cs.CV

Typography-Based Monocular Distance Estimation Framework for Vehicle Safety Systems

基于字体的单目距离估计框架用于车辆安全系统

Manognya Lokesh Reddy, Zheng Liu

机构 * Department of Computer and Information Science, University of Michigan-Dearborn(密歇根大学迪尔伯恩分校计算机与信息科学系) Department of Industrial and Manufacturing Systems Engineering, University of Michigan-Dearborn(密歇根大学迪尔伯恩分校工业与制造系统工程系)

AI总结 本文提出基于车牌标准化字体的单目距离估计方法,利用字符高度和针孔相机模型计算距离,结合多种鲁棒性增强技术,实现高精度实时距离估计。

Comments 25 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00415 2026-03-25 cs.AI

Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigm

迈向自演化基准:通过验证-再生产范式下的测试时间探索合成代理轨迹

Dadi Guo, Tianyi Zhou, Dongrui Liu, Chen Qian, Qihan Ren, Shuai Shao, Zhiyuan Fan, Yi R. Fung, Kun Wang, Linfeng Zhang, Jing Shao

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Hong Kong University of Science and Technology(香港科技大学) University of Michigan(密歇根大学) Renmin University of China(中国人民大学) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学)

AI总结 本文提出TRACE框架,通过测试时间探索使代理在原有任务基础上自演化更复杂的任务,提升基准的动态性和可靠性,适应并改进了AIME-2024等推理数据集。

Comments This is a work in progress due to methodology refinement and further evaluation

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12164 2026-03-25 cs.CL cs.DB cs.LG

Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Generator-Validator Fine-tuning

Table-LLM-Specialist: 用于表格任务的语言模型专家使用迭代生成-验证微调

Junjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong, Shi Han, Dongmei Zhang, Surajit Chaudhuri

机构 * University of Michigan(密歇根大学) Microsoft Corporation(微软公司)

AI总结 本文提出Table-LLM-Specialist,通过生成-验证框架实现表格任务的有效微调,无需人工标注数据,提升性能并降低部署成本。

Comments Full version of a paper in EMNLP 2025; code is available at: https://github.com/microsoft/Table-Specialist

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22248 2026-03-24 cs.LG cs.AI cs.IT math.IT stat.ML

Confidence-Based Decoding is Provably Efficient for Diffusion Language Models

基于置信度的解码在扩散语言模型中具有可证明的效率

Changxiao Cai, Gen Li

机构 * Department of Industrial and Operations Engineering, University of Michigan(工业与运营工程系,密歇根大学) Department of Statistics and Data Science, The Chinese University of Hong Kong(统计与数据科学系,香港中文大学)

AI总结 本文提出了一种基于熵和的解码策略,证明其在KL散度下具有ε精度,并在迭代次数上达到O(H(X0)/ε)的效率,适用于低熵数据分布。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21810 2026-03-24 eess.SY cs.MA cs.RO cs.SY

Partial Attention in Deep Reinforcement Learning for Safe Multi-Agent Control

深度强化学习中多智能体控制的安全部分注意力

Turki Bin Mohaya, Peter Seiler

机构 * Department of Electrical Engineering and Computer Science at the University of Michigan(密歇根大学电气工程与计算机科学系)

AI总结 本文提出在多智能体安全控制中应用注意力机制,设计神经网络控制高速公路汇入场景的自动驾驶车辆,通过部分注意力提升安全性和效率。

Comments This work has been accepted for publication in the proceedings of the 2026 American Control Conference (ACC), New Orleans, Louisiana, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19217 2026-03-24 cs.CL

Modality Matching Matters: Calibrating Language Distances for Cross-Lingual Transfer in URIEL+

模态匹配至关重要:在URIEL+中校准语言距离以实现跨语言迁移

York Hay Ng, Aditya Khan, Xiang Lu, Matteo Salloum, Michael Zhou, Phuong H. Hoang, A. Seza Doğruöz, En-Shiun Annie Lee

机构 * University of Toronto, Canada(多伦多大学) University of Michigan, USA(密歇根大学) Harvard University, USA(哈佛大学) Carnegie Mellon University, USA(卡内基梅隆大学) LT3, IDLab, Universiteit Gent, Belgium(IDLab,根特大学) Ontario Tech University, Canada(安大略技术大学)

AI总结 本文提出一种类型匹配的语言距离框架,通过结构感知的表示方法提升跨语言迁移性能,特别是在任务相关距离类型下表现更优。

Comments Accepted to EACL 2026 SRW

Journal ref In Proceedings of EACL 2026 (Volume 4: Student Research Workshop), pages 110 to 130. ACL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02406 2026-03-24 stat.ML cs.AI cs.CL cs.IT cs.LG math.IT

A Training-free Method for LLM Text Attribution

无需训练的LLM文本归因方法

Tara Radvand, Mojtaba Abdolmaleki, Mohamed Mostagir, Ambuj Tewari

机构 * Ross School of Business, University of Michigan, United States(密歇根大学罗斯商学院) Department of Statistics, University of Michigan, United States(密歇根大学统计学系)

AI总结 本文提出无需训练的LLM文本归因方法,通过零样本统计测试区分不同LLM生成文本,并证明测试误差随文本长度指数下降,同时验证理论结果和对抗性后编辑的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02868 2026-03-24 cs.AI

PrecLLM: A Privacy-Preserving Framework for Efficient Clinical Annotation Extraction from Unstructured EHRs using Small-Scale LLMs

PrecLLM: 一种用于从非结构化电子健康记录中高效提取临床注释的隐私保护框架,使用小型语言模型

Yixiang Qu, Yifan Dai, Shilin Yu, Pradham Tanikella, Malvika Pillai, Walter Chen, Jialiu Xie, Yishan Ren, Duan Wang, Yikai Wang, Sid Sheth, Guanting Chen, Yufeng Liu, Travis Schrank, Trevor Hackman, Didong Li, Di Wu

机构 * Department of Biostatistics, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校生物统计学系) Department of Genetics, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校遗传学系) Curriculum for Bioinformatics and Computational Biology, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校生物信息学与计算生物学课程) Carolina Health Informatics Program, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校健康信息学计划) Department of Statistics and Operations Research, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校统计学与运筹学系) Department of Otolaryngology/Head and Neck Surgery, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校耳鼻喉科及头颈外科系) Department of Statistics, University of Michigan(密歇根大学统计学系) Department of Biomedical Sciences, Adams School of Dentistry, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校阿德姆牙科学院生物医学科学系) Computational Medicine Program, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校计算医学计划) Lineberger Comprehensive Cancer Center, University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校林伯格综合癌症中心)

AI总结 本文提出PrecLLM框架,利用小型语言模型高效处理非结构化电子健康记录,通过正则表达式和RAG技术提升隐私保护下的临床注释提取性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20525 2026-03-24 cs.RO cs.SY eess.SY

High-Speed, All-Terrain Autonomy: Ensuring Safety at the Limits of Mobility

高速全地形自主性:在移动极限下的安全性保证

James R. Baxter, Bogdan I. Epureanu, Paramsothy Jayakumar, Tulga Ersal

机构 * Department of Mechanical Engineering, University of Michigan(密歇根大学机械工程系) U.S. Army Ground Vehicle Systems Center(美国陆军地面车辆系统中心)

AI总结 本文提出一种新型局部轨迹规划器,通过能量约束实现高速越野车辆在复杂地形中的安全行驶,通过仿真和实验证明其在极端场景下的有效性。

Comments 19 pages, 16 figures, submitted to IEEE Transactions on Robotics

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20443 2026-03-24 cs.RO

TRGS-SLAM: IMU-Aided Gaussian Splatting SLAM for Blurry, Rolling Shutter, and Noisy Thermal Images

TRGS-SLAM:基于IMU的高斯点云SLAM用于模糊、滚动快门和噪声热图像

Spencer Carmichael, Katherine A. Skinner

机构 * University of Michigan(密歇根大学)

AI总结 本文提出TRGS-SLAM,一种基于3DGS的热惯性SLAM系统,能处理热图像中的模糊、滚动快门和噪声问题,通过改进的3DGS渲染方法和创新技术实现高精度跟踪。

Comments Project page: https://umautobots.github.io/trgs_slam

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02259 2026-03-24 cs.GT cs.AI

Stochastically Dominant Peer Prediction

随机占优同伴预测

Yichi Zhang, Shengwei Xu, David Pennock, Grant Schoenebeck

机构 * DIMACS, Rutgers University(罗格斯大学DIMACS研究中心) University of Michigan, Ann Arbor(密歇根大学安娜堡分校)

AI总结 本文提出随机占优真实性机制,以增强同伴预测机制的诚实性,通过二元彩票评分和新机制在保持敏感度的同时提高效率。

Comments 29 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12339 2026-03-24 cs.RO cs.CV

SPOT: Point Cloud Based Stereo Visual Place Recognition for Similar and Opposing Viewpoints

SPOT:基于点云的立体视觉位置识别用于相似和对立视角

Spencer Carmichael, Rahul Agrawal, Ram Vasudevan, Katherine A. Skinner

机构 * Department of Robotics, University of Michigan(机器人学系,密歇根大学) Department of Robotics and the Department of Mechanical Engineering, University of Michigan(机器人学系和机械工程系,密歇根大学)

AI总结 本文提出SPOT技术,利用立体视觉里程计估计的结构进行对立视角视觉位置识别,通过双距离矩阵序列匹配方法提升识别精度,实验表明在不同光照条件下,SPOT在对立视角识别中达到91.7%的召回率,且存储和运行效率优于现有方法。

Comments Expanded version with added appendix. Published in ICRA 2024. Project page: https://umautobots.github.io/spot

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20129 2026-03-23 cs.RO

KUKAloha: A General, Low-Cost, and Shared-Control based Teleoperation Framework for Construction Robot Arm

KUKAloha:一种通用、低成本且基于共享控制的施工机器人臂远程操作框架

Yifan Xu, Qizhang Shen, Vineet Kamat, Carol Menassa

机构 * Department of Civil and Environmental Engineering, University of Michigan, United States of America(密歇根大学土木与环境工程系) Department of Robotics, University of Michigan, United States of America(密歇根大学机器人系)

AI总结 KUKAloha通过轻量引导臂和自主感知模块实现人机协作,提升施工机器人操作的安全性和重复性,实验表明其能降低操作负担并提高任务效率。

Comments 9 pages, 4 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17246 2026-03-23 cs.LG

On the Cone Effect and Modality Gap in Medical Vision-Language Embeddings

关于医学视觉-语言嵌入中的锥效应与模态间隙

David Restrepo, Miguel L Martins, Chenwei Wu, Luis Filipe Nakayama, Diego M Lopez, Stergios Christodoulidis, Maria Vakalopoulou, Enzo Ferrante

机构 * CentraleSupélec, Université Paris-Saclay, France(中央理工巴黎高等学院,巴黎萨克雷大学,法国) University of Porto, Portugal(葡萄牙波尔图大学) University of Michigan, USA(美国密歇根大学) Federal University of São Paulo, Brazil(巴西圣保罗联邦大学) Universidad del Cauca, Colombia(哥伦比亚考卡大学) Universidad de Buenos Aires, Argentina(阿根廷布宜诺斯艾利斯大学)

AI总结 本文研究了视觉语言模型中的锥效应及其对多模态性能的影响,通过轻量级机制调控模态间隙,发现医学数据集对间隙调节更敏感,但过度消除间隙并非最优,任务相关分离更佳。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05502 2026-03-23 math.OC cs.RO cs.SY eess.SY

Feasibility Analysis and Constraint Selection in Optimization-Based Controllers

基于优化控制器的可行性分析与约束选择

Panagiotis Rousseas, Haejoon Lee, Dimos V. Dimarogonas, Dimitra Panagou

机构 * Division of Decision and Control Systems, School of Electrical Engineering and Computer Science, KTH Royal Institute of Technology(决策与控制系统系,电气工程与计算机科学学院,皇家理工学院) Department of Robotics, University of Michigan(机器人系,密歇根大学) Department of Aerospace Engineering, University of Michigan(航空航天工程系,密歇根大学)

AI总结 本文提出了一种新的理论分析方法,用于评估线性约束的可行性并选择可行约束,通过仿真验证了算法在性能和计算效率上的优势。

Comments 13 pages, 4 figures, submitted to IEEE Transactions on Automatic Control

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02364 2026-03-20 cs.LG cs.CV stat.ML

Linearly Separable Features in Shallow Nonlinear Networks: Width Scales Polynomially with Intrinsic Data Dimension

浅层非线性网络中的线性可分特征:宽度与内在数据维度呈多项式关系

Alec S. Xu, Can Yaras, Peng Wang, Qing Qu

机构 * University of Michigan(密歇根大学) University of Macao(澳门大学)

AI总结 本文研究浅层非线性网络的线性可分能力,证明网络宽度可多项式缩放于数据内在维度而非 ambient 维度,实验验证了理论结论。

Comments 33 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.16740 2026-03-20 cs.RO cs.AI cs.LG

PLM-Net: Perception Latency Mitigation Network for Vision-Based Lateral Control of Autonomous Vehicles

PLM-Net:基于视觉的自动驾驶横向控制中感知延迟缓解网络

Aws Khalil, Jaerock Kwon

机构 * Department of Electrical and Computer Engineering, University of Michigan-Dearborn(电气与计算机工程系,密歇根大学-迪尔伯恩分校)

AI总结 PLM-Net通过模块化深度学习框架缓解视觉模仿学习车道保持系统中的感知延迟,通过插件架构保持原始控制流程,实现实时延迟缓解,降低转向误差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16936 2026-03-19 cs.CV cs.AI

TDMM-LM: Bridging Facial Understanding and Animation via Language Models

TDMM-LM:通过语言模型弥合面部理解与动画之间的差距

Luchuan Song, Pinxin Liu, Haiyang Liu, Zhenchao Jin, Yolo Yunlong Tang, Zichong Xu, Susan Liang, Jing Bi, Jason J Corso, Chenliang Xu

机构 * University of Rochester(罗切斯特大学) University of Tokyo(东京大学) University of Michigan(密歇根大学) Voxel51

AI总结 本文通过语言模型实现面部动作的双向建模,提出TDMM-LM框架,利用生成模型合成面部行为数据集,并通过Motion2Language和Language2Motion任务实现面部动画与理解的统一路径。

Comments 12 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16830 2026-03-19 cs.CL

PEPPER: Perception-Guided Perturbation for Robust Backdoor Defense in Text-to-Image Diffusion Models

PEPPER:基于感知的扰动用于文本到图像扩散模型的鲁棒后门防御

Oscar Chew, Po-Yi Lu, Jayden Lin, Kuan-Hao Huang, Hsuan-Tien Lin

机构 * Texas A&M University(德克萨斯A&M大学) National Taiwan University(台湾国立大学) University of Michigan(密歇根大学)

AI总结 PEPPER通过重写提示生成语义上不同但视觉上相似的描述,干扰输入提示中的触发器,提高文本到图像扩散模型的鲁棒性,尤其有效对抗基于文本编码器的攻击。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27183 2026-03-19 cs.CL

Simple Additions, Substantial Gains: Expanding Scripts, Languages, and Lineage Coverage in URIEL+

简单增加,显著提升:在URIEL+中扩展脚本、语言和谱系覆盖

Mason Shipton, York Hay Ng, Aditya Khan, Phuong Hanh Hoang, Xiang Lu, A. Seza Doğruöz, En-Shiun Annie Lee

机构 * Ontario Tech University(安大略技术大学) University of Toronto(多伦多大学) University of Michigan(密歇根大学) LT3, IDLab, Universiteit Gent(LT3,IDLab,根特大学)

AI总结 URIEL+通过引入脚本向量、整合Glottolog和扩展谱系推断,提升了语言覆盖和特征稀疏性,使低资源语言的跨语言迁移性能提升6%。

Comments Accepted to LREC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13939 2026-03-18 cs.CL cs.AI cs.CY

Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers

读者更倾向于AI在受版权保护的书籍上训练输出的文本而非专家人类作家

Tuhin Chakrabarty, Jane C. Ginsburg, Paramveer Dhillon

机构 * Department of Computer Science, Stony Brook University(石溪大学计算机科学系) Columbia Law School(哥伦比亚法学院) School of Information Science, University of Michigan(密歇根大学信息科学学院) MIT Initiative on the Digital Economy(麻省理工学院数字经济倡议)

AI总结 研究比较了AI与专家作家在模仿50位获奖作者风格的文本表现,发现微调AI在风格和质量上更受读者青睐,且效果稳健。

Comments Preprint Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16166 2026-03-18 cs.RO cs.CV

SignNav: Leveraging Signage for Semantic Visual Navigation in Large-Scale Indoor Environments

SignNav:利用标识信息在大规模室内环境中实现语义视觉导航

Jian Sun, Yuming Huang, He Li, Shuqi Xiao, Shenyan Guo, Maani Ghaffari, Qingbiao Li, Chengzhong Xu, Hui Kong

机构 * Faculty of Science and Technology, University of Macau(澳门大学科技学院) Department of Robotics, University of Michigan(密歇根大学机器人系)

AI总结 本文提出SignNav任务,通过解读标识信息进行语义导航,构建LSI-Dataset并提出START模型,实现端到端决策,达到80%的成功率和0.74 NDTW。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11370 2026-03-18 cs.LG

Relaxed Efficient Acquisition of Context and Temporal Features

放松的高效上下文与时间特征获取

Yunni Qu, Dzung Dinh, Grant King, Whitney Ringwald, Bing Cai Kok, Kathleen Gates, Aidan Wright, Junier Oliva

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Michigan(密歇根大学) University of Minnesota Twin Cities(明尼苏达大学双城分校)

AI总结 REACT框架通过联合优化上下文描述和时间适应性特征获取,提高预测性能并降低获取成本,解决生物医学中测量资源受限的问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11141 2026-03-18 cs.CV

Learning complete and explainable visual representations from itemized text supervision

从条目化文本监督中学习完整且可解释的视觉表示

Yiwei Lyu, Chenhui Zhao, Soumyanil Banerjee, Shixuan Liu, Akshay Rao, Akhil Kondepudi, Honglak Lee, Todd C. Hollon

机构 * University of Michigan(密歇根大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 本文提出ItemizedCLIP框架,通过条目化文本监督学习完整且可解释的视觉表示,在四个自然条目化领域和一个合成数据集上实现了零样本性能和细粒度可解释性的显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09347 2026-03-18 stat.ML cs.LG math.ST stat.TH

Inference for Deep Neural Network Estimators in Generalized Nonparametric Models

深度神经网络估计器在广义非参数模型中的推断

Xuran Meng, Yi Li

机构 * Department of Biostatistics, University of Michigan(密歇根大学生物统计学系)

AI总结 本文提出在广义非参数回归模型中使用深度神经网络估计器,并开发了严谨的推断框架,通过Ensemble Subsampling Method方法构建置信区间,展示了在非参数逻辑、泊松和二项回归模型中的有效性,并应用于eICU数据集进行ICU再入院风险预测。

Comments 91 pages, 14 figures, 20 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15418 2026-03-17 cs.RO cs.AI

MA-VLCM: A Vision Language Critic Model for Value Estimation of Policies in Multi-Agent Team Settings

MA-VLCM:一种用于多智能体团队设置中策略价值估计的视觉语言批评模型

Shahil Shaik, Aditya Parameshwaran, Anshul Nayak, Jonathon M. Smereka, Yue Wang

机构 * University of Michigan, Ann Arbor(密歇根大学安娜堡分校) Clemson University(克莱姆斯大学)

AI总结 本文提出MA-VLCM,通过预训练的视觉语言模型替代传统集中批评者,提升多智能体强化学习的样本效率和泛化能力,适用于资源受限的机器人部署。

Comments 7 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15237 2026-03-17 cs.CV

Multi-turn Physics-informed Vision-language Model for Physics-grounded Anomaly Detection

多轮物理引导的视觉-语言模型用于物理基础的异常检测

Yao Gu, Xiaohao Xu, Yingna Wu

机构 * Shanghaitech University(上海科技大学) University of Michigan, Ann Arbor(密歇根大学安娜堡分校)

AI总结 本文提出一种多轮物理引导的视觉-语言模型,通过编码物体属性和动态约束提升物理基础异常检测性能,实验显示其在视频级检测中达到96.7%的AUROC,优于现有最佳方法。

Comments Accepted by IEEE ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏