arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

International Conference on Learning Representations · 会议 · Machine Learning

共收录 9452
2412.15176 2026-04-21 cs.LG

Rethinking Uncertainty Estimation in LLMs: A Principled Single-Sequence Measure

重新思考大语言模型中的不确定性估计:一种原理性的单序列度量

Lukas Aichberger, Kajetan Schweighofer, Sepp Hochreiter

机构 * ELLIS Unit Linz and LIT AI Lab(林茨ELLIS单元和LIT AI实验室) Institute for Machine Learning, Johannes Kepler University Linz(机器学习研究所,林茨约瑟夫·克里格大学) NXAI GmbH(NXAI公司)

AI总结 本文提出基于proper scoring rules框架的G-NLL方法,通过单序列贪心解码实现高效可靠的不确定性估计,挑战传统多序列方法的必要性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17803 2026-04-21 cs.AI cs.LG

Adversarial Arena: Crowdsourcing Data Generation through Interactive Competition

对抗领域:通过互动竞争进行众包数据生成

Prasoon Goyal, Sattvik Sahai, Michael Johnston, Hangjie Shi, Yao Lu, Shaohua Liu, Anna Rumshisky, Rahul Gupta, Anna Gottardi, Desheng Zhang, Lavina Vaz, Leslie Ball, Lucy Hu, Luke Dai, Samyuth Sagi, Maureen Murray, Sankaranarayanan Ananthakrishnan

机构 * Amazon Nova Responsible AI(亚马逊Nova负责任人工智能)

AI总结 本文提出对抗领域框架,通过攻击者和防御者之间的互动竞争生成高质量对话数据,提升低资源领域和多轮对话的质量。

Comments 10 pages, 3rd DATA-FM workshop @ ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17568 2026-04-21 cs.LG math.ST stat.ML stat.TH

Diverse Dictionary Learning

多样字典学习

Yujia Zheng, Zijian Li, Shunxing Fan, Andrew Gordon Wilson, Kun Zhang

机构 * CMU(卡内基梅隆大学) MBZUAI(马克斯·普朗克人工智能研究所) NYU(纽约大学)

AI总结 本文提出多样字典学习问题,探讨在无法完全辨识的情况下,如何通过集合运算恢复潜在变量的结构,并展示其在合成和真实数据中的有效性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17224 2026-04-21 cs.LG stat.ML

LASER: Low-Rank Activation SVD for Efficient Recursion

LASER:低秩激活SVD用于高效递归

Ege Çakar, Ketan Ali Raghu, Lia Zheng

机构 * Harvard University(哈佛大学)

AI总结 本文研究了递归架构中激活流形的几何结构,提出LASER框架通过动态压缩实现激活记忆节省,同时保持模型精度。

Comments Accepted to the Latent and Implicit Thinking Workshop at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03692 2026-04-21 cs.CV cs.AI

Error as Signal: Stiffness-Aware Diffusion Sampling via Embedded Runge-Kutta Guidance

误差作为信号:通过嵌入式龙格-库塔引导的刚性感知扩散采样

Inho Kong, Sojin Lee, Youngjoon Hong, Hyunwoo J. Kim

机构 * Korea University(韩国大学) KAIST(韩国科学技术院) Seoul National University(首尔国立大学) Korea Institute for Advanced Study(韩国高级研究院)

AI总结 本文提出嵌入式龙格-库塔引导方法,通过识别刚性区域减少局部截断误差,提升扩散模型采样稳定性与质量。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19555 2026-04-21 cs.CR cs.AI

SOK: A Taxonomy of Attack Vectors and Defense Strategies for Agentic Supply Chain Runtime

SOK:代理供应链运行时的攻击向量和防御策略分类

Xiaochong Jiang, Shiqi Yang, Wenting Yang, Yichen Liu, Cheng Ji

机构 * Independent Researcher(独立研究者)

AI总结 本文系统化了现有研究,提出统一的运行时框架,区分数据和工具供应链攻击,并识别出病毒代理循环现象,主张采用零信任运行时架构。

Comments Published at ICLR 2026 Workshop on AI for Mechanism Design and Strategic Decision Making

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21613 2026-04-21 cs.CL cs.AI cs.LG

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining

超越URLs:元数据多样性与位置对高效LLM预训练的影响

Dongyang Fan, Diba Hashemi, Sai Praneeth Karimireddy, Martin Jaggi

机构 * EPFL(苏黎世联邦理工学院) University of Southern California(南加州大学)

AI总结 本文研究了元数据类型对LLM预训练的影响,发现细粒度文档质量指标能提升训练效率,并提出元数据追加方法以加速预训练过程。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23542 2026-04-21 cs.CL cs.AI cs.LG

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization

大语言模型判官的保质期:未来证明、向后兼容性和问题泛化

Janvijay Singh, Austin Xu, Yilun Zhou, Yefan Zhou, Dilek Hakkani-Tur, Shafiq Joty

机构 * Salesforce AI Research(Salesforce AI研究院) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Dartmouth College(达特茅斯学院)

AI总结 本文研究了细调判官模型在现实部署中的挑战,探讨了未来证明、向后兼容性和问题泛化三个核心问题,并通过实验发现持续学习在适应响应分布变化方面表现更优。

Comments Updated after ICLR 2026 Acceptance; 29 pages;

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01082 2026-04-21 cs.LG cs.PL

RefineStat: Efficient Exploration for Probabilistic Program Synthesis

RefineStat: 高效探索概率程序合成

Madhav Kanda, Shubham Ugare, Sasa Misailovic

机构 * University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 RefineStat通过引入语言模型驱动框架,在概率程序合成中强制语义约束并应用诊断感知细化,提升生成程序的语法正确性和统计可靠性。

Comments RefineStat constrains LM decoding with statistical validity checks and uses diagnostic-guided resampling (priors/likelihoods) to transform small LMs' drafts into correct, reliable probabilistic programs that can match or surpass closed-source models

Journal ref ICLR 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01038 2026-04-21 q-bio.BM cs.LG

Learning residue level protein dynamics with multiscale Gaussians

通过多尺度高斯学习残基水平蛋白质动力学

Mihir Bafna, Bowen Jing, Bonnie Berger

机构 * CSAIL, Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室) Dept. of Mathematics, Massachusetts Institute of Technology(麻省理工学院数学系)

AI总结 本文提出DynaProt框架,通过多变量高斯方法从静态结构预测蛋白质动力学,实现残基灵活性预测和协方差矩阵重建,具有更少参数和高精度。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19564 2026-04-21 cs.LG cs.AI

Bi-LoRA: Efficient Sharpness-Aware Minimization for Fine-Tuning Large-Scale Models

Bi-LoRA:高效锐度感知最小化用于大规模模型微调

Yuhang Liu, Tao Li, Zhehao Huang, Zuopeng Yang, Xiaolin Huang

机构 * Institute of Image Processing and Pattern Recognition, School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(图像处理与模式识别研究所,自动化与智能感知学院,上海交通大学) MoE Key Laboratory of System Control and Information Processing (Shanghai)(系统控制与信息处理MOE重点实验室(上海))

AI总结 Bi-LoRA通过引入辅助LoRA模块,有效解决传统SAM在LoRA参数上的限制,实现更高效的锐度优化,提升模型泛化能力。

Comments 32 pages,ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20879 2026-04-21 cs.CV

DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking

DriveAgent-R1: 通过主动感知和混合思维推进基于视觉语言模型的自动驾驶

Weicheng Zheng, Xiaofei Mao, Nanfei Ye, Pengxiang Li, Kun Zhan, Xianpeng Lang, Hang Zhao

机构 * Shanghai Qi Zhi Institute(上海启智研究院) LiAuto Tongji University(同济大学) Tsinghua University(清华大学)

AI总结 DriveAgent-R1通过主动感知和混合思维框架,提升自动驾驶中的视觉推理能力,采用三阶段训练策略,在复杂场景中实现高效决策,展现与顶级模型相当的性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13674 2026-04-21 cs.CL cs.AI

PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention

PrefixMemory-Tuning: 通过解耦前缀与注意力来现代化前缀微调

Haonan Wang, Brian Chen, Siquan Li, Xinhe Liang, Hwee Kuan Lee, Kenji Kawaguchi, Tianyang Hu

机构 * National University of Singapore(新加坡国立大学) Bioinformatics Institute, A*STAR(A*STAR生物信息研究所) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

AI总结 本文提出PrefixMemory-Tuning,通过将前缀模块移出注意力头以提升表达能力,实验证明其在多个基准上优于传统前缀微调方法,并在通用任务中表现接近现代PEFT技术。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12622 2026-04-21 cs.LG cs.AI math.OC

DR-SAC: Distributionally Robust Soft Actor-Critic for Reinforcement Learning under Uncertainty

DR-SAC:在不确定环境下用于强化学习的分布鲁棒软Actor-Critic

Mingxuan Cui, Duo Zhou, Yuxuan Han, Grani A. Hanasusanto, Qiong Wang, Huan Zhang, Zhengyuan Zhou

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) New York University(纽约大学)

AI总结 本文提出DR-SAC,首个基于Actor-Critic的分布鲁棒强化学习算法,用于连续动作空间的离线学习。通过最大化熵正则化奖励,对抗最坏可能的转移模型,在KL散度约束的不确定性集内实现收敛保证,并结合生成建模方法估计未知的名义转移模型。

Comments 31 Pages. Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12176 2026-04-21 cs.LG stat.ML

"Faithful to What?" On the Limits of Fidelity-Based Explanations

忠实于什么?关于基于忠实性的解释的局限性

Jackson Eshbaugh

机构 * Lafayette College(拉法叶学院)

AI总结 研究指出基于忠实性的解释方法无法准确反映模型预测性能,通过线性得分λ(f)发现高忠实度代理模型可能无法恢复数据结构的预测优势。

Comments 6 pages, 3 figures, 3 tables. Accepted at the Workshop on Scientific Methods for Understanding Deep Learning (Sci4DL) at ICLR 2026. Code available at https://github.com/jacksoneshbaugh/lambda-linearity-score/tree/main

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22226 2026-04-21 cs.CV

Expressive yet Efficient Feature Expansion with Adaptive Cross-Hadamard Products

具有适应性交叉哈达玛积的表达性高效特征扩展

Xuyang Zhang, Xi Zhang, Liang Chen, Hao Shi, Qingshan Guo

机构 * Beijing Institute of Technology(北京理工大学) Chongqing Innovation Center(重庆创新中心)

AI总结 本文提出适应性交叉哈达玛模块,通过可微离散采样和动态软号归一化实现高效特征重用,提升视觉模型的效率与精度。

Comments Accepted by ICLR 2026. Camera-ready Version. 24 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02904 2026-04-21 cs.CL cs.AI cs.LG

Enhancing Trust in Large Language Models via Uncertainty-Calibrated Fine-Tuning

通过不确定性校准微调增强大语言模型的信任

Ranganath Krishnan, Piyush Khanna, Omesh Tickoo

机构 * Capital One, AI Labs(Capital One人工智能实验室) Wayve Technologies(Wayve技术公司) Intel Corporation(英特尔公司)

AI总结 本文提出一种不确定性感知微调方法,用于提升大语言模型在自然语言生成任务中的不确定性估计能力,从而提高生成响应的可信度并减少幻觉现象。

Comments ICLR 2026 Trustworthy AI workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17148 2026-04-21 cs.AI

Graph-of-Agents: A Graph-based Framework for Multi-Agent LLM Collaboration

图-智能体:一种基于图的多智能体LLM协作框架

Sukwon Yun, Jie Peng, Pingzhi Li, Wendong Fan, Jie Chen, James Zou, Guohao Li, Tianlong Chen

机构 * UNC Chapel Hill(UNC夏洛茨维尔分校) Eigent AI(Eigent人工智能) MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室) Stanford University(斯坦福大学)

AI总结 本文提出Graph-of-Agents框架,通过节点采样、边构建和消息传递提升多智能体协作效率,仅用3个智能体在多个基准测试中优于使用全部6个智能体的基线方法。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16919 2026-04-21 cs.LG cs.AI cs.CV

Noise-Adaptive Diffusion Sampling for Inverse Problems Without Task-Specific Tuning

具有任务无关调优的噪声自适应扩散采样用于逆问题

Yingzhi Xia, Setthakorn Tanomkiattikun, Liangli Zhen, Zaiwang Gu

机构 * Institute of High Performance Computing, Agency for Science, Technology and Research, Singapore(高性能计算研究所,科技研究局,新加坡) Institute for Infocomm Research, Agency for Science, Technology and Research, Singapore(信息通信研究所,科技研究局,新加坡) Johns Hopkins University(约翰·霍普金斯大学)

AI总结 本文提出噪声空间Hamilton蒙特卡洛方法,用于逆问题求解,避免局部极值并提升鲁棒性,通过噪声空间探索提升重建质量。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16780 2026-04-21 cs.CV cs.AI cs.LG

FairNVT: Improving Fairness via Noise Injection in Vision Transformers

FairNVT:通过噪声注入提升视觉变换器的公平性

Qiaoyue Tang, Sepidehsadat Hosseini, Mengyao Zhai, Thibaut Durand, Greg Mori

机构 * University of British Columbia(不列颠哥伦比亚大学) RBC Borealis Simon Fraser University(西蒙 Fraser大学)

AI总结 FairNVT通过在视觉变换器中注入噪声,提升表示和预测层面的公平性,同时保持任务准确性,采用轻量适配器学习任务相关和敏感嵌入,并通过校准高斯噪声和正交约束减少敏感属性泄露。

Comments ICLR 2026 Algorithmic Fairness Across Alignment Procedures and Agentic Systems (AFAA) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16776 2026-04-21 cs.AI

SAVE: A Generalizable Framework for Multi-Condition Single-Cell Generation with Gene Block Attention

SAVE:一种多条件单细胞生成的通用框架,基于基因块注意力

Jiahao Li, Jiayi Dong, Peng Ye, Xiaochi Zhou, Haohai Lu, Fei Wang

机构 * College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院) Shanghai Key Laboratory of Intelligent Information Processing, Fudan University(复旦大学上海智能信息处理重点实验室)

AI总结 本文提出SAVE框架,通过基因块注意力和流匹配机制提升多条件单细胞生成的泛化能力,优于现有方法,尤其在低资源场景中表现突出。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16733 2026-04-21 cs.CV

Active World-Model with 4D-informed Retrieval for Exploration and Awareness

具有4D信息检索的主动世界模型用于探索与意识

Elaheh Vaezpour, Amirhosein Javadi, Tara Javidi

机构 * KavAI UCSD(加州大学圣迭戈分校)

AI总结 本文提出AW4RE模型,通过结合4D信息检索、几何支持与时间一致性,提升在动态环境中感知决策的准确性与一致性。

Comments 11 pages, 4 figures, submitted to ICLR 2026 2nd Workshop on World Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16678 2026-04-21 cs.LG

UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels

UniCon:通过核方法实现高效的对比对齐统一框架

Hangke Sui, Yuqing Wang, Minh N Do

机构 * Department of Electrical & Computer Engineering, The Grainger College of Engineering, UIUC(电气与计算机工程系,格拉inger工程学院,伊利诺伊大学香槟分校) Siebel School of Computing and Data Science, The Grainger College of Engineering, UIUC(计算与数据科学学院,格拉inger工程学院,伊利诺伊大学香槟分校) Coordinated Science Laboratory, University of Illinois Urbana-Champaign (UIUC)(协调科学实验室,伊利诺伊大学香槟分校) VinUni-Illinois Smart Health Center(VinUni-伊利诺伊智能健康中心)

AI总结 UniCon通过引入对比相似度权重矩阵S(γ),提供了一种统一的对比对齐框架,实现了闭式全局解,替代了 minibatch反向传播,提升了效率并保持了通用性和性能。

Comments 33 pages, 8 figures, 8 tables. Accepted by The Fourteenth International Conference on Learning Representations (ICLR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16516 2026-04-21 cs.CV cs.LG cs.MM

Operationalizing Fairness in Text-to-Image Models: A Survey of Bias, Fairness Audits and Mitigation Strategies

在文本到图像模型中实现公平性:对偏见、公平性审计及缓解策略的综述

Megan Smith, Venkatesh Thirugnana Sambandham, Florian Richter, Laura Crompton, Matthias Uhl, Torsten Schön

机构 * AImotion Bavaria, Technische Hochschule Ingolstadt(AImotion巴伐利亚、因戈尔施塔特技术大学) School of Transformation and Sustainability, Catholic University of Eichstätt-Ingolstadt(转型与可持续性学院,埃施塔特-因戈尔施塔特天主教大学) Chair of Economic and Social Ethics, University of Hohenheim(经济与社会伦理系,霍亨海姆大学)

AI总结 本文综述了文本到图像模型中的公平性研究,分析了偏见类型和公平性概念的分类,指出现有研究在目标公平与阈值公平之间的差距,并提出新的公平性操作框架。

Comments ICLR 2026 Algorithmic Fairness Across Alignment Procedures and Agentic Systems (AFAA) Workshop, reviews can be found at: https://openreview.net/forum?id=8DOkyBGWwP

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16506 2026-04-21 cs.CV cs.CL

Medical thinking with multiple images

多图像医学推理

Zonghai Yao, Benlu Wang, Yifan Zhang, Junda Wang, Iris Xia, Zhipeng Tang, Shuo Han, Feiyun Ouyang, Zhichao Yang, Arman Cohan, Hong Yu

机构 * Manning College of Information and Computer Sciences, UMass Amherst(UMass阿默斯特信息与计算机科学学院) Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝福德医疗保健中心) Department of Computer Science, Yale University(耶鲁大学计算机科学系) Miner School of Computer and Information Sciences, UMass Lowell(UMass洛威计算机与信息科学学院) Department of Electrical and Computer Engineering, UMass Lowell(UMass洛威电气与计算机工程系)

AI总结 本文提出MedThinkVQA多图像医学推理基准,通过专家标注数据验证多图像整合对临床推理的重要性,发现多图像推理瓶颈在于证据提取与对齐,且增加推理时间计算效果有限。

Comments Equal contribution for the first two authors. To appear in the proceedings of the Fourteenth International Conference on Learning Representations (ICLR 2026). Code is in https://github.com/benluwang/MedThinkVQA. Dataset is in https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.16391 2026-04-21 cs.RO cs.CV

Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining

通过分离前向和逆向动力学预训练实现解耦的机器人学习

Wenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng, Xin Jin, Li Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Fudan University(复旦大学) Eastern Institute of Technology, Ningbo(宁波东部技术研究所) Shanghai Innovation Institute(上海创新研究院) Tsinghua University(清华大学)

AI总结 本文提出DeFI框架,通过分离前向和逆向动力学预训练,利用不同数据源解决视觉-动作耦合问题,实现端到端微调,实验表明在CALVIN ABC-D和SimplerEnv上取得最佳性能。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23987 2026-04-21 cs.LG

Can we generate portable representations for clinical time series data using LLMs?

能否利用LLMs为临床时间序列数据生成可移植的表示?

Zongliang Ji, Yifei Sun, Andre Amaral, Anna Goldenberg, Rahul G. Krishnan

机构 * University of Toronto, Canada(加拿大多伦多大学) Sunnybrook Health Sciences Centre, Canada(加拿大阳光医疗科学中心) Vector Institute, Canada(加拿大向量研究所)

AI总结 本文研究了利用LLMs生成可移植患者表示的方法,通过将不规则ICU时间序列转换为自然语言摘要,再嵌入固定长度向量,实现跨医院的预测任务,展示了其在多个临床任务中的竞争力和可移植性。

Comments Accepted to the 14th International Conference on Learning Representations (ICLR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19926 2026-04-21 cs.LG cs.AI

Rethinking LoRA for Privacy-Preserving Federated Learning in Large Models

重新思考LoRA用于大模型隐私保护联邦学习

Jin Liu, Yinbin Miao, Ning Xi, Junkang Liu

机构 * School of Cyber Engineering, Xidian University(西安电子科技大学电子工程学院) College of Intelligence and Computing, Tianjin University(天津大学智能与计算学院)

AI总结 本文提出LA-LoRA方法,解决DPFL中LoRA的性能下降问题,通过解耦梯度交互和对齐更新方向,提升在严格隐私约束下的鲁棒性,实验显示在Swin Transformer和RoBERTa模型上达到SOTA性能。

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11528 2026-04-21 cs.CR cs.AI cs.CL

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs

别跟踪我!对抗大语言模型中属性推断攻击的主动防御

Dong Yan, Jian Liang, Ran He, Tieniu Tan

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

AI总结 本文提出TRACE-RPS框架,通过细粒度匿名化和推断阻止优化,有效降低大语言模型中属性推断准确性,实现隐私保护与模型效用的平衡。

Comments Accepted at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02830 2026-04-21 cs.CV

Densemarks: Learning Canonical Embeddings for Human Heads Images via Point Tracks

Densemarks: 通过点轨迹学习人类头部图像的规范嵌入

Dmitrii Pozdeev, Alexey Artemov, Ananta R. Bhattarai, Artem Sevastopolsky

机构 * Technical University of Munich (TUM)(慕尼黑技术大学) University of Bielefeld(比勒菲尔德大学)

AI总结 本文提出DenseMarks,通过点轨迹学习人类头部图像的规范嵌入,实现高质量的密集对应。利用Vision Transformer预测每个像素的3D嵌入,结合对比损失训练网络,通过多任务学习和空间连续性约束,实现可解释的规范空间,用于语义部分匹配、面部跟踪和立体重建。

Comments ICLR 2026. Project page: https://diddone.github.io/densemarks/ .Video: https://youtu.be/o8DOOYFW0gI .21 pages, 13 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏