arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Toronto(多伦多大学)

共收录 1022
2605.30617 2026-06-01 cs.RO math.OC

Exploiting Chordal Sparsity for Globally Optimal Estimation with Factor Graphs

利用弦稀疏性实现因子图的全局最优估计

Avinash Subramanian, Connor Holmes, Timothy D. Barfoot, Frank Dellaert, Frederike Dümbgen

机构 * College of Computing, Georgia Institute of Technology(佐治亚理工学院计算机学院) Robotics Institute, University of Toronto(多伦多大学机器人研究所) Department of Mechanical Engineering, Carnegie Mellon University(卡内基梅隆大学机械工程系)

AI总结 本文提出在GTSAM框架中自动构建凸半定规划松弛,并利用贝叶斯树分解加速求解,实现因子图的全局最优估计。

Journal ref ICRA 2026 WORKSHOP ON FRONTIERS OF OPTIMIZATION FOR ROBOTICS

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30481 2026-06-01 cs.CL

When English Rewrites Local Knowledge: Global Narrative Dominance in Large Language Models

当英语重写地方知识:大语言模型中的全球叙事主导

Md Arid Hasan, Ruwad Naswan, Farhan Samir, Sharifa Sultana, Syed Ishtiaque Ahmed

机构 * University of Toronto(多伦多大学) BUET(巴特利特大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

AI总结 本研究通过构建孟加拉语文化数据集CulturalNB,评估大语言模型在低资源文化背景下的跨语言知识一致性,发现英语提问会系统性地增加全球替代和制度框架,减少地方视角覆盖,表明文化失败不仅是知识缺失,更是根基和叙事优先级问题。

Comments Submitted to ARR

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30215 2026-06-01 cs.CV

Déjà View: Looping Transformers for Multi-View 3D Reconstruction

Déjà View: 用于多视图3D重建的循环Transformer

Alessandro Burzio, Tobias Fischer, Sven Elflein, Qunjie Zhou, Riccardo de Lutio, Jiawei Ren, Jiahui Huang, Shengyu Huang, Marc Pollefeys, Laura Leal-Taixé, Zan Gojcic, Haithem Turki

机构 * NVIDIA University of Modena and Reggio Emilia, AImageLab(摩德纳和雷焦艾米利亚大学,AImageLab) University of Toronto, Vector Institute(多伦多大学,向量研究所) ETH Zürich(苏黎世联邦理工学院)

AI总结 提出DéjàView模型,通过循环应用单个Transformer块进行迭代细化,以更少的参数和计算量在多个3D重建基准上达到或超越大规模前馈模型。

Comments Project Page: https://research.nvidia.com/labs/dvl/projects/dvlt

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29146 2026-06-01 cs.CL cs.AI

SafeRx-Agent: A Knowledge-Grounded Multi-Agent Framework for Safe and Explainable Medication Recommendation

SafeRx-Agent: 基于知识的多智能体框架用于安全且可解释的药物推荐

Xinyu Wang, Hanwei Wu, Zhenghan Tai, Sicheng Lyu, Qincheng Lu, Ziyu Zhao, Jijun Chi, Jingrui Tian, Xiao-Wen Chang, Ziyang Song

机构 * McGill University(麦吉尔大学) McMaster University(麦马斯特大学) University of Toronto(多伦多大学) Ohio University(俄亥俄大学)

AI总结 提出SafeRx-Agent,一种基于知识的多智能体框架,通过患者上下文、外部临床知识和安全验证来推荐可追溯的药物集合,在MIMIC-III和MIMIC-IV数据集上提高了细粒度药物预测准确性,同时控制了药物相互作用、禁忌症和药物集合大小。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22967 2026-06-01 cs.LG

Learned Relay Representations for Forward-Thinking Discrete Diffusion Models

学习的中继表示用于前向思考的离散扩散模型

Benjamin Rozonoyer, Jacopo Minniti, Dhruvesh Patel, Neil Band, Avishek Joey Bose, Tim G. J. Rudner, Andrew McCallum

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) University of Toronto(多伦多大学) Stanford University(斯坦福大学) Imperial College London(伦敦帝国学院) Mila Vijil

AI总结 提出Learned Relay Representations (Relay)方法,通过可微通道传递潜在信息,使掩码扩散模型在去噪步骤间前向思考,减少推理延迟并提升性能。

Comments 16 pages, 3 figures. Equal contribution: Benjamin Rozonoyer, Jacopo Minniti, and Dhruvesh Patel. Code: https://github.com/jacopo-minniti/relay

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.26262 2026-06-01 cs.CV

Semantic Foam: Unifying Spatial and Semantic Scene Decomposition

Semantic Foam:统一空间与语义场景分解

Amr Sharafeldin, Shrisudhan Govindarajan, Thomas Walker, Aryan Mikaeili, Daniel Rebain, Kwang Moo Yi, Andrea Tagliasacchi

机构 * Simon Fraser University(西蒙弗雷泽大学) University of Toronto(多伦多大学) Wayve Technologies(Wayve技术公司) University of British Columbia(不列颠哥伦比亚大学) University of Edinburgh(爱丁堡大学)

AI总结 提出Semantic Foam,通过扩展Radiant Foam表示,结合Voronoi网格的空间分解和显式语义特征场,实现高质量、一致性的语义分割。

Comments 15 pages, 10 figures, Accepted to CVPR 2026 (Highlight) , Project page: http://semanticfoam.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19966 2026-06-01 cond-mat.mtrl-sci cs.LG physics.chem-ph physics.comp-ph

Global Plane Waves From Local Gaussians: Periodic Charge Densities in a Blink

从局部高斯到全局平面波:眨眼间的周期电荷密度

Jonas Elsborg, Felix Ærtebjerg, Luca Thiede, Alán Aspuru-Guzik, Tejs Vegge, Arghya Bhowmik

机构 * Technical University of Denmark(技术大学) University of Toronto(多伦多大学) CAPeX Pioneer Center for Accelerating P2X Materials Discovery(CAPeX先锋中心) Canadian Institute for Advanced Research (CIFAR)(加拿大高级研究 institute) Vector Institute for Artificial Intelligence(人工智能研究所)

AI总结 提出ELECTRAFI模型,利用实空间各向异性高斯的解析傅里叶变换和泊松求和公式,通过单次逆FFT快速重建周期电荷密度,在保持高精度的同时速度提升高达633倍。

Comments ICML 2026, 29 pages including appendix, 11 Figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.30054 2026-05-29 cs.SE cs.AI

Projectional Decoding: Towards Semantic-Aware LLM Generation

投影式解码:迈向语义感知的LLM生成

Boqi Chen, José Antonio Hernández López, Aren A. Babikian

机构 * University of Ottawa(渥太华大学) University of Murcia(穆尔西亚大学) University of Toronto(多伦多大学)

AI总结 提出投影式解码框架,通过维护部分图模型作为主要工件表示,实现增量语义验证和错误检测,以提升LLM生成工件的语义有效性。

Comments 5 pages, 3 figures. Accepted at FSE 2026 IVR track

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29622 2026-05-29 cs.LG physics.chem-ph

MōLe-Λ: Learning the Coupled-Cluster Response State for Energies, Gradients, and Properties

MōLe-Λ: 学习耦合簇响应态以获取能量、梯度和性质

Andreas Burger, Luca Thiede, Abdulrahman Aldossary, Jorge A. Campos-Gonzalez-Angulo, Alex Zook, Jérôme Florian Gonthier, Alán Aspuru-Guzik

机构 * University of Toronto(多伦多大学) Vector Institute for Artificial Intelligence(人工智能向量研究所) NVIDIA(英伟达) Canadian Institute for Advanced Research (CIFAR)(加拿大高级研究研究院)

AI总结 提出MōLe-Λ模型,通过联合学习左右手振幅预测耦合簇响应态,高效计算能量、梯度及多类分子性质。

Comments ICML 2026 AI4Physics

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01456 2026-05-29 cs.LG cs.CV

Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations

Rectified LpJEPA:具有稀疏和最大熵表示的联合嵌入预测架构

Yilun Kuang, Yash Dagade, Tim G. J. Rudner, Randall Balestriero, Yann LeCun

机构 * New York University(纽约大学) Duke University(杜克大学) University of Toronto(多伦多大学) Brown University(布朗大学)

AI总结 提出Rectified Distribution Matching Regularization (RDMReg)损失,通过将表示对齐到Rectified Generalized Gaussian分布,实现稀疏且最大熵的表示,从而改进联合嵌入预测架构(JEPA)的性能。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17798 2026-05-29 cs.RO

SM2ITH: Safe Mobile Manipulation with Interactive Human Prediction via Task-Hierarchical Bilevel Model Predictive Control

SM2ITH:通过任务分层双层模型预测控制实现安全移动操作与人机交互预测

Francesco D'Orazio, Sepehr Samavi, Xintong Du, Siqi Zhou, Giuseppe Oriolo, Angela P. Schoellig

机构 * Department of Computer, Control and Management Engineering, of Sapienza University of Rome(意大利萨皮恩扎大学计算机、控制与管理工程系) University of Toronto Institute for Aerospace Studies (UTIAS) and the Vector Institute for Artificial Intelligence(多伦多大学航空航天研究所(UTIAS)和向量人工智能研究所) Learning Systems and Robotics lab at the Technical University of Munich and the Munich Institute for Robotics and Machine Intelligence (MIRMI)(慕尼黑技术大学学习系统与机器人实验室及慕尼黑机器人与机器智能研究所(MIRMI)) School of Computing Science, Faculty of Applied Sciences, Simon Fraser University(西蒙·弗雷泽大学应用科学学院计算机科学系)

AI总结 提出SM$^2$ITH框架,结合分层任务模型预测控制与双层优化的人机交互预测,实现动态人机环境中的安全高效移动操作。

Comments Accepted to the IEEE International Conference on Robotics and Automation (ICRA) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29318 2026-05-29 cs.GR cs.CV

FreeForm: Reduced-Order Deformable Simulation from Particle-Based Skinning Eigenmodes

FreeForm: 基于粒子蒙皮特征模态的降阶可变形仿真

Donglai Xiang, Vismay Modi, Rishit Dagli, Ty Trusty, Gilles Daviet, Anka He Chen, Nicholas Sharp, David I. W. Levin

机构 * NVIDIA University of Toronto(多伦多大学)

AI总结 提出一种基于再生核粒子法的无网格降阶超弹性物体仿真方法,通过求解弹性能量Hessian矩阵的广义特征系统构建降阶蒙皮权重,实现40倍训练加速并降低仿真误差。

Comments CVPR 2026, project website: https://research.nvidia.com/labs/sil/projects/freeform/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27995 2026-05-29 cs.AI

AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios

AsyncTool: 多任务场景下异步函数调用能力的评估

Kou Shi, Ziao Zhang, Shiting Huang, Avery Nie, Zhen Fang, Qiuchen Wang, Lin Chen, Huaian Chen, Zehui Chen, Feng Zhao

机构 * University of Science and Technology of China(中国科学技术大学) University of Toronto(多伦多大学)

AI总结 提出AsyncTool基准,通过模拟工具响应延迟的多任务环境,评估基于大语言模型的智能体在异步工具调用中的任务协调与效率。

Comments https://github.com/StoKou/repo-asynctool

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22080 2026-05-29 cs.CV cs.AI

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

JMed48k:用于视觉语言模型评估的多专业日本医疗执照基准

Yue Xun, Junyu Liu, Qian Niu, Xinyi Wang, Zheng Yuan, Zirui Li, Zequn Zhang, Bowen Zhao, Shujun Wang, Irene Li, Kan Hatakeyama-Sato, Yusuke Iwasawa, Yutaka Matsuo

机构 * The Hong Kong Polytechnic University(香港理工大学) Kyoto University(京都大学) The University of Tokyo(东京大学) Hohai University(淮海大学) University of Science and Technology of China(中国科学技术大学) University of Toronto(多伦多大学)

AI总结 本文提出JMed48k,一个包含48,862道试题和20,142张图像的多专业日本医疗执照基准,通过评估21个模型并引入配对图像移除审计,发现专有和开源模型显著受益于图像,而医学专用模型对视觉证据利用有限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26623 2026-05-29 cs.RO

A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation

用于在线连续时间连续体机器人状态估计的滑动窗口滤波器

Spencer Teetaert, Sven Lilge, Jessica Burgner-Kahrs, Timothy D. Barfoot

机构 * University of Toronto Robotics Institute(多伦多大学机器人研究所)

AI总结 提出一种专为连续体机器人设计的随机滑动窗口滤波器,在保持超实时运行速度的同时,通过连续时间方法提升滤波精度并实现在线操作。

Comments 8 pages, 6 figures. Submitted to IEEE-RAS International Conference on Soft Robotics 2026

Journal ref 2026 IEEE 9th International Conference on Soft Robotics (RoboSoft), 239-246

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16873 2026-05-29 cs.CV cs.SI

Multimodal LLMs See Sentiment

多模态大语言模型感知情感

Neemias B. da Silva, John Harrison, Rodrigo Minetto, Myriam R. Delgado, Bogdan T. Nassu, Thiago H. Silva

机构 * Universidade Tecnológica Federal do Paraná(联邦技术大学帕拉纳州大学) University of Toronto(多伦多大学)

AI总结 本文通过系统评估研究,探讨多模态大语言模型在图像情感分析中的三种方法,发现基于MLLM描述的两阶段流水线在微调后性能显著优于传统基线。

Comments 24 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.28816 2026-05-28 cs.CV

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

Gamma-World: 超越双玩家的生成式多智能体世界建模

Fangfu Liu, Kai He, Tianchang Shen, Tianshi Cao, Sanja Fidler, Yueqi Duan, Jun Gao, Igor Gilitschenski, Zian Wang, Xuanchi Ren

机构 * NVIDIA Tsinghua University(清华大学) University of Toronto(多伦多大学) Vector Institute(向量研究所)

AI总结 提出一种生成式多智能体世界模型,通过Simplex Rotary Agent Encoding实现智能体置换等价性,并采用Sparse Hub Attention降低跨智能体注意力成本,支持多玩家交互视频生成。

Comments Project Page: https://research.nvidia.com/labs/sil/projects/gamma-world

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27957 2026-05-28 cs.CL

DisasterBench: Benchmarking LLM Planning under Typed Tool Interface Constraints

DisasterBench: 在类型化工具接口约束下基准测试LLM规划

Zhitong Chen, Kai Yin, Weifeng Zhang, Zhiyuan Wang, Xiangjue Dong, Chengkai Liu, Zhewei Liu, Yiming Xiao, Ali Mostafavi, James Caverlee

机构 * Texas A&M University(德克萨斯A&M大学) University of Toronto(多伦多大学)

AI总结 提出DisasterBench基准,通过类型化工具接口评估LLM在灾害响应中的结构化多智能体规划能力,并引入首次故障点(FPoF)方法进行步骤级故障归因,揭示语义推理与执行约束之间的差距。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27736 2026-05-28 cs.LG cs.CV

Explicit Critic Guidance for Aligning Diffusion Models

显式评论家引导的对齐扩散模型

Zhengyang Liang, Qihang Zhang, Ceyuan Yang

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) The Chinese University of Hong Kong(香港中文大学)

AI总结 提出一种状态对齐的潜在演员-评论家框架,通过将扩散模型自身作为时间步条件价值函数,实现轨迹级PPO训练和推理时引导,在单/多奖励基准上优于先前方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27686 2026-05-28 cs.CV cs.AI

Tensor Memory: Fixed-Size Recurrent State for Long-Horizon Transformers

张量记忆:用于长程Transformer的固定大小循环状态

Kabir Swain, Sijie Han, Daniel Karl I. Weidele, Mauro Martino, Antonio Torralba

机构 * Massachusetts Institute of Technology, Cambridge, MA, USA(麻省理工学院) IBM Research, Cambridge, MA, USA(IBM研究院) University of Toronto, Toronto, Canada(多伦多大学)

AI总结 提出张量记忆模块,通过固定大小的3D循环张量状态增强Transformer,以解耦状态容量与输入长度,并保持空间归纳偏置,适用于长程视频理解。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27646 2026-05-28 cs.LG cs.AI

Hurwitz Quaternion Multiplicative Quantization for KV Cache Compression

Hurwitz四元数乘法量化用于KV缓存压缩

Kabir Swain, Sijie Han, Daniel Karl I. Weidele, Mauro Martino, David Cox, Antonio Torralba

机构 * Massachusetts Institute of Technology, Cambridge, MA, USA(麻省理工学院) IBM Research, Cambridge, MA, USA(IBM研究院) University of Toronto, Toronto, Canada(多伦多大学)

AI总结 提出一种免校准的Hurwitz四元数乘法量化方法,通过将K/V的4元素块视为四元数并用量化乘积编码,在约5比特下匹配fp16困惑度,实现高达5.05倍KV缓存压缩。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27541 2026-05-28 cs.LG

SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse Training

SparseOpt:解决稀疏训练中归一化引起的梯度倾斜

Mohammed Adnan, Rohan Jain, Tom Jacobs, Ekansh Sharma, Rahul G. Krishnan, Rebekka Burkholz, Yani Ioannou

机构 * University of Calgary(卡尔加里大学) University of Toronto(多伦多大学) Vector Institute(向量研究所) CISPA Helmholtz Center for Information Security(CISPA海德堡信息安全中心)

AI总结 针对动态稀疏训练收敛慢的问题,通过分析批归一化对稀疏训练的不利影响,提出稀疏感知优化器SparseOpt,实现更快的收敛和更好的泛化。

Comments Accepted International Conference on Machine Learning (ICML) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27281 2026-05-27 cs.LG stat.ML

Causal Risk Minimization for High-Dimensional Treatments

高维处理变量的因果风险最小化

Nikita Dhawan, Arnav Paruthi, Andrew Kim, Lovedeep Gondara, Jekaterina Novikova, Chris J. Maddison

机构 * University of Toronto(多伦多大学) Vector Institute(向量研究所) Vanguard(先锋)

AI总结 针对高维处理空间(如文本)的因果推断,提出通过分解因果误差为矩平衡误差序列并优化高阶平衡目标,以及将高维处理投影到低维属性的方法,实现无需属性特定训练的因果估计。

Comments 18 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.27068 2026-05-27 cs.CL cs.AI cs.MA

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

QUACK: 多模态社交推理智能体中的沟通知识质疑、理解与审计

Ye Yuan, Rui Song, Weien Li, Zeyu Li, Haochen Liu, Xiangyu Kong, Changjiang Han, Yonghan Yang, Zichen Zhao, Zixuan Dong, Fuyuan Lyu, Bowei He, Haolun Wu, Jikun Kang, Xue Liu

机构 * McGill University(麦吉尔大学) Mila - Quebec AI Institute(魁北克人工智能研究所) University of Cambridge(剑桥大学) MBZUAI - Mohamed bin Zayed University of Artificial Intelligence(MBZUAI - 摩苏尔·本·扎耶德人工智能大学) University of Toronto(多伦多大学) Salesforce

AI总结 提出QUACK框架,通过游戏结果、行为轨迹和话语一致性三级评估,自动审计多模态社交推理智能体语言与感知行为的一致性,发现最强智能体仍有15.1%的空间幻觉和过半无据指控。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26316 2026-05-27 cs.CV cs.AI

E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control

E$^3$C: 具有3D环境记忆和自我-外部人体姿态控制的视频生成

Qiao Gu, Lingni Ma, Adam W Harley, Richard Newcombe, Florian Shkurti, Julian Straub

机构 * Meta Reality Labs(Meta现实实验室) University of Toronto(多伦多大学)

AI总结 提出E$^3$C可控视频扩散框架,通过3D点云记忆和双通道人体控制(自我与外骨骼),实现物理一致的自我中心视频生成。

Comments Preprint. Project Page: https://e3c-videogen.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15283 2026-05-27 cs.CV cs.GR

LuxRemix: Lighting Decomposition and Remixing for Indoor Scenes

LuxRemix: 室内场景的光照分解与重新混合

Ruofan Liang, Norman Müller, Ethan Weber, Duncan Zauss, Nandita Vijaykumar, Peter Kontschieder, Christian Richardt

机构 * Meta Reality Labs(Meta现实实验室) University of Toronto(多伦多大学)

AI总结 提出一种基于图像的光照分解模型,从多视图场景捕获中分解室内光照为独立光源,并通过多视图光照协调集成到可重光照的3D高斯溅射表示中,实现交互式光源编辑。

Comments CVPR 2026. Project page: https://luxremix.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04310 2026-05-27 cs.AI

EvoEmo: Towards Evolved Emotional Policies for Adversarial LLM Agents in Multi-Turn Price Negotiation

EvoEmo:面向多轮价格谈判中对抗性LLM智能体的进化情感策略

Yunbo Long, Liming Xu, Lukas Beckenbauer, Yuhan Liu, Alexandra Brintrup

机构 * Department of Engineering, University of Cambridge(剑桥大学工程系) Rotman School of Management, University of Toronto(多伦多大学罗特曼管理学院) TUM School of Management, Technical University of Munich(慕尼黑技术大学管理学院) The Alan Turing Institute, London, UK(伦敦阿尔安·图灵研究院)

AI总结 提出EvoEmo进化强化学习框架,通过将情感状态转移建模为马尔可夫决策过程并采用种群遗传优化,动态优化多轮谈判中的情感表达,显著提升LLM智能体的谈判成功率、效率和买家节省。

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.11997 2026-05-27 cs.LG cs.AI cs.RO

Continual Model-Based Reinforcement Learning with Hypernetworks

基于超网络的连续模型强化学习

Yizhou Huang, Kevin Xie, Homanga Bharadhwaj, Florian Shkurti

机构 * Division of Engineering Science, University of Toronto, Canada(多伦多大学工程科学系) Department of Computer Science, University of Toronto, Canada(多伦多大学计算机科学系)

AI总结 提出HyperCRL方法,利用任务条件超网络在序列任务中持续学习动力学模型,避免重新训练并固定存储开销,在机器人 locomotion 和 manipulation 任务中优于现有持续学习方法。

Comments Updated link to project website in the abstract. 7 pages (+2 pages in appendix), 8 figures. In proceedings of the 2021 IEEE International Conference on Robotics and Automation

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.26111 2026-05-26 cs.CV cs.AI cs.GR cs.LG cs.MM

Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation

从多模态大语言模型中榨取能力用于主题驱动生成

Shuhong Zheng, Aashish Kumar Misraa, Yu-Teng Li, Yu-Jhe Li, Igor Gilitschenski

机构 * University of Toronto & Vector Institute(多伦多大学及向量研究所) Adobe(Adobe公司) Google(谷歌公司)

AI总结 提出一种结合多模态大语言模型和VAE身份条件的方法,通过双层级聚合模块和多阶段去噪策略,在主题驱动图像生成中实现多模态理解与身份保持的平衡,优于现有方法。

Comments 33 pages, 18 figures, Project Page: https://zsh2000.github.io/squeeze-mllm-subject-gen/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.25998 2026-05-26 cs.LG

Causal methods for LLM development and evaluation

因果方法在LLM开发与评估中的应用

Dennis Frauen, Marie Brockschmidt, Konstantin Hess, Haorui Ma, Yuchen Ma, Abdurahman Maarouf, Maresa Schröder, Jonas Schweisthal, Yuxin Wang, Athiya Deviyani, Sonali Parbhoo, Rahul G. Krishnan, Stefan Feuerriegel

机构 * Imperial College London(帝国理工学院伦敦分校) University of Toronto(多伦多大学)

AI总结 本文提出因果方法可解决LLM开发与评估中的关键因果问题,并系统梳理其在预训练、对齐、路由等环节的应用机会。

Comments Published in KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏