arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2509.26127 2026-03-04 cs.CV 57%

EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model

EchoGen: 通过前馈主体驱动自回归模型在任意场景中生成视觉回声

Ruixiao Dong, Zhendong Wang, Keli Liu, Li Li, Ying Chen, Kai Li, Daowen Li, Houqiang Li

机构 * University of Science and Technology of China(中国科学技术大学) Alibaba Group(阿里巴巴集团)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 EchoGen通过前馈主体驱动自回归模型在任意场景中生成高质量视觉回声,实现高效生成与高保真度的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02138 2026-03-03 cs.CV 57%

OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens

OmniLottie:通过参数化Lottie令牌生成向量动画

Yiying Yang, Wei Cheng, Sijin Chen, Honghao Fu, Xianfang Zeng, Yujun Cai, Gang Yu, Xingjun Ma

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 OmniLottie通过参数化Lottie令牌生成高质量向量动画,结合多模态指令与大规模数据集提升动画生成能力。

Comments Accepted by CVPR 2026. Project Page: https://openvglab.github.io/OmniLottie/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01552 2026-03-03 cs.CV 57%

Align-cDAE: Alzheimer's Disease Progression Modeling with Attention-Aligned Conditional Diffusion Auto-Encoder

Align-cDAE: 利用注意力对齐的条件扩散自编码器进行阿尔茨海默病进展建模

Ayantika Das, Keerthi Ram, Mohanasankar Sivaprakasam

机构 * Department of Electrical Engineering, Indian Institute of Technology Madras, Chennai, India(电子工程系,印度理工学院马德拉斯,钦奈,印度) Sudha Gopalakrishnan Brain Centre, Indian Institute of Technology Madras, Chennai, India(苏达·戈帕拉克里希南脑中心,印度理工学院马德拉斯,钦奈,印度) Department of Electrical Engineering, Indian Institute of Technology Madras Chennai, India(电子工程系,印度理工学院马德拉斯钦奈,印度)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 Align-cDAE通过引入注意力对齐和结构化潜在空间,提升扩散自编码器在阿尔茨海默病进展建模中的精度和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01068 2026-03-03 cs.CV cs.LG 57%

LLaDA-o: An Effective and Length-Adaptive Omni Diffusion Model

LLaDA-o:一种高效且长度自适应的多模态扩散模型

Zebin You, Xiaolu Zhang, Jun Zhou, Chongxuan Li, Ji-Rong Wen

机构 * Gaoling School of Artificial Intelligence, Renmin University of China, Beijing, China.(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models(北京大型模型研究关键实验室) Engineering Research Center of Next-Generation Intelligent Search(下一代智能搜索工程研究中心)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 LLaDA-o通过混合扩散框架和数据驱动的长度适应策略,实现了高效且灵活的多模态扩散建模,展示了在文本到图像生成任务中的卓越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11758 2026-03-03 q-bio.QM cs.AI 57%

Protein Structure Tokenization via Geometric Byte Pair Encoding

通过几何字对编码进行蛋白质结构分词

Michael Sun, Weize Yuan, Gang Liu, Wojciech Matusik, Marinka Zitnik

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Harvard Medical School(哈佛医学院) Apple(苹果公司) Notre Dame(诺特大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 GeoBPE通过几何字对编码实现蛋白质结构分词,提供压缩、数据效率和泛化能力,支持多架构应用并增强功能解释性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00521 2026-03-03 cs.LG cs.AI 57%

Phys-Diff: A Physics-Inspired Latent Diffusion Model for Tropical Cyclone Forecasting

Phys-Diff:一种基于物理的潜在扩散模型用于热带气旋预测

Lei Liu, Xiaoning Yu, Kang Chen, Jiahui Huang, Tengyuan Liu, Hongwei Zhao, Bin Li

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 Phys-Diff通过结合物理启发的归纳偏置和多模态数据整合,提升热带气旋预测的物理一致性与性能

Comments 5 pages, 4 figures. Accepted to IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00266 2026-03-03 cs.CV 57%

Adversarial Patch Generation for Visual-Infrared Dense Prediction Tasks via Joint Position-Color Optimization

为视觉-红外密集预测任务生成对抗性补丁的联合位置-颜色优化

He Li, Wenyue He, Weihang Kong, Xingchen Zhang

机构 * Yanshan University(燕山大学) University of Exeter(埃克塞特大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出AP-PCO框架,通过联合优化位置和颜色生成对抗性补丁,提升视觉-红外密集预测任务中的攻击性能和隐蔽性。

Comments 12 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10980 2026-03-03 cs.CV 57%

TrueSkin: Towards Fair and Accurate Skin Tone Recognition and Generation

TrueSkin: 向公平和准确的皮肤色调识别与生成迈进

Haoming Lu

机构 * Topaz Labs(Topaz实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 TrueSkin通过系统化的皮肤色调数据集,提升模型在公平性和准确性上的表现,验证了其在识别和生成任务中的有效性。

Comments The dataset is available for download at https://drive.google.com/file/d/1_ndw5uyY4h4DLL5iGTL4bVDKdE_g_H4B/view?usp=sharing

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00116 2026-03-03 cs.CV 57%

VoxelDiffusionCut: Non-destructive Internal-part Extraction via Iterative Cutting and Structure Estimation

VoxelDiffusionCut:通过迭代切割和结构估计实现非破坏性内部部件提取

Takumi Hachimine, Yuhwan Kwon, Cheng-Yu Kuo, Tomoya Yamanokuchi, Takamitsu Matsubara

机构 * Division of Information Science, Graduate School of Science and Technology, Nara Institute of Science and Technology(信息科学系,科学技术研究生学校,奈良科学技术研究所)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 VoxelDiffusionCut通过迭代切割和结构估计方法,利用扩散模型估计体素结构以实现非破坏性内部部件提取。

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19326 2026-03-02 cs.MA cs.AI 57%

City Editing: Hierarchical Agentic Execution for Dependency-Aware Urban Geospatial Modification

城市编辑:面向依赖意识的层级代理执行

Rui Liu, Steven Jige Quan, Zhong-Ren Peng, Zijun Yao, Han Wang, Zhengzhang Chen, Kunpeng Liu, Yanjie Fu, Dongjie Wang

机构 * University of Kansas, Lawrence, KS, USA(堪萨斯大学) Seoul National University, Seoul, South Korea(首尔国立大学) University of Florida, Gainesville, FL, USA(佛罗里达大学) NEC Laboratories America, Princeton, NJ, USA(NEC美国实验室) Clemson University, Clemson, SC, USA(克莱姆森大学) Arizona State University, Tempe, AZ, USA(亚利桑那州立大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出了一种基于层级代理的框架,用于高效、准确地修改城市规划,通过分层几何意图和迭代验证机制提升空间修改的效率与一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03467 2026-02-27 cs.CV 57%

ThinkRL-Edit: Thinking in Reinforcement Learning for Reasoning-Centric Image Editing

ThinkRL-Edit: 基于强化学习的图像编辑推理框架

Hengjia Li, Liming Jiang, Qing Yan, Yizhi Song, Hao Kang, Zichuan Liu, Xin Lu, Boxi Wu, Deng Cai

机构 * Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 ThinkRL-Edit通过引入链式推理和无偏奖励策略,提升图像编辑的推理能力,实现更精准和稳定的编辑效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21631 2026-02-26 cs.CV 57%

UniHand: A Unified Model for Diverse Controlled 4D Hand Motion Modeling

UniHand:一种统一模型,用于多样可控的4D手部运动建模

Zhihao Sun, Tong Wu, Ruirui Tu, Daoguo Dong, Zuxuan Wu

机构 * Institute of Trustworthy Embodied AI (TEAI)(可信具身人工智能研究所) Fudan University(复旦大学) Stanford University(斯坦福大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 UniHand通过统一的扩散框架,将手部运动估计与生成统一为条件运动合成,有效整合异构输入并提升鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21596 2026-02-26 cs.CV 57%

A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers

扩散变换器条件嵌入中的隐藏语义瓶颈

Trung X. Pham, Kang Zhang, Ji Woo Hong, Chang D. Yoo

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 研究揭示扩散变换器条件嵌入中的语义瓶颈,发现语义信息集中于少量维度,通过剪枝提升生成质量。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21461 2026-02-26 cs.CL 57%

VecGlypher: Unified Vector Glyph Generation with Language Models

VecGlypher: 一种基于语言模型的统一向量图形单元生成方法

Xiaoke Huang, Bhavul Gauri, Kam Woh Ng, Tony Ng, Mengmeng Xu, Zhiheng Liu, Weiming Ren, Zhaochong An, Zijian Zhou, Haonan Qiu, Yuyin Zhou, Sen He, Ziheng Wang, Tao Xiang, Xiao Han

机构 * Meta AI

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL

AI总结 VecGlypher通过多模态语言模型直接生成高质量向量图形,无需位图中间步骤,显著提升字体创建效率和可编辑性。

Comments Accepted to CVPR'26. Project page: https://xk-huang.github.io/VecGlypher/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21416 2026-02-26 cs.CV 57%

WildSVG: Towards Reliable SVG Generation Under Real-Word Conditions

WildSVG:迈向真实世界条件下可靠的SVG生成

Marco Terral, Haotian Zhang, Tianyang Zhang, Meng Lin, Xiaoqing Xie, Haoran Dai, Darsh Kaushik, Pai Peng, Nicklas Scharpff, David Vazquez, Joan Rodriguez

机构 * QuiverAI Columbia University(哥伦比亚大学) Illinois Institute of Technology(伊利诺伊理工学院) Mila - Quebec Artificial Intelligence Institute(魁北克人工智能研究所) University of Wisconsin-Madison(威斯康星大学麦迪逊分校) ServiceNow Research(ServiceNow研究)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 WildSVG通过引入现实世界基准测试,揭示了现有多模态模型在真实场景下生成SVG的不足,并指出了改进方向。

Comments 10 pages, 6 pages of additional material

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03255 2026-02-26 cs.LG cs.AI 57%

SciTS: Scientific Time Series Understanding and Generation with LLMs

SciTS: 基于大语言模型的科学时间序列理解与生成

Wen Wu, Ziyang Zhang, Liwei Liu, Xuenan Xu, Jimin Zhuang, Ke Fan, Qitan Lv, Junlin Liu, Chen Zhang, Zheqi Yuan, Siyuan Hou, Tianyi Lin, Kai Chen, Bowen Zhou, Chao Zhang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Tsinghua University(清华大学) Harbin Institute of Technology(哈尔滨工业大学) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 SciTS提出了一种基于LLM的科学时间序列理解与生成框架,通过基准测试发现通用LLM在泛化能力上优于专门模型,并引入TimeOmni框架提升性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16156 2026-02-25 cs.CV 57%

Pluggable Pruning with Contiguous Layer Distillation for Diffusion Transformers

可插拔剪枝与连续层蒸馏用于扩散变换器

Jian Ma, Qirong Peng, Xujie Zhu, Peixing Xie, Chen Chen, Haonan Lu

机构 * OPPO AI Center(OPPO人工智能中心) Sun Yat-sen University(中山大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 本研究提出PPCL方法,通过连续层蒸馏和灵活剪枝技术,在保持图像生成质量的同时实现50%的参数压缩,适用于资源受限环境。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08019 2026-02-25 cs.CV 57%

Coherent and Multi-modality Image Inpainting via Latent Space Optimization

通过潜在空间优化实现一致性和多模态图像修复

Lingzhi Pan, Tong Zhang, Bingyuan Chen, Qi Zhou, Wei Ke, Sabine Süsstrunk, Mathieu Salzmann

机构 * Xi’an Jiaotong University(西安交通大学) EPFL(苏黎世联邦理工学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 PILOT通过潜在空间优化提升图像修复的一致性和多模态处理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20057 2026-02-24 cs.RO cs.AI 57%

AdaWorldPolicy: World-Model-Driven Diffusion Policy with Online Adaptive Learning for Robotic Manipulation

AdaWorldPolicy:基于世界模型的扩散策略与在线自适应学习的机器人操控

Ge Yuan, Qiyuan Qiao, Jing Zhang, Dong Xu

机构 * The University of Hong Kong(香港大学) Beihang University(北航大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

AI总结 AdaWorldPolicy通过结合世界模型和在线自适应学习,实现机器人操控的高效动态适应。

Comments Homepage: https://AdaWorldPolicy.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19542 2026-02-24 cs.CV 57%

Vinedresser3D: Agentic Text-guided 3D Editing

Vinedresser3D: 基于代理的文本引导3D编辑

Yankuan Chi, Xiang Li, Zixuan Huang, James M. Rehg

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 Vinedresser3D通过多模态大语言模型和潜在空间编辑技术实现高质量文本引导的3D编辑,提升编辑精度与一致性。

Comments CVPR 2026, Project website:https://vinedresser3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19193 2026-02-24 cs.RO cs.AI 57%

Visual Prompt Guided Unified Pushing Policy

基于视觉提示的统一推送策略

Hieu Bui, Ziyan Gao, Yuya Hosoda, Joo-Ho Lee

机构 * Graduate School of Information Science and Engineering, Ritsumeikan University, Japan(立命馆大学信息科学与工程研究生院) Japan Advanced Institute of Science and Technology (JAIST)(日本先进科学研究院) College of Information Science and Engineering, Ritsumeikan University, Japan(立命馆大学信息科学与工程学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出一种基于视觉提示的统一推送策略,通过整合轻量级提示机制提升多模态推送动作的生成效率,适用于广泛规划问题,并在桌面清洁任务中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04808 2026-02-24 q-bio.NC cs.AI 57%

Setting up for failure: automatic discovery of the neural mechanisms of cognitive errors

失败的设定:自动发现认知错误的神经机制

Puria Radmard, Paul M. Bays, Máté Lengyel

机构 * Department of Engineering, University of Cambridge(工程系,剑桥大学) Department of Psychology, University of Cambridge(心理学系,剑桥大学) Department of Cognitive Science, Central European University(认知科学系,中央欧亚大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出通过训练RNNs复制行为特征,自动发现认知错误的神经机制,解决了传统方法在数据有限和行为优化上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21896 2026-02-24 cs.AI 57%

GenesisGeo: Technical Report

GenesisGeo:技术报告

Minfeng Zhu, Zi Wang, Sizhe Ji, Zhengtong Du, Shengqiang Tai, Junming Ke, Xiao Deng, Zanlang Yin, Xiuqi Huang, Heyu Wang, Wei Chen

机构 * State Key Lab of CAD&CG, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室) Polytechnic Institute, Zhejiang University(浙江大学多科大学院) Hangzhou Research Institute of AI and Holographic Technology(杭州人工智能与全息技术研究 institutes) Volkswagen Group Innovation(大众集团创新) School of Mathematical Science, Zhejiang University(浙江大学数学科学学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出GenesisGeo-1M数据集及基于多任务学习的几何学习框架,通过大规模合成数据提升模型在几何推理任务中的性能,实现金牌级表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18711 2026-02-24 cs.CV 57%

HIME: Mitigating Object Hallucinations in LVLMs via Hallucination Insensitivity Model Editing

HIME: 通过幻觉不敏感模型编辑缓解LVLMs中的物体幻觉

Ahmed Akl, Abdelwahed Khamis, Ali Cheraghian, Zhe Wang, Sara Khalifa, Kewen Wang

机构 * School of Information and Communication Technology, Griffith University, Australia(信息与通信技术学院,格里菲斯大学) Data61, CSIRO, Australia(Data61,澳大利亚联邦科学与工业研究组织) School of Engineering, Macquarie University, Sydney, Australia(工程学院,麦觉大学) School of Information Systems, Queensland University of Technology, Australia(信息系统学院,昆士兰技术大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 HIME通过分层加权编辑方法有效抑制LVLMs中的物体幻觉,减少61.8%的幻觉问题,无需额外参数或计算开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18451 2026-02-24 cs.CY cs.AI 57%

Developing a Multi-Agent System to Generate Next Generation Science Assessments with Evidence-Centered Design

开发一个多智能体系统以生成下一代科学评估并采用证据中心设计

Yaxuan Yang, Jongchan Park, Yifan Zhou, Xiaoming Zhai

机构 * AI4STEM Education Center, University of Georgia(AI4STEM教育中心,佐治亚大学) Department of Educational Psychology, University of Georgia(教育心理学系,佐治亚大学) School of Computing, University of Georgia(计算学院,佐治亚大学) Department of Mathematics, Science, and Social Studies Education, University of Georgia(数学、科学与社会科学教育系,佐治亚大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本研究提出将证据中心设计整合到多智能体系统中,以自动生成符合NGSS的评估项目,发现AI生成的项目在包容性方面表现良好,但存在清晰性和多模态设计的局限。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14514 2026-02-23 cs.CV 57%

Efficient Text-Guided Convolutional Adapter for the Diffusion Model

高效的文本引导卷积适配器用于扩散模型

Aryan Das, Koushik Biswas, Swalpa Kumar Roy, Badri Narayana Patro, Vinay Kumar Verma

机构 * VIT Bhopal(维特大学博帕尔分校) IIIT Delhi(德里印度理工学院) Tezpur University Assam(泰朱普大学阿萨姆分校) IIT Kanpur(坎普尔印度理工学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出Nexus Prime和Slim适配器,通过文本引导提升扩散模型的结构保持条件生成性能,显著减少参数量并保持高效果。

Comments Accepted in WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18016 2026-02-23 cs.CV 57%

Towards LLM-centric Affective Visual Customization via Efficient and Precise Emotion Manipulating

面向高效精准情感操控的基于大语言模型的有情感视觉定制

Jiamin Luo, Xuqian Gu, Jingjing Wang, Jiahong Lu

机构 * School of Computer Science and Technology, Soochow University(计算机科学与技术学院,苏州大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出基于大语言模型的有情感视觉定制任务,通过高效精准的情感操控方法实现图像主观情感的生成与修改。

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13567 2026-02-23 cs.AI 57%

On the Value of Labeled Data and Symbolic Methods for Hidden Neuron Activation Analysis

关于标记数据和符号方法在隐藏神经元激活分析中的价值

Abhilekha Dalal, Rushrukh Rayan, Adrita Barua, Eugene Y. Vasserman, Md Kamruzzaman Sarker, Pascal Hitzler

机构 * Kansas State University(堪萨斯州立大学) Bowie State University(比弗州立大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出了一种基于符号方法和背景知识的可解释人工智能方法,能够为卷积神经网络中的神经元提供有意义的解释,并在定量和定性方面优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16006 2026-02-19 cs.CV 57%

BTReport: A Framework for Brain Tumor Radiology Report Generation with Clinically Relevant Features

BTReport: 一种用于脑肿瘤放射学报告生成的框架,结合临床相关特征

Juampablo E. Heras Rivera, Dickson T. Chen, Tianyi Ren, Daniel K. Low, Asma Ben Abacha, Alberto Santamaria-Pang, Mehmet Kurt

机构 * University of Washington(华盛顿大学) University of Washington School of Medicine(华盛顿大学医学院) Microsoft Health AI(微软健康人工智能) Johns Hopkins School of Medicine(约翰霍普金斯医学院)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

AI总结 BTReport通过确定性特征提取和大语言模型结合,生成可解释的脑肿瘤放射学报告,并提供配套数据集提升临床应用

Comments Accepted to Medical Imaging with Deep Learning (MIDL) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15776 2026-02-18 cs.AI 57%

GlobeDiff: State Diffusion Process for Partial Observability in Multi-Agent Systems

GlobeDiff:多智能体系统中部分可观测性的状态扩散过程

Yiqin Yang, Xu Yang, Yuhua Jiang, Ni Mu, Hao Hu, Runpeng Xie, Ziyou Zhang, Siyuan Li, Yuan-Hua Ni, Qianchuan Zhao, Bo Xu

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) Tsinghua University(清华大学) Moonshot AI Nankai University(南开大学) Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

AI总结 GlobeDiff通过多模态扩散过程解决多智能体系统中部分可观测性问题,实现高保真度的全局状态推断。

Journal ref ICLR-2026

详情

展开后加载摘要…

URL PDF HTML 收藏