arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4946 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4946 篇

2505.15922 2026-02-12 cs.CL 79%

Aligning Dialogue Agents with Global Feedback via Large Language Model Multimodal Reward Decomposition

通过大语言模型多模态奖励分解对齐对话代理

Dong Won Lee, Hae Won Park, Cynthia Breazeal, Louis-Philippe Morency

机构 * MIT(麻省理工学院) CMU(卡内基梅隆大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出了一种基于大语言模型的多模态奖励分解方法,通过分解会话级反馈来提升对话生成质量,无需人工反馈。

Comments 9 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10659 2026-02-12 cs.CV 79%

Multimodal Priors-Augmented Text-Driven 3D Human-Object Interaction Generation

多模态先验增强的文本驱动3D人-物交互生成

Yin Wang, Ziyao Zhang, Zhiying Leng, Haitian Liu, Frederick W. B. Li, Mu Li, Xiaohui Liang

机构 * State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室) Beihang University(北京航空航天大学) Department of Computer Science, University of Durham(杜伦大学计算机科学系) Zhongguancun Laboratory(中关村实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出MP-HOI框架,通过多模态先验、增强物体表示、模态感知MoE模型和级联扩散技术,提升文本驱动的3D人-物交互生成效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09528 2026-02-11 cs.CV 79%

SchröMind: Mitigating Hallucinations in Multimodal Large Language Models via Solving the Schrödinger Bridge Problem

SchröMind: 通过求解薛定谔桥问题减轻多模态大语言模型中的幻觉

Ziqiang Shi, Rujie Liu, Shanshan Yu, Satoshi Munakata, Koichi Shirahata

机构 * Fujitsu Research \& Development Center Co.,LTD., Beijing, China Fujitsu Limited, Tokyo, Japan

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 SchröMind通过求解薛定谔桥问题,有效减轻多模态大语言模型中的幻觉问题,提升模型在医疗等高风险领域的应用能力。

Comments ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08249 2026-02-10 eess.IV cs.CV 79%

A Unified Framework for Multimodal Image Reconstruction and Synthesis using Denoising Diffusion Models

基于去噪扩散模型的多模态图像重建与合成统一框架

Weijie Gan, Xucheng Wang, Tongyao Wang, Wenshang Wang, Chunwei Ying, Yuyang Hu, Yasheng Chen, Hongyu An, Ulugbek S. Kamilov

机构 * Department of Computer Science and Engineering, Washington University in St. Louis(华盛顿大学圣路易斯分校计算机科学与工程系) Mallinckrodt Institute of Radiology, Washington University in St. Louis(华盛顿大学圣路易斯分校马林克罗德特放射医学研究所) Department of Electrical and Systems Engineering, Washington University in St. Louis(华盛顿大学圣路易斯分校电气与系统工程系) Department of Neurology, Washington University in St. Louis(华盛顿大学圣路易斯分校神经病学系) Department of Biomedical Engineering, Washington University in St. Louis(华盛顿大学圣路易斯分校生物医学工程系) Division of Biology and Biomedical Sciences, Washington University in St. Louis(华盛顿大学圣路易斯分校生物学与生物医学科学 division) Department of Electrical and Computer Engineering, University of Wisconsin–Madison(威斯康星大学麦迪逊分校电气与计算机工程系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 Any2all通过统一框架实现多模态图像重建与合成,利用去噪扩散模型提升性能与质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.09851 2026-02-10 cs.RO cs.CV 79%

Simultaneous Tactile-Visual Perception for Learning Multimodal Robot Manipulation

同时触觉-视觉感知用于学习多模态机器人操作

Yuyang Li, Yinghan Chen, Zihang Zhao, Puhao Li, Tengyu Liu, Siyuan Huang, Yixin Zhu

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) School of Psychological and Cognitive Sciences, Peking University(北京大学心理与认知科学学院) Beijing Key Lab of Behavior and Mental Health, Peking University(北京大学行为与心理健康北京市重点实验室) Beijing Institute for General Artificial Intelligence(北京一般人工智能研究院) State Key Lab for General Artificial Intelligence(一般人工智能国家重点实验室) Embodied Intelligence Lab, PKU-Wuhan Institute for Artificial Intelligence(具身智能实验室,北京大学武汉人工智能研究院) Department of Computer Science and Technology, University of Cambridge(剑桥大学计算机科学与技术系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 TacThru-UMI通过结合同时触觉-视觉感知与现代学习框架,实现了高精度多模态机器人操作。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.23278 2026-02-10 cs.CV 79%

UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing

UniLiP: 适配CLIP以实现统一的多模态理解、生成与编辑

Hao Tang, Chenwei Xie, Xiaoyi Bao, Tingyu Weng, Pandeng Li, Yun Zheng, Liwei Wang

机构 * Center for Data Science, Peking University(北京大学数据科学中心) Alibaba Group(阿里巴巴集团) CASIA Center for Machine Learning Research, Peking University(北京大学机器学习研究中心) State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 UniLIP通过两阶段训练和双条件架构,提升CLIP在多模态理解、生成和编辑任务中的性能,实现高效参数下的高表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25682 2026-02-10 cs.CL 79%

PairUni: Pairwise Training for Unified Multimodal Language Models

PairUni: 用于统一多模态语言模型的成对训练

Jiani Zheng, Zhiyang Teng, Kunpeng Qiu, Xiangtai Li, Anran Wang, Yu Tian, Ye Tian, Haochen Wang, Zhuochen Wang

机构 * ByteDance(字节跳动)

专题命中 多模态生成 :multimodal(title);cross-modal(abstract);分类 cs.CL

AI总结 PairUni通过成对训练提升统一多模态语言模型的理解与生成能力,采用配对数据集和PairGRPO算法实现更有效的策略学习。

Comments 22 pages, 11 figures, and 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11924 2026-02-09 cs.CV 79%

Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation

通过跨模态注意力安装实现对齐的新视角图像和几何合成

Min-Seop Kwak, Junho Kim, Sangdoo Yun, Dongyoon Han, Taekyung Kim, Seungryong Kim, Jin-Hwa Kim

机构 * NAVER AI Lab(NAVER AI实验室) KAIST AI(韩国科学技术院人工智能学院) SNU AIIS(首尔国立大学人工智能研究所)

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于扩散的框架,通过跨模态注意力蒸馏实现新视角图像和几何的对齐生成,提升几何鲁棒性和重建质量。

Comments Project page at https://cvlab-kaist.github.io/MoAI

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06166 2026-02-09 cs.CV 79%

M3: High-fidelity Text-to-Image Generation via Multi-Modal, Multi-Agent and Multi-Round Visual Reasoning

M3:通过多模态、多智能体和多轮视觉推理实现高保真的文本到图像生成

Bangji Yang, Ruihan Guo, Jiajun Fan, Chaoran Cheng, Ge Liu

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

AI总结 M3通过多智能体和多轮视觉推理提升开源模型的文本到图像生成能力,超越商业系统并实现最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.12203 2026-02-04 cs.CV 79%

LazyDrag: Enabling Stable Drag-Based Editing on Multi-Modal Diffusion Transformers via Explicit Correspondence

LazyDrag: 通过显式对应关系在多模态扩散变换器上实现稳定的拖拽编辑

Zixin Yin, Xili Dai, Duomin Wang, Xianfang Zeng, Lionel M. Ni, Gang Yu, Heung-Yeung Shum

机构 * The Hong Kong University of Science and Technology(香港科技大学) StepFun The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

AI总结 LazyDrag通过显式对应关系实现多模态扩散变换器的稳定拖拽编辑,消除了隐式点匹配依赖,提升生成能力与精确控制。

Comments https://zxyin.github.io/LazyDrag

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20418 2026-02-04 eess.IV cs.CV 79%

Diff4MMLiTS: Advanced Multimodal Liver Tumor Segmentation via Diffusion-Based Image Synthesis and Alignment

Diff4MMLiTS: 通过基于扩散的图像合成与对齐的先进多模态肝肿瘤分割

Shiyun Chen, Li Lin, Pujin Cheng, ZhiCheng Jin, JianJian Chen, HaiDong Zhu, Kenneth K. Y. Wong, Xiaoying Tang

机构 * Department of Electronic and Electrical Engineering, Southern University of Science and Technology, Shenzhen, China(电子与电气工程系,南方科技大学,深圳,中国) Department of Electrical and Electronic Engineering, The University of Hong Kong, Hong Kong SAR, China(电气与电子工程系,香港大学,香港特别行政区,中国) Department of Radiology, Zhongda Hospital, Medical School, Southeast University, Nanjing, China(放射科,中大医院,医学院,东南大学,南京,中国) Jiaxing Research Institute, Southern University of Science and Technology, Jiaxing, China(嘉兴研究所,南方科技大学,嘉兴,中国)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 Diff4MMLiTS通过基于扩散的图像合成与对齐技术,实现肝肿瘤的多模态分割,无需严格对齐的多模态数据,提升了分割性能。

Comments International Workshop on Machine Learning in Medical Imaging, 668-678

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01901 2026-02-03 cs.CV 79%

Q Cache: Visual Attention is Valuable in Less than Half of Decode Layers for Multimodal Large Language Model

Q Cache:视觉注意力在少于一半的解码层中具有价值用于多模态大语言模型

Jiedong Zhuang, Lu Lu, Ming Dai, Rui Hu, Jian Chen, Qiang Liu, Haoji Hu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 Q Cache通过跨层共享相似注意力模式,减少多模态大语言模型中的KV缓存使用量,提升吞吐量并保持性能

Comments Accepted by AAAI26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00960 2026-02-03 cs.LG cs.AI cs.CE stat.CO stat.ML 79%

Multimodal Scientific Learning Beyond Diffusions and Flows

多模态科学学习超越扩散与流

Leonardo Ferreira Guilhoto, Akshat Kaushal, Paris Perdikaris

机构 * Graduate Group in Applied Mathematics and Computational Science(应用数学与计算科学联合研究生组) Department of Computer and Information Science(计算机与信息科学系) Department of Mechanical Engineering and Applied Mechanics(机械工程与应用力学系)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出混合密度网络作为多模态科学学习中更高效、更稳定的替代方法,实现对科学问题中多模态不确定性的有效建模与解决。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00849 2026-02-03 cs.LG cs.AI cs.NA math.NA 79%

RMFlow: Refined Mean Flow by a Noise-Injection Step for Multimodal Generation

RMFlow:通过噪声注入步骤细化均流以实现多模态生成

Yuhao Huang, Shih-Hsin Wang, Andrea L. Bertozzi, Bao Wang

机构 * Department of Mathematics and Scientific Computing and Imaging (SCI) Institute University of Utah(数学与科学计算及成像学院(SCI)院,犹他大学) Department of Mathematics, UCLA(数学系,加州大学洛杉矶分校)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

AI总结 RMFlow通过引入噪声注入步骤,改进均流模型,实现高效多模态生成,仅需单次功能评估即可达到接近最先进的性能。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00762 2026-02-03 cs.CL cs.HC 79%

WordCraft: Scaffolding the Keyword Method for L2 Vocabulary Learning with Multimodal LLMs

WordCraft: 通过多模态大语言模型 scaffolding 关键词方法用于L2词汇学习

Yuheng Shao, Junjie Xiong, Chaoran Wu, Xiyuan Wang, Ziyu Zhou, Yang Ouyang, Qinyi Tao, Quan Li

机构 * School of Information Science and Technology, ShanghaiTech University(信息科学与技术学院,上海科技大学) School of Creativity and Art, ShanghaiTech University(创意与艺术学院,上海科技大学) Shanghai Fengxian Dai Wen Middle School(上海奉贤戴文中学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

AI总结 WordCraft通过多模态大语言模型辅助L2词汇学习,提升关键词方法的使用效果和学习者参与度。

Comments Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI' 26), April 13--17, 2026, Barcelona, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00107 2026-02-03 cs.CV cs.RO eess.IV 79%

Efficient UAV trajectory prediction: A multi-modal deep diffusion framework

高效无人机轨迹预测:一种多模态深度扩散框架

Yuan Gao, Xinyu Guo, Wenjing Xie, Zifan Wang, Hongwen Yu, Gongyang Li, Shugong Xu

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出一种多模态深度融合框架,通过融合激光雷达和毫米波雷达数据提升无人机轨迹预测精度,实验显示其比基线模型提升40%。

Comments in Chinese language

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21821 2026-01-30 cs.CV 79%

MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods

MMFineReason: 通过以开放数据为中心的方法缩小多模态推理差距

Honglin Lin, Zheng Liu, Yun Zhu, Chonghan Qin, Juekai Lin, Xiaoran Shang, Conghui He, Wentao Zhang, Lijun Wu

机构 * Shanghai Artificial Intelligence Laboratory, OpenDataLab(上海人工智能实验室、OpenDataLab) Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) The University of Hong Kong(香港大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 MMFineReason通过大规模多模态推理数据集和基于难度的过滤策略,提升模型推理能力与参数效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21076 2026-01-30 cs.AI 79%

Multi-modal Imputation for Alzheimer's Disease Classification

多模态缺失数据填补用于阿尔茨海默病分类

Abhijith Shaji, Tamoghna Chattopadhyay, Sophia I. Thomopoulos, Greg Ver Steeg, Paul M. Thompson, Jose-Luis Ambite

机构 * Information Sciences Institute(信息科学研究所) University of Southern California(美国南加州大学) University of California(加州大学) Stevens Neuroimaging and Informatics Institute(史蒂文斯神经影像与信息学研究所)

专题命中 多模态生成 :multi-modal(title);multimodal(abstract);分类 cs.AI

AI总结 本文提出利用条件去噪扩散概率模型填补缺失的DWI扫描,以提高多模态深度学习模型在阿尔茨海默病三类分类中的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18543 2026-01-29 cs.CV 79%

GenAgent: Scaling Text-to-Image Generation via Agentic Multimodal Reasoning

GenAgent:通过代理多模态推理扩展文本到图像生成

Kaixun Jiang, Yuzheng Wang, Junjie Zhou, Pandeng Li, Zhihang Liu, Chen-Wei Xie, Zhaoyu Chen, Yun Zheng, Wenqiang Zhang

机构 * Fudan University(复旦大学) Tongyi Lab(通义实验室) Nanjing University(南京大学) University of Science and Technology of China(中国科学技术大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 GenAgent通过代理多模态推理提升文本到图像生成性能,采用两阶段训练策略实现自主多轮交互与任务自适应推理。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19839 2026-01-28 cs.RO cs.AI cs.HC 79%

HARMONI: Multimodal Personalization of Multi-User Human-Robot Interactions with LLMs

HARMONI: 基于大语言模型的多用户人机交互多模态个性化

Jeanne Malécot, Hamed Rahimi, Jeanne Cattoni, Marie Samson, Mouad Abrini, Mahdi Khoramshahi, Maribel Pino, Mohamed Chetouani

机构 * Institut Curie, Université Paris-Saclay(巴黎-萨克勒大学Curie研究所) Institute of Intelligent Systems and Robotics (ISIR), Sorbonne University(索邦大学智能系统与机器人研究所) Assistance Publique – Hôpitaux de Paris (AP-HP), Université Paris Cité(巴黎公共医院(AP-HP)与巴黎城市大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

AI总结 HARMONI通过多模态个性化框架,利用大语言模型提升多用户人机交互的持续个性化和动态适应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19834 2026-01-28 cs.AI 79%

Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models

视觉生成通过多模态世界模型解锁类人推理

Jialong Wu, Xiaoying Zhang, Hongyi Yuan, Xiangcheng Zhang, Tianhao Huang, Changjing He, Chaoyi Deng, Renrui Zhang, Youbin Wu, Mingsheng Long

机构 * Tsinghua University(清华大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出视觉生成在特定任务中优于纯语言推理,通过构建VisWorld-Eval评估套件验证了多模态世界模型提升类人推理的能力。

Comments Project page: https://thuml.github.io/Reasoning-Visual-World

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04548 2026-01-23 cs.CV 79%

Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model

Skywork UniPic 2.0: 通过在线强化学习构建 Kontext 模型以实现统一多模态模型

Hongyang Wei, Baixin Xu, Hongbo Liu, Size Wu, Jie Liu, Yi Peng, Peiyu Wang, Zexiang Liu, Jingwen He, Yidan Xietian, Chuanxin Tang, Zidong Wang, Yichen Wei, Liang Hu, Boyi Jiang, Wei Li, Ying He, Yang Liu, Xuchen Song, Yangguang Li, Yahui Zhou

机构 * Skywork Multimodality Team(Skywork多模态团队)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 Skywork UniPic 2.0 通过在线强化学习构建 Kontext 模型,实现统一多模态模型,提升图像生成与编辑能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20362 2026-01-22 cs.CV 79%

CRAFT: Continuous Reasoning and Agentic Feedback Tuning for Multimodal Text-to-Image Generation

CRAFT:连续推理与代理反馈调优用于多模态文本到图像生成

V. Kovalev, A. Kuvshinov, A. Buzovkin, D. Pokidov, D. Timonin

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 CRAFT通过显式推理和反馈调优提升多模态生成模型的可靠性,实现可控且高效的文本到图像生成。

Comments 37 pages, 42 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11614 2026-01-21 cs.CV cs.LG q-bio.NC 79%

Multi-modal MRI-Based Alzheimer's Disease Diagnosis with Transformer-based Image Synthesis and Transfer Learning

基于多模态MRI的阿尔茨海默病诊断:基于Transformer的图像合成与迁移学习

Jason Qiu

机构 * Marvin Ridge High School(马文岭高中)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

AI总结 本研究提出基于Transformer的图像合成方法,利用T1w MRI预测FA和MD,提升AD诊断准确率和MCI检测能力。

Comments 19 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06049 2026-01-13 cs.CY cs.AI 79%

The Violation State: Safety State Persistence in a Multimodal Language Model Interface

违规状态:多模态语言模型接口中的安全状态持续

Bentley DeVilling

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

AI总结 研究发现多模态语言模型在面对版权拒绝后,会持续拒绝无关图像生成请求,揭示了会话级安全状态的持续性问题。

Comments 19 pages, 1 figure, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05572 2026-01-12 cs.CV 79%

Towards Generalized Multi-Image Editing for Unified Multimodal Models

面向统一多模态模型的通用多图像编辑

Pengcheng Xu, Peng Tang, Donghao Luo, Xiaobin Hu, Weichu Cui, Qingdong He, Zhennan Chen, Jiangning Zhang, Charles Ling, Boyu Wang

机构 * Western University(西方大学) Tencent YouTu Lab(腾讯优图实验室) Nanjing University(南京大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种面向统一多模态模型的通用多图像编辑框架,通过可学习的潜在分离器和正弦索引编码提升多图像编辑任务中的视觉一致性和泛化能力。

Comments Project page: https://github.com/Pengchengpcx/MIE-UMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04056 2026-01-08 cs.CL 79%

Bridging the Discrete-Continuous Gap: Unified Multimodal Generation via Coupled Manifold Discrete Absorbing Diffusion

弥合离散-连续鸿沟:通过耦合流形离散吸收扩散实现统一多模态生成

Yuanfeng Xu, Yuhao Chen, Liang Lin, Guangrun Wang

机构 * Sun Yat-sen University(中山大学) Guangdong Key Lab of Big Data Analysis & Processing(广东省大数据分析与处理重点实验室) X-Era AI Lab(X-Era AI实验室)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL

AI总结 CoM-DAD通过耦合流形离散吸收扩散框架,实现离散与连续数据的统一多模态生成,解决多模态对齐与训练稳定性问题。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03250 2026-01-07 cs.CV 79%

A Versatile Multimodal Agent for Multimedia Content Generation

多功能多模态代理用于多媒体内容生成

Daoan Zhang, Wenlin Yao, Xiaoyang Wang, Yebowen Hu, Jiebo Luo, Dong Yu

机构 * University of Rochester(罗切斯特大学) Tencent AI Lab, Bellevue(腾讯AI实验室) University of Central Florida(中央佛罗里达大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种多功能多模态代理,通过技能获取理论和两阶段相关策略,实现复杂多媒体内容生成任务的自动化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01363 2026-01-06 cs.AI 79%

A unified multimodal understanding and generation model for cross-disciplinary scientific research

面向跨学科科学研究的统一多模态理解和生成模型

Xiaomeng Yang, Zhiyu Tan, Xiaohui Zhong, Mengping Yang, Qiusheng Huang, Lei Chen, Libo Wu, Hao Li

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

AI总结 FuXi-Uni是一种统一多模态模型,能跨学科领域理解和生成科学数据,通过自然语言和科学数值预测提升跨学科研究效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00051 2026-01-05 cs.CV 79%

TeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World Model

TeleWorld:面向动态多模态合成的4D世界模型

Yabo Chen, Yuanzhi Liang, Jiepeng Wang, Tingxi Chen, Junfei Cheng, Zixiao Gu, Yuyang Huang, Zicheng Jiang, Wei Li, Tian Li, Weichen Li, Zuoxin Li, Guangce Liu, Jialun Liu, Junqi Liu, Haoyuan Wang, Qizhen Weng, Xuan'er Wu, Xunzhi Xiang, Xiaoyan Yang, Xin Zhang, Shiwen Zhang, Junyu Zhou, Chengcheng Zhou, Haibin Huang, Chi Zhang, Xuelong Li

机构 * TeleWorld Team(TeleWorld团队)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 TeleWorld提出了一种实时多模态4D世界建模框架,通过生成-重建-引导范式实现动态场景重建与长期记忆,提升世界模型的交互性和计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏