arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-10 至 2026-03-10 共收录 151 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 18 篇

2509.21060 2026-03-10 eess.AS 57%

Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models

测量音频对正确性的影响:面向大音频语言模型的音频贡献意识后训练

Haolin He, Xingjian Du, Renhe Sun, Zheqi Dai, Yujia Xiao, Mingru Yang, Jiayi Zhou, Xiquan Li, Zhengxi Liu, Zining Liang, Chunyat Wu, Qianhua He, Tan Lee, Xie Chen, Wei-Long Zheng, Weiqiang Wang, Mark Plumbley, Jian Liu, Qiuqiang Kong

专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS

AI总结 本文提出AudioMCQ数据集和Audio-Contribution Filtering方法,通过弱到强和混合到强的后训练范式,提升大音频语言模型的音频贡献意识和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06854 2026-03-10 cs.SD cs.AI 57%

Are Audio-Language Models Listening? Audio-Specialist Heads for Adaptive Audio Steering

音频-语言模型在倾听吗?音频专家头用于自适应音频引导

Neta Glazer, Lenny Aharon, Ethan Fetaya

机构 * Bar-Ilan University(巴伊兰大学) Columbia University(哥伦比亚大学)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文通过识别音频专家注意力头,提出一种方法来增强音频信息在音频-语言模型中的作用,从而提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08571 2026-03-10 cs.HC cs.IR cs.SD 50%

LoopLens: Supporting Search as Creation in Loop-Based Music Composition

LoopLens:在基于循环的音乐创作中支持搜索作为创作

Sheng Long, Atsuya Kobayashi, Kei Tateno

机构 * Northwestern University(西北大学) Sony Group Corporation(索尼集团)

专题命中 音频语音多模态 :multimodal(abstract)

AI总结 LoopLens通过可视化音频搜索结果,支持在基于循环的音乐创作中进行创意搜索与拼接,揭示了不同专业背景用户在探索与利用之间的行为差异。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 视频多模态 10 篇

2409.03597 2026-03-10 cs.SD cs.AI eess.AS 81%

Multimodal Laryngoscopic Video Analysis for Assisted Diagnosis of Vocal Fold Paralysis

多模态喉镜视频分析用于辅助声带麻痹诊断

Yucong Zhang, Xin Zou, Jinshan Yang, Wenjun Chen, Juan Liu, Faya Liang, Ming Li

机构 * School of Computer Science(计算机科学学院) Suzhou Municipal Key Laboratory of Multimodal Intelligent Systems(多模态智能系统苏州市重点实验室) Sun Yat-sen Memorial Hospital of Sun Yat-sen University(中山大学孙逸仙纪念医院) School of Artificial Intelligent(人工智能学院) Department of Otorhinolaryngology(耳鼻喉科部门)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI、eess.AS

AI总结 本文提出多模态喉镜视频分析系统,结合音频和视频数据自动提取关键片段,用于辅助诊断声带麻痹。

Comments Accepted by Computer Speech and Language

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11954 2026-03-10 cs.AI 79%

UniCast: A Unified Framework for Instance-Conditioned Multimodal Time-Series Forecasting

UniCast: 一个实例条件多模态时间序列预测的统一框架

Sehyuk Park, Soyeon Caren Han, Eduard Hovy

机构 * The Universito of Melbourne(墨尔本大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 UniCast通过实例条件提示和动态模态路由,实现多模态时间序列预测的统一框架,有效提升预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22295 2026-03-10 cs.LG 78%

Aurora: Towards Universal Generative Multimodal Time Series Forecasting

Aurora:迈向通用生成多模态时间序列预测

Xingjian Wu, Jianxin Jin, Wanghui Qiu, Peng Chen, Yang Shu, Bin Yang, Chenjuan Guo

机构 * East China Normal University(东华师范大学)

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 Aurora通过多模态输入和零样本推理,提升跨领域时间序列预测的泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05437 2026-03-10 cs.CV cs.AI 62%

SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning

SAIL:基于相似性感知引导和跨字幕增强学习的弱监督密集视频字幕生成

Ye-Chan Kim, SeungJu Cha, Si-Woo Kim, Minju Jeon, Hyungee Kim, Dong-Jin Kim

机构 * Hanyang University(翰阳大学)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 SAIL通过跨模态对齐和大语言模型增强策略,提升弱监督密集视频字幕生成的准确性和语义感知能力。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07660 2026-03-10 cs.CV 57%

Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence

Holi-Spatial: 将视频流转化为整体3D空间智能

Yuanyuan Gao, Hao Li, Yifei Liu, Xinhao Ji, Yuning Gong, Yuanjun Liao, Fangfu Liu, Manyuan Zhang, Yuchen Yang, Dan Xu, Xue Yang, Huaxi Huang, Hongjie Zhang, Ziwei Liu, Xiao Sun, Dingwen Zhang, Zhihang Zhong

机构 * Shanghai AI Lab(上海人工智能实验室) Northwestern Polytechnical University(西北工业大学) Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) Nanyang Technological University(南洋理工大学) Beihang University(北京航空航天大学) Sichuan University(四川大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学) Fudan University(复旦大学) Hong Kong University of Science and Technology(香港科技大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 Holi-Spatial通过自动化方法构建大规模多模态3D空间数据集,提升空间推理任务的模型性能。

Comments project page: https://visionary-laboratory.github.io/holi-spatial/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.08942 2026-03-10 cs.LG cs.AI 57%

Language in the Flow of Time: Time-Series-Paired Texts Weaved into a Unified Temporal Narrative

时间之流中的语言:将时间序列配对文本编织成统一的时间叙事

Zihao Li, Xiao Lin, Zhining Liu, Jiaru Zou, Ziwei Wu, Lecheng Zheng, Dongqi Fu, Yada Zhu, Hendrik Hamann, Hanghang Tong, Jingrui He

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Meta IBM Research(IBM研究院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出TaTS框架,将时间序列配对文本作为辅助变量,提升多模态时间序列预测性能。

Comments ICLR 2026, 47 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06732 2026-03-10 cs.CV 57%

HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos

HERO: 分层嵌入-细化用于视频中开放词汇时间句接地

Tingting Han, Xinsong Tao, Yufei Yin, Min Tan, Sicheng Zhao, Zhou Yu

机构 * Laboratory of Complex Systems Modeling and Simulation, School of Computer Science and Technology, Hangzhou Dianzi University(电子科技大学复杂系统建模与仿真实验室,计算机科学与技术学院) Zhejiang Key Laboratory of Space Information Sensing and Transmission, Hangzhou Dianzi University(浙江省空间信息感知与传输重点实验室,电子科技大学) Department of Psychological and Cognitive Sciences, Tsinghua University(清华大学心理与认知科学系)

专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV

AI总结 HERO提出一种分层嵌入-细化框架,用于视频中开放词汇时间句接地,通过多级语义建模和跨模态细化提升视频与语言的对齐能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08060 2026-03-10 cs.HC 50%

CinemaWorld: Generative Augmented Reality with LLMs and 3D Scene Generation for Movie Augmentation

CinemaWorld: 基于 LLM 和 3D 场景生成的生成增强现实用于电影增强

Keiichi Ihara, DaeHo Lee, Manato Abe, Hye-Young Jo, Ryo Suzuki

专题命中 视频多模态 :multimodal(abstract)

AI总结 CinemaWorld 利用 LLM 和 3D 场景生成技术,通过增强现实将电影内容与现实环境结合,提升观影沉浸感和娱乐性。

Comments 13 pages, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07279 2026-03-10 q-bio.QM 50%

Learning When to Look: On-Demand Keypoint-Video Fusion for Animal Behavior Analysis

学习何时观察:按需关键点-视频融合用于动物行为分析

Weihan Li, Jingyang Ke, Yule Wang, Chengrui Li, Anqi Wu

专题命中 视频多模态 :multimodal(abstract)

AI总结 LookAgain通过按需视觉定位结合关键点与视频数据,实现高效且低计算成本的动物行为分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00919 2026-03-10 cs.RO 50%

Green-VLA: Staged Vision-Language-Action Model for Generalist Robots

Green-VLA: 通用机器人中的分阶段视觉-语言-动作模型

I. Apanasevich, M. Artemyev, R. Babakyan, P. Fedotova, D. Grankin, E. Kupryashin, A. Misailidi, D. Nerus, A. Nutalapati, G. Sidorov, I. Efremov, M. Gerasyov, D. Pikurov, Y. Senchenko, S. Davidenko, D. Kulikov, M. Sultankin, K. Askarbek, O. Shamanin, D. Statovoy, E. Zalyaev, I. Zorin, A. Letkin, E. Rusakov, A. Silchenko, V. Vorobyov, S. Sobolnikov, A. Postnikov

专题命中 视频多模态 :multimodal(abstract)

AI总结 Green-VLA是一种面向真实环境的分阶段视觉-语言-动作模型,通过多阶段训练和强化学习对齐,提升机器人在多样化形态上的泛化能力和性能。

Comments 22 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 跨模态检索 14 篇

2509.11165 2026-03-10 cs.CV 83%

Traffic-MLLM: Curiosity-Regularized Supervised Learning for Traffic Scenario Case-Based Reasoning

Traffic-MLLM: 基于好奇心正则化的交通场景案例推理监督学习

Waikit Xiu, Qiang Lu, Bingchen Liu, Chen Sun, Xiying Li

机构 * The University of Hong Kong(香港大学) Sun Yat-sen University(中山大学) University of Bristol(布里斯托大学)

专题命中 跨模态检索 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 Traffic-MLLM通过好奇心驱动的结构化案例学习,提升多模态交通场景推理的鲁棒性和适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08202 2026-03-10 cs.CV cs.AI 81%

MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data

MM-TS: 多模态温度和边距调度用于长尾数据的对比学习

Siarhei Sheludzko, Dhimitrios Duka, Bernt Schiele, Hilde Kuehne, Anna Kukleva

机构 * University of Bonn(波恩大学) MPI for Informatics, SIC(信息研究所) Tuebingen AI Center/University of Tuebingen(图宾根人工智能中心/图宾根大学) MIT-IBM Watson AI Lab(麻省理工-IBM Watson人工智能实验室)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 MM-TS通过动态温度和边距调度提升多模态对比学习在长尾数据上的性能,实现InfoNCE和最大边距方法的统一。

Comments 18 pages, 11 figures. Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07997 2026-03-10 cs.AI 79%

CMMR-VLN: Vision-and-Language Navigation via Continual Multimodal Memory Retrieval

CMMR-VLN:通过持续多模态记忆检索实现视觉与语言导航

Haozhou Li, Xiangyu Dong, Huiyan Jiang, Yaoming Zhou, Xiaoguang Ma

机构 * Foshan Graduate School of Innovation at Northeastern University(东北大学创新研究生院) Faculty of Robot Science and Engineering at Northeastern University(东北大学机器人科学与工程学院) College of Software at Northeastern University(东北大学软件学院) School of Aeronautic Science and Engineering at Beihang University(北航航空科学与工程学院)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 CMMR-VLN通过引入持续多模态记忆检索机制,提升视觉与语言导航任务中对先前经验的选择性利用能力,显著提高导航成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07039 2026-03-10 cs.AI 79%

Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding

具有4D空间-时间嵌入的自监督多模态世界模型

Lance Legel, Qin Huang, Brandon Voelker, Daniel Neamati, Patrick Alan Johnson, Favyen Bastani, Jeff Rose, James Ryan Hennessy, Robert Guralnick, Douglas Soltis, Pamela Soltis, Shaowen Wang

机构 * Ecological Intelligence Lab(生态智能实验室) School of Complex Adaptive Systems(复杂适应系统学院) University of Houston(休斯顿大学) Geosensing Systems Engineering & Sciences Lab(传感系统工程与科学实验室) Stanford University(斯坦福大学) Allen Institute for Artificial Intelligence(人工智能研究院) Spatial Intelligence Lab(空间智能实验室) Department of Computer Science(计算机科学系) Georgia Institute of Technology(佐治亚理工学院) Florida Museum of Natural History(佛罗里达自然历史博物馆) University of Florida(佛罗里达大学) NSF Institute for Geospatial Understanding(国家科学基金会地理理解研究所) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.AI

AI总结 DeepEarth通过4D空间-时间嵌入实现自监督多模态世界模型,在生态预测中取得最佳性能。

Comments 8 pages, 5 figures, 1 table. Presented at 2026 World Modeling Workshop, Mila Quebec

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06982 2026-03-10 cs.CV cs.IR 79%

Optimizing Multi-Modal Models for Image-Based Shape Retrieval: The Role of Pre-Alignment and Hard Contrastive Learning

优化多模态模型用于基于图像的形状检索:预对齐和硬对比学习的作用

Paul Julius Kühn, Cedric Spengler, Michael Weinmann, Arjan Kuijper, Saptarshi Neil Sinha

机构 * Fraunhofer IGD(弗劳恩霍夫图像研究中心) Delft University of Technology(代尔夫特理工大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出通过预对齐和硬对比学习优化多模态模型,提升基于图像的形状检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06704 2026-03-10 cs.CV cs.LG 70%

On the Generalization Capacities of MLLMs for Spatial Intelligence

关于多模态大语言模型在空间智能中的泛化能力

Gongjie Zhang, Wenhao Li, Quanhao Qian, Jiuniu Wang, Deli Zhao, Shijian Lu, Ran Xu

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) HuPan Lab(虎朋实验室) Nanyang Technological University(南洋理工大学)

专题命中 跨模态检索 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出Camera-Aware MLLM框架,通过注入相机内参、数据增强和蒸馏几何先验,提升多模态大语言模型在空间任务中的泛化能力与鲁棒性。

Comments ICLR 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21194 2026-03-10 cs.CV cs.AI 62%

BotaCLIP: Contrastive Learning for Botany-Aware Representation of Earth Observation Data

BotaCLIP:基于对比学习的植物学感知地球观测数据表示

Selene Cerna, Sara Si-Moussi, Wilfried Thuiller, Hadrien Hendrikx, Vincent Miele

机构 * univ-grenoble-alpes(格勒诺布尔阿尔卑斯大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 BotaCLIP通过对比学习将植物学知识注入地球观测数据表示,提升生物多样性建模任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.11148 2026-03-10 cs.CL cs.AI 62%

Improving the Efficiency of Visually Augmented Language Models

提升视觉增强语言模型的效率

Paula Ontalvilla, Aitor Ormazabal, Gorka Azkune

机构 * HiTZ Center - Ixa, University of the Basque Country (UPV/EHU)(巴斯克大学HiTZ中心 - Ixa,巴斯克大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出BLIND-VALM模型,通过使用CLIP获取的视觉 grounding 文本表示,提升视觉增强语言模型的效率和性能。

Comments COLING 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07497 2026-03-10 cs.CV 57%

AMR-CCR: Anchored Modular Retrieval for Continual Chinese Character Recognition

AMR-CCR: 以锚定模块检索实现连续中文字符识别

Yuchuan Wu, Yinglian Zhu, Haiyang Yu, Ke Niu, Bin Li, Xiangyang Xue

机构 * Fudan University(复旦大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 AMR-CCR通过锚定模块检索框架实现连续中文字符识别,解决持续类增长和脚本多样性问题,提出轻量级脚本注入模块和多原型字典以提升识别性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07493 2026-03-10 cs.CV 57%

RayD3D: Distilling Depth Knowledge Along the Ray for Robust Multi-View 3D Object Detection

RayD3D: 沿射线蒸馏深度知识以实现鲁棒的多视角3D目标检测

Rui Ding, Zhaonian Kuang, Zongwei Zhou, Meng Yang, Xinhu Zheng, Gang Hua

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 RayD3D通过沿射线蒸馏深度知识,提升多视角3D目标检测的鲁棒性,无需增加计算成本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19112 2026-03-10 cs.CV 57%

Universal 3D Shape Matching via Coarse-to-Fine Language Guidance

通过粗到细的语言引导实现通用3D形状匹配

Qinfeng Xiao, Guofeng Mei, Bo Yang, Liying Zhang, Jian Zhang, Kit-lun Yick

机构 * Hong Kong Polytechnic University, HK SAR(香港理工大学) Fondazione Bruno Kessler, Italy(布鲁诺·凯斯勒基金会) University of Technology Sydney, Australia(悉尼科技大学)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 UniMatch通过粗到细的语言引导方法,实现跨类别非等距形状的通用3D匹配。

Comments Accepted by CVPR 2026

Journal ref CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06722 2026-03-10 cs.LG cs.AI 57%

ProtAlign: Contrastive learning paradigm for Sequence and structure alignment

ProtAlign: 用于序列和结构对齐的对比学习范式

Aditya Ranganath, Hasin Us Sami, Kowshik Thopalli, Bhavya Kailkhura, Wesam Sakla

机构 * Center for Applied Scientific Computing, Lawrence Livermore National Laboratory(应用科学计算中心,劳伦斯利弗莫尔国家实验室)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.AI

AI总结 ProtAlign通过对比学习范式,统一了蛋白质序列与结构的表示,提升跨模态检索和下游预测性能。

Comments 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07611 2026-03-10 cs.HC 50%

Beyond Semantic Similarity: Open Challenges for Embedding-Based Creative Process Analysis Across AI Design Tools

超越语义相似性:基于嵌入的AI设计工具创意过程分析的开放挑战

Seung Won Lee, Semin Jin, Kyung Hoon Hyun

专题命中 跨模态检索 :multimodal(abstract)

AI总结 本文探讨了基于嵌入的AI设计工具创意过程分析中超越语义相似性的开放挑战,提出通过上下文感知干预提升分析敏感性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06927 2026-03-10 cs.RO 50%

A Contrastive Fewshot RGBD Traversability Segmentation Framework for Indoor Robotic Navigation

一种用于室内外机器人导航的对比式少样本RGBD可 traversability 分割框架

Qiyuan An, Tuan Dang, Fillia Makedon

机构 * Department of Computer Science and Engineering, University of Texas at Arlington(计算机科学与工程系,德克萨斯大学阿灵顿分校) Uber(优步) Cognitive Robotics Lab, Department of Electrical Engineering and Computer Science, University of Arkansas(认知机器人实验室,电气工程与计算机科学系,阿肯色大学)

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 本文提出了一种基于对比学习的少样本RGBD分割框架,利用负例原型和稀疏深度信息提升室内机器人导航的可 traversability 分割性能。

Journal ref IEEE International Conference on Robotics & Automation 2026 (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 多模态生成 16 篇

2603.08069 2026-03-10 cs.CV 83%

Synthetic Defect Image Generation for Power Line Insulator Inspection Using Multimodal Large Language Models

利用多模态大语言模型生成合成缺陷图像用于电力线路绝缘子检测

Xuesong Wang, Caisheng Wang

机构 * Department of Electrical and Computer Engineering, Wayne State University(电气与计算机工程系,韦恩州立大学)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 利用多模态大语言模型生成合成缺陷图像,提升电力线路绝缘子缺陷识别的准确率和数据效率。

Comments Submitted to Engineering Applications of Artificial Intelligence, Feb. 16, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07540 2026-03-10 cs.CV cs.AI 81%

How Long Can Unified Multimodal Models Generate Images Reliably? Taming Long-Horizon Interleaved Image Generation via Context Curation

统一多模态模型能可靠生成图像多长?通过上下文筛选驯服长周期交错图像生成

Haoyu Chen, Qing Liu, Yuqian Zhou, He Zhang, Zhaowen Wang, Mengwei Ren, Jingjing Ren, Xiang Wang, Zhe Lin, Lei Zhu

机构 * The Hong Kong University of Science(香港科学与技术大学) Adobe Research(Adobe研究) Huazhong University of Science(华中科技大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究提出UniLongGen,通过动态筛选模型记忆以提升长周期图像生成的稳定性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24099 2026-03-10 cs.CV 79%

Unified Multi-Modal Interactive & Reactive 3D Motion Generation via Rectified Flow

通过修正流实现统一的多模态交互与反应3D运动生成

Prerit Gupta, Shourya Verma, Ananth Grama, Aniket Bera

机构 * Department of Computer Science Purdue University(计算机科学系 哈佛大学)

专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV

AI总结 DualFlow通过修正流实现多模态双人运动生成,提升运动质量与节奏同步性。

Comments Under review at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏