arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-23 至 2026-03-23 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 6 篇

2603.19337 2026-03-23 cs.CV cs.AI 84%

Diffusion-Guided Semantic Consistency for Multimodal Heterogeneity

扩散引导的语义一致性用于多模态异质性

Jing Liu, Zhengliang Guo, Yan Wang, Xiaoguang Zhu, Yao Du, Zehua Wang, Victor C. M. Leung

机构 * The University of British Columbia(不列颠哥伦比亚大学) Fudan University(复旦大学) Duke Kunshan University(杜克昆山大学) East China Normal University(华东师范大学) University of California, Davis(加州大学戴维斯分校) China University of Mining and Technology(中国矿业大学) SMBU Shenzhen University(深圳大学)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出SemanticFL框架,利用预训练扩散模型的语义表示,通过跨模态对比学习提升联邦学习在多模态异质数据中的鲁棒性,实验表明其在CIFAR-10等数据集上准确率提升达5.49%。

Comments Accepted by IEEE ICME 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19598 2026-03-23 cs.CV 79%

FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow

FlowScene: 多模态图校正流驱动的风格一致室内场景生成

Zhifei Yang, Guangyao Zhai, Keyang Lu, YuYang Yin, Chao Zhang, Zhen Xiao, Jieyi Long, Nassir Navab, Yikai Wang

机构 * School of Computer Science, Peking University(北京大学计算机科学系) Technical University of Munich(慕尼黑技术大学) Beijing Jiaotong University(北京交通大学) Beijing Digital Native Digital City Research Center(北京数字原生数字城市研究院) Theta Labs, Inc.(Theta Labs公司) Beijing Normal University(北京师范大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 FlowScene通过多模态图校正流模型生成风格一致的室内场景,实现对布局、形状和纹理的精细控制,优于现有基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20192 2026-03-23 cs.CV cs.AI 62%

LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation

LumosX:通过身份与属性的关系实现个性化视频生成

Jiazheng Xing, Fei Du, Hangjie Yuan, Pengwei Liu, Hongbin Xu, Hai Ci, Ruigang Niu, Weihua Chen, Fan Wang, Yong Liu

机构 * Zhejiang University(浙江大学) DAMO Academy, Alibaba Group(阿里云达摩院) Hupan Lab(虎斑实验室) National University of Singapore(新加坡国立大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 LumosX通过结合身份与属性关系,提升多主体个性化视频生成的精细度和一致性,采用数据与模型双改进策略,构建了全面的基准测试平台。

Comments ICLR 2026 Camera Ready Version. Code and Models: https://jiazheng-xing.github.io/lumosx-home/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.20188 2026-03-23 cs.CV 57%

Wildfire Spread Scenarios: Increasing Sample Diversity of Segmentation Diffusion Models with Training-Free Methods

野火蔓延场景:通过无训练方法增加分割扩散模型的样本多样性

Sebastian Gerard, Josephine Sullivan

机构 * KTH Royal Institute of Technology(皇家理工学院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出通过无训练方法提升分割扩散模型样本多样性,验证了粒子引导和SPELL等技术在野火蔓延场景中的有效性,提升了HM IoU指标。

Comments Accepted at NLDL 2026. This version contains small corrections compared to the initial publication, see appendix for details

Journal ref Proceedings of the 7th Northern Lights Deep Learning Conference (NLDL), PMLR, Jan. 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19970 2026-03-23 cs.LG cs.AI 57%

Graph2TS: Structure-Controlled Time Series Generation via Quantile-Graph VAEs

Graph2TS: 通过分位数图变分自编码器实现结构控制的时间序列生成

Shaoshuai Du, Joze M. Rozanec, Andy Pimentel, Ana-Lucia Varbanescu

机构 * Informatics Institute, University of Amsterdam, Amsterdam, The Netherlands(阿姆斯特丹大学信息学院) Computer Architecture for Embedded Systems, University of Twente, Enschede, The Netherlands(特文特大学嵌入式系统计算机架构) Laboratory of Artificial Intelligence, Institute Jozef Stedan, Ljubljana, Slovenia(乔泽夫·斯特达安研究所人工智能实验室)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI

AI总结 本文提出Graph2TS,通过分位数图变分自编码器实现时间序列生成,通过结构与残差分离提升分布保真度和时序对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.19752 2026-03-23 cs.CV 57%

PhysNeXt: Next-Generation Dual-Branch Structured Attention Fusion Network for Remote Photoplethysmography Measurement

PhysNeXt:下一代双分支结构注意力融合网络用于远程光体积脉搏波测记测量

Junzhe Cao, Bo Zhao, Zhiyi Niu, Dan Guo, Yue Sun, Haochen Liang, Yong Xu, Zitong YU

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Great Bay University(大湾区大学) Hefei University of Technology(合肥工业大学) Macao Polytechnic University(澳门理工学院)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

AI总结 PhysNeXt通过融合视频帧和STMap表示,结合时空差分建模单元和跨模态交互模块,提升rPPG信号提取的鲁棒性,实验表明其在挑战性条件下具有更稳定和精细的信号恢复能力。

详情

展开后加载摘要…

URL PDF HTML 收藏