arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-03-02 至 2026-03-02 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 6 篇

2602.23739 2026-03-02 cs.CV 83%

U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation

U-Mind:一种实时光学交互的统一框架

Xiang Deng, Feng Gao, Yong Zhang, Youxin Pang, Xu Xiaoming, Zhuoliang Kang, Xiaoming Wei, Yebin Liu

机构 * Tsinghua University(清华大学) Meituan(美团)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 U-Mind 提出了一种统一框架,通过实时生成和多模态联合建模,实现高效且连贯的多模态交互,推动智能对话代理的发展。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23366 2026-03-02 cs.HC cs.IR 82%

Doc To The Future: Infomorphs for Interactive, Multimodal Document Transformation and Generation

文档面向未来:用于交互式、多模态文档转换与生成的Infomorphs

Balasaravanan Thoravi Kumaravel

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)

AI总结 Doc To The Future提出Infomorphs,通过模块化、用户可控的AI增强转换,实现交互式多模态文档转换与生成,提升生成式AI在信息工作中的透明度和模块化交互。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.24183 2026-03-02 cs.CV cs.LG 79%

A multimodal slice discovery framework for systematic failure detection and explanation in medical image classification

一种用于医学图像分类中系统性故障检测和解释的多模态切片发现框架

Yixuan Liu, Kanwal K. Bhatia, Ahmed E. Fetit

机构 * Department of Computing, Imperial College London, UK(帝国理工学院 computing 部,英国) Aival, London, UK(Aival,伦敦,英国)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 该研究提出了一种多模态切片发现框架,用于医学图像分类中的系统性故障检测与解释,展示了其在故障发现和解释生成方面的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23711 2026-03-02 cs.CV 70%

Can Unified Generation and Understanding Models Maintain Semantic Equivalence Across Different Output Modalities?

统一的生成与理解模型能否在不同输出模态间保持语义等价性?

Hongbo Jiang, Jie Li, Yunhang Shen, Pingyang Dai, Xing Sun, Haoyu Cao, Liujuan Cao

机构 * Tencent Youtu Lab(腾讯云图实验室) Xiamen University(厦门大学) Computing Lab, Department of Artificial Intelligence, School of Informatics(计算实验室,人工智能系,信息学院)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本研究探讨统一生成与理解模型在不同输出模态间保持语义等价性的能力,发现其在视觉回答任务中存在性能下降问题,原因在于跨模态语义对齐的失效。

Comments Equal contribution by Jie Li

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23969 2026-03-02 cs.MM cs.CV 62%

MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation

MSVBench: 向多镜头视频生成的人机水平评估迈进

Haoyuan Shi, Yunxin Li, Nanhao Deng, Zhenran Xu, Xinyu Chen, Longyue Wang, Baotian Hu, Min Zhang

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Alibaba International Digital Commerce(阿里巴巴国际数字商业)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.MM

AI总结 MSVBench通过引入分层脚本和参考图像,提出混合评估框架,验证了视频生成模型的连贯性和吸引力,并展示了其在多镜头视频生成中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19326 2026-03-02 cs.MA cs.AI 57%

City Editing: Hierarchical Agentic Execution for Dependency-Aware Urban Geospatial Modification

城市编辑:面向依赖意识的层级代理执行

Rui Liu, Steven Jige Quan, Zhong-Ren Peng, Zijun Yao, Han Wang, Zhengzhang Chen, Kunpeng Liu, Yanjie Fu, Dongjie Wang

机构 * University of Kansas, Lawrence, KS, USA(堪萨斯大学) Seoul National University, Seoul, South Korea(首尔国立大学) University of Florida, Gainesville, FL, USA(佛罗里达大学) NEC Laboratories America, Princeton, NJ, USA(NEC美国实验室) Clemson University, Clemson, SC, USA(克莱姆森大学) Arizona State University, Tempe, AZ, USA(亚利桑那州立大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出了一种基于层级代理的框架,用于高效、准确地修改城市规划,通过分层几何意图和迭代验证机制提升空间修改的效率与一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏