arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-24 至 2026-02-24 共收录 17 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 17 篇

2602.19348 2026-02-24 cs.CV cs.AI 84%

MultiDiffSense: Diffusion-Based Multi-Modal Visuo-Tactile Image Generation Conditioned on Object Shape and Contact Pose

MultiDiffSense: 基于扩散的多模态视觉-触觉图像生成,基于物体形状和接触姿态

Sirine Bhouri, Lan Wei, Jian-Qing Zheng, Dandan Zhang

机构 * Department of Bioengineering, Imperial-X Initiative, Imperial College London(生物工程系、Imperial-X计划、帝国理工学院伦敦分校) CAMS-Oxford Institute, University of Oxford(CAMS-牛津研究所、牛津大学)

专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 MultiDiffSense是一种基于扩散的多模态视觉-触觉图像生成模型,通过双条件化实现可控且物理一致的多模态生成,提升了触觉传感数据集的生成效率和跨模态学习能力。

Comments Accepted by 2026 ICRA

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19497 2026-02-24 cs.CV 83%

MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models

MICON-Bench: 多图像上下文图像生成的基准测试与增强

Mingrui Wu, Hang Liu, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(多媒体可信感知与高效计算重点实验室,中国教育部,厦门大学) Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)

专题命中 多模态生成 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 MICON-Bench通过六个任务评估多图像上下文生成能力,并提出DAR机制提升生成质量与连贯性。

Comments CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19409 2026-02-24 cs.SD 82%

AuditoryHuM: Auditory Scene Label Generation and Clustering using Human-MLLM Collaboration

AuditoryHuM: 利用人机协同生成和聚类听觉场景标签

Henry Zhong, Jörg M. Buchholz, Julian Maclaren, Simon Carlile, Richard F. Lyon

机构 * Australian Hearing Hub, Macquarie University, Sydney, Australia(澳大利亚听力中心、麦觉里大学、悉尼、澳大利亚) Google Research Australia, Sydney, Australia(谷歌澳大利亚研究、悉尼、澳大利亚)

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract)

AI总结 AuditoryHuM通过人机协同方法实现听觉场景标签的自动生成与聚类,提供了一种高效且低成本的标准化分类解决方案,适用于边缘设备部署。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19822 2026-02-24 cs.CV cs.AI 81%

Efficient endometrial carcinoma screening via cross-modal synthesis and gradient distillation

通过跨模态合成与梯度蒸馏实现高效的子宫内膜癌筛查

Dongjing Shan, Yamei Luo, Jiqing Xuan, Lu Huang, Jin Li, Mengchu Yang, Zeyu Chen, Fajin Lv, Yong Tang, Chunxiang Zhang

机构 * School of Medical Information and Engineering, Southwest Medical University(西南医科大学医学信息与工程学院) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) Department of Ultrasound, Affiliated Hospital of Southwest Medical University(西南医科大学附属医院超声科) Department of Functional Examination Unit, Zibo Hospital of Traditional Chinese Medicine(淄博中医药医院功能检查科) Key Laboratory of Medical Electrophysiology, Ministry of Education& Medical Electrophysiological Key Laboratory of Sichuan Province, Institute of Cardiovascular Research, Southwest Medical University(教育部医学电生理重点实验室、四川省医学电生理重点实验室、西南医科大学心血管研究所) Department of Radiology, the First Affiliated Hospital of Chongqing Medical University(重庆医科大学第一附属医院放射科) Department of Cardiology, Affiliated Hospital of Southwest Medical University(西南医科大学附属医院心内科) International Research Center for Complexity Sciences, Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院复杂科学研究中心) Institute of Intelligent Chinese Medicine, Chongqing University of Chinese Medicine(重庆中医药大学智能中药研究院) Basic Medicine Research Innovation Center for Cardiometabolic Diseases, Ministry of Education, Southwest Medical University(教育部心脑血管疾病基础医学研究创新中心、西南医科大学)

专题命中 多模态生成 :cross-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种高效的两阶段深度学习框架,通过跨模态合成与梯度蒸馏技术,在低计算成本下实现高灵敏度和特异度的子宫内膜癌筛查。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09609 2026-02-24 cs.CV 79%

Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing

Tele-Omni: 一种用于视频生成与编辑的统一多模态框架

Jialun Liu, Tian Li, Xiao Cao, Yukuo Ma, Gonghu Shang, Haibin Huang, Chi Zhang, Xiangzhen Chang, Zhiyong Huang, Jiakui Hu, Zuoxin Li, Yuanzhi Liang, Cong Liu, Junqi Liu, Robby T. Tan, Haitong Tang, Qizhen Weng, Yifan Xu, Liying Yang, Xiaoyan Yang, Peng Yu, Shiwen Zhang, Xuelong Li

机构 * TeleAI

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 Tele-Omni是一种统一多模态框架,通过解析文本、图像和参考视频指令,实现视频生成与编辑的灵活控制,提升时间一致性和视觉一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15857 2026-02-24 cs.IR 78%

Large-scale Benchmarks for Multimodal Recommendation with Ducho

基于Ducho的多模态推荐系统大规模基准测试

Matteo Attimonelli, Danilo Danese, Angela Di Fazio, Daniele Malitesta, Claudio Pomo, Tommaso Di Noia

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 本文首次提出基于Ducho的多模态推荐系统大规模基准测试,通过统一实验环境评估多模态特征提取器的性能。

Comments Accepted in Expert Systems with Applications

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18903 2026-02-24 cs.CV cs.HC 74%

SCHEMA for Gemini 3 Pro Image: A Structured Methodology for Controlled AI Image Generation on Google's Native Multimodal Model

Gemini 3 Pro图像的SCHEMA:为Google原生多模态模型的受控AI图像生成的结构化方法

Luca Cazzaniga

机构 * Independent Researcher(独立研究者)

专题命中 多模态生成 :multimodal(title);分类 cs.CV

AI总结 SCHEMA为Google Gemini 3 Pro图像提供结构化提示工程方法,通过三级系统提升AI图像生成的可控性,实现高合规率和跨领域应用。

Comments 24 pages, 8 tables. Based on SCHEMA Method v1.0 (deposited December 11, 2025). Previously published on Zenodo: doi:10.5281/zenodo.18721380

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19163 2026-02-24 cs.CV cs.MM cs.SD 73%

JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation

JavisDiT++:联合音频视频生成的统一建模与优化

Kai Liu, Yanhao Zheng, Kai Wang, Shengqiong Wu, Rongjunchen Zhang, Jiebo Luo, Dimitrios Hatzinakos, Ziwei Liu, Hao Fei, Tat-Seng Chua

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学) University of Toronto(多伦多大学) HiThink Research(HiThink研究) University of Rochester(罗切斯特大学) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM

AI总结 JavisDiT++通过统一建模与优化方法,在联合音频视频生成任务中实现高质量的同步与语义对齐。

Comments Accepted by ICLR 2026. Homepage: https://JavisVerse.github.io/JavisDiT2-page

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20057 2026-02-24 cs.RO cs.AI 57%

AdaWorldPolicy: World-Model-Driven Diffusion Policy with Online Adaptive Learning for Robotic Manipulation

AdaWorldPolicy:基于世界模型的扩散策略与在线自适应学习的机器人操控

Ge Yuan, Qiyuan Qiao, Jing Zhang, Dong Xu

机构 * The University of Hong Kong(香港大学) Beihang University(北航大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI

AI总结 AdaWorldPolicy通过结合世界模型和在线自适应学习,实现机器人操控的高效动态适应。

Comments Homepage: https://AdaWorldPolicy.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19542 2026-02-24 cs.CV 57%

Vinedresser3D: Agentic Text-guided 3D Editing

Vinedresser3D: 基于代理的文本引导3D编辑

Yankuan Chi, Xiang Li, Zixuan Huang, James M. Rehg

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 Vinedresser3D通过多模态大语言模型和潜在空间编辑技术实现高质量文本引导的3D编辑,提升编辑精度与一致性。

Comments CVPR 2026, Project website:https://vinedresser3d.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19193 2026-02-24 cs.RO cs.AI 57%

Visual Prompt Guided Unified Pushing Policy

基于视觉提示的统一推送策略

Hieu Bui, Ziyan Gao, Yuya Hosoda, Joo-Ho Lee

机构 * Graduate School of Information Science and Engineering, Ritsumeikan University, Japan(立命馆大学信息科学与工程研究生院) Japan Advanced Institute of Science and Technology (JAIST)(日本先进科学研究院) College of Information Science and Engineering, Ritsumeikan University, Japan(立命馆大学信息科学与工程学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出一种基于视觉提示的统一推送策略,通过整合轻量级提示机制提升多模态推送动作的生成效率,适用于广泛规划问题,并在桌面清洁任务中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04808 2026-02-24 q-bio.NC cs.AI 57%

Setting up for failure: automatic discovery of the neural mechanisms of cognitive errors

失败的设定:自动发现认知错误的神经机制

Puria Radmard, Paul M. Bays, Máté Lengyel

机构 * Department of Engineering, University of Cambridge(工程系,剑桥大学) Department of Psychology, University of Cambridge(心理学系,剑桥大学) Department of Cognitive Science, Central European University(认知科学系,中央欧亚大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出通过训练RNNs复制行为特征,自动发现认知错误的神经机制,解决了传统方法在数据有限和行为优化上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21896 2026-02-24 cs.AI 57%

GenesisGeo: Technical Report

GenesisGeo:技术报告

Minfeng Zhu, Zi Wang, Sizhe Ji, Zhengtong Du, Shengqiang Tai, Junming Ke, Xiao Deng, Zanlang Yin, Xiuqi Huang, Heyu Wang, Wei Chen

机构 * State Key Lab of CAD&CG, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室) Polytechnic Institute, Zhejiang University(浙江大学多科大学院) Hangzhou Research Institute of AI and Holographic Technology(杭州人工智能与全息技术研究 institutes) Volkswagen Group Innovation(大众集团创新) School of Mathematical Science, Zhejiang University(浙江大学数学科学学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本文提出GenesisGeo-1M数据集及基于多任务学习的几何学习框架,通过大规模合成数据提升模型在几何推理任务中的性能,实现金牌级表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18711 2026-02-24 cs.CV 57%

HIME: Mitigating Object Hallucinations in LVLMs via Hallucination Insensitivity Model Editing

HIME: 通过幻觉不敏感模型编辑缓解LVLMs中的物体幻觉

Ahmed Akl, Abdelwahed Khamis, Ali Cheraghian, Zhe Wang, Sara Khalifa, Kewen Wang

机构 * School of Information and Communication Technology, Griffith University, Australia(信息与通信技术学院,格里菲斯大学) Data61, CSIRO, Australia(Data61,澳大利亚联邦科学与工业研究组织) School of Engineering, Macquarie University, Sydney, Australia(工程学院,麦觉大学) School of Information Systems, Queensland University of Technology, Australia(信息系统学院,昆士兰技术大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 HIME通过分层加权编辑方法有效抑制LVLMs中的物体幻觉,减少61.8%的幻觉问题,无需额外参数或计算开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18451 2026-02-24 cs.CY cs.AI 57%

Developing a Multi-Agent System to Generate Next Generation Science Assessments with Evidence-Centered Design

开发一个多智能体系统以生成下一代科学评估并采用证据中心设计

Yaxuan Yang, Jongchan Park, Yifan Zhou, Xiaoming Zhai

机构 * AI4STEM Education Center, University of Georgia(AI4STEM教育中心,佐治亚大学) Department of Educational Psychology, University of Georgia(教育心理学系,佐治亚大学) School of Computing, University of Georgia(计算学院,佐治亚大学) Department of Mathematics, Science, and Social Studies Education, University of Georgia(数学、科学与社会科学教育系,佐治亚大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 本研究提出将证据中心设计整合到多智能体系统中,以自动生成符合NGSS的评估项目,发现AI生成的项目在包容性方面表现良好,但存在清晰性和多模态设计的局限。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19276 2026-02-24 cs.SE 50%

ComUICoder: Component-based Reusable UI Code Generation for Complex Websites via Semantic Segmentation and Element-wise Feedback

ComUICoder:基于语义分割和元素反馈的组件式可重用UI代码生成用于复杂网站

Jingyu Xiao, Jiantong Qin, Shuoqi Li, Man Ho Lam, Yuxuan Wan, Jen-tse Huang, Yintong Huo, Michael R. Lyu

专题命中 多模态生成 :multimodal(abstract)

AI总结 ComUICoder通过语义分割和元素反馈技术,提升复杂网站中可重用UI代码的生成质量与重用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26308 2026-02-24 cs.RO 50%

Anomaly detection for generic failure monitoring in robotic assembly, screwing and manipulation

通用故障监控中机器人装配、拧螺钉和操作中的异常检测

Niklas Grambow, Lisa-Marie Fenner, Felipe Kempkes, Philip Hotz, Dingyuan Wan, Jörg Krüger, Kevin Haninger

机构 * Department of Automation at Fraunhofer IPK(弗劳恩霍夫研究所自动化部门) Department of Industrial Automation Technology at TU Berlin(柏林技术大学工业自动化技术部门)

专题命中 多模态生成 :multi-modal(abstract)

AI总结 本文提出了一种适用于多种机器人任务的异常检测方法,通过比较不同自编码器方法,验证了其在不同任务和控制策略中的泛化能力,并展示了在布线和拧螺钉任务中高可靠性的检测效果。

Comments 8 pages, 5 figures, 4 tables, the paper has been accepted for publication in the IEEE Robotics and Automation Letters

详情

展开后加载摘要…

URL PDF HTML 收藏