arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4951 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4951 篇

2107.04724 2021-07-15 q-bio.NC cs.LG eess.IV 78%

Longitudinal Correlation Analysis for Decoding Multi-Modal Brain Development

Qingyu Zhao, Ehsan Adeli, Kilian M. Pohl

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.03993 2021-05-10 physics.med-ph physics.bio-ph physics.chem-ph physics.optics 78%

Multimodal microscopy for characterization of amyloid-${\unicode[Times]{x3B2}}$ plaques biomarkers in animal model of Alzheimer's disease

Renan Cunha, Lucas Lafeta, Emerson A. Fonseca, Alexandre Barbosa, Marco A. Romano-Silva, Rafael Vieira, Ado Jorio, Leandro M. Malard

专题命中 多模态生成 :multimodal(title,abstract)

Comments 17 pages (10 pages article, 7 pages supplementary information), 10 figures

Journal ref Analyst, 2021,146, 2945-2954

详情

展开后加载摘要…

URL PDF HTML 收藏
2011.03389 2020-12-30 physics.flu-dyn physics.ao-ph 78%

Multi-modal excitation to model the Quasi-Biennial Oscillation

Pierre Léard, Daniel Lecoanet, Michael Le Bars

专题命中 多模态生成 :multi-modal(title,abstract)

Comments Main text: 6 pages, 4 figures. Supplemental: 7 pages, 5 figures. To be published in Physical Review Letters

详情

展开后加载摘要…

URL PDF HTML 收藏
2004.12545 2020-08-07 cs.NI eess.SP 78%

Wireless VR/Haptic Open Platform for Multimodal Teleoperation

Tae Hun Jung, Hanju Yoo, Yuna Jin, Chae Eun Rhee, Chan-Byoung Chae

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.03640 2020-04-28 cs.GR cs.LG 78%

Unsupervised multi-modal Styled Content Generation

Omry Sendik, Dani Lischinski, Daniel Cohen-Or

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.10925 2019-08-30 stat.ME 78%

Multimodal Neuroimaging Data Integration and Pathway Analysis

Yi Zhao, Lexin Li, Brian S. Caffo

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.09225 2019-02-26 cs.LG stat.ML 78%

Harmonizing Maximum Likelihood with GANs for Multimodal Conditional Generation

Soochan Lee, Junsoo Ha, Gunhee Kim

专题命中 多模态生成 :multimodal(title,abstract)

Comments Accepted as a conference paper at ICLR 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1902.06273 2019-02-19 cs.RO 78%

"Touching to See" and "Seeing to Feel": Robotic Cross-modal SensoryData Generation for Visual-Tactile Perception

Jet-Tsyn Lee, Danushka Bollegala, Shan Luo

专题命中 多模态生成 :cross-modal(title,abstract)

Comments 7 pages, IEEE International Conference on Robotics and Automation 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.02736 2018-12-03 cs.LG cs.DS math.PR stat.ML 78%

Beyond Log-concavity: Provable Guarantees for Sampling Multi-modal Distributions using Simulated Tempering Langevin Monte Carlo

Rong Ge, Holden Lee, Andrej Risteski

专题命中 多模态生成 :multi-modal(title,abstract)

Comments 53 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1808.07312 2018-08-23 eess.SP 78%

Recovering Hidden Components in Multimodal Data with Composite Diffusion Operators

Tal Shnitzer, Mirela Ben-Chen, Leonidas Guibas, Ronen Talmon, Hau-Tieng Wu

专题命中 多模态生成 :multimodal(title);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.11129 2018-07-13 cs.IT math.IT 78%

MIMO Over-the-Air Computation for High-Mobility Multi-Modal Sensing

Guangxu Zhu, Kaibin Huang

专题命中 多模态生成 :multi-modal(title,abstract)

Comments An extended version of a shorter conference submission

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.00532 2018-03-02 cs.RO 78%

Reconfigurable Manipulator Simulation for Robotics and Multimodal Machine Learning Application: Aaria

Arttu Hautakoski, Mohammad M. Aref, Jouni Mattila

专题命中 多模态生成 :multimodal(title,abstract)

Comments preprint before submission to conference: 2018 IEEE International Conference on Automation Science and Engineering , 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1708.04149 2017-08-15 physics.med-ph physics.optics 78%

High-resolution multimodal flexible coherent Raman endoscope

Alberto Lombardini, Vasyl Mytskaniuk, Siddharth Sivankutty, Esben Ravn Andresen, Xueqin Chen, Jérôme Wenger, Marc Fabert, Nicolas Joly, Frédéric Louradour, Alexandre Kudlinski, Hervé Rigneault

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1701.01137 2017-07-05 gr-qc 78%

An architecture for efficient gravitational wave parameter estimation with multimodal linear surrogate models

Richard O'Shaughnessy, Jonathan Blackman, Scott E. Field

专题命中 多模态生成 :multimodal(title);multi-modal(abstract)

Comments 10 pages, 3 figures, and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.04997 2017-04-18 stat.ML cs.LG 78%

Multimodal Prediction and Personalization of Photo Edits with Deep Generative Models

Ardavan Saeedi, Matthew D. Hoffman, Stephen J. DiVerdi, Asma Ghandeharioun, Matthew J. Johnson, Ryan P. Adams

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1704.00972 2017-04-05 cs.HC 78%

MIS: Multimodal Interaction Services in a cloud perspective

Patrizia Grifoni, Fernando Ferri, Maria Chiara Caschera, Arianna D'Ulizia, Mauro Mazzei

专题命中 多模态生成 :multimodal(title,abstract)

Comments 10 pages, 4 figures

Journal ref Journal of Next Generation Information Technology (JNIT), Volume 5, Issue 4, Pages 1-10, November 2014

详情

展开后加载摘要…

URL PDF HTML 收藏
1702.05441 2017-02-20 cs.NE 78%

Toward Abstraction from Multi-modal Data: Empirical Studies on Multiple Time-scale Recurrent Models

Junpei Zhong, Angelo Cangelosi, Tetsuya Ogata

专题命中 多模态生成 :multi-modal(title,abstract)

Comments Accepted by IJCNN 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1508.03541 2016-07-13 physics.optics 78%

Super-resolved multimodal multiphoton microscopy with spatial frequency-modulated imaging

Jeffrey J. Field, Keith W. Wernsing, Scott R. Domingue, Alyssa M. Allende Motz, Keith F. DeLuca, Jennifer G. DeLuca, Darius Kuciauskas, Dean H. Levi, Jeff A. Squier, Randy A. Bartels

专题命中 多模态生成 :multimodal(title,abstract)

Comments 16 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
1601.06473 2016-01-27 cs.RO 78%

Teaching Robots to Do Object Assembly using Multi-modal 3D Vision

Weiwei Wan, Feng Lu, Zepei Wu, Kensuke Harada

专题命中 多模态生成 :multi-modal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1011.4019 2015-05-20 cond-mat.soft 78%

Three dimensional optical manipulation and structural imaging of soft materials by use of laser tweezers and multimodal nonlinear microscopy

Rahul P. Trivedi, Taewoo Lee, Kris A. Bertness, Ivan I. Smalyukh

专题命中 多模态生成 :multimodal(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1009.2285 2015-05-19 cond-mat.soft cond-mat.mtrl-sci physics.optics 78%

Multimodal nonlinear optical polarizing microscopy of long-range molecular order in liquid crystals

Taewoo Lee, Rahul P. Trivedi, Ivan I. Smalyukh

专题命中 多模态生成 :multimodal(title,abstract)

Comments Total 12 pages, 4 figures, submitted to Optics Letters on August 2010

详情

展开后加载摘要…

URL PDF HTML 收藏
1408.3969 2014-08-19 astro-ph.IM stat.CO 78%

Efficient Exploration of Multi-Modal Posterior Distributions

Yi-Ming Hu, Martin Hendry, Ik Siong Heng

专题命中 多模态生成 :multi-modal(title,abstract)

Comments 6 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
1304.0925 2013-04-04 stat.ME 78%

A new approach to multi-modal diffusions with applications to protein folding

Julie Forman, Michael Sørensen

专题命中 多模态生成 :multi-modal(title,abstract)

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1002.2527 2010-02-26 cs.CR 78%

Secured Cryptographic Key Generation From Multimodal Biometrics Feature Level Fusion Of Fingerprint And Iris

A. Jagadeesan, K. Duraiswamy

专题命中 多模态生成 :multimodal(title,abstract)

Comments Pages IEEE format, International Journal of Computer Science and Information Security, IJCSIS January 2010, ISSN 1947 5500, http://sites.google.com/site/ijcsis/

Journal ref International Journal of Computer Science and Information Security, IJCSIS, Vol. 7, No. 1, pp. 296-305, January 2010, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.01513 2026-06-02 cs.DC cs.AI cs.CL cs.LG 77%

Compliance-Scored Best-of-N Guardrail Orchestration for Multimodal Document Generation in Payments Dispute Defense

基于合规评分的Best-of-N护栏编排用于支付争议防御中的多模态文档生成

Nataraj Agaram Sundar, Tejas Morabia

机构 * eBay Inc.(eBay公司)

专题命中 多模态生成 :multimodal(title,comments);分类 cs.CL、cs.AI

AI总结 提出一种结合多候选生成与合规评分早退机制的护栏编排层,通过并行生成、加权评分和最佳输出选择,在支付争议防御场景中实现高合规率与低延迟。

Comments 8 pages, 7 figures, 4 tables. Preprint. Applied systems paper on compliance-scored guardrail orchestration for multimodal LLM document generation. Contains aggregate operational readouts; not a randomized A/B test

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29025 2026-08-06 cs.CV 版本更新 77%

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

用于一致多参考图像编辑的评估-验证奖励

Yingmao Miao, Pengfei Zhang, Xiaochen Lv, Meng Yu, Lei Sun, Xiangxiang Chu, Chao Shen, Chenhao Lin

机构 * Xi’an Jiaotong University(西安交通大学) Amap, Alibaba(高德,阿里巴巴) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 针对多参考图像编辑的视觉一致性与奖励模型缺失问题,提出多维度评估-验证奖励EVR,实现无需架构改动的现成编辑器RL微调,性能优于Qwen-Image-Edit并达到或超NanoBanana。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27113 2026-07-30 cs.CV 新提交 77%

Veritas++: Value-aware On-Policy Distillation for Perception-Enhanced AIGI Detection

Veritas++:面向感知增强的AIGI检测的值感知在线策略蒸馏

Hao Tan, Jun Lan, Zichang Tan, Ajian Liu, Zijian Yu, Chuanbiao Song, Huijia Zhu, Weiqiang Wang, Jun Wan, Zhen Lei

机构 * School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences (UCAS)(中国科学院大学先进交叉科学学院) Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(中国科学院大学人工智能学院) Ant Group(蚂蚁集团) Sangfor Technologies Inc.(深信服科技股份有限公司)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multi-modal(abstract);分类 cs.CV

AI总结 该研究针对现有AIGI检测模型的感知瓶颈,提出Veritas++框架,通过PoRL与VaOPD机制增强感知能力,提升了检测泛化性与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.11886 2026-07-14 cs.CV 新提交 77%

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

读回:预训练的多模态语言模型是文本到图像生成的零样本奖励模型

Runhui Huang, Qihui Zhang, Zhe Liu, Yu Gao, Jie Wu, Hengshuang Zhao

机构 * The University of Hong Kong(香港大学) ByteDance Seed(字节跳动Seed) Peking University(北京大学)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);image-text(abstract);分类 cs.CV

AI总结 研究提出SpectraReward将预训练多模态语言模型转为图像生成强化学习奖励模型,用图像条件提示对数似然作奖励,还引入Self-SpectraReward形成闭环框架。经广泛实验验证,二者能提升生成性能,表明奖励-策略对齐是关键。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.29655 2026-07-08 cs.CV cs.GR 版本更新 77%

SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation

SuperVoxelGPT: 自适应有序3D令牌化用于自回归形状生成

Yuan Li, Congyi Zhang, Xifeng Gao, Xiaohu Guo

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校) Tencent America(腾讯美国)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 提出SuperVoxelGPT框架,通过自适应且有序的超体素令牌化解决自回归3D生成中序列长度与空间顺序的矛盾,实现高质量、高效率的形状生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22607 2026-07-07 cs.CV 版本更新 77%

Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off

Dress-ED:基于指令的虚拟试穿和试脱编辑

Davide Lobba, Fulvio Sanguigni, Bin Ren, Marcella Cornia, Rita Cucchiara, Nicu Sebe

机构 * University of Modena(摩德纳大学) University of Trento(特伦托大学) University of Pisa(比萨大学)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV

AI总结 本文提出Dress-ED数据集,首次统一了虚拟试穿、试脱和文本引导的服装编辑,包含146k个样本,涵盖外观和结构修改,提供统一的多模态扩散框架作为指令驱动的VTON和VTOFF基线。

Comments Accepted at ECCV 2026. Project page at https://aimagelab.github.io/Dress-ED/

详情

展开后加载摘要…

URL PDF HTML 收藏