arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2502.10475 2025-04-24 cs.CR cs.AI cs.CV 62%

X-SG$^2$S: Safe and Generalizable Gaussian Splatting with X-dimensional Watermarks

Zihang Cheng, Huiping Zhuang, Chun Li, Xin Meng, Ming Li, Fei Richard Yu, Liqiang Nie

机构 * South China University of Technology(南方科技大学) Shenzhen MSU-BIT University(深圳MSU-BIT大学) Peking University(北京大学) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.14871 2025-04-24 cs.CV cs.AI 62%

BrainVis: Exploring the Bridge between Brain and Visual Signals via Image Reconstruction

Honghao Fu, Zhiqi Shen, Jing Jih Chin, Hao Wang

机构 * Honghao Fu 1, 2(Fu Honghao 1, 2) Zhiqi Shen 2(Shen Zhiqi 2) Jing Jih Chin 2(Chin Jing Jih 2) Hao Wang 1(Wang Hao 1)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14260 2025-04-22 cs.CV cs.CL 62%

Cross-attention for State-based model RWKV-7

Liu Xiao, Li Zhiyuan, Lin Yueyu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00238 2025-04-18 cs.AI cs.CV cs.LG q-bio.NC 62%

Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem

Declan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata, Kia Ghods, Amogh Joshi, Alexander Ku, Steven M. Frankland, Thomas L. Griffiths, Jonathan D. Cohen, Taylor W. Webb

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08033 2025-04-11 cs.CV cs.AI cs.GR 62%

GaussianAnything: Interactive Point Cloud Flow Matching For 3D Object Generation

Yushi Lan, Shangchen Zhou, Zhaoyang Lyu, Fangzhou Hong, Shuai Yang, Bo Dai, Xingang Pan, Chen Change Loy

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments ICLR 2025 project page: https://nirvanalan.github.io/projects/GA/

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07046 2025-04-10 cs.CV cs.CL 62%

A Unified Agentic Framework for Evaluating Conditional Image Generation

Jifang Wang, Xue Yang, Longyue Wang, Zhenran Xu, Yiyu Wang, Yaowei Wang, Weihua Luo, Kaifu Zhang, Baotian Hu, Min Zhang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Work in progress. GitHub: https://github.com/HITsz-TMG/Agentic-CIGEval

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01081 2025-04-10 cs.CV cs.CL eess.IV 62%

ShieldGemma 2: Robust and Tractable Image Content Moderation

Wenjun Zeng, Dana Kurniawan, Ryan Mullins, Yuchi Liu, Tamoghna Saha, Dirichi Ike-Njoku, Jindong Gu, Yiwen Song, Cai Xu, Jingjing Zhou, Aparna Joshi, Shravan Dheep, Mani Malek, Hamid Palangi, Joon Baek, Rick Pereira, Karthik Narasimhan

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17550 2025-04-10 cs.LG cs.MM cs.SD eess.AS 62%

A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation

Masato Ishii, Akio Hayakawa, Takashi Shibuya, Yuki Mitsufuji

专题命中 多模态生成 :cross-modal(abstract);分类 cs.MM、eess.AS

Comments IJCNN 2025. The source code is available: https://github.com/SonyResearch/SVG_baseline

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05800 2025-04-09 cs.CV cs.LG cs.MM 62%

Storybooth: Training-free Multi-Subject Consistency for Improved Visual Storytelling

Jaskirat Singh, Junshen Kevin Chen, Jonas Kohler, Michael Cohen

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00915 2025-04-04 cs.CV cs.AI 62%

Empower Vision Applications with LoRA LMM

Liang Mi, Weijun Wang, Wenming Tu, Qingfeng He, Rui Kong, Xinyu Fang, Yazhu Dong, Yikang Zhang, Yunchun Li, Meng Li, Haipeng Dai, Guihai Chen, Yunxin Liu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments EuroSys'2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22796 2025-04-01 cs.CV cs.AI 62%

DiTFastAttnV2: Head-wise Attention Compression for Multi-Modality Diffusion Transformers

Hanling Zhang, Rundong Su, Zhihang Yuan, Pengtao Chen, Mingzhu Shen Yibo Fan, Shengen Yan, Guohao Dai, Yu Wang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22200 2025-03-31 cs.SD cs.CV eess.AS 62%

Enhance Generation Quality of Flow Matching V2A Model via Multi-Step CoT-Like Guidance and Combined Preference Optimization

Haomin Zhang, Sizhe Shan, Haoyu Wang, Zihao Chen, Xiulong Liu, Chaofan Ding, Xinhan Di

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、eess.AS

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.08814 2025-03-31 cs.AI cs.CL cs.CY 62%

SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector

Kyeongryul Lee, Heehyeon Kim, Joyce Jiyoung Whang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

Comments 6 pages, 2 figures, 1 tables. AI for Public Missions (AIPM) Workshop at the 39th AAAI Conference on Artificial Intelligence (AAAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18627 2025-03-25 cs.CV cs.AI 62%

Dig2DIG: Dig into Diffusion Information Gains for Image Fusion

Bing Cao, Baoshuo Cai, Changqing Zhang, Qinghua Hu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16770 2025-03-25 cs.CL cs.AI cs.LG 62%

Persona-Coded Poly-Encoder: Persona-Guided Multi-Stream Conversational Sentence Scoring

Junfeng Liu, Christopher Symons, Ranga Raju Vatsavai

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL、cs.AI

Comments The 35th IEEE International Conference on Tools with Artificial Intelligence (ICTAI)

Journal ref 2023 IEEE 35th International Conference on Tools with Artificial Intelligence (ICTAI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15855 2025-03-21 cs.CV cs.AI 62%

VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling

Hyojun Go, Byeongjun Park, Hyelin Nam, Byung-Hoon Kim, Hyungjin Chung, Changick Kim

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Project page: https://gohyojun15.github.io/VideoRFSplat/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11989 2025-03-21 cs.CL cs.AI 62%

Applications of Large Language Model Reasoning in Feature Generation

Dharani Chandra

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

Comments I just updated the format of the references in the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15176 2025-03-20 cs.HC cs.CL cs.CV 62%

A Review on Large Language Models for Visual Analytics

Navya Sonal Agarwal, Sanjay Kumar Sonbhadra

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14503 2025-03-19 cs.CV cs.AI cs.LG 62%

The Power of Context: How Multimodality Improves Image Super-Resolution

Kangfu Mei, Hossein Talebi, Mojtaba Ardakani, Vishal M. Patel, Peyman Milanfar, Mauricio Delbracio

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13730 2025-03-19 cs.CV cs.CL 62%

TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark

Forouzan Fallah, Maitreya Patel, Agneet Chatterjee, Vlad I. Morariu, Chitta Baral, Yezhou Yang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21301 2025-03-19 eess.IV cs.AI cs.CV cs.LG 62%

Evaluating the Posterior Sampling Ability of Plug&Play Diffusion Methods in Sparse-View CT

Liam Moroy, Guillaume Bourmaud, Frédéric Champagnat, Jean-François Giovannelli

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12018 2025-03-18 cs.CV cs.AI 62%

Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art

Zhe Jin, Tat-Seng Chua

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11958 2025-03-18 cs.CV cs.AI cs.LG cs.RO 62%

CHOrD: Generation of Collision-Free, House-Scale, and Organized Digital Twins for 3D Indoor Scenes with Controllable Floor Plans and Optimal Layouts

Chong Su, Yingbin Fu, Zheyuan Hu, Jing Yang, Param Hanji, Shaojun Wang, Xuan Zhao, Cengiz Öztireli, Fangcheng Zhong

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Chong Su and Yingbin Fu contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11096 2025-03-17 cs.CV cs.AI cs.HC 62%

Augmenting Image Annotation: A Human-LMM Collaborative Framework for Efficient Object Selection and Label Generation

He Zhang, Xinyi Fu, John M. Carroll

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments This paper will appear at ICLR 2025 Workshop on Bidirectional Human-AI Alignment

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08280 2025-03-12 cs.CV cs.AI 62%

OminiControl2: Efficient Conditioning for Diffusion Transformers

Zhenxiong Tan, Qiaochu Xue, Xingyi Yang, Songhua Liu, Xinchao Wang

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08269 2025-03-12 cs.CV cs.AI 62%

Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial Attacks

Junying Wang, Hongyuan Zhang, Yuan Yuan

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by CVPR-25

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08014 2025-03-12 cs.CV cs.AI 62%

MAGIC: Mastering Physical Adversarial Generation in Context through Collaborative LLM Agents

Yun Xing, Nhat Chung, Jie Zhang, Yue Cao, Ivor Tsang, Yang Liu, Lei Ma, Qing Guo

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19973 2025-03-11 cs.CV cs.AI 62%

Can Large Language Models Unveil the Mysteries? An Exploration of Their Ability to Unlock Information in Complex Scenarios

Chao Wang, Luning Zhang, Zheng Wang, Yang Zhou

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 11pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15232 2025-03-11 cs.CV cs.CL 62%

DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

Run Luo, Yunshui Li, Longze Chen, Wanwei He, Ting-En Lin, Ziqiang Liu, Lei Zhang, Zikai Song, Xiaobo Xia, Tongliang Liu, Min Yang, Binyuan Hui

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 25 pages. arXiv admin note: text overlap with arXiv:2401.10208 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13982 2025-03-06 cs.CV cs.AI 62%

Safety Without Semantic Disruptions: Editing-free Safe Image Generation via Context-preserving Dual Latent Reconstruction

Jordan Vice, Naveed Akhtar, Mubarak Shah, Richard Hartley, Ajmal Mian

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments This research is supported by the NISDRG project #20100007, funded by the Australian Government

详情

展开后加载摘要…

URL PDF HTML 收藏