arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2412.09656 2024-12-16 cs.CV cs.AI 62%

From Noise to Nuance: Advances in Deep Generative Image Models

Benji Peng, Chia Xin Liang, Ziqian Bi, Ming Liu, Yichao Zhang, Tianyang Wang, Keyu Chen, Xinyuan Song, Pohsun Feng

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06742 2024-12-11 cs.CV cs.AI 62%

ContRail: A Framework for Realistic Railway Image Synthesis using ControlNet

Andrei-Robert Alexandrescu, Razvan-Gabriel Petec, Alexandru Manole, Laura-Silvia Diosan

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 5 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06643 2024-12-10 cs.CV cs.AI 62%

Detecting Facial Image Manipulations with Multi-Layer CNN Models

Alejandro Marco Montejano, Angela Sanchez Perez, Javier Barrachina, David Ortiz-Perez, Manuel Benavent-Lledo, Jose Garcia-Rodriguez

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17572 2024-12-10 cs.CV cs.AI 62%

CityX: Controllable Procedural Content Generation for Unbounded 3D Cities

Shougao Zhang, Mengqi Zhou, Yuxi Wang, Chuanchen Luo, Rongyu Wang, Yiwei Li, Zhaoxiang Zhang, Junran Peng

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05694 2024-12-10 cs.MM cs.GR cs.SD eess.AS 62%

Combining Genre Classification and Harmonic-Percussive Features with Diffusion Models for Music-Video Generation

Leonardo Pina, Yongmin Li

专题命中 多模态生成 :audio-visual(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.13853 2024-12-10 cs.RO cs.AI cs.CV 62%

RealDex: Towards Human-like Grasping for Robotic Dexterous Hand

Yumeng Liu, Yaxun Yang, Youzhuo Wang, Xiaofei Wu, Jiamin Wang, Yichen Yao, Sören Schwertfeger, Sibei Yang, Wenping Wang, Jingyi Yu, Xuming He, Yuexin Ma

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Project Page: https://4dvlab.github.io/RealDex_page/

Journal ref Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.04086 2024-12-09 cs.CV cs.AI 62%

BodyMetric: Evaluating the Realism of Human Bodies in Text-to-Image Generation

Nefeli Andreou, Varsha Vivek, Ying Wang, Alex Vorobiov, Tiffany Deng, Raja Bala, Larry Davis, Betty Mohler Tesch

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00769 2024-12-09 cs.CV cs.AI 62%

GameGen-X: Interactive Open-world Game Video Generation

Haoxuan Che, Xuanhua He, Quande Liu, Cheng Jin, Hao Chen

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Homepage: https://gamegen-x.github.io/ Github: https://github.com/GameGen-X/GameGen-X

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03837 2024-12-06 cs.AI cs.CV 62%

Movie Gen: SWOT Analysis of Meta's Generative AI Foundation Model for Transforming Media Generation, Advertising, and Entertainment Industries

Abul Ehtesham, Saket Kumar, Aditi Singh, Tala Talaei Khoei

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.03011 2024-12-05 cs.CV cs.AI 62%

Human Multi-View Synthesis from a Single-View Model:Transferred Body and Face Representations

Yu Feng, Shunsi Zhang, Jian Shu, Hanfeng Zhao, Guoliang Pang, Chi Zhang, Hao Wang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00053 2024-12-03 cs.LG cs.AI cs.CL 62%

LeMoLE: LLM-Enhanced Mixture of Linear Experts for Time Series Forecasting

Lingzheng Zhang, Lifeng Shen, Yimin Zheng, Shiyuan Piao, Ziyue Li, Fugee Tsung

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19786 2024-12-02 cs.CV cs.CL cs.LG 62%

MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks

Yiming Wu, Wei Ji, Kecheng Zheng, Zicheng Wang, Dong Xu

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Five figures, six tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.02791 2024-12-02 cs.CV cs.AI 62%

Efficient Text-driven Motion Generation via Latent Consistency Training

Mengxian Hu, Minghao Zhu, Xun Zhou, Qingqing Yan, Shu Li, Chengju Liu, Qijun Chen

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13579 2024-11-27 cs.CV cs.AI 62%

LTOS: Layout-controllable Text-Object Synthesis via Adaptive Cross-attention Fusions

Xiaoran Zhao, Tianhao Wu, Yu Lai, Zhiliang Tian, Zhen Huang, Yahui Liu, Zejiang He, Dongsheng Li

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16080 2024-11-26 cs.CV cs.AI cs.GR cs.LG 62%

Boosting 3D Object Generation through PBR Materials

Yitong Wang, Xudong Xu, Li Ma, Haoran Wang, Bo Dai

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted to SIGGRAPH Asia 2024 Conference Papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02256 2024-11-26 cs.CV cs.AI cs.GR 62%

EMDM: Efficient Motion Diffusion Model for Fast and High-Quality Motion Generation

Wenyang Zhou, Zhiyang Dou, Zeyu Cao, Zhouyingcheng Liao, Jingbo Wang, Wenjia Wang, Yuan Liu, Taku Komura, Wenping Wang, Lingjie Liu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments ECCV 2024. Project Page: https://frank-zy-dou.github.io/projects/EMDM/index.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08666 2024-11-19 cs.CV cs.AI 62%

A Survey on Vision Autoregressive Model

Kai Jiang, Jiaxing Huang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments This work will be integrated into another project

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10004 2024-11-18 eess.IV cs.AI cs.CV 62%

EyeDiff: text-to-image diffusion model improves rare eye disease diagnosis

Ruoyu Chen, Weiyi Zhang, Bowen Liu, Xiaolan Chen, Pusheng Xu, Shunming Liu, Mingguang He, Danli Shi

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 28 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08642 2024-11-14 cs.CV cs.AI 62%

Towards More Accurate Fake Detection on Images Generated from Advanced Generative and Neural Rendering Models

Chengdong Dong, Vijayakumar Bhagavatula, Zhenyu Zhou, Ajay Kumar

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 13 pages, 8 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08424 2024-11-14 cs.CV cs.AI 62%

A Heterogeneous Graph Neural Network Fusing Functional and Structural Connectivity for MCI Diagnosis

Feiyu Yin, Yu Lei, Siyuan Dai, Wenwen Zeng, Guoqing Wu, Liang Zhan, Jinhua Yu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01240 2024-11-14 cs.SE cs.CL cs.CV cs.HC 62%

AURORA: Navigating UI Tarpits via Automated Neural Screen Understanding

Safwat Ali Khan, Wenyu Wang, Yiran Ren, Bin Zhu, Jiangfan Shi, Alyssa McGowan, Wing Lam, Kevin Moran

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Published at 17th IEEE International Conference on Software Testing, Verification and Validation (ICST) 2024, 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.03739 2024-11-08 cs.CV cs.AI cs.LG cs.RO 62%

Aligning Text-to-Image Diffusion Models with Reward Backpropagation

Mihir Prabhudesai, Anirudh Goyal, Deepak Pathak, Katerina Fragkiadaki

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

Comments This paper is subsumed by a later paper of ours: arXiv:2407.08737

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02791 2024-11-06 cs.CL cs.AI cs.IR cs.LG cs.NE stat.ML 62%

Language Models and Cycle Consistency for Self-Reflective Machine Translation

Jianqiao Wangni

专题命中 多模态生成 :any-to-any(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00086 2024-11-06 cs.CV cs.AI 62%

ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer

Zhen Han, Zeyinzi Jiang, Yulin Pan, Jingfeng Zhang, Chaojie Mao, Chenwei Xie, Yu Liu, Jingren Zhou

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20494 2024-10-31 cs.CV cs.AI cs.LG 62%

Slight Corruption in Pre-training Data Makes Better Diffusion Models

Hao Chen, Yujin Han, Diganta Misra, Xiang Li, Kai Hu, Difan Zou, Masashi Sugiyama, Jindong Wang, Bhiksha Raj

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

Comments NeurIPS 2024 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.19247 2024-10-30 cs.RO cs.AI cs.CV cs.LG 62%

Non-rigid Relative Placement through 3D Dense Diffusion

Eric Cai, Octavian Donca, Ben Eisner, David Held

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Conference on Robot Learning (CoRL), 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18823 2024-10-30 cs.CV cs.AI 62%

Towards Visual Text Design Transfer Across Languages

Yejin Choi, Jiwan Chung, Sumin Shim, Giyeong Oh, Youngjae Yu

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20220 2024-10-29 cs.RO cs.AI cs.CV cs.LG 62%

Neural Fields in Robotics: A Survey

Muhammad Zubair Irshad, Mauro Comi, Yen-Chen Lin, Nick Heppert, Abhinav Valada, Rares Ambrus, Zsolt Kira, Jonathan Tremblay

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 20 pages, 20 figures. Project Page: https://robonerf.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02845 2024-10-24 cs.SD cs.MM eess.AS 62%

Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model

Tornike Karchkhadze, Mohammad Rasool Izadi, Ke Chen, Gerard Assayag, Shlomo Dubnov

专题命中 多模态生成 :cross-modal(abstract);分类 cs.MM、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.16477 2024-10-22 cs.CV cs.CL 62%

DaLPSR: Leverage Degradation-Aligned Language Prompt for Real-World Image Super-Resolution

Aiwen Jiang, Zhi Wei, Long Peng, Feiqiang Liu, Wenbo Li, Mingwen Wang

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏