arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86539 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3480 篇

2410.04207 2024-10-16 cs.LG stat.ML 67%

Learning on LoRAs: GL-Equivariant Processing of Low-Rank Weight Spaces for Large Finetuned Models

Theo Putterman, Derek Lim, Yoav Gelberg, Stefanie Jegelka, Haggai Maron

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00173 2024-10-02 cs.LG 67%

GaNDLF-Synth: A Framework to Democratize Generative AI for (Bio)Medical Imaging

Sarthak Pati, Szymon Mazurek, Spyridon Bakas

专题命中 文生图 :diffusion(abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08675 2024-07-23 cs.AI 67%

CAD-Prompted Generative Models: A Pathway to Feasible and Novel Engineering Designs

Leah Chong, Jude Rayan, Steven Dow, Ioanna Lykourentzou, Faez Ahmed

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

Comments 11 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.09507 2024-07-18 eess.IV 67%

Can Generative AI Replace Immunofluorescent Staining Processes? A Comparison Study of Synthetically Generated CellPainting Images from Brightfield

Xiaodan Xing, Siofra Murdoch, Chunling Tang, Giorgos Papanastasiou, Jan Cross-Zamirski, Yunzhe Guo, Xianglu Xiao, Carola-Bibiane Schönlieb, Yinhai Wang, Guang Yang

专题命中 文生图 :diffusion(abstract);image synthesis(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08100 2024-07-12 cs.LG math.OC math.PR 67%

Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates

Steffen Dereich, Robin Graeber, Arnulf Jentzen

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

Comments 54 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07672 2024-07-11 cs.HC 67%

StoryDiffusion: How to Support UX Storyboarding With Generative-AI

Zhaohui Liang, Xiaoyu Zhang, Kevin Ma, Zhao Liu, Xipei Ren, Kosa Goucher-Lambert, Can Liu

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.13534 2024-07-08 cs.HC 67%

Prompting AI Art: An Investigation into the Creative Skill of Prompt Engineering

Jonas Oppenlaender, Rhema Linder, Johanna Silvennoinen

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

Comments 42 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.08739 2024-06-14 cs.CY cs.LG 67%

At the edge of a generative cultural precipice

Diego Porres, Alex Gomez-Villa

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

Comments Accepted at the CVPR Fourth Workshop on Ethical Considerations in Creative applications of Computer Vision

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.18385 2024-05-01 cs.HC cs.AI 67%

Equivalence: An analysis of artists' roles with Image Generative AI from Conceptual Art perspective through an interactive installation design practice

Yixuan Li, Dan C. Baciu, Marcos Novak, George Legrady

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00018 2024-04-02 cs.HC cs.AI cs.SI 67%

Can AI Outperform Human Experts in Creating Social Media Creatives?

Eunkyung Park, Raymond K. Wong, Junbum Kwon

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.01673 2024-01-26 eess.AS cs.CL cs.SD 67%

Disentanglement in a GAN for Unconditional Speech Synthesis

Matthew Baas, Herman Kamper

专题命中 文生图 :diffusion(abstract);image synthesis(abstract)

Comments 12 pages, 5 tables, 4 figures. Accepted to IEEE TASLP. arXiv admin note: substantial text overlap with arXiv:2210.05271

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10434 2023-10-24 cs.CL cs.AI cs.LG 67%

Learning the Visualness of Text Using Large Vision-Language Models

Gaurav Verma, Ryan A. Rossi, Christopher Tensmeyer, Jiuxiang Gu, Ani Nenkova

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

Comments Accepted at EMNLP 2023 (Main, long); 9 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.10120 2023-10-18 cs.LG cs.AI 67%

Selective Amnesia: A Continual Learning Approach to Forgetting in Deep Generative Models

Alvin Heng, Harold Soh

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.12915 2023-09-19 cs.HC cs.AI 67%

Language as Reality: A Co-Creative Storytelling Game Experience in 1001 Nights using Generative AI

Yuqian Sun, Zhouyi Li, Ke Fang, Chang Hee Lee, Ali Asadipour

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

Comments The paper was accepted by The 19th AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE 23)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.05940 2023-09-13 cs.CR 67%

Catch You Everything Everywhere: Guarding Textual Inversion via Concept Watermarking

Weitao Feng, Jiyan He, Jie Zhang, Tianwei Zhang, Wenbo Zhou, Weiming Zhang, Nenghai Yu

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.04997 2023-09-12 cs.CY 67%

Gender Bias in Multimodal Models: A Transnational Feminist Approach Considering Geographical Region and Culture

Abhishek Mandal, Suzanne Little, Susan Leavy

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

Comments Selected for publication at the Aequitas 2023: Workshop on Fairness and Bias in AI | co-located with ECAI 2023, Kraków, Poland

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.08961 2023-08-28 cs.CL cs.AI cs.HC 67%

Grimm in Wonderland: Prompt Engineering with Midjourney to Illustrate Fairytales

Martin Ruskov

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

Comments 19th Conference on Information and Research science Connecting to Digital and Library Science, February 23-24, 2023, Bari, Italy

Journal ref in 19th IRCDL, CEUR WS vol. 3365, pp. 180-191 (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.14151 2023-07-27 cs.LG stat.ML 67%

Learning Disentangled Discrete Representations

David Friede, Christian Reimers, Heiner Stuckenschmidt, Mathias Niepert

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.03668 2023-06-02 cs.LG cs.CL 67%

Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery

Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, Tom Goldstein

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

Comments 15 pages, 12 figures, Code is available at https://github.com/YuxinWenRick/hard-prompts-made-easy

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11176 2023-05-25 cs.RO cs.AI 67%

Instruct2Act: Mapping Multi-modality Instructions to Robotic Actions with Large Language Model

Siyuan Huang, Zhengkai Jiang, Hao Dong, Yu Qiao, Peng Gao, Hongsheng Li

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12013 2023-05-23 cs.HC cs.AI cs.CY 67%

Constructing Dreams using Generative AI

Safinah Ali, Daniella DiPaola, Randi Williams, Prerna Ravi, Cynthia Breazeal

专题命中 文生图 :image generation(abstract);text-to-image(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.12253 2023-03-23 cs.HC 67%

The Prompt Artists

Minsuk Chang, Stefania Druga, Alex Fiannaca, Pedro Vergani, Chinmay Kulkarni, Carrie Cai, Michael Terry

专题命中 文生图 :text-to-image(abstract);image editing(abstract)

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.06832 2023-03-17 cs.LG cs.AI 67%

ODIN: On-demand Data Formulation to Mitigate Dataset Lock-in

SP Choi, Jihun Lee, Hyeongseok Ahn, Sanghee Jung, Bumsoo Kang

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

Comments 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03765 2023-02-16 cs.CL cs.AI 67%

Visualize Before You Write: Imagination-Guided Open-Ended Text Generation

Wanrong Zhu, An Yan, Yujie Lu, Wenda Xu, Xin Eric Wang, Miguel Eckstein, William Yang Wang

专题命中 文生图 :text-to-image(abstract);image synthesis(abstract)

Comments EACL 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.03998 2022-10-06 cs.HC 67%

I Learn to Diffuse, or Data Alchemy 101: a Mnemonic Manifesto

Victor Schetinger, Velitchko Filipov, Ignacio Pérez-Messina, Ethan Smith, Rodrigo Oliveira de Oliveira

专题命中 文生图 :text-to-image(abstract);diffusion(abstract)

Comments Submission to IEEE alt.vis 2022. Short paper containing a 4-page comic, and an immense amount of content in a linked miro board

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.04443 2019-09-11 cs.LG stat.ML 67%

Learning Priors for Adversarial Autoencoders

Hui-Po Wang, Wen-Hsiao Peng, Wei-Jan Ko

专题命中 文生图 :text-to-image(abstract);image synthesis(abstract)

Comments Accepted by APSIPA ASC, 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1905.06586 2019-05-17 cs.LG stat.ML 67%

On Conditioning GANs to Hierarchical Ontologies

Hamid Eghbal-zadeh, Lukas Fischer, Thomas Hoch

专题命中 文生图 :image generation(abstract);image synthesis(abstract)

Comments Under review at MLKgraphs2019: http://www.dexa.org/mlkgraphs2019

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.13541 2026-08-14 cs.CV cs.GR 新提交 62%

SCULPT: Subtractive Composition for 3D Part Generation

SCULPT:用于3D部件生成的减法组合方法

Sikuang Li, Chen Yang, Jiemin Fang, Jiazhong Cen, Yuhe Wei, Jichen Pang, Wei Shen, Qi Tian

机构 * Shanghai Jiao Tong University(上海交通大学) Huawei(华为)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 SCULPT是一种3D部件生成框架,通过减法组合方式生成部件,解决了现有方法的边界问题,在PartObjaverse上实现了最优几何性能,还能完成细粒度纹理部件分解。

Comments Project page: https://sculpt-part.github.io/ Code: https://github.com/sculpt-part/SCULPT

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08831 2026-08-04 cs.GR cs.CV cs.RO 版本更新 62%

DiffPhysCam: Differentiable Physics-Based Camera Simulation for Inverse Rendering and Embodied AI

DiffPhysCam:面向逆渲染与具身智能的可微分基于物理的相机仿真

Bo-Hsun Chen, Nevindu M. Batagoda, Dan Negrut

机构 * Simulation-Based Engineering Lab, University of Wisconsin-Madison(模拟基于工程实验室,威斯康星大学麦迪逊分校)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 本文提出DiffPhysCam,一款可微分基于物理的相机模拟器,解决现有虚拟相机局限,支持正向与逆渲染,经实验验证可提升机器人感知性能,还用于自主地面车辆导航的虚拟实验。

Comments 37 pages, 24 figures, and 5 tables. Code of DiffPhysCam-CamCaliExp: https://github.com/DanielYamChen/DiffPhysCam-CamCaliExp Code of DiffPhysCam-NovelViewSynthesis: https://github.com/DanielYamChen/DiffPhysCam-NovelViewSynthesis Data of DiffPhysCam_Data: https://huggingface.co/datasets/DanielYamChen/DiffPhysCam_Data Simulation video: https://youtu.be/gQwSMrdmHJI

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.05290 2026-06-05 cs.CV cs.AI cs.MM 62%

Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

模型是否共享安全表示?面向安全视觉生成的跨模型引导

Tobia Poppi, Silvia Cappelletti, Sara Sarto, Florian Schiffers, Garin Kessler, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) University of Pisa(比萨大学) Amazon Prime Video(亚马逊prime视频)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

AI总结 本文提出首个跨模型安全引导框架,通过源语言模型估计安全方向并迁移至目标生成器,无需目标侧不安全数据即可实现安全控制,且不牺牲生成质量。

Comments Project page: https://aimagelab.github.io/cross-model-safety-representations/

详情

展开后加载摘要…

URL PDF HTML 收藏