arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 86539 信号源:cs.CV, cs.GR, cs.MM

1. 文生图 3480 篇

2605.26062 2026-05-26 cs.GR cs.CV 62%

Look Both Ways Before You Cross: Lifting Cross Fields From 2D Visual Priors

过马路前左右看:从2D视觉先验中提取交叉场

Dale Decatur, Jacob Serfaty, Oded Stein, Amir Vaxman, Rana Hanocka

机构 * University of Chicago(芝加哥大学) University of Southern California(南加州大学) University of Edinburgh(爱丁堡大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 提出CrossLift方法,利用文本到图像先验从2D图像中提取方向信号,通过两次平滑插值将其反投影到网格表面,生成语义对齐的交叉场和四边形网格。

Comments Project page at: https://crosslift.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.24398 2026-05-26 cs.CV cs.AI cs.GR 62%

VectorArk: Learning Practical Image Vectorization with Rounded Polygon Representation

VectorArk: 学习基于圆角多边形表示的实际图像矢量化

Tarun Gehlaut, Difan Liu, Charu Bansal, Krutik Malani, Souymodip Chakraborty, Ankit Phogat, Matthew Fisher, Vineet Batra

机构 * Adobe

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

AI总结 提出VectorArk模型,采用圆角多边形表示和退化模型,实现鲁棒且实用的图像矢量化,在多个数据集上取得优越的几何完整性和伪影抑制效果。

Comments CVPR 2026. Project page: https://vectorark.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03558 2026-04-30 cs.CV cs.AI cs.MM 62%

ELIQ: A Label-Free Framework for Quality Assessment of Evolving AI-Generated Images

ELIQ:一种无需标签的进化AI生成图像质量评估框架

Xinyue Li, Zhiming Xu, Min Tang, Zhaolin Cai, Sijing Wu, Xiongkuo Min, Yitong Chen, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Xi'an Jiaotong University(西安交通大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

AI总结 ELIQ通过自动构建正负样本对,利用预训练多模态模型提升质量评估,实现无需人工标注的高质量评估,优于现有无标签方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22783 2026-04-07 cs.IR cs.CV cs.LG cs.MM cs.SD 62%

Compact Hypercube Embeddings for Fast Text-based Wildlife Observation Retrieval

紧凑超立方体嵌入用于快速基于文本的野生动物观察检索

Ilyass Moummad, Marius Miron, David Robinson, Kawtar Zaher, Hervé Goëau, Olivier Pietquin, Pierre Bonnet, Emmanuel Chemla, Matthieu Geist, Alexis Joly

机构 * Inria, LIRMM, UM(法国国家信息与自动化研究所,蒙彼利埃计算机科学实验室,蒙彼利埃大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

AI总结 本文提出紧凑超立方体嵌入框架,通过二进制表示实现高效文本检索,提升大规模野生动物图像和音频数据库的检索效率与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12459 2026-03-17 cs.CV cs.GR 62%

From Particles to Fields: Reframing Photon Mapping with Continuous Gaussian Photon Fields

从粒子到场:用连续高斯光场重新框架光子映射

Jiachen Tao, Benjamin Planche, Van Nguyen Nguyen, Junyi Wu, Yuchun Liu, Haoxuan Wang, Zhongpai Gao, Gengyu Zhang, Meng Zheng, Feiran Wang, Anwesa Choudhuri, Zhenghao Zhao, Weitai Kang, Terrence Chen, Yan Yan, Ziyan Wu

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 本文提出高斯光场,通过学习表示将光子分布编码为各向异性3D高斯体,提升多视角渲染效率,实现光子级精度与神经场景表示的结合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17623 2026-03-04 cs.MM cs.CV 62%

Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?

合成感知:生成图像能否解锁潜在的视觉先验以用于以文本为中心的推理?

Yuesheng Huang, Peng Zhang, Xiaoxin Wu, Riliang Liu, Jiaqi Liang

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

AI总结 本文探讨生成图像能否解锁潜在视觉先验以提升文本中心推理,通过多模态融合架构和提示工程策略实现性能提升。

Comments Accepted as a poster at the International Conference on Machine Learning (ICML 2025) NewInML Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19750 2026-01-29 cs.MM cs.CV cs.IR 62%

Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues

在电子商务目录中评估多模态大语言模型用于缺失模态补全

Junchen Fu, Wenhao Deng, Kaiwen Zheng, Ioannis Arapakis, Yu Ye, Yongxin Ni, Joemon M. Jose, Xuri Ge

机构 * University of Glasgow(格拉斯哥大学) Telefónica Scientific Research(电信科研机构) National University of Singapore(新加坡国立大学) Shandong University(山东大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

AI总结 本文研究了多模态大语言模型在电子商务目录中补全缺失模态的能力,通过MMPCBench基准测试发现MLLMs在细粒度对齐上存在不足,并探索了GRPO方法以提升补全效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03052 2025-12-04 cs.GR cs.CV 62%

LATTICE: Democratize High-Fidelity 3D Generation at Scale

LATTICE:大规模实现高保真3D生成

Zeqiang Lai, Yunfei Zhao, Zibo Zhao, Haolin Liu, Qingxiang Lin, Jingwei Huang, Chunchao Guo, Xiangyu Yue

机构 * MMLab, CUHK(CUHK人工智能实验室) Tencent Hunyuan(腾讯文言)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

AI总结 LATTICE通过VoxSet半结构化表示实现高效高保真3D生成,支持任意分辨率解码和灵活推理,达到最先进的性能。

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05229 2025-10-29 cs.CV cs.MM 62%

Does CLIP perceive art the same way we do?

Andrea Asperti, Leonardo Dessì, Maria Chiara Tonetti, Nico Wu

机构 * Dept. of Informatics (DISI) University of Bologna(信息学院(DISI)博洛尼亚大学)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.MM

Journal ref Proceedings of IEEE International Conference on Content-Based Multimedia Indexing (IEEE CBMI 2025), Dublin, Ireland, 22-24 October 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26055 2025-10-01 cs.GR cs.CV cs.LG 62%

GaussEdit: Adaptive 3D Scene Editing with Text and Image Prompts

Zhenyu Shu, Junlong Yu, Kai Chao, Shiqing Xin, Ligang Liu

机构 * School of Computer and Data Engineering, NingboTech University(计算机与数据工程学院,宁波科技学院) Ningbo Institute, Zhejiang University(浙江大学宁波学院) School of Software Technology, Zhejiang University(软件技术学院,浙江大学) School of Big Data and Artificial Intelligence Management, Xi’an Jiaotong University(大数据与人工智能管理学院,西安交通大学) School of Computer Science and Technology, ShanDong University(计算机科学与技术学院,山东大学) Graphics & Geometric Computing Laboratory, School of Mathematical Sciences, University of Science and Technology of China(图形与几何计算实验室,数学科学学院,中国科学技术大学)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Journal ref IEEE Transactions on Visualization and Computer Graphics. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15773 2025-08-22 cs.CV cs.GR cs.LG 62%

Scaling Group Inference for Diverse and High-Quality Generation

Gaurav Parmar, Or Patashnik, Daniil Ostashev, Kuan-Chieh Wang, Kfir Aberman, Srinivasa Narasimhan, Jun-Yan Zhu

机构 * Carnegie Mellon University(卡内基梅隆大学) Snap Research(Snap研究公司) Tel Aviv University(特拉维夫大学)

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

Comments Project website: https://www.cs.cmu.edu/~group-inference, GitHub: https://github.com/GaParmar/group-inference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14624 2025-07-22 cs.GR cs.CV 62%

Real-Time Scene Reconstruction using Light Field Probes

Yaru Liu, Derek Nowrouzezahri, Morgan Mcguire

机构 * University of Cambridge(剑桥大学) McGill University(麦吉尔大学)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20875 2025-06-27 cs.GR cs.CV 62%

3DGH: 3D Head Generation with Composable Hair and Face

Chengan He, Junxuan Li, Tobias Kirschstein, Artem Sevastopolsky, Shunsuke Saito, Qingyang Tan, Javier Romero, Chen Cao, Holly Rushmeier, Giljoo Nam

机构 * Yale University(耶鲁大学) Meta Codec Avatars Lab(Meta 编码人脸实验室) Technical University of Munich(慕尼黑技术大学)

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Accepted to SIGGRAPH 2025. Project page: https://c-he.github.io/projects/3dgh/

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15891 2025-05-20 cs.GR cs.CV 62%

TexPro: Text-guided PBR Texturing with Procedural Material Modeling

Ziqiang Dang, Wenqi Dong, Zesong Yang, Bangbang Yang, Liang Li, Yuewen Ma, Zhaopeng Cui

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

Comments Accepted by CVM 2025 and CVMJ (Computational Visual Media Journal)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19296 2025-03-26 cs.CV cs.MM 62%

Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval

Haoqiang Lin, Haokun Wen, Xuemeng Song, Meng Liu, Yupeng Hu, Liqiang Nie

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.18588 2025-01-31 cs.HC cs.AI cs.CV cs.MM 62%

Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching

David Chuan-En Lin, Hyeonsu B. Kang, Nikolas Martelaro, Aniket Kittur, Yan-Ying Chen, Matthew K. Hong

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

Comments Accepted to CHI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16096 2024-11-26 cs.CV cs.AI cs.MM 62%

ENCLIP: Ensembling and Clustering-Based Contrastive Language-Image Pretraining for Fashion Multimodal Search with Limited Data and Low-Quality Images

Prithviraj Purushottam Naik, Rohit Agarwal

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16254 2024-07-25 cs.CV cs.AI cs.CL cs.MM 62%

Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models

Samuele Poppi, Tobia Poppi, Federico Cocchi, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

Comments ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15155 2024-07-23 cs.CV cs.AI cs.MM 62%

Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification

Yunyi Xuan, Weijie Chen, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.MM

Comments Accepted by ACMMM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.02430 2024-07-03 cs.CV cs.AI cs.GR cs.LG 62%

Meta 3D TextureGen: Fast and Consistent Texture Generation for 3D Objects

Raphael Bensadoun, Yanir Kleiman, Idan Azuri, Omri Harosh, Andrea Vedaldi, Natalia Neverova, Oran Gafni

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.16759 2024-03-21 cs.CV cs.GR 62%

StyleHumanCLIP: Text-guided Garment Manipulation for StyleGAN-Human

Takato Yoshikawa, Yuki Endo, Yoshihiro Kanamori

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments VISIAPP 2024, project page: https://www.cgg.cs.tsukuba.ac.jp/~yoshikawa/pub/style_human_clip/

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.09249 2023-12-15 cs.CV cs.GR 62%

ZeroRF: Fast Sparse View 360° Reconstruction with Zero Pretraining

Ruoxi Shi, Xinyue Wei, Cheng Wang, Hao Su

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Project page: https://sarahweiii.github.io/zerorf/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18491 2023-12-01 cs.CV cs.AI cs.GR cs.LG 62%

ZeST-NeRF: Using temporal aggregation for Zero-Shot Temporal NeRFs

Violeta Menéndez González, Andrew Gilbert, Graeme Phillipson, Stephen Jolly, Simon Hadfield

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments VUA BMVC 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.03335 2023-11-07 cs.CV cs.GR 62%

Cross-Image Attention for Zero-Shot Appearance Transfer

Yuval Alaluf, Daniel Garibi, Or Patashnik, Hadar Averbuch-Elor, Daniel Cohen-Or

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

Comments Project page: https://garibida.github.io/cross-image-attention

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.12302 2023-09-22 cs.CV cs.GR 62%

Text-Guided Vector Graphics Customization

Peiying Zhang, Nanxuan Zhao, Jing Liao

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.GR

Comments Accepted by SIGGRAPH Asia 2023. Project page: https://intchous.github.io/SVGCustomization

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.09637 2023-08-16 cs.CV cs.AI cs.GR cs.LG 62%

InfiniCity: Infinite-Scale City Synthesis

Chieh Hubert Lin, Hsin-Ying Lee, Willi Menapace, Menglei Chai, Aliaksandr Siarohin, Ming-Hsuan Yang, Sergey Tulyakov

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.04343 2023-08-09 cs.CV cs.IR cs.MM 62%

Unifying Two-Stream Encoders with Transformers for Cross-Modal Retrieval

Yi Bin, Haoxuan Li, Yahui Xu, Xing Xu, Yang Yang, Heng Tao Shen

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

Comments Accepted at ACM Multimedia 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.00148 2023-08-02 cs.CV cs.GR 62%

Controlling Geometric Abstraction and Texture for Artistic Images

Martin Büßemeyer, Max Reimann, Benito Buchheim, Amir Semmo, Jürgen Döllner, Matthias Trapp

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.04399 2023-04-11 cs.CV cs.AI cs.LG cs.MM 62%

CAVL: Learning Contrastive and Adaptive Representations of Vision and Language

Shentong Mo, Jingfei Xia, Ihor Markevych

专题命中 文生图 :text-to-image(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.04868 2023-02-10 cs.CV cs.GR 62%

MEGANE: Morphable Eyeglass and Avatar Network

Junxuan Li, Shunsuke Saito, Tomas Simon, Stephen Lombardi, Hongdong Li, Jason Saragih

专题命中 文生图 :image synthesis(abstract);分类 cs.CV、cs.GR

Comments Project page: https://junxuan-li.github.io/megane/

详情

展开后加载摘要…

URL PDF HTML 收藏