arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2026-05-19 至 2026-05-19 共收录 214 信号源:cs.CV, cs.GR, cs.MM

1. 其他图像生成 4 篇

2501.04163 2026-05-19 cs.HC 71%

HistoryPalette: Supporting Exploration and Reuse of Past Alternatives in Image Generation and Editing

HistoryPalette: 支持在图像生成与编辑中探索和重用过去的选择

Karim Benharrak, Amy Pavel

专题命中 其他图像生成 :image generation(title)

AI总结 HistoryPalette通过组织历史设计替代方案,帮助创作者快速预览和重用先前工作,提升创意任务的效率与协作体验。

Comments Accepted to CHI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17564 2026-05-19 cs.CV 57%

A Conditional U-Net Pipeline with Pre- and Post-Processing for Aerial RGB-to-Thermal Image Translation

具有预处理和后处理的条件U-Net管道用于航空RGB到热图像转换

Tseten Sherpa, Sikandar Ali, Shubham Parab, Haoyun Feng, Matthew Dennis, Keenan Gibbons, Verrah Otiende, Geoffrey H. Siwo

机构 * Department of Data Science, University of Michigan, Ann Arbor, MI, USA(数据科学系,密歇根大学,安阿伯,MI,美国) Department of Information Science, University of Michigan, Ann Arbor, MI, USA(信息科学系,密歇根大学,安阿伯,MI,美国) Department of Computer Science, University of Michigan, Ann Arbor, MI, USA(计算机科学系,密歇根大学,安阿伯,MI,美国) Arcknow, New York, USA(Arcknow,纽约,美国) School of Environmental Sustainability, University of Michigan, Ann Arbor, MI, USA(可持续环境学院,密歇根大学,安阿伯,MI,美国) SmithGroup, Ann Arbor, MI, USA(SmithGroup,安阿伯,MI,美国) Michigan Institute for Data and AI in Society (MIDAS), University of Michigan, Ann Arbor, MI, USA(密歇根数据与人工智能社会研究院(MIDAS),密歇根大学,安阿伯,MI,美国) United States International University (USIU), Nairobi, Kenya(美国国际大学(USIU),内罗毕,肯尼亚) Department of Learning Health Sciences, University of Michigan Medical School, Ann Arbor, MI, USA(学习健康科学系,密歇根大学医学院,安阿伯,MI,美国) Department of Pharmacology, University of Michigan Medical School, Ann Arbor, MI, USA(药理学系,密歇根大学医学院,安阿伯,MI,美国) Center for Global Health Equity, University of Michigan, Ann Arbor, MI, USA(全球健康公平中心,密歇根大学,安阿伯,MI,美国)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出了一种基于条件U-Net的简单架构,结合天气数据和针对性预处理与后处理技术,以提高航空RGB到热图像转换的性能,实验结果显示其在PSNR、SSIM和LPIPS指标上优于现有方法。

Comments 8 pages, 7 figures, NeurIPS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24763 2026-05-19 cs.CV 57%

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Tuna-2:像素嵌入在多模态理解和生成中优于视觉编码器

Zhiheng Liu, Weiming Ren, Xiaoke Huang, Shoufa Chen, Tianhong Li, Mengzhao Chen, Yatai Ji, Sen He, Jonas Schult, Belinda Zeng, Tao Xiang, Wenhu Chen, Ping Luo, Luke Zettlemoyer, Yuren Cong

机构 * Meta AI The University of Hong Kong(香港大学) University of Waterloo(滑铁卢大学)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出Tuna-2,一种基于像素嵌入的统一多模态模型,通过直接使用像素嵌入进行多模态理解和生成,展示了统一像素空间建模在高质量图像生成中可以与潜在空间方法竞争,并证明了预训练视觉编码器在多模态建模中并非必要。

Comments Project page: https://tuna-ai.org/tuna-2

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08704 2026-05-19 cs.CV cs.LG 57%

Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?

重新思考生成图像预训练:我们离扩大下一步像素预测还有多远?

Xinchen Yan, Chen Liang, Lijun Yu, Adams Wei Yu, Yifeng Lu, Quoc V. Le

机构 * Google Deepmind(谷歌DeepMind)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文研究了自回归下一步像素预测的扩展特性,探讨了统一视觉模型中简单且端到端但尚未充分探索的框架。通过在32x32分辨率的图像上训练Transformer模型,评估了三个目标指标:下一步像素预测目标、ImageNet分类准确率和基于生成的完成度(通过Fr'echet距离测量)。研究发现,最优扩展策略高度依赖任务,且随着图像分辨率的增加,模型大小必须比数据量增长得更快。通过预测发现,计算能力是主要瓶颈,而非训练数据量。随着计算能力每年增长四到五倍,预计在五年内可实现像素级图像建模。

Comments Accepted by ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏