arXivDaily arXiv每日学术速递 周一至周五更新

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

共收录 762 信号源:cs.CV, cs.GR, cs.MM

1. 其他图像生成 762 篇

2606.21304 2026-06-23 cs.CV 新提交 57%

A Test-time Actor-Critic Approach to News Images Generation

一种测试时的演员-评论家方法用于新闻图像生成

Damianos Galanopoulos, Vasileios Mezaris

机构 * Information Technologies Institute (ITI), Centre of Research and Technology Hellas (CERTH)(信息技术研究所(ITI),希腊研究与技术中心(CERTH))

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 提出ACIG方法,基于强化学习的Actor-Critic范式,在测试时通过反馈循环生成、评估和优化新闻图像提示,在MediaEval NewsImages 2026挑战中取得最佳结果。

Comments MediaEval 2026 Workshop, Amsterdam, NL

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.04421 2026-06-23 cs.CV 57%

Conditional Generative Adversarial Networks for Emoji Synthesis with Word Embedding Manipulation

基于词嵌入操作的条件生成对抗网络用于表情符号合成

Nelu Radpour, Vivek Bheda

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出一种优化训练的深度卷积GAN,通过整合Google的word2vec词嵌入,生成高度逼真、与真实表情符号几乎相同的合成表情符号。

Comments 5 pages, 3 figures, 2 graphs

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14877 2026-06-18 cs.CV 版本更新 57%

HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling

HeatKV:针对视觉自回归建模的头部调制KV缓存压缩

Jonathan Cederlund, Axel Berg, William Isaksson, Durmus Alp Emre Acar, Chuteng Zhou, Pontus Giselsson

机构 * Dept. of Automatic Control, Lund University(自动控制系,吕勒欧大学) Arm(Arm公司)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出HeatKV方法,通过根据每个头部对先前生成尺度的注意力进行调整,实现更高效的KV缓存压缩,提升内存利用率并保持图像生成质量。

Comments 18 pages total including appendix; 6 main-paper figures, 2 appendix figures; 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13679 2026-06-15 cs.CV 新提交 57%

InterleaveThinker: Reinforcing Agentic Interleaved Generation

InterleaveThinker: 强化智能体交错生成

Dian Zheng, Harry Lee, Manyuan Zhang, Kaituo Feng, Zoey Guo, Ray Zhang, Hongsheng Li

机构 * CUHK MMLab(香港中文大学多媒体实验室) Meituan(美团) CUHK IMIXR(香港中文大学IMIXR实验室)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 提出首个多智能体管线InterleaveThinker,通过规划器和评论家智能体使现有图像生成器具备交错生成能力,并利用GRPO强化单步指令修正,显著提升生成性能。

Comments Project Page: https://zhengdian1.github.io/InterleaveThinker-proj/ Code: https://github.com/zhengdian1/InterleaveThinker

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13625 2026-06-12 cs.CV 新提交 57%

Revisiting Vehicle Color Recognition in Long-Tailed Surveillance Scenarios

重新审视长尾监控场景中的车辆颜色识别

Vinícius Orrú, Bruno H. Foggiatto, Gabriel E. Lima, David Menotti, Rayson Laroca

机构 * Pontifical Catholic University of Paraná(巴拉那天主教大学) National High Court of Brazil(巴西国家高等法院) Federal University of Paraná(巴拉那联邦大学)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 针对监控场景中车辆颜色分布高度不平衡的问题,本文提出结合生成式数据增强、视觉表征、损失重加权等方法的综合方案,在UFPR-VeSV数据集上实现94.6%微平均和79.7%宏平均准确率,宏平均比近期文献提升8.2个百分点。

Comments Accepted for presentation at the 2026 International Conference on Pattern Recognition (ICPR) - V3SC Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.09400 2026-06-09 cs.CV 新提交 57%

vesselFM-CT: Segmenting All Blood Vessels in CT Images for System-Level Cardiovascular Analysis

vesselFM-CT:在CT图像中分割所有血管以实现系统级心血管分析

Bastian Wittmann, Chinmay Prabhakar, Suprosanna Shit, Bjoern Menze

机构 * Department of Quantitative Biomedicine, University of Zurich(苏黎世大学定量生物医学系)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 提出vesselFM-CT模型,通过迭代多步训练和TubeLoss损失函数,实现CT图像中从大血管到微小肠系膜血管的全分割,优于基线方法,支持系统级心血管分析。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.20549 2026-05-21 cs.CV 57%

MAPS: A Synthetic Dataset for Probing Vision Models in a Controlled 3D Scene Space

MAPS:用于在受控3D场景空间中探测视觉模型的合成数据集

Santiago Galella, Pamela Osuna-Vargas, Maren Wehrheim, Martina G. Vilas, Gemma Roig, Matthias Kaschube

机构 * FIAS & Institute of Computer Science Goethe University Frankfurt(FIAS与计算机科学研究所弗赖堡大学) Mila & Department of Biology York University(Mila与生物学系约克大学) Institute of Computer Science Goethe University Frankfurt(计算机科学研究所弗赖堡大学)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出MAPS数据集,用于在受控3D场景空间中研究视觉模型的行为,通过回归敏感性分析评估20种模型对场景因素的依赖性,发现相机距离和高度是导致识别失败的主要因素,且现代CNN和Transformer模型在敏感性上表现出相似性。

Comments 33 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.17564 2026-05-19 cs.CV 57%

A Conditional U-Net Pipeline with Pre- and Post-Processing for Aerial RGB-to-Thermal Image Translation

具有预处理和后处理的条件U-Net管道用于航空RGB到热图像转换

Tseten Sherpa, Sikandar Ali, Shubham Parab, Haoyun Feng, Matthew Dennis, Keenan Gibbons, Verrah Otiende, Geoffrey H. Siwo

机构 * Department of Data Science, University of Michigan, Ann Arbor, MI, USA(数据科学系,密歇根大学,安阿伯,MI,美国) Department of Information Science, University of Michigan, Ann Arbor, MI, USA(信息科学系,密歇根大学,安阿伯,MI,美国) Department of Computer Science, University of Michigan, Ann Arbor, MI, USA(计算机科学系,密歇根大学,安阿伯,MI,美国) Arcknow, New York, USA(Arcknow,纽约,美国) School of Environmental Sustainability, University of Michigan, Ann Arbor, MI, USA(可持续环境学院,密歇根大学,安阿伯,MI,美国) SmithGroup, Ann Arbor, MI, USA(SmithGroup,安阿伯,MI,美国) Michigan Institute for Data and AI in Society (MIDAS), University of Michigan, Ann Arbor, MI, USA(密歇根数据与人工智能社会研究院(MIDAS),密歇根大学,安阿伯,MI,美国) United States International University (USIU), Nairobi, Kenya(美国国际大学(USIU),内罗毕,肯尼亚) Department of Learning Health Sciences, University of Michigan Medical School, Ann Arbor, MI, USA(学习健康科学系,密歇根大学医学院,安阿伯,MI,美国) Department of Pharmacology, University of Michigan Medical School, Ann Arbor, MI, USA(药理学系,密歇根大学医学院,安阿伯,MI,美国) Center for Global Health Equity, University of Michigan, Ann Arbor, MI, USA(全球健康公平中心,密歇根大学,安阿伯,MI,美国)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出了一种基于条件U-Net的简单架构,结合天气数据和针对性预处理与后处理技术,以提高航空RGB到热图像转换的性能,实验结果显示其在PSNR、SSIM和LPIPS指标上优于现有方法。

Comments 8 pages, 7 figures, NeurIPS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.24763 2026-05-19 cs.CV 57%

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Tuna-2:像素嵌入在多模态理解和生成中优于视觉编码器

Zhiheng Liu, Weiming Ren, Xiaoke Huang, Shoufa Chen, Tianhong Li, Mengzhao Chen, Yatai Ji, Sen He, Jonas Schult, Belinda Zeng, Tao Xiang, Wenhu Chen, Ping Luo, Luke Zettlemoyer, Yuren Cong

机构 * Meta AI The University of Hong Kong(香港大学) University of Waterloo(滑铁卢大学)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出Tuna-2,一种基于像素嵌入的统一多模态模型,通过直接使用像素嵌入进行多模态理解和生成,展示了统一像素空间建模在高质量图像生成中可以与潜在空间方法竞争,并证明了预训练视觉编码器在多模态建模中并非必要。

Comments Project page: https://tuna-ai.org/tuna-2

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.08704 2026-05-19 cs.CV cs.LG 57%

Rethinking Generative Image Pretraining: How Far Are We From Scaling Up Next-Pixel Prediction?

重新思考生成图像预训练:我们离扩大下一步像素预测还有多远?

Xinchen Yan, Chen Liang, Lijun Yu, Adams Wei Yu, Yifeng Lu, Quoc V. Le

机构 * Google Deepmind(谷歌DeepMind)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文研究了自回归下一步像素预测的扩展特性,探讨了统一视觉模型中简单且端到端但尚未充分探索的框架。通过在32x32分辨率的图像上训练Transformer模型,评估了三个目标指标:下一步像素预测目标、ImageNet分类准确率和基于生成的完成度(通过Fr'echet距离测量)。研究发现,最优扩展策略高度依赖任务,且随着图像分辨率的增加,模型大小必须比数据量增长得更快。通过预测发现,计算能力是主要瓶颈,而非训练数据量。随着计算能力每年增长四到五倍,预计在五年内可实现像素级图像建模。

Comments Accepted by ICML2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.16165 2026-05-18 cs.CV cs.AI 57%

Second-Order Multi-Level Variance Correction for Modality Competition in Multimodal Models

二阶多级方差校正用于多模态模型中的模态竞争

Yishun Lu, Wes Armour

机构 * University of Oxford, Oxford, United Kingdom(牛津大学,英国)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出ML-FOP-SOAP框架,通过多级方差校正提升多模态对齐稳定性,实验显示在Janus和Emu3数据集上,该方法提高了样本效率和训练速度,适用于大规模多模态基础模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22663 2026-05-13 cs.CV 57%

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model

AIA: 重新思考统一多模态模型中的架构解耦策略

Dian Zheng, Manyuan Zhang, Hongyu Li, Kai Zou, Hongbo Liu, Ziyu Guo, Kaituo Feng, Yexin Liu, Ying Luo, Hongsheng Li

机构 * MMLab, CUHK(CUHK多媒体实验室) Meituan(美团) USTC(中国科学技术大学) TJU(天津大学)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出AIA损失函数,通过学习任务特定的多模态交互模式,缓解任务冲突而不依赖模型解耦,提升生成与理解性能。

Comments Project page: https://zhengdian1.github.io/AIA-project/ Code: https://github.com/zhengdian1/AIA

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25809 2026-05-11 cs.CV 57%

Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning

基于指令和证据的对比双流解码用于 grounded 视觉语言推理

Yashwant Pravinrao Bangde, Debaditya Roy

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Indian Institute of Technology Kharagpur(印度理工学院Kharagpur分校)

专题命中 其他图像生成 :generative vision(abstract);分类 cs.CV

AI总结 本文提出 IECD$^2$ 方法,通过双流解码平衡语言信息和视觉真实性,提升视觉语言推理任务的准确性和减少幻觉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.27804 2026-05-01 cs.CV cs.CR cs.LG 57%

Machine Unlearning for Class Removal through SISA-based Deep Neural Network Architectures

通过基于SISA的深度神经网络架构实现类别移除的机器无学习

Ishrak Hamim Mahi, Siam Ferdous, Md Sakib Sadman Badhon, Nabid Hasan Omi, Md Habibun Nabi Hemel, Farig Yousuf Sadeque, Md. Tanzim Reza

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Brac University(布拉大学) Dhaka, Bangladesh(孟加拉国达卡)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出一种改进的SISA框架,用于在卷积神经网络中实现类别级无学习,通过增强的重放机制和门控网络提升选择性遗忘效率,实验表明该方法能有效移除特定类别数据并保持模型性能。

Comments 10 pages, 9 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14574 2026-04-17 cs.CV 57%

M3D-Net: Multi-Modal 3D Facial Feature Reconstruction Network for Deepfake Detection

M3D-Net:面向深度伪造检测的多模态3D面部特征重建网络

Haotian Wu, Yue Cheng, Shan Bian

机构 * College of Mathematics and Informatics(数学与信息学院) South China Agricultural University(华南农业大学) Wushan Road, Tianhe District(天河南路) Guangzhou(广州) China(中国)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出M3D-Net,通过多模态特征融合和3D重建模块提升深度伪造检测的准确性和鲁棒性,实验表明其在多个数据集上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12318 2026-04-15 cs.CV 57%

Cell Instance Segmentation via Multi-Task Image-to-Image Schrödinger Bridge

基于多任务图像到图像Schrödinger桥的细胞实例分割

Hayato Inoue, Shota Harada, Shumpei Takezaki, Ryoma Bise

机构 * Dept. of Information Science and Technology(信息科学与技术系)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出基于Schrödinger桥的多任务图像到图像生成框架,通过反向距离图实现边界感知监督,无需预训练或后处理即可在PanNuke和MoNuSeg数据集上实现竞争性或优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06129 2026-04-08 cs.CV cs.AI 57%

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer

PoM:一种线性时间的注意力替代方案:多项式混合器

David Picard, Nicolas Dufour, Lucas Degeorge, Arijit Ghosh, Davide Allegro, Tom Ravaud, Yohann Perron, Corentin Sautier, Zeynep Sonat Baltaci, Fei Meng, Syrine Kalleli, Marta López-Rauhut, Thibaut Loiseau, Ségolène Albouy, Raphael Baena, Elliot Vincent, Loic Landrieu

机构 * LIGM, CNRS, Univ Gustave Eiffel, ENPC, Institut Polytechnique de Paris(LIGM,法国国家科学研究中心,古斯塔夫·埃菲尔大学,巴黎高科路桥学院,巴黎综合理工学院) LIX, École Polytechnique, CNRS, IP Paris(LIX,巴黎综合理工学院,法国国家科学研究中心,巴黎综合理工学院) Department of Information Engineering, Università degli Studi di Padova(帕多瓦大学信息工程系) AMIAD, Pole recherche, EFEO(AMIAD,法国远东学院研究部) LASTIG, Univ Gustave Eiffel, IGN, Géodata Paris(LASTIG,古斯塔夫·埃菲尔大学,法国国家地理与森林信息研究所,巴黎地理数据)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出PoM,一种线性复杂度的替代注意力机制,通过学习多项式函数聚合输入令牌,实现高效序列处理,减少长序列计算成本。

Comments Accepted to CVPR Findings 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00955 2026-04-02 cs.CV 57%

Enhancing Gradient Inversion Attacks in Federated Learning via Hierarchical Feature Optimization

通过分层特征优化增强联邦学习中的梯度反向攻击

Hao Fang, Wenbo Yu, Bin Chen, Xuan Wang, Shu-Tao Xia, Qing Liao, Ke Xu

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出GIFD方法,通过分解GAN模型并搜索中间层的分层特征,提升梯度反向攻击的表达能力和泛化能力,同时引入正则化约束和OOD场景下的标签映射技术,实现像素级重建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10805 2026-04-01 cs.LG cs.CV 57%

Interpretable and Steerable Concept Bottleneck Sparse Autoencoders

可解释且可操控的概念瓶颈稀疏自编码器

Akshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy, Shusen Liu, Wesam A. Sakla, Kowshik Thopalli

机构 * University of California, San Diego(加州大学圣迭戈分校) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出CB-SAE框架,通过剪枝低效神经元和增加概念瓶颈提升LVLMs和图像生成任务的可解释性和可操控性,效果提升分别为32.1%和14.5%。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28333 2026-03-31 cs.CV cs.AI 57%

Integrating Multimodal Large Language Model Knowledge into Amodal Completion

将多模态大语言模型知识整合到无模态补全中

Heecheol Yun, Eunho Yang

机构 * KAIST(韩国科学技术院) AITRICS

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出AmodalCG框架,利用多模态大语言模型知识指导无模态补全,通过评估遮挡程度选择性调用MLLM指导,结合视觉生成模型迭代优化补全结果,实验表明优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.04317 2026-03-24 cs.CV cs.CL 57%

WiFi-GEN: High-Resolution Indoor Imaging from WiFi Signals Using Generative AI

WiFi-GEN:利用生成式人工智能进行高分辨率室内成像

Jianyang Shi, Bowen Zhang, Amartansh Dubey, Ross Murch, Liwen Jing

机构 * Pengcheng Laboratory(鹏城实验室) College of Big Data and Internet(大数据与互联网学院) Shenzhen Technology University(深圳技术大学) Department of Electrical Engineering(电子工程系) Indian Institute of Technology Delhi(印度理工学院德里分校) Department of Electronic and Computer Engineering(电子与计算机工程系) HKUST(香港科技大学) School of Computer Science and Technology(计算机科学与技术学院) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出WiFi-GEN,通过生成式人工智能将测量到的WiFi功率转换为高分辨率室内图像,其形状重建精度比物理模型方法高275%,且Frechet Inception Distance得分降低82%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24267 2026-03-17 cs.CV 57%

FakeScope: Large Multimodal Expert Model for Transparent AI-Generated Image Forensics

FakeScope:大型多模态专家模型用于透明的AI生成图像取证

Yixuan Li, Yu Tian, Yipo Huang, Wei Lu, Shiqi Wang, Weisi Lin, Anderson Rocha

机构 * College of Computing, City University of Hong Kong(城市大学 computing 学院) School of Data Science and Artificial Intelligence, Chang’an University(长安大学数据科学与人工智能学院) School of Computer Science and Engineering, Ministry of Education Key Laboratory of Information Technology, Guangdong Province Key Laboratory of Information Security Technology, Sun Yat-Sen University(中山大学计算机科学与工程学院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学 computing 与数据科学学院) Artificial Intelligence Lab. (Recod.ai) at the University of Campinas(坎皮纳斯大学人工智能实验室)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 FakeScope通过多模态专家模型提升AI生成图像的检测能力,提供可解释的取证分析,实现高精度识别与深入解释,具备零样本检测和跨生成器泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04980 2026-03-06 cs.CV 57%

A Simple Baseline for Unifying Understanding, Generation, and Editing via Vanilla Next-token Prediction

一种通过 vanilla next-token 预测统一理解、生成和编辑的简单基线

Jie Zhu, Hanghang Ma, Jia Wang, Yayong Guan, Yanbing Zeng, Lishuai Gao, Junqiang Wu, Jie Hu, Leye Wang

机构 * Key Lab of High Confidence Software Technologies (Peking University), Ministry of Education, China(高可信软件技术重点实验室(北京大学),教育部,中国) School of Computer Science, Peking University, Beijing, China(北京大学计算机学院,北京,中国) School of Computer Science and Technology, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学计算机科学与技术学院,北京,中国)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出 Wallaroo,一种通过 vanilla next-token 预测统一多模态理解、生成和编辑的简单基线模型,展示了其在多任务中的竞争力。

Comments Technical report. This work serves as a straightforward autoregressive baseline for unifying understanding, generation, and editing

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18970 2026-02-26 cs.CV 57%

Pay Attention to Where You Looked

关注你所注视的地方

Alex Berian, JhihYang Wu, Daniel Brignac, Natnael Daba, Abhijit Mahalanobis

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出了一种自适应视角加权机制,通过调整源视角的重要性来提升少样本NVS的合成质量与真实感。

Comments ICIP 2025 Workshop on Generative AI for World Simulations and Communications

Journal ref International Conference on Image Processing 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.20034 2026-02-16 cs.CV 57%

MaskInversion: Localized Embeddings via Optimization of Explainability Maps

MaskInversion: 通过可解释性图的优化生成局部嵌入

Walid Bousselham, Sofian Chaybouti, Christian Rupprecht, Vittorio Ferrari, Hilde Kuehne

机构 * Tuebingen AI Center University of Tuebingen(图宾根人工智能中心 图宾根大学) University of Oxford(牛津大学) Meta MIT-IBM Watson AI Lab(麻省理工-IBM Watson人工智能实验室)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 MaskInversion通过优化可解释性图生成特定图像区域的嵌入,适用于多种视觉-语言任务。

Comments Project page: https://walidbousselham.com/MaskInversion

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18851 2026-01-28 cs.CV 57%

SelfieAvatar: Real-time Head Avatar reenactment from a Selfie Video

SelfieAvatar: 从自拍视频实现实时头部虚拟形象重现

Wei Liang, Hui Yu, Derui Ding, Rachael E. Jack, Philippe G. Schyns

机构 * Department of Control Science and Engineering, University of Shanghai for Science and Technology(上海理工大学控制科学与工程系) School of Psychology and Neuroscience, University of Glasgow(格拉斯哥大学心理学与神经科学学院)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本研究提出基于3DMM和StyleGAN的方法,利用自拍视频实现实时头部虚拟形象重现,提升细节还原与高保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11952 2026-01-21 cs.CV 57%

Decoder Gradient Shields: A Family of Provable and High-Fidelity Methods Against Gradient-Based Box-Free Watermark Removal

解码器梯度防护:一种可证明且高保真的对抗梯度基于的框外水印移除方法

Haonan An, Guang Hua, Wei Du, Hangcheng Cao, Yihang Tao, Guowen Xu, Susanto Rahardja, Yuguang Fang

机构 * Hong Kong JC STEM Lab of Smart City and Department of Computer Science, City University of Hong Kong(香港JC STEM实验室及城市大学计算机科学系) Infocomm Technology Cluster, Singapore Institute of Technology(新加坡理工学院信息通信技术集群) College of Information Science & Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科学与技术大学计算机科学与工程学院) Department of Electronic and Electrical Engineering, University College London(伦敦大学学院电子与电气工程系)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本文提出了解码器梯度防护(DGSs)方法,通过防止梯度传播来增强对抗框外水印移除的鲁棒性,有效保护模型知识产权。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04785 2026-01-09 cs.CV cs.AI 57%

SRU-Pix2Pix: A Fusion-Driven Generator Network for Medical Image Translation with Few-Shot Learning

SRU-Pix2Pix:一种融合驱动的生成器网络用于医学图像翻译的少样本学习

Xihe Qiu, Yang Dai, Xiaoyu Tan, Sijia Li, Fenghao Sun, Lu Gan, Liang Liu

机构 * School of Electronic and Electrical Engineering, Shanghai University of Engineering Science(上海工程技术大学电子电气工程学院) Department of Thoracic Surgery, Zhongshan Hospital of Fudan University(复旦大学中山医院胸外科) Clinical Research Unit, Institute of Clinical Science, Zhongshan Hospital of Fudan University(复旦大学中山医院临床科研部) Department of Medical Oncology, Cancer Center, and Fudan Zhangjiang Institute, Zhongshan Hospital of Fudan University(复旦大学中山医院医学肿瘤科、癌症中心及复旦张江研究院)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 SRU-Pix2Pix通过整合SEResNet和U-Net++提升医学图像翻译的生成质量与结构保真度,实现少样本下的高效翻译

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22766 2025-12-30 eess.IV cs.CV nucl-ex 57%

SwinCCIR: An end-to-end deep network for Compton camera imaging reconstruction

SwinCCIR:一种用于康普顿相机成像重建的端到端深度网络

Minghao Dong, Xinyang Luo, Xujian Ouyang, Yongshun Xiao

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 SwinCCIR通过端到端深度学习框架提升康普顿相机成像重建效果,有效解决伪影和形变问题。

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.10067 2025-12-23 cs.CV cs.LG 57%

Independent Density Estimation

独立密度估计

Jiahao Liu, Senhao Cao

机构 * Orcava Inc.(Orcava公司)

专题命中 其他图像生成 :image generation(abstract);分类 cs.CV

AI总结 本研究提出独立密度估计(IDE)方法,通过学习句子中单个词与图像特征之间的联系,提升模型在组合泛化任务上的性能。

Comments 10 pages, 1 table, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏