arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉与机器人

图像生成

图像生成、文生图、图像编辑、扩散模型和可控生成。

2026-03-10 至 2026-03-10 共收录 10 信号源:cs.CV, cs.GR, cs.MM

1. 可控生成 10 篇

2508.21363 2026-03-10 cs.CV 79%

Efficient Diffusion-Based 3D Human Pose Estimation with Hierarchical Temporal Pruning

高效扩散模型基于层次时间剪枝的3D人体姿态估计

Yuquan Bi, Hongsong Wang, Xinli Shi, Zhipeng Gui, Jie Gui, Yuan Yan Tang

机构 * School of Cyber Science and Engineering, Southeast University(东南大学计算机科学与工程学院) School of Computer Science and Engineering, Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Ministry of Education, Southeast University(东南大学计算机科学与工程学院、新一代人工智能技术及其交叉应用国家重点实验室) National Center for Applied Mathematics, Southeast University(东南大学应用数学国家中心) Purple Mountain Laboratories, Nanjing(紫金山实验室) Engineering Research Center of Blockchain Application, Supervision and Management (Southeast University), Ministry of Education(区块链应用、监督与管理工程研究中心(东南大学)) Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系) Faculty of Science and Technology, UOW College Hong Kong(香港大学学院科技学院)

专题命中 可控生成 :diffusion(title,abstract);分类 cs.CV

AI总结 本文提出一种高效扩散模型基于层次时间剪枝的3D人体姿态估计方法,通过分阶段剪枝策略显著降低计算成本并提升推理效率。

Comments Accepted by IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHOLOGY

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.11161 2026-03-10 cs.CV cs.GR cs.MM 67%

altiro3D: Scene representation from single image and novel view synthesis

altiro3D: 从单张图像和新视角合成生成场景表示

E. Canessa, L. Tenze

机构 * The Abdus Salam International Centre for Theoretical Physics, ICTP, Trieste 34151, Italy(阿布杜斯·萨拉姆国际理论物理学中心,ICTP,特里埃斯特)

专题命中 可控生成 :inpainting(abstract);分类 cs.CV、cs.GR、cs.MM

AI总结 altiro3D通过单张图像生成多视角和视频,结合深度估计与快速算法实现逼真3D体验。

Comments In press (2023) Springer International Journal of Information Technology (IJIT) 10 pages, 3 figures

Journal ref International Journal of Information Technology 16 (2024) 33

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08028 2026-03-10 cs.CV cs.MM 62%

Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades

通过文本到骨骼级联实现可控的复杂人类运动视频生成

Ashkan Taghipour, Morteza Ghahremani, Zinuo Li, Hamid Laga, Farid Boussaid, Mohammed Bennamoun

机构 * Department of Computer Science and Software Engineering, The University of Western Australia(计算机科学与软件工程系,西澳大学) Munich Center for Machine Learning (MCML) and Technical University of Munich (TUM)(慕尼黑机器学习中心(MCML)和技术大学慕尼黑(TUM)) School of Information Technology, Murdoch University(信息科技学院,墨尔本大学) Department of Electrical, Electronics and Computer Engineering, The University of Western Australia(电子、电子与计算机工程系,西澳大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV、cs.MM

AI总结 本文提出通过文本到骨骼级联框架生成可控复杂人类运动视频,解决文本条件模糊和姿态控制成本高的问题,引入合成数据集并展示模型在多个指标上的优越性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07691 2026-03-10 cs.RO cs.CV 57%

RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation

RoboPCA:基于姿态的从人类示范中学习空间可及性的方法用于机器人操作

Zhanqi Xiao, Ruiping Wang, Xilin Chen

机构 * Key Laboratory of AI Safety of CAS, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室、计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

AI总结 RoboPCA通过联合预测接触区域和姿态,提升机器人操作中从人类示范中学习空间可及性的性能。

Comments Accepted to ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11952 2026-03-10 cs.CV 57%

UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding

UniUGG: 通过几何-语义编码实现统一的3D理解与生成

Yueming Xu, Jiahui Zhang, Ze Huang, Yurui Chen, Yanpeng Zhou, Zhenyu Chen, Yu-Jie Yuan, Pengxiang Xia, Guowei Huang, Xinyue Cai, Zhongang Qi, Xingyue Quan, Jianye Hao, Hang Xu, Li Zhang

机构 * Fudan University(复旦大学) Huawei Noah’s Ark Lab(华为诺亚实验室)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

AI总结 UniUGG通过几何-语义编码实现3D理解和生成的统一框架,结合大语言模型和潜在扩散模型,提升空间理解和生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06672 2026-03-10 cs.CV cs.AI 57%

Does Semantic Noise Initialization Transfer from Images to Videos? A Paired Diagnostic Study

语义噪声初始化能否从图像迁移到视频?一项配对诊断研究

Yixiao Jing, Chaoyu Zhang, Zixuan Zhong, Peizhou Huang

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

AI总结 该研究探讨了语义噪声初始化在文本到视频生成中的有效性,发现其在时间相关维度有小幅度提升,但整体效果与基线相当,建议采用提示级评估和噪声空间分析作为标准方法。

Comments 8 pages, 1 figure. Accepted to the ICLR 2026 Workshop on Multimodal Intelligence: Next Token Prediction & Beyond

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17908 2026-03-10 cs.CV cs.AI cs.LG 57%

ReDepth Anything: Test-Time Depth Refinement via Self-Supervised Re-lighting

ReDepth Anything: 通过自监督重照明进行测试时深度细化

Ananta R. Bhattarai, Helge Rhodin

机构 * Bielefeld University(比勒菲尔德大学)

专题命中 可控生成 :diffusion(abstract);分类 cs.CV

AI总结 ReDepth Anything通过自监督重照明技术,在测试时提升深度估计的准确性和真实感,优于现有方法并达到最新水平。

Comments Accepted at CVPR 2026 (Findings). Project Page: https://anantarb.github.io/redepth Code: https://github.com/anantarb/Re-Depth-Anything

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11364 2026-03-10 cs.RO 56%

ActivePose: Active 6D Object Pose Estimation and Tracking for Robotic Manipulation

ActivePose: 用于机器人操作的主动6D物体姿态估计与跟踪

Sheng Liu, Zhe Li, Weiheng Wang, Han Sun, Heng Zhang, Hongpeng Chen, Yusen Qin, Arash Ajoudani, Yizhao Wang

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工大学) Shanghai Jiao Tong University(上海交通大学) Istituto Italiano di Tecnologia(意大利理工学院) The Hong Kong Polytechnic University(香港理工大学) D-Robotics(D-机器人)

专题命中 可控生成 :diffusion(abstract,comments)

AI总结 ActivePose通过结合视觉-语言模型与机器人想象力,实现主动6D物体姿态估计与跟踪,提升机器人操作的可靠性。

Comments 6D Pose, Diffusion Policy, Robot Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01587 2026-03-10 math.NA cs.NA math.OC 50%

The State-Dependent Riccati Equation in Nonlinear Optimal Control: Analysis, Error Estimation and Numerical Approximation

非线性最优控制中的状态依赖性里卡蒂方程:分析、误差估计与数值近似

Luca Saluzzi

专题命中 可控生成 :diffusion(abstract)

AI总结 本文研究了非线性最优控制中状态依赖性里卡蒂方程的理论分析、误差估计及数值近似方法,提出最优半线性分解策略并评估了两种数值方法的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06831 2026-03-10 cs.RO math.OC 50%

Learning-Based Robust Control: Unifying Exploration and Distributional Robustness for Reliable Robotics via Free Energy

基于学习的鲁棒控制:通过自由能统一探索与分布鲁棒性以实现可靠的机器人控制

Hozefa Jesawada, Giovanni Russo, Abdalla Swikir, Fares Abu-Dakka

机构 * Mechanical Engineering Program, New York University Abu Dhabi(纽约大学阿布扎赫尔分校机械工程项目) Department of Information and Electrical Engineering & Applied Mathematics, University of Salerno(萨勒诺大学信息与电气工程及应用数学系) Mohammed bin Zayed University for Artificial Intelligence, Abu Dhabi, UAE(穆罕默德·本·扎耶德人工智能大学)

专题命中 可控生成 :diffusion(abstract)

AI总结 本文提出了一种基于自由能原理的鲁棒控制方法,通过联合学习环境动态和奖励,提升机器人在真实环境中的鲁棒性和可重复操作性能。

详情

展开后加载摘要…

URL PDF HTML 收藏