arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Nanjing University(南京大学)

2026-05-04 至 2026-05-04 共收录 5
2605.00658 2026-05-04 cs.CV

UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors

UniVidX:一种基于扩散先验的统一多模态框架,用于 versatile 视频生成

Houyuan Chen, Hong Li, Xianghao Kong, Tianrui Zhu, Shaocong Xu, Weiqing Xiao, Yuwei Guo, Chongjie Ye, Lvmin Zhang, Hao Zhao, Anyi Rao

机构 * Beihang University(北京航空航天大学) Nanjing University(南京大学) Stanford University(斯坦福大学) Tsinghua University(清华大学)

AI总结 UniVidX 提出一种统一多模态框架,通过扩散先验实现 versatile 视频生成,采用随机条件掩码、解耦门控LoRA和跨模态自注意力等设计,提升多模态一致性与生成性能。

Comments Project page: https://houyuanchen111.github.io/UniVidX.github.io/ Accepted to ACM Transactions on Graphics (Proceedings of SIGGRAPH 2026)

Journal ref ACM Trans. Graph. 45, 4, Article 51 (July 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00513 2026-05-04 cs.CL cs.LG

ControBench: An Interaction-Aware Benchmark for Controversial Discourse Analysis on Social Networks

ControBench:一种考虑交互的争议性 discourse 分析社交网络基准

Ta Thanh Thuy, Jiaqi Zhu, Xuan Liu, Lin Shang, Reihaneh Rabbany, Guillaume Rabusseau, Lihui Chen, Zheng Yilun, Sitao Luan

机构 * Nanyang Technological University(南洋理工大学) NVIDIA(NVIDIA公司) Nanjing University(南京大学) Mila - Quebec AI Institute(魁北克AI研究所) McGill University(麦吉尔大学) University of Montreal(蒙特利尔大学)

AI总结 ControBench结合异质社交交互图与丰富文本语义,通过Reddit讨论数据构建,用于研究争议性 discourse 分析中的政治极化、虚假信息和内容审核问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00510 2026-05-04 cs.LG cs.CV physics.comp-ph

Scale-Aware Adversarial Analysis: A Diagnostic for Generative AI in Multiscale Complex Systems

基于尺度的对抗分析:生成AI在多尺度复杂系统中的诊断

Mengke Zhao, Guang-Xing Li, Duo Xu, Keping Qiu

机构 * School of Astronomy and Space Science, Nanjing University(南京大学天文与空间科学学院) South-Western Institute for Astronomy Research, Yunnan University(云南西南天文研究所) Key Laboratory of Modern Astronomy and Astrophysics (Nanjing University), Ministry of Education(南京大学现代天文与天体物理重点实验室) Canadian Institute for Theoretical Astrophysics, University of Toronto(多伦多大学加拿大理论天体物理研究所)

AI总结 本文提出基于约束扩散分解的诊断框架,用于评估生成AI在多尺度复杂系统中的性能,揭示模型在物理约束下的稳定性与连续性问题。

Comments submitted, comments welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00498 2026-05-04 cs.CV

GOR-IS: 3D Gaussian Object Removal in the Intrinsic Space

GOR-IS: 在内在空间中进行3D高斯物体移除

Yonghao Zhao, Yupeng Gao, Jian Yang, Jin Xie, Beibei Wang

机构 * Nankai University(南开大学) Nanjing University(南京大学)

AI总结 本文提出GOR-IS框架,通过分解场景为内在组件并建模光传输,实现物理一致且视觉连贯的3D物体移除,实验表明其在感知相似度和PSNR上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22285 2026-05-04 cs.CV

VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding

VideoDetective:通过外在查询和内在相关性进行长视频理解的线索搜索

Ruoliu Yang, Chu Wu, Caifeng Shan, Ran He, Chaoyou Fu

机构 * Nanjing University(南京大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

AI总结 本文提出VideoDetective框架,通过结合查询与片段的相关性及片段间亲和力,有效定位长视频问答中的关键片段,提升多模态大语言模型在长视频理解中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏