CommentsProceedings of the 39th Annual Conference on Neural Information Processing Systems, ARLET Workshop (Aligning Reinforcement Learning Experimentalists and Theorists)
Journal refTransactions on Machine Learning Research, Vol. 2026, June 2026
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
SMART: 基于音频增强多模态大模型的镜头感知视频时刻检索
An Yu, Weiheng Lu, Jian Li, Zhenfei Zhang, Yunhang Shen, Felix X. -F. Ye, Ming-Ching Chang
机构
*
Department of Computer Science, University at Albany - SUNY(University at Albany - SUNY 计算机科学系)
;
School of Software & Microelectronics, Peking University(北京大学软件与微电子学院)
;
Nanjing University(南京大学)
;
Xiamen University(厦门大学)
;
Department of Mathematics and Statistics, University at Albany - SUNY(University at Albany - SUNY 数学与统计学系)
DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
DySink:动态帧 sinks 用于自回归长视频生成
Bo Ye, Xinyu Cui, Jian Zhao, Tong Wei, Min-Ling Zhang
机构
*
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
;
Key Lab. of Computer Network and Information Integration, Southeast University(东南大学计算机网络与信息集成重点实验室)
;
Zhongguancun Academy(中关村学院)
;
Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)
;
Institute of Automation, CAS(中国科学院自动化研究所)
MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling
MMPhysVideo: 通过联合多模态建模提升视频生成的物理合理性
Shubo Lin, Xuanyang Zhang, Wei Cheng, Weiming Hu, Gang Yu, Jin Gao
机构
*
State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(中国科学院自动化研究所多模态人工智能系统国家重点实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
StepFun(阶跃星辰)
;
Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information(多模态信息超级智能安全北京市重点实验室)
;
School of Information Science and Technology, Shanghai Tech University(上海科技大学信息科学与技术学院)