Transformer self-attention encoder-decoder with multimodal deep learning for response time series forecasting and digital twin support in wind structural health monitoring
Resolving Spatio-Temporal Entanglement in Video Prediction via Multi-Modal Attention
通过多模态注意力解决视频预测中的时空纠缠
Shreyam Gupta, P. Agrawal, Priyam Gupta
机构
*
Indian Institute of Technology (BHU), Varanasi(印度理工学院(巴纳拉斯印度教大学),瓦拉纳西)
;
University of Colorado, Boulder(科罗拉多大学博尔德分校)
;
Erasmus+, Intelligent Field Robotic Systems (IFRoS), University of Girona(伊拉斯谟+,智能现场机器人系统(IFRoS),赫罗纳大学)
机构
*
Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州)人工智能研究所)
;
Hubei Key Laboratory of Inland Shipping Technology (Wuhan University of Technology)(湖北内河航运技术重点实验室(武汉理工大学))
;
School of Navigation, Wuhan University of Technology(武汉理工大学航海学院)
;
School of Advanced Technology, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学先进科技学院)
;
School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
;
School of Information Engineering, Yancheng Institute of Technology(盐城职业技术学院信息工程学院)
;
School of Engineering, Stanford University(斯坦福大学工程学院)
;
Centre for AI and Data Science Innovation and the School of Science and Engineering, James Cook University(詹姆斯库克大学人工智能与数据科学创新中心及科学与工程学院)
Next Point-of-interest (POI) Recommendation Model Based on Multi-modal Spatio-temporal Context Feature Embedding
基于多模态时空上下文特征嵌入的下一个兴趣点(POI)推荐模型
Lingyu Zhang, Pengfei Xu, Rui Ban, Zhenchao Zhang, Songtao Liu, Yan Wang, Yunhai Wang
机构
*
Institute of Trustworthy Autonomous Systems, Southern University of Science and Technology(SUSTech)(可信自主系统研究院,南方科技大学)
;
School of Information Sciences and Technology, Northwest University(信息科学与技术学院,西北大学)
;
China Information Technology Designing & Consulting Institute Co., Ltd.(中国信息科技设计与咨询研究院有限公司)
;
Renmin University of China(中国人民大学)
OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions
OmniVCus: 基于多模态控制条件的前馈主体驱动视频定制
Yuanhao Cai, He Zhang, Xi Chen, Jinbo Xing, Yiwei Hu, Yuqian Zhou, Kai Zhang, Zhifei Zhang, Soo Ye Kim, Tianyu Wang, Yulun Zhang, Xiaokang Yang, Zhe Lin, Alan Yuille
机构
*
Johns Hopkins University(约翰霍普金斯大学)
;
Adobe Research(Adobe研究)
;
The University of Hong Kong(香港大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Jiao Tong University(上海交通大学)
专题命中
视频多模态
:multimodal(title);分类 cs.CV
AI总结
OmniVCus通过多模态控制条件和改进的嵌入机制实现高效的多主体视频定制。
CommentsNeurIPS 2025; A data construction pipeline and a diffusion Transformer framework for controllable subject-driven video customization
Designing a Multimodal Viewer for Piano Performance Analysis -- a Pedagogy-First Approach
为钢琴表演分析设计多模态查看器——一种以教学为导向的方法
Joonhyung Bae, Hyeyoon Cho, Kirak Kim, Dawon Park, Taegyun Kwon, Yoon-Seok Choi, Hyeon Hur, Shigeru Kai, Yohei Wada, Satoshi Obata, Akira Maezawa, Jaebum Park, Jonghwa Park, Juhan Nam
机构
*
Indian Institute of Technology Gandhinagar(印度理工学院甘地纳加尔)
;
Indian Institute of Technology Madras(印度理工学院马德拉斯)
;
Dr. A. P. J. Abdul Kalam Technical University(阿卜杜勒·卡拉姆技术大学)
Generative Video Semantic Communication via Multimodal Semantic Fusion with Large Model
Hang Yin, Li Qiao, Yu Ma, Shuo Sun, Kan Li, Zhen Gao, Dusit Niyato
机构
*
School of Information and Electronics, Beijing Institute of Technology(信息与电子学院,北京理工大学)
;
School of Computer Science and Engineering, Nanyang Technological University(计算机科学与工程学院,南洋理工大学)