TraMP-LLaMA: Generative Interpretability with Decoupled Instruction Tuning for Facial Expression Quality Assessment
TraMP-LLaMA: 基于解耦指令微调的生成式可解释性用于面部表情质量评估
Shuchao Duan, Alan Whone, Hossein Rahmani, Jun Liu, Majid Mirmehdi
机构
*
School of Computer Science University of Bristol(布里斯托大学计算机科学学院)
;
Translational Health Sciences University of Bristol(布里斯托大学转化健康科学学院)
;
School of Computing and Communications Lancaster University(兰卡斯特大学计算与通信学院)
机构
*
Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究所)
;
Robbyant, Ant Group(蚂蚁集团 Robbyant)
;
Hongkong University of Science and Technology(香港科技大学)
Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction
在预测之前想象:用于视频事件预测的交错潜在视觉推理
Tianxiang Jiang, Linquan Wu, Sheng Xia, Songze Li, Ziang Yan, Haoyu Yang, Yu Qiao, Yi Wang
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
City University of Hong Kong(香港城市大学)
;
Nanjing University(南京大学)
;
Fudan University(复旦大学)
;
Zhejiang University(浙江大学)
;
University of Electronic Science and Technology of China(电子科技大学)
Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation
Mamba增强的隐式运动学习用于音频驱动肖像动画
Xuan Wei, Jiahui Chen, Kaiheng Li, Mingyu Shao, Qingqi Hong
机构
*
Fujian Provincial Natural Science Foundation of China(福建省自然科学基金委员会)
;
Giant Interactive Group Inc.(巨匠互动集团有限公司)
;
National Natural Science Foundation of China(国家自然科学基金委员会)
机构
*
University of the Chinese Academy of Sciences(中国科学院大学)
;
National University of Singapore(新加坡国立大学)
;
Zhejiang University(浙江大学)
;
State Key Laboratory of Communication Content Cognition, People’s Daily Online(人民日報網通信內容認知重點實驗室)
High-Speed Vision Improves Zero-Shot Semantic Understanding of Human Actions
高速视觉提升人类动作的零样本语义理解
Yongpeng Cao, Yuji Yamakawa
机构
*
Institute of Industrial Science, The University of Tokyo(东京大学工业科学研究所)
;
Interfaculty Initiative in Information Studies, Graduate School of Interdisciplinary Information Studies, The University of Tokyo(东京大学跨学科信息研究 graduate school)
Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation
Prompt-to-Gesture:测量图像到视频指称手势生成的能力
Hassan Ali, Doreen Jirak, Luca Müller, Stefan Wermter
机构
*
Knowledge Technology Group, Department of Informatics, University of Hamburg(汉堡大学信息学院知识技术组)
;
Behavioral Lab, Department of Product Development, University of Antwerp(安特卫普大学产品开发学院行为实验室)
机构
*
Department of Electrical and Computer Engineering, Northeastern University(东北大学电气与计算机工程系)
;
Khoury College of Computer Science, Northeastern University(东北大学科赫里计算机科学学院)