Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
先关注再注意:通过自回归注视实现高效的可扩展视频理解
Baifeng Shi, Stephanie Fu, Long Lian, Hanrong Ye, David Eigen, Aaron Reite, Boyi Li, Jan Kautz, Song Han, David M. Chan, Pavlo Molchanov, Trevor Darrell, Hongxu Yin
机构
*
Carnegie Mellon University(卡内基梅隆大学)
;
Kyung Hee University(Kyung Hee大学)
;
Korea Advanced Institute of Science and Technology(韩国科学技术院)
;
University of Pittsburgh(匹兹堡大学)
CommentsThis version corrects the author affiliation to reflect the accurate institutional information at the time of publication. No technical content of the paper has been changed
EventFlash: Towards Efficient MLLMs for Event-Based Vision
EventFlash: 向基于事件的视觉高效MLLMs迈进
Shaoyu Liu, Jianing Li, Guanghui Zhao, Yunjian Zhang, Wen Jiang, Ming Li, Xiangyang Ji
机构
*
Xidian University(西安电子科技大学)
;
Tsinghua University(清华大学)
;
Beijing Institute of Technology(北京理工大学)
;
Guangdong Laboratory of Artificial Intelligence and Digital Economy(SZ)(广东人工智能与数字经济实验室(深圳))
GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction
GTPred:评估多模态大语言模型在可解释地理定位和拍摄时间预测中的基准测试
Jinnao Li, Zijian Chen, Tingzhu Chen, Changbo Wang
机构
*
School of Computer Science and Technology, East China Normal University(东华大学计算机科学与技术学院)
;
Institute of Image Communication and Information Processing, Shanghai Jiao Tong University(上海交通大学图像通信与信息处理研究院)
;
School of Humanities, Shanghai Jiao Tong University(上海交通大学人文学院)
;
Shanghai AI Laboratory(上海人工智能实验室)
机构
*
Computer Science Department, University of Crete(塞萨洛尼基大学计算机科学系)
;
Institute of Computer Science (ICS), Foundation for Research & Technology – Hellas (FORTH)(希腊基础研究与技术机构计算机科学研究所)
机构
*
ByteDance Intelligent Creation(字节跳动智能创作)
;
Tsinghua University(清华大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Shanghai Jiao Tong University(上海交通大学)
;
Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
CommentsOur paper was initially titled "Video-SSR1: Self-Supervised Reinforcement Video Reasoning." Upon noticing its close resemblance to the title of a recently released paper, we have decided to rename our work as "ViSS-R1."