UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks
UpstreamQA:一个用于视频问答任务显式推理的模块化框架
Jason Nguyen, Ameet Rao, Alexander Chang, Ishaan Kumar, Erin Tan
机构
*
Lincoln North Star High School(林肯北星高中)
;
The Charter School of Wilmington(威尔明顿 charter 学校)
;
Greenwich High School(格林威治高中)
;
Santa Susana High School(圣塔苏娜高中)
;
UC Berkeley(伯克利大学)
机构
*
RMIT University(皇家墨尔本理工大学)
;
Monash University(墨尔本大学)
;
Adelaide University(阿德莱德大学)
;
The University of Hong Kong(香港大学)
;
ESPOL University(ESPOL大学)
机构
*
Department of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong(香港中文大学计算机科学与工程系)
;
SkyReels, Skywork AI(SkyReels与Skywork AI)
;
Independent Researcher(独立研究者)
Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
机器人中的视觉-语言-动作:数据集、基准和数据引擎的综述
Ziyao Wang, Bingying Wang, Hanrong Zhang, Tingting Du, Tianyang Chen, Guoheng Sun, Yexiao He, Zheyu Shen, Wanghao Ye, Ang Li
机构
*
University of Maryland, College Park(马里兰大学学院公园分校)
;
University of Utah(犹他大学)
;
Northeastern University(东北大学)
;
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Ant Group(蚂蚁集团)
;
Visual Geometry Group, University of Oxford(牛津大学视觉几何组)
;
Ningbo Institute of Digital Twin, Eastern Institute of Technology, Ningbo(宁波数字孪生研究院,东部技术研究院,宁波)
;
Zhejiang Key Laboratory of Industrial Intelligence and Digital Twin(浙江工业智能与数字孪生重点实验室)
CommentsAccepted to The 33rd ACM International Conference on Advances in Geographic Information Systems(SIGSPATIAL '25) as a short paper in the Short Paper Track
Journal refIn Proceedings of the 33rd ACM International Conference on Advances in Geographic Information Systems (SIGSPATIAL '25), 2025, pp. 1186-1189