Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
通过VLM引导的迭代自优化提升物理导向的视频生成
Yang Liu, Xilin Zhao, Peisong Wen, Siran Dai, Qingming Huang
机构
*
School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)
;
School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
机构
*
Institute of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)
;
College of AI, Tsinghua University(清华大学人工智能学院)
;
Shanghai Qi Zhi Institute(上海启智研究所)
;
State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室)
机构
*
South China Normal University(华南师范大学)
;
Nanyang Technological University(南洋理工大学)
;
City University of Hong Kong(香港城市大学)
;
Shenzhen Academy of Inspection and Quarantine(深圳检验检疫局)
;
Shenzhen University(深圳大学)
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
Manasi Sharma, Chen Bo Calvin Zhang, Chaithanya Bandi, Clinton Wang, Ankit Aich, Huy Nghiem, Tahseen Rabbani, Ye Htet, Brian Jang, Sumana Basu, Aishwarya Balwani, Denis Peskoff, Marcos Ayestaran, Sean M. Hendryx, Brad Kenstler, Bing Liu
机构
*
Scale AI
;
University of Maryland(马里兰大学)
;
University of Chicago(芝加哥大学)
;
Washington University, St. Louis(圣路易斯华盛顿大学)
;
McGill University(麦吉尔大学)
;
University of California, Berkeley(加州大学伯克利分校)
On the generalization of language models from in-context learning and finetuning: a controlled study
Andrew K. Lampinen, Arslan Chaudhry, Stephanie C. Y. Chan, Cody Wild, Diane Wan, Alex Ku, Jörg Bornschein, Razvan Pascanu, Murray Shanahan, James L. McClelland
机构
*
Google DeepMind(谷歌DeepMind)
;
Google DeepMind & Stanford University(谷歌DeepMind与斯坦福大学)
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Peking University(北京大学)
;
Tencent(腾讯)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Tsinghua University(清华大学)
;
Nanyang Technological University(南洋理工大学)
;
Westlake University(西湖大学)
Multimodal Fusion with LLMs for Engagement Prediction in Natural Conversation
Cheng Charles Ma, Kevin Hyekang Joo, Alexandria K. Vail, Sunreeta Bhattacharya, Álvaro Fernández García, Kailana Baker-Matsuoka, Sheryl Mathew, Lori L. Holt, Fernando De la Torre
机构
*
Computer Science Department, Carnegie Mellon University(卡内基梅隆大学计算机科学系)
;
Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)
;
Neuroscience Institute, Carnegie Mellon University(卡内基梅隆大学神经科学研究所)
;
Department of Psychology, The University of Texas at Austin(德克萨斯大学奥斯汀分校心理学系)
;
Center for Perceptual Systems, The University of Texas at Austin(德克萨斯大学奥斯汀分校感知系统中心)
MedXplain-VQA: Multi-Component Explainable Medical Visual Question Answering
Hai-Dang Nguyen, Minh-Anh Dang, Minh-Tan Le, Minh-Tuan Le
机构
*
Faculty Of Information Technology VNU University of Engineering(信息科技学院越南工程大学)
;
IT-BT Convergence Technology Division Vietnam-Korea Institute of Science(IT-BT融合技术部门越南-韩国科学技术院)
;
TADI Global Lab TADI Global Company Limited(TADI全球实验室TADI全球有限公司)
;
Faculty of Finance Banking Academy of Vietnam(金融学院越南银行学院)
Peering Inside the Black Box: Uncovering LLM Errors in Optimization Modelling through Component-Level Evaluation
Dania Refai, Moataz Ahmed
机构
*
Computer Science Department, KFUPM, Dhahran 31261, Saudi Arabia
;
SDAIA-KFUPM Joint Research Center for Artificial Intelligence, King Fahd University of Petroleum \& Minerals, Dhahran, 31261, Saudi Arabia