A Benchmark and Multi-Agent System for Instruction-driven Cinematic Video Compilation
一个用于指令驱动电影视频编译的基准和多智能体系统
Peixuan Zhang, Chang Zhou, Ziyuan Zhang, Hualuo Liu, Chunjie Zhang, Jingqi Liu, Xiaohui Zhou, Xi Chen, Shuchen Weng, Si Li, Boxin Shi
机构
*
School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院)
;
AI Technology Center, Online Video Business Unit, Tencent PCG(腾讯PCG在线视频事业部AI技术中心)
;
Tsinghua University(清华大学)
;
Beijing Academy of Artificial Intelligence(北京智源人工智能研究院)
;
State Key Lab of Multimedia Info. Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室)
;
Nat’l Eng. Research Ctr. of Visual Technology, School of Computer Science, Peking University(北京大学计算机学院国家视觉技术工程研究中心)
;
School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)
机构
*
Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
University of Science and Technology of China(中国科学技术大学)
;
Peking University(北京大学)
Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training
数据混合代理:学习重新加权领域以实现持续预训练
Kailai Yang, Xiao Liu, Lei Ji, Hao Li, Xiao Liang, Zhiwei Liu, Yeyun Gong, Peng Cheng, Mao Yang
机构
*
The University of Manchester(曼彻斯特大学)
;
Microsoft Research(微软研究院)
;
Imperial College London(伦敦帝国学院)
;
University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
EmbodiedGovBench: A Benchmark for Governance, Recovery, and Upgrade Safety in Embodied Agent Systems
EmbodiedGovBench: 一种用于具身代理系统治理、恢复和升级安全性的基准
Xue Qin, Simin Luan, John See, Cong Yang, Zhijun Li
机构
*
School of Software, Harbin Institute of Technology(哈尔滨工业大学软件学院)
;
School of Computer Science and Technology, Harbin Institute of Technology(哈尔滨工业大学计算机科学与技术学院)
;
School of Mathematical and Computer Sciences, Heriot-Watt University, Malaysia Campus(赫瑞-瓦特大学马来西亚校区数学与计算机科学学院)
;
School of Future Science and Engineering, Soochow University(苏州大学未来科学与工程学院)
A Dual-Positive Monotone Parameterization for Multi-Segment Bids and a Validity Assessment Framework for Reinforcement Learning Agent-based Simulation of Electricity Markets
多段报价的双正单调参数化及强化学习代理仿真电力市场的有效性评估框架
Zunnan Xu, Zhaoxia Jing, Zhanhua Pan
机构
*
School of Electric Power Engineering, South China University of Technology(华南理工大学电力学院)
;
Department of Engineering, University of Exeter(埃克塞特大学工程学院)
LABBench2: An Improved Benchmark for AI Systems Performing Biology Research
LABBench2:一种改进的AI系统进行生物研究的基准测试
Jon M Laurent, Albert Bou, Michael Pieler, Conor Igoe, Alex Andonian, Siddharth Narayanan, James Braza, Alexandros Sanchez Vassopoulos, Jacob L Steenwyk, Blake Lash, Andrew D White, Samuel G Rodriques
机构
*
Broad Institute, Cambridge, MA, USA(博德研究所,美国马萨诸塞州剑桥市)
If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs
如果一个大语言模型是一个角色,它会知道自己故事吗?评估大语言模型的终身学习
Siqi Fan, Xiusheng Huang, Yiqun Yao, Xuezhi Fang, Kang Liu, Peng Han, Shuo Shang, Aixin Sun, Yequan Wang
机构
*
University of Electronic Science and Technology of China(电子科技大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所复杂系统认知与决策智能重点实验室)
;
Nanyang Technological University(南洋理工大学)
;
Peking University(北京大学)
;
Spin Matrix, China(中国Spin Matrix公司)
EVGeoQA: Benchmarking LLMs on Dynamic, Multi-Objective Geo-Spatial Exploration
EVGeoQA:基于动态多目标地理空间探索的LLM基准测试
Jianfei Wu, Zhichun Wang, Zhensheng Wang, Zhiyu He
机构
*
School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)
;
Beijing Key Laboratory of Artificial Intelligence for Education(北京市教育人工智能重点实验室)
;
Engineering Research Center of Intelligent Technology and Educational Application, Ministry of Education(教育部智能技术与教育应用工程研究中心)
;
College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机学院)