MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos
MMR-V:未言明的是什么?视频中多模态深度推理的基准测试
Kejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li, Shangqing Tu, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao
机构
*
The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院,北京,中国)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)
;
Tsinghua University(清华大学)
Enhancing LLMs' Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry
利用全国脓毒症登记处的真实世界数据增强大语言模型的临床推理能力
Junu Kim, Chaeeun Shim, Sungjin Park, Su Yeon Lee, Gee Young Suh, Chae-Man Lim, Seong Jin Choi, Song Mi Moon, Kyoung-Ho Song, Eu Suk Kim, Hong Bin Kim, Sejoong Kim, Chami Im, Dong-Wan Kang, Yong Soo Kim, Hee-Joon Bae, Sung Yoon Lim, Han-Gil Jeong, Edward Choi
机构
*
Korea Advanced Institute of Science and Technology(韩国科学技术院)
;
Microsoft(微软)
;
Asan Medical Center, University of Ulsan College of Medicine(釜山大学医学院阿桑医疗中心)
;
Samsung Medical Center, Sungkyunkwan University School of Medicine(成均馆大学医学院三星医疗中心)
;
Seoul National University Bundang Hospital, Seoul National University College of Medicine(首尔国立大学医学院首尔国立大学医院)
PEARL: Auditable Repair for Scientific Reasoning Graph Extraction
PEARL:科学推理图提取的可审计修复
Bohan Su, Pengze Li, Yuchen Lu, Xi Chen
机构
*
School of Computer Science, Wuhan University(武汉大学计算机科学学院)
;
Artificial Intelligence Innovation and Incubation, Fudan University(复旦大学人工智能创新与孵化)
;
Department of Computer Science, Faculty of Science, University of Bath(巴斯大学理学院计算机科学系)
;
College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)
Orientation Reading by Production Vision-Language Models on Optotype Charts: A Controlled Multi-Model Evaluation Across Reasoning Modes, Prompts, and Access Modalities
Benchmarking Resource-Efficient LLMs for Research Topic Ontology Generation in the Biomedical Field
用于生物医学领域研究主题本体生成的资源高效语言模型基准测试
Tanay Aggarwal, Angelo Salatino, Francesco Osborne, Enrico Motta
机构
*
Knowledge Media Institute, The Open University, Milton Keynes, UK(开放大学知识媒体研究所)
;
Department of Business and Law, University of Milano-Bicocca, Milan, IT(米兰-比科卡大学商业与法律系)
Comments20 pages, 1 figure, 9 tables. v2 adds Round 2: Russian-market coding agents (SourceCraft CLI, Koda CLI), Antigravity with Gemini 3.1 Pro / 3.5 Flash, and Codex CLI with GPT-5.6 on the same frozen task set, plus a tool-call contamination re-audit (network + disk layers). Data, full trajectories and harness: https://github.com/eugeneshilow/rubench
机构
*
Department of Chemical and Materials Engineering, New Mexico State University(新墨西哥州立大学化学与材料工程系)
;
Department of Electrical and Communications Engineering, New Jersey Institute of Technology(新泽西理工学院电气与通信工程系)
机构
*
Independent Researchers(独立研究者)
;
Tencent(腾讯)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Tsinghua University(清华大学)
FETS Benchmark: Foundation Models Enable Scalable and Generalizable Energy Time Series Forecasting
FETS基准:基础模型在能源时间序列预测中优于数据集特定的机器学习
Marco Obermeier, Marco Pruckner, Florian Haselbeck, Andreas Zeiselmair
机构
*
Julius-Maximilians-Universität Würzburg, Modeling and Simulation Lab(乌尔姆-马克斯·普朗克大学,建模与仿真实验室)
;
Weihenstephan-Triesdorf University of Applied Sciences, Smart Farming(魏因施泰因-特里尔夫应用科学大学,智能农业)
;
Weihenstephan-Triesdorf University of Applied Sciences, Digital Energy Transition(魏因施泰因-特里尔夫应用科学大学,数字能源转型)