I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning
我在视频中寻找你:面向以人为中心的视频推理的身份条件查询
Shibo Gao, Chongxiao Wang, Chenglong Huang, Jie Ma, Haolin Shi, Fei Ding, Jing Li, Qiang Lyu, Yangyang Liu, Yang Liu, Jun Liu, Linlin Huang, Peipei Yang
机构
*
Beijing Jiaotong University(北京交通大学)
;
HUJING Digital Media & Entertainment Group(汇晶数字媒体与娱乐集团)
;
MAIS Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS)
GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models
GraphVerse:面向多模态大语言模型的综合性视觉图推理基准
Yuanfu Sun, Yuanhang Ren, Kang Li, Chuanhao Ji, Jiaxi Li, Jiajin Liu, Ninghao Liu, Qiaoyu Tan
机构
*
New York University(纽约大学)
;
Sensetime Research(商汤科技研究院)
;
Tsinghua University(清华大学)
;
New York University Shanghai(上海纽约大学)
;
University of Georgia(佐治亚大学)
;
The Hong Kong Polytechnic University(香港理工大学)
Comments17 pages, 5 figures, 9 tables. v2 corrects scorer and taxonomy defects, adds no-model baselines showing label leakage, re-runs the Lithuanian cells on de-leaked text, and withdraws the claim that few-shot helps on judgment-form classification everywhere; all tables and figures regenerated. Dataset: https://huggingface.co/datasets/overthelex/multi-legal-bench
ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
ResidencyRL:在模拟临床环境中开展的强化学习
Valentin Liévin, Samuel Schmidgall, Tim Strother, Alex Bijamov, Akshay Goel, Anil Palepu, Chunjong Park, Vahid Balazadeh, Min Woo Sun, Marius Guerard, Justin Chen, Dave Steiner, Vikram Dhillon, Ibrahim Azar, Akhil Mehta, Nicholas Spetsieris, Shilpan Shah, Maen Abdelrahim, Amit Dahiya, Yun Liu, Katherine Chou, Yossi Matias, Avinatan Hassidim, Dale R. Webster, Quoc V. Le, Raia Hadsell, Joelle Barral, Carey Radebaugh, Aleksandra Faust, Shekoofeh Azizi, Mike Schaekermann, Po-Hsuan Cameron Chen, Tao Tu, David Racz, Lin Yang
机构
*
Google DeepMind(谷歌DeepMind)
;
Google Research(谷歌研究院)
;
Houston Methodist Hospital(休斯顿卫理公会医院)
;
Trinity Health Group(三一健康集团)
;
Stanford Oncology Partners(斯坦福肿瘤学伙伴)
;
St. Luke Hospital(圣卢克医院)
专题命中
推理评测
:reasoning(abstract);分类 cs.CL、cs.AI
AI总结
本研究提出 ResidencyRL,通过多轮强化学习训练临床 AI 智能体,在模拟临床环境中提升诊断准确性、降低漏报率,且能力可迁移至多个医学基准测试,为临床 AI 发展提供了新路径。
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
多模态大语言模型(MLLMs)能否解码创造性飞跃?推出面向跨概念理解的C4框架
Ming Wang, Yuqing Zhang, Tingna Xie, Xiangju Li, Xiaocui Yang, Daling Wang, Shi Feng, Yifei Zhang
机构
*
School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
;
School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算与信息系统学院)
;
School of Computer Science and Engineering, Shandong University of Science and Technology(山东科技大学计算机科学与工程学院)
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training
DiDPO:用于编码智能体训练的差异内差异策略优化
Xucong Wang, Zhe Zhao, Liheng Yu, Di Wu, Xiaofeng Cao, Pengkun Wang
机构
*
University of Science and Technology of China (USTC)(中国科学技术大学)
;
Stanford University(斯坦福大学)
;
Suzhou Institute for Advanced Research, USTC(中国科学技术大学苏州高等研究院)
;
Tongji University(同济大学)
机构
*
Alibaba Group(阿里巴巴集团)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tsinghua University(清华大学)
;
University of Alberta(阿尔伯塔大学)
;
Zhejiang University(浙江大学)
Comments16 pages, 9 figures. v2: updated author list (added Yang Liu and Yuxin Li; marked core contributors and project lead) and added a release-date note on the first page
MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pressure in Large Language Models
MEDLEY-BENCH:在AI元认知中规模购买评估但不控制
Farhad Abtahi, Abdolamir Karbalaie, Eduardo Illueca-Fernandez, Fernando Seoane
机构
*
Department of Clinical Science, Intervention and Technology (CLINTEC), Karolinska Institutet(临床科学、干预与技术部门(CLINTEC),Karolinska研究所)
;
Department of Clinical Physiology, Karolinska University Hospital(临床生理学部门,Karolinska大学医院)
;
Department of Biomedical Engineering and Health Systems, KTH Royal Institute of Technology(生物医学工程与健康系统部门,KTH皇家理工学院)
;
Department of Textile Technology, University of Borås(纺织技术部门,Borås大学)
;
Department of Medical Technologies, Karolinska University Hospital(医学技术部门,Karolinska大学医院)
"LLM Agent Performance" Is Not a Single Evaluation Target
基于LLM的智能体评估统一框架的必要性
Pengyu Zhu, Li Sun, Philip S. Yu, Sen Su
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
;
Chongqing University of Posts and Telecommunications(重庆邮电大学)