Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
重新思考大语言模型作为评判器:通过语义能力不对称利用小语言模型进行表示作为评判器
Zhuochun Li, Yong Zhang, Ming Li, Yuelyu Ji, Yiming Zeng, Ning Cheng, Yun Zhu, Yanmeng Wang, Shaojun Wang, Jing Xiao, Daqing He
机构
*
Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司)
;
University of Pittsburgh(匹兹堡大学)
;
University of Maryland, College Park(马里兰大学学院公园分校)
;
University of Connecticut(康涅狄格大学)
AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring
AVA-VLM:用于野外建筑工地监测的自适应视觉注意力视觉语言模型
Younggun Kim, Taeheon Kim, Youngseo Kim, Seunghee Park
机构
*
University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
;
AI+KCIR Global Resilience Research Center(人工智能与韩国气候变化影响韧性研究中心)
;
SmartInside AI Co., Ltd.(思玛特因赛德人工智能有限公司)
;
Sungkyunkwan University(成均馆大学)
URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment
URSA:用于功利性逆合成评估的化学感知基准测试
Bogdan Zagribelnyy, Ivan Ilin, Nikita Bondarev, Anton Morgunov, Arkadii Lin, Maksim Kuznetsov, Rim Shayakhmetov, Vladimir Aladinskiy, Alex Aliper, Alex Zhavoronkov
机构
*
Insilico Medicine AI Limited(Insilico Medicine AI有限公司)
;
Independent researcher(独立研究者)
;
Insilico Medicine Canada Inc.(Insilico Medicine加拿大公司)
;
Insilico Medicine Hong Kong Ltd.(Insilico Medicine香港有限公司)
CausalGame: Benchmarking Causal Thinking of LLM Agents in Games
因果游戏:在游戏中对大语言模型智能体的因果思维进行基准测试
Zhenhao Chen, Yongqiang Chen, Chenxi Liu, Junchi Yu, Xiangchen Song, Zijian Li, Jialin Li, Philip Torr, Bo Han, Kun Zhang
机构
*
MBZUAI(穆罕默德·本·扎耶德人工智能大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
Hong Kong Baptist University(香港浸会大学)
;
University of Oxford(牛津大学)
;
New York University, Abu Dhabi(纽约大学阿布扎比分校)
CommentsZhenhao, Yongqiang, and Chenxi contributed equally to the project. A short version is accepted at the Forty-Third International Conference on Machine Learning (ICML) 2026 as an Oral presentation. Project website https://causalgame.github.io/
Evaluation of Medical Vision Language Models HuluMed and MedGemma, and general purpose chatbots Gemma 3, ChatGPT Plus, and Claude Pro on real previously unseen wound images
医学视觉语言模型 HuluMed 和 MedGemma 以及通用聊天机器人 Gemma 3、ChatGPT Plus 和 Claude Pro 在真实未见伤口图像上的评估
Yunzhe Xue, Mohammed Saim Ahmed Quadri, Neal Panse, Justin W. Ady, Usman Roshan
机构
*
Department of Computer Science, New Jersey Institute of Technology(新泽西理工学院计算机科学系)
;
Vascular and Endovascular Surgery, Robert Wood Johnson Hospital(罗伯特·伍德·约翰逊医院血管外科)
;
Department of Data Science, New Jersey Institute of Technology(新泽西理工学院数据科学系)
专题命中
推理评测
:reasoning(abstract);planning(abstract)
AI总结
本研究评估了六种视觉语言模型在慢性伤口分析任务上的表现,发现通用模型 ChatGPT 和 Claude 显著优于医学专用模型,表明广泛的多模态推理能力比领域知识更重要。
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(信息多媒体国家重点实验室,计算机学院,北京大学)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Institute for Brain and Intelligence, Fudan University(脑与智能研究院,复旦大学)
;
University of Science and Technology Beijing(北京科技大学)
;
Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)
Comments07 pages, 01 figure, accepted for presentation at the IEEE International Conference on Communication, Computing, Networking, and Control in Cyber-Physical Systems (CCNCPS 2026)
From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents
从知道到行动:基准测试LLM代理的自我意识能力
Yifan Li, Shengbin Yue, Boyu Feng, Jinhu Qi, Bo Ke, Zixing Song, Hongru Wang, Zhongyu Wei, Irwin King
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Fudan University(复旦大学)
;
University of Edinburgh(爱丁堡大学)
;
Tencent(腾讯)
;
University of Bristol(布里斯托大学)