A Reasoning-Focused Legal Retrieval Benchmark
机构 * Stanford University(斯坦福大学) ; Princeton University(普林斯顿大学)
专题命中 推理评测 :reasoning(title);分类 cs.CL
Comments CS&Law 2025. For data, see https://reglab.github.io/legal-rag-benchmarks/
AI 大模型
大模型数学、逻辑、规划、多步推理和测试时计算能力。
机构 * Stanford University(斯坦福大学) ; Princeton University(普林斯顿大学)
专题命中 推理评测 :reasoning(title);分类 cs.CL
Comments CS&Law 2025. For data, see https://reglab.github.io/legal-rag-benchmarks/
专题命中 推理评测 :CoT(title);分类 cs.AI
Comments Project Page: https://ucsc-vlaa.github.io/Complex-Edit/, Dataset: https://huggingface.co/datasets/UCSC-VLAA/Complex-Edit
专题命中 推理评测 :reasoning(title);分类 cs.AI
专题命中 推理评测 :reasoning(title);分类 cs.CL
Comments ACL 2021
专题命中 推理评测 :reasoning(title);分类 cs.LG
Comments 29 pages, 16 figures
专题命中 推理评测 :reasoning(title);分类 cs.CL
Comments GenBench Workshop by EMNLP 2024: Camera-ready version
专题命中 推理评测 :reasoning(title);分类 cs.CL
Comments 18 pages, 7 figures, accepted to COLM 2024. Data available here: https://github.com/a-brassard/ACORN
专题命中 推理评测 :reasoning(title);分类 cs.CL
专题命中 推理评测 :reasoning(title);分类 cs.CL
专题命中 推理评测 :reasoning(title);分类 cs.AI
Comments 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)
专题命中 推理评测 :reasoning(title);分类 cs.AI
Comments Preprint with additional experiments
专题命中 推理评测 :reasoning(title);分类 cs.AI
专题命中 推理评测 :reasoning(title);分类 cs.CL
专题命中 推理评测 :reasoning(title);分类 cs.LG
Comments Published on NeurIPS2023
专题命中 推理评测 :reasoning(title);分类 cs.CL
Comments CoNLL 2023 BabyLM Challenge
专题命中 推理评测 :reasoning(title);分类 cs.CL
Comments Accepted to ACL 2023(Short Paper)
专题命中 推理评测 :reasoning(title);分类 cs.CL
专题命中 推理评测 :reasoning(title);分类 cs.CL
Comments EACL 2023
专题命中 推理评测 :reasoning(title);分类 cs.CL
Comments Camera ready, to appear at the Natural Legal Language Processing Workshop 2022 co-located with EMNLP
专题命中 推理评测 :reasoning(title);分类 cs.CL
Comments The 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP 2022)
专题命中 推理评测 :reasoning(title);分类 cs.CL
Comments this is a long paper, the short version was accepted by SemDial 2022
专题命中 推理评测 :reasoning(title);分类 cs.CL
Comments Accepted to EMNLP 2021 Findings
专题命中 推理评测 :reasoning(title);分类 cs.AI
Comments 46 pages, 1 Postscript figure
Journal ref TPLP Vol 3(4&5) (2003) 425-461
基于美国核管理委员会反应堆操作员执照考试的多模态语言模型微调与检索策略基准测试
机构 * organization= Department of Nuclear Engineering, Hanyang University , addressline= 222 Wangsimni-ro , postcode= 04763 , state= Seongdong-gu , city= Seoul , country= South Korea ; organization= The Grainger College of Engineering, Nuclear, Plasma \& Radiological Engineering, University of Illinois Urbana-Champaign , city= Urbana , state= IL , country= USA
专题命中 推理评测 :chain-of-thought(abstract);CoT(abstract_cn);分类 cs.CL、cs.AI
AI总结 该研究针对美国核管理委员会反应堆操作员执照考试,评估310亿参数多模态模型应用核知识的能力,通过对比基础模型与多种微调及检索配置,发现固定大小分块RAG的SFT配置表现最佳,并揭示了分块策略规律及RAFT与SFT的性能差异。
ReLoop:结构建模与行为验证用于可靠的基于LLM的优化
机构 * McCormick School of Engineering, Northwestern University(西北大学工程学院) ; Wenzhou Buyi Pharmacy Chain Co., Ltd.(温州-buyi药链有限公司) ; College of Computer Science and Artificial Intelligence, Wenzhou University(温州大学计算机科学与人工智能学院) ; Department of Decision Analytics and Operations, City University of Hong Kong(香港城市大学决策分析与运营部门) ; Institute of Operations Research and Analytics, National University of Singapore(新加坡国立大学运筹学与分析研究所)
专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG
AI总结 ReLoop通过结构生成和行为验证机制,解决LLM生成优化代码的可行性与正确性差距问题,提升代码执行准确性和鲁棒性。
Comments Code and benchmark: https://github.com/junbolian/ReLoop
LLMs能否像自动定理证明器一样进行Rust验证?VCoT-Bench:通过验证思维链进行评估
机构 * University of Virginia(弗吉尼亚大学)
专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.AI、cs.LG
AI总结 本文提出VCoT-Bench,通过验证思维链评估LLMs在Rust验证中的能力,揭示其在不同证明类型和缺失证明情况下的脆弱性。
Comments Accepted at ICML 2026
Avalon-ToM-Bench:通过非对称游戏机制评估细粒度心理理论
机构 * National Taiwan University(台湾大学) ; CyCraft AI Lab(CyCraft人工智能实验室)
专题命中 推理评测 :reasoning(abstract);chain-of-thought(abstract);分类 cs.CL、cs.AI
AI总结 本文提出Avalon-ToM-Bench基准,基于阿瓦隆游戏机制评估大语言模型的细粒度心理理论,发现模型ToM能力不足源于推理策略而非知识或表征,推理训练比测试时思维链增益更显著。
ComboShoppingBench:评估大型语言模型智能体在带优惠券的预算约束组合购物篮任务中的表现
专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI
AI总结 本研究推出ComboShoppingBench组合购物基准,通过探索智能体、LLM评判者及确定性验证,发现各类LLM智能体在该基准上表现不佳,凸显其在约束感知组合购物上仍有巨大改进空间。
WuYuEval:面向固体废物管理的大语言模型多级基准测试
专题命中 推理评测 :reasoning(abstract);planning(abstract);分类 cs.CL、cs.AI
AI总结 WuYuEval是面向固体废物管理的多级基准测试,含基础与专家模块,评估33个LLM的SWM能力,发现其在专业任务表现差,为开发专用基础模型提供依据。
大型语言模型有多独立?一种用于审计行为纠缠和重加权验证者集合的统计框架
机构 * Texas A&M University(德克萨斯农工大学)
专题命中 推理评测 :reasoning(abstract);verifier(abstract);分类 cs.CL、cs.AI
AI总结 本文提出一种统计框架,用于审计大型语言模型之间的行为纠缠,通过多分辨率层次结构和信息理论度量,分析行为纠缠对验证性能的影响,并通过重加权验证者集合减少相关偏差。