Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing
复现与压力测试两种提升大语言模型(LLM)推理可靠性的方法:测试时概率聚合与逻辑表示编辑
机构 * Sungkyunkwan University(成均馆大学)
专题命中 逻辑推理 :reasoning(title,abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 本研究复现了RPC与LCF两种提升LLM推理可靠性的方法,在多任务域与多模型上压力测试,发现RPC优势未达显著且样本量增大后效果逆转,LCF效果弱且无统计显著性。
Comments 16 pages, 3 figures, 9 tables. Code, data, and experiment logs: https://github.com/rabqatab/llm-reasoning-reliability-reproduction