机构
*
State Key Laboratory of Automotive Simulation and Control, Jilin University(吉林大学汽车仿真与控制国家重点实验室)
;
Center of Research for Cyber Security and Network (CSNET), Faculty of Computer Science and Information Technology, Universiti Malaya(马来西亚大学计算机科学与信息技术学院网络安全与网络研究中心)
;
Interdisciplinary Centre for Security, Reliability and Trust (SnT), University of Luxembourg(卢森堡大学安全、可靠性与信任跨学科中心)
;
IMT, Department of Humanities and Technology, Roskilde University(罗斯基勒大学人文与技术系IMT)
;
Center for Security, Theory and Algorithmic Research, International Institute of Information Technology(国际信息技术研究所安全、理论与算法研究中心)
;
Department of Computer Science and Engineering, College of Informatics, Korea University(韩国大学信息学院计算机科学与工程系)
专题命中
推理与问题求解
:large language model(title,abstract);language model(title,abstract);LLM(abstract)
Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models
推理下的校准漂移:思维链预算如何导致大型语言模型过度自信
Prakul Sunil Hiremath, Harshit R. Hiremath
机构
*
Department of Computer Science and Engineering, Visvesvaraya Technological University, Belagavi(维斯瓦拉亚科技大学计算机科学与工程系,贝拉加维)
;
Department of Computer Science and Business System, SG Balekundri Institute of Technology, Belagavi(SG巴莱昆德里理工学院计算机科学与商业系统系,贝拉加维)
专题命中
推理与问题求解
:large language model(title,abstract);language model(title,abstract);分类 cs.CL、cs.AI、cs.LG
Comments31 pages, 4 figures, 3 tables. Introduces Calibration Drift Under Reasoning (CDUR) with theoretical analysis and preliminary experiments; includes CABStop; code and data available
机构
*
Department of Computer Science and Technology, Jilin University(吉林大学计算机科学与技术学院)
;
Department of Computer Science, National Taiwan University(国立台湾大学计算机科学系)
专题命中
推理与问题求解
:large language model(title,abstract);language model(title,abstract);prompting(abstract)
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
多样性越低,安全性越差:大规模语言模型测试时缩放的间接但广泛的风险
Shahriar Kabir Nahin, Hadi Askari, Muhao Chen, Anshuman Chhabra
机构
*
Bellini College of AI, Cybersecurity, and Computing(人工智能、网络安全与计算学院)
;
University of South Florida, Tampa, Florida, USA(佛罗里达州塔帕斯大学)
;
University of California, Davis, California, USA(加州大学戴维斯分校)
专题命中
推理与问题求解
:large language model(title);language model(title);LLM(abstract,abstract_cn);分类 cs.CL、cs.AI、cs.LG
MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding
MLLM-HWSI: 一种用于分层全滑动图像理解的多模态大语言模型
Basit Alawode, Arif Mahmood, Muaz Khalifa Al-Radi, Shahad Albastaki, Asim Khan, Muhammad Bilal, Moshira Ali Abdalla, Mohammed Bennamoun, Sajid Javed
机构
*
Department of Computer Science, Khalifa University of Science and Technology(卡利法科技大学计算机科学系)
;
Information Technology University(信息技术大学)
;
KAU(卡乌大学)
;
University of the Western Australia(西澳大学)
专题命中
推理与问题求解
:large language model(title,abstract);language model(title,abstract);LLM(abstract)
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
LLM概率集中:对齐如何缩小生成范围
Chenghao Yang, Sida Li, Ari Holtzman
专题命中
推理与问题求解
:LLM(title,abstract);large language model(abstract);language model(abstract);prompting(abstract)
AI总结
研究发现对齐微调通过减少生成多样性,使LLM生成更一致,从而影响复杂推理稳定性。
CommentsCodebase: https://github.com/yangalan123/LLMBranchingFactor. V3: Significantly rewrite the whole paper for a clearer structure. Correct problems in the theory parts (Remove emphasis on AEP, discussions on variable LLM generation lengths) and strengthen asymptotic analysis. Add Qwen and OLMo2 experiments. Preliminary SFT v.s. RL comparison to better understand the alignment effects on BF