From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States
从方向到大小:多模态指令微调如何在Transformer隐藏状态中重新组织身份指定提示的几何编码
Jorge A. Castillo, Marco Torres Yévenes, Juan Carlos Lanas
Comments18 pages, 7 tables. v2: corrections to the classical sources (Quintilian, Cicero) and to several cited figures, and Appendix B corpus statistics aligned to the delivered dataset; measurements, results and conclusions unchanged. Data, code, and the trained LoRA adapter: https://federicoboggia.binatomy.com/pubblicazioni/
Gauge dependence and structured-output corruption in sign-branched repetition penalties: measurements across models, inference stacks, and alternative repetition controls
From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
从行为到理解:LLM代理中时间概念的符合可解释性
Trilok Padhi, Ramneet Kaur, Krishiv Agarwal, Adam D. Cobb, Daniel Elenius, Manoj Acharya, Colin Samplawski, Alexander M. Berenbeim, Nathaniel D. Bastian, Susmit Jha, Ugur Kursuncu, Anirban Roy
机构
*
Georgia State University(佐治亚州立大学)
;
Computer Science Lab, SRI(SRI计算机科学实验室)
;
Army Cyber Institute(陆军网络学院)
;
United States Military Academy(美国军事学院)
;
University of Florida(佛罗里达大学)
Aligning Language Model Benchmarks with Pairwise Preferences
将语言模型基准与成对偏好对齐
Marco Gutierrez, Xinyi Leng, Hannah Cyberey, Jonathan Richard Schwarz, Ahmed Alaa, Thomas Hartvigsen
机构
*
School of Data Science, University of Virginia(弗吉尼亚大学数据科学学院)
;
Imperial College London(伦敦帝国理工学院)
;
Thomson Reuters Foundational Research(汤姆森路透基础研究)
;
Department of Electrical Engineering and Computer Science, UC Berkeley and UCSF(伯克利大学电气工程与计算机科学系及旧金山大学)
Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning
分布偏好优化:大语言模型遗忘的细粒度视角
Kai Qin, Jiaqi Wu, Jianxiang He, Haoyuan Sun, Yifei Zhao, Xu Wang, Bin Liang, Yongzhe Chang, Cheng Li, Tiantian Zhang, Houde Liu
机构
*
Tsinghua University(清华大学)
;
The Hong Kong University of Science and Technology(香港科技大学)
;
Jianghuai Advanced Technology Center(江淮先进技术中心)
;
University of Technology Sydney(悉尼大学)
Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact
大语言模型的心理特征很大程度上是测量假象
Jelena Meyer, David Garcia, Dirk U. Wulff
机构
*
Max Planck Institute for Human Development(马克斯·普朗克人类发展研究所)
;
University of Konstanz(康斯坦茨大学)
;
Barcelona Supercomputing Center(巴塞罗那超级计算中心)
;
University of Basel(巴塞尔大学)