A Model Context Protocol Server for Astrophysical RAG: Unified Access to HI, Dwarf, Globular Cluster, IntZ, and ALPINE Kinematic Corpora with FAISS Semantic Search
CommentsThis revision substantially expands the empirical evaluation to eleven open-weight and three frontier models, adding matched query-cost, team-reward, group-size, and private-share incentive analyses. It also extends the weight-level and GEPA results, frozen-prompt information-structure interventions, statistical uncertainty analyses, and qualitative prompt/reasoning-trace studies
CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks
CoLA: 跨模态低秩适配用于多模态下游任务
Wish Suharitdamrong, Tony Alex, Muhammad Awais, Sara Atito
机构
*
Centre for Vision, Speech and Signal Processing (CVSSP)(视觉、语音和信号处理中心)
;
University of Surrey(塞维利亚大学)
;
Surrey Institute for People-Centred AI(以人为本的人工智能研究所)
Comments136 pages, 12 practical works, preprint. Textbook for senior undergraduates and graduate students. Original contributions on low-resource languages (Tajik, Tatar and other). Companion repository available
机构
*
Tencent HY LLM Frontier(腾讯HY大模型前沿团队)
;
University of Georgia(佐治亚大学)
;
University of Maryland, College Park(马里兰大学帕克分校)
;
University of Pennsylvania(宾夕法尼亚大学)
;
University of Minnesota, Twin Cities(明尼苏达大学双城分校)
;
Indiana University(印第安纳大学)
;
National University of Singapore(新加坡国立大学)
;
Hong Kong Polytechnic University(香港理工大学)
专题命中
后训练与偏好优化
:LLM(abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI、cs.LG
In-context superposition: human-like working memory interference in large language models
类人工作记忆干扰在大语言模型中
Hua-Dong Xiong, Li Ji-An, Jiaqi Huang, Robert C. Wilson, Kwonjoon Lee, Xue-Xin Wei
机构
*
School of Psychological and Brain Sciences, Georgia Tech(佐治亚理工学院心理与脑科学学院)
;
Department of Psychology, New York University(纽约大学心理学系)
;
Department of Cognitive Science, Indiana University Bloomington(印第安纳大学布卢明顿分校认知科学系)
;
Honda Research Institute(本田研究所)
;
Center of Excellence for Computational Cognition, Georgia Tech(佐治亚理工学院计算认知卓越中心)
;
Departments of Neuroscience and Psychology, The University of Texas at Austin(德克萨斯大学奥斯汀分校神经科学和心理学系)
专题命中
长上下文与记忆
:large language model(title,abstract);language model(title,abstract);分类 cs.AI、cs.LG
Human-like fleeting memory improves language learning but impairs reading time prediction in transformer language models
类人短暂记忆提升语言学习但损害变压器语言模型的阅读时间预测
Abishek Thamma, Micha Heilbron
机构
*
University of Amsterdam, Amsterdam Brain and Cognition(阿姆斯特丹大学,阿姆斯特丹脑与认知中心)
;
Vrije Universiteit Amsterdam, Department of Informatics(阿姆斯特丹自由大学,信息学院)
;
Max Planck Institute for Psycholinguistics(马克斯·普朗克心理学语言学研究所)
Commentsv2: Revised after peer review. Accepted for publication in Transactions of the Association for Computational Linguistics v3: Added link to code repository. Code: https://github.com/drhanjones/fmt-llm
Journal refTransactions of the Association for Computational Linguistics 14 (2026) 877-892
Comments12 pages, 4 figures, 9 tables. v2: adds Adam-mini discussion, learning-rate sweeps with repeated seeds for the AdamW and Adafactor baselines, and a tier ablation; corrects the attribution of the perplexity advantage between momentum and the factored estimator. Code and per-run training logs: https://github.com/nuemaan/skewadam
机构
*
Department of Computer Science, Rensselaer Polytechnic Institute(里士满理工学院计算机科学系)
;
Department of Mechanical, Aerospace, and Nuclear Engineering, Rensselaer Polytechnic Institute(里士满理工学院机械、航空航天与核工程系)
;
Department of Computer Science and Engineering, University of California San Diego(加州大学圣地亚哥分校计算机科学与工程系)
;
College of Education, Zhejiang Normal University(浙江师范大学教育学院)
;
School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院)
专题命中
推理与问题求解
:large language model(title,abstract);language model(title,abstract);分类 cs.AI