CommentsFine-tuned Flan-T5-xl outperforms the top #1 results of transformer-based classifier in RuSentNE-2023 competition, to appear in Lobachevskii Journal of Mathematics No.8/2024 proceedings
Interpretable Cross-Lingual Alignment in Small Language Models: Probing Cultural and Pragmatic Reasoning in Japanese-English Bilingual LLMs
小语言模型中的可解释跨语言对齐:探究日英双语大语言模型的文化与语用推理
Florian Braun
机构
*
Sakaguchi–Inui Laboratory (Tohoku University)(东北大学坂口–乾实验室)
;
Natural Language Understanding Team at RIKEN AIP(理化学研究所先进智能项目中心自然语言理解团队)
;
Swallow Project at the Institute of Science Tokyo(东京科学大学Swallow项目)
;
Sakana AI
;
Tohoku University(东北大学)
;
RIKEN AIP(理化学研究所先进智能项目中心)
;
Institute of Science Tokyo(东京科学大学)
Comments17 pages, 5 figures, 9 tables. v2 corrects scorer and taxonomy defects, adds no-model baselines showing label leakage, re-runs the Lithuanian cells on de-leaked text, and withdraws the claim that few-shot helps on judgment-form classification everywhere; all tables and figures regenerated. Dataset: https://huggingface.co/datasets/overthelex/multi-legal-bench
机构
*
University of New South Wales, NSW, Sydney, Australia(新南威尔士大学,新州,悉尼,澳大利亚)
;
Suzhou Institute for Advanced Research, University of Science(苏州先进研究院,科学大学)
;
Shanghai Artificial Intelligence Laboratory, Shanghai, China(上海人工智能实验室,上海,中国)
;
Cornell University(康奈尔大学)
专题命中
推理评测
:reasoning(title);分类 cs.CL、cs.AI
AI总结
提出AtomWorld基准,通过十种基本原子结构操作评估LLM在材料科学中的空间推理能力,发现Claude Opus 4.6表现最佳但复杂空间关系操作成功率低,表明LLM更适合作为辅助工具而非完全自主的科研代理。
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
EconCausal: 面向大语言模型的上下文感知经济推理基准
Donggyu Lee, Hyeok Yun, Meeyoung Cha, Sungwon Park, Sangyoon Park, Jihee Kim
机构
*
Graduate School of Data Science, KAIST(韩国科学技术院数据科学研究生院)
;
College of Business, KAIST(韩国科学技术院商学院)
;
Data Science for Humanity Group, MPI-SP(马克斯·普朗克所际数据科学为人类集团)
;
School of Computing, KAIST(韩国科学技术院计算学院)
;
Division of Social Science, HKUST(香港科技大学社会科学系)