Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling
通过输出预填充改进多选问答中的LLM首词预测
Silvia Cappelletti, Tobia Poppi, Samuele Poppi, Zheng-Xin Yong, Diego Garcia-Olano, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
机构
*
University of Modena and Reggio Emilia(摩德纳大学与雷焦艾米利亚大学)
;
University of Pisa(比萨大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Brown University(布朗大学)
;
Meta
专题命中
评测与基准
:LLM(title);large language model(abstract);language model(abstract);分类 cs.CL
CommentsAccepted to FSE 2026; An anonymous link containing the dataset, construction scripts, and experimental code is publicly available for reproducibility: https://figshare.com/s/4f202bc0921e26b41dc2
High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination
高波动性和行动偏差区分LLMs与人类在群体协调中的表现
Sahaj Singh Maini, Robert L. Goldstone, Zoran Tiganj
机构
*
Department of Computer Science, Indiana University Bloomington(印第安纳大学布卢明顿分校计算机科学系)
;
Cognitive Science Program, Indiana University Bloomington(印第安纳大学布卢明顿分校认知科学项目)
;
Department of Psychological and Brain Sciences, Indiana University Bloomington(印第安纳大学布卢明顿分校心理与脑科学系)
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
Task-Guided Prompting for Unified Remote Sensing Image Restoration
基于任务引导的统一遥感图像修复
Wenli Huang, Yang Wu, Xiaomeng Xin, Zhihong Liu, Jinjun Wang, Ye Deng
机构
*
School of Electronic and Information Engineering, Ningbo University of Technology(宁波工程学院电子与信息工程学院)
;
University of Exeter(埃克塞特大学)
;
Engineering Research Center of Intelligent Finance, Ministry of Education(教育部智能金融工程研究中心)
;
School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics(西南财经大学计算机与人工智能学院)
Vibe Coding XR: Accelerating AI + XR Prototyping with XR Blocks and Gemini
Vibe Coding XR: 通过XR Blocks和Gemini加速AI+XR原型设计
Ruofei Du, Benjamin Hersh, David Li, Nels Numan, Xun Qian, Yanhe Chen, Zhongyi Zhou, Xingyue Chen, Jiahao Ren, Robert Timothy Bettridge, Xiang 'Anthony' Chen, Faraz Faruqi, Steve Toh, David Kim
专题命中
评测与基准
:LLM(abstract);large language model(abstract);language model(abstract)
机构
*
University of Waterloo(滑铁卢大学)
;
University of Toronto(多伦多大学)
;
HKUST(香港科技大学)
;
Shanghai University(上海大学)
;
Independent Contributor(独立贡献者)
;
Vector Institute(向量研究所)
;
University of British Columbia(不列颠哥伦比亚大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
PolyReal: A Benchmark for Real-World Polymer Science Workflows
PolyReal:一个用于现实世界聚合物科学工作流程的基准
Wanhao Liu, Weida Wang, Jiaqing Xie, Suorong Yang, Jue Wang, Benteng Chen, Guangtao Mei, Zonglin Yang, Shufei Zhang, Yuchun Mo, Lang Cheng, Jin Zeng, Houqiang Li, Wanli Ouyang, Yuqiang Li
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Fudan University(复旦大学)
;
Northwestern Polytechnical University(西北工业大学)
;
Tongji University(同济大学)
;
The University of Hong Kong(香港大学)
;
National University of Singapore(新加坡国立大学)
;
Nanyang Technological University(南洋理工大学)
专题命中
评测与基准
:large language model(abstract);language model(abstract)
Zheng-Xin Yong, Parv Mahajan, Andy Wang, Ida Caspary, Yernat Yestekov, Zora Che, Mosh Levy, Elle Najt, Dennis Murphy, Prashant Kulkarni, Lev McKinney, Kei Nishimura-Gasparian, Ram Potham, Aengus Lynch, Michael L. Chen
机构
*
Constellation
;
Anthropic Fellows Program(Anthropic研究员项目)
;
Brown University(布朗大学)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
;
Imperial College London(伦敦帝国学院)
;
University of Maryland, College Park(马里兰大学帕克分校)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Bar Ilan University(巴伊兰大学)
;
University of Toronto(多伦多大学)
;
University of Oxford(牛津大学)
专题命中
评测与基准
:LLM(abstract);分类 cs.CL、cs.AI
AI总结
Kimi K2.5作为开源大模型,在安全评估中显示出潜在风险,包括CBRNE滥用、网络安全漏洞及政治偏见,但其在拒绝恶意请求方面表现较弱,凸显开源模型的安全挑战。