Conv-to-Bench: Evaluating Language Models Via User-Assistant Dialogues In Code Tasks
Conv-to-Bench: 通过代码任务中的用户-助手对话评估语言模型
Victor M. dos Santos, Andre C. Castro, Samuel L. de S. Toledo, Bruno M. L. Calura, Lisandra C. de M. Menezes, Raul C. R. Mata, Telma W. de L. Soares, Bryan L. M. de Oliveira
机构
*
Institute of Mathematics and Computer Science, University of São Paulo(圣保罗大学数学与计算机科学学院)
;
Institute of Informatics, Federal University of Goiás(戈亚斯联邦大学信息学院)
;
HUG Labs(HUG实验室)
;
Advanced Knowledge Center for Immersive Technologies (AKCIT)(沉浸式技术高级知识中心)
CommentsThis submission is being withdrawn because it was submitted without the knowledge and authorization of all co-authors. The authors need to resolve this authorship/authorization issue before any public posting
Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning
Med-CoReasoner: 通过语言感知的协同推理减少医学推理中的语言差异
Fan Gao, Sherry T. Tong, Jiwoong Sohn, Jiahao Huang, Junfeng Jiang, Ding Xia, Piyalitt Ittichaiwong, Kanyakorn Veerakanjana, Hyunjae Kim, Qingyu Chen, Edison Marrese Taylor, Kazuma Kobayashi, Akiko Aizawa, Irene Li
机构
*
The University of Tokyo(东京大学)
;
ETH Zürich(苏黎世联邦理工学院)
;
National Institute of Informatics(日本信息处理学会)
;
Siriraj Informatics and Data Innovation Center(Siriraj信息与数据创新中心)
;
Yale University(耶鲁大学)
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
FTibSuite:面向藏语视觉语言建模的综合资源套件
Guixian Xu, Yide Liang, Zeli Su, Xuexian Song, Ziyin Zhang, Yushuang Dong, Ting Zhang, Xu Han
机构
*
Hainan International College, Minzu University of China(民族大学海南国际学院)
;
School of Information Engineering, Minzu University of China(民族大学信息工程学院)
;
Shanghai Jiao Tong University(上海交通大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
ODOV: Benchmark the Open-Domain Open-Vocabulary Object Detection
ODOV:开放域开放词汇目标检测基准
Yupeng Zhang, Ruize Han, Fangnan Zhou, Wei Feng, Liang Wan
机构
*
College of Intelligence and Computing, Tianjin University(天津大学智能计算学院)
;
Key Research Center for Surface Monitoring and Analysis of Relics, State Administration of Cultural Heritage(文物表面监测与分析国家重点研究中心)
;
Faculty of Computer Science and Artificial Intelligence, Shenzhen University of Advanced Technology(深圳先进技术大学计算机科学与人工智能学院)
A Technical Policy Blueprint for Trustworthy Decentralized AI
可信去中心化人工智能的技术政策蓝图
Hasan Kassem, Orion Banks, Omar Benjelloun, Sergen Cansiz, Brandon Edwards, Patrick Foley, Inken Hagestedt, Taeho Jung, Peter Kairouz, Marco Lorenzi, Peter Mattson, Prakash Moorthy, Ann K Novakowski, Michael O'Connor, Bruno Rodrigues, Holger Roth, Micah Sheller, Dimitris Stripelis, Renato Umeton, Marc Vesin, Wenbin Zhang, Mic Bowman, Alexandros Karargyris
CommentsThis work appears in the Proceedings of the 43rd International Conference on Machine Learning (ICML 2026) and was selected as an Oral paper at the ICML 2025 DataWorld Workshop
Comments19 pages, 2 figures, 3 main tables; supplementary appendix with 6 tables, 2 figures, and a reproducibility methods section. Describes 17 configured agents in a persistent research environment and introduces the PARE-M (Persistent Agentic Research Environment Measurement) framework
FAST-GOAL: Fast and Efficient Global-local Object Alignment Learning
FAST-GOAL: 快速高效的全局-局部对象对齐学习
Hyungyu Choi, Young Kyun Jang, Chanho Eom
机构
*
Department of Virtual Convergence, Graduate School of Advanced Imaging Science, Multimedia & Films (GSAIM), Chung-Ang University(虚拟融合系,高级影像科学研究生院,多媒体与电影系(GSAIM), Chung-Ang 大学)
The Labyrinth and the Thread: Rethinking Regularizations in Sequential Knowledge Editing for Large Language Models
迷宫与线索:重新思考大语言模型顺序知识编辑中的正则化方法
Zheng Wang, Kaixuan Zhang, Wanfang Chen, Jingwen Zhang, Xiaonan Lu
机构
*
Bosch Center for Artificial Intelligence (BCAI)(博世人工智能中心(BCAI))
;
Bosch (China) Investment Ltd.(博世(中国)投资有限公司)
;
School of Statistics, East China Normal University(东华大学统计学院)
Post-training makes large language models less human-like
后训练使大型语言模型更不像人类
Marcel Binz, Elif Akata, Abdullah Almaatouq, Mohammed Alsobay, Oleksii Ariasov, Franziska Brändle, David Broska, Jason W. Burton, Nuno Busch, Frederick Callaway, Vanessa Cheung, Brian Christian, Julian Coda-Forno, Can Demircan, Vittoria Dentella, Maria K. Eckstein, Noémi Éltető, Michael Franke, Thomas L. Griffiths, Fritz Günther, Susanne Haridi, Sebastian Hellmann, Stefan Herytash, Linus Hof, Eleanor Holton, Isabelle Hoxha, Zak Hussain, Akshay Jagadish, Elif Kara, Valentin Kriegmair, Evelina Leivada, Li Ji-An, Tobias Ludwig, Maximilian Maier, Marcelo G. Mattar, Marvin Mathony, Alireza Modirshanechi, Robin Na, Mariia Nadverniuk, Antonios Nasioulas, Surabhi S. Nath, Helen Niemeyer, Kate Nussenbaum, Sebastian Olschewski, Thorsten Pachur, Stefano Palminteri, Aliona Petrenco, Camille V. Phaneuf-Hadd, Angelo Pirrone, Manuel Rausch, Laura Raveling, Shashank Reddy, Milena Rmus, Evan M. Russek, Tankred Saanum, Kai Sandbrink, Louis Schiekiera, Johannes A. Schubert, Luca M. Schulze Buschoff, Nishad Singhi, Leah H. Somerville, Mikhail S. Spektor, Xin Sui, Christopher Summerfield, Mirko Thalmann, Anna I. Thoma, Taisiia Tikhomirova, Vuong Truong, Polina Tsvilodub, Konstantinos Voudouris, Kristin Witte, Shuchen Wu, Dirk U. Wulff, Hua-Dong Xiong, Songlin Xu, Lance Ying, Xinyu Zhang, Jian-Qiao Zhu, Eric Schulz
机构
*
Helmholtz Munich(海德堡-慕尼黑亥姆霍兹中心)
;
Massachusetts Institute of Technology(麻省理工学院)
;
University of Tübingen(图宾根大学)
;
University of Oxford(牛津大学)
;
Stanford(斯坦福大学)
A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration
一个通用悬崖与一个设计指纹:LLM编排下的跨段缺陷检测
Hiroki Fukui
机构
*
Research Institute of Criminal Psychiatry(刑事精神病研究机构)
;
Sex Offender Medical Center(性犯罪医疗中心)
;
Department of Neuropsychiatry, Graduate School of Medicine, Kyoto University(京都大学医学研究生院神经精神病学部门)
机构
*
Department of Electrical and Computer Engineering, McMaster University(麦基尔大学电气与计算机工程系)
;
Department of Biomedical Engineering, University of Southern California(南加州大学生物医学工程系)
;
Department of Electrical and Computer Engineering, University of Rochester(罗切斯特大学电气与计算机工程系)
;
BIOSEN Group(BIOSEN集团)