MLaGA: Multimodal Large Language and Graph Assistant
MLaGA: 多模态大语言与图助手
Dongzhe Fan, Yi Fang, Jiajin Liu, Djellel Difallah, Qiaoyu Tan
机构
*
New York University(纽约大学)
;
New York University Shanghai(纽约大学上海)
;
New York University Brooklyn(纽约大学布鲁克林)
;
Virginia Polytechnic Institute and State University(弗吉尼亚理工大学)
;
New York University Abu Dhabi(纽约大学阿布扎克)
Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models
推理下的校准漂移:思维链预算如何导致大型语言模型过度自信
Prakul Sunil Hiremath, Harshit R. Hiremath
机构
*
Department of Computer Science and Engineering, Visvesvaraya Technological University, Belagavi(维斯瓦拉亚科技大学计算机科学与工程系,贝拉加维)
;
Department of Computer Science and Business System, SG Balekundri Institute of Technology, Belagavi(SG巴莱昆德里理工学院计算机科学与商业系统系,贝拉加维)
Comments31 pages, 4 figures, 3 tables. Introduces Calibration Drift Under Reasoning (CDUR) with theoretical analysis and preliminary experiments; includes CABStop; code and data available
OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models
OpenMedReason: 医学视觉语言模型的科学推理监督
Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
机构
*
York University(约克大学)
;
Vector Institute(向量研究所)
;
University of British Columbia(不列颠哥伦比亚大学)
;
University of Toronto(多伦多大学)
;
Unity Health Toronto / St. Michael’s Hospital(多伦多联合健康/圣迈克尔医院)
;
University Health Network(大学健康网络)
;
Arc Institute(弧研究所)
;
Queen's University(女王大学)
Beyond representational alignment with brain-guided language models for robust reasoning
超越表征对齐:基于大脑引导的语言模型实现稳健推理
Mingqing Xiao, Kai Du, Zhouchen Lin
机构
*
State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(北京大学通用人工智能国家重点实验室、智能科学与技术学院)
;
Department of Psychological and Cognitive Sciences, Tsinghua University(清华大学心理与认知科学系)
;
Microsoft Research Asia(微软亚洲研究院)
Comments9 pages main text, 31 pages total (including references and appendix). 5 figures, 16 tables. Preprint under review. Code and data will be made available upon publication
机构
*
School of Computer Science, Chongqing University(重庆大学计算机学院)
;
AI Research Institution, Mashang Financial Institution(马上金融人工智能研究院)
;
Department of Information, Third Military Medical University(陆军军医大学信息系)
机构
*
School of Philosophy (Political Philosophy) Renmin University of China(哲学学院(政治哲学)中国人民大学)
;
School of Government and Policy Johns Hopkins University(政府与政策学院约翰霍普金斯大学)
Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction
Embodied-BenchClaw:用于具身空间智能基准构建的自主多智能体系统
Baoyang Jiang, Fengchun Zhang, Leyuan Wang, Haotian Li, Yida Wang, Zhe Ji, Jinshan Lai, Xi Ren, Jianwei Hu, Qiang Ma
机构
*
QiYuan Lab(启元实验室)
;
School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
;
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
Automated Creativity Evaluation of Language Models Across Open-Ended Tasks
语言模型在开放式任务中的自动化创造力评估
Min Sen Tan, Zachary Kit Chun Choy, Syed Ali Redha Alsagoff, Nadya Yuki Wangsajaya, Mohor Banerjee, Swaagat Bikash Saikia, Alvin Chan
机构
*
Raffles Institution(莱佛士书院)
;
College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院)
;
Lee Kong Chian School of Medicine, Nanyang Technological University(南洋理工大学李光前医学院)
;
Centre of AI in Medicine (C-AIM), Nanyang Technological University(南洋理工大学人工智能医学中心)
Comments65 pages, 3 figures, 5 tables. Reference architecture with a reference implementation of the policy-engine core and microbenchmark results; full-system evaluation identified as future work
Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?
重新评估高性能大语言模型在波兰医学考试中的表现:真实能力还是偏差驱动?
Antoni Lasik, Jakub Pokrywka, Łukasz Grzybowski, Jeremi Ignacy Kaczmarek, Gabriela Korzańska, Janusz Świeczkowski-Feiz, Oskar Pastuszek, Paulina Hoffman, Jakub Tomasz Dąbrowski, Wojciech Kusa
机构
*
NASK National Research Institute(NASK国家研究所)
;
Adam Mickiewicz University(亚当·密茨凯维奇大学)
;
ARAAI Poland(ARAAI波兰)
;
Poznań University of Medical Sciences(波兹南医科大学)
;
Centre of Postgraduate Medical Education, Poland(波兰研究生医学教育中心)
;
T. Marciniak Lower Silesian Specialist Hospital(T. 马尔奇尼亚克下西里西亚专科医院)
;
Medical University of Warsaw(华沙医科大学)
RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark
RAIL: 基于CHC框架重新思考大型音频语言模型中的听觉智能
Hongyu Jin, Siyi Wang, Yang Xiao, Jiaheng Dong, Shihong Tan, Kaiyuan peng, Georgiana Juravle, Shanquan Chen, Gongping Huang, Hong Jia, Eun-Jung Holden, James Bailey, Ting Dang
机构
*
School of Computing and Information Systems, The University of Melbourne(墨尔本大学计算与信息系统学院)
;
Faculty of Psychology and Educational Sciences, Alexandru Ioan Cuza University of Iași(亚历山德鲁伊万库扎大学心理学与教育科学学院)
;
School of Electronic Information, Wuhan University(武汉大学电子信息学院)
;
School of Public Health, The University of Hong Kong(香港大学公共卫生学院)
;
School of Computer Science, The University of Auckland(奥克兰大学计算机科学学院)
;
Department of Data Science and Artificial Intelligence, Monash University(莫纳什大学数据科学与人工智能系)