CommentsAccepted for publication in the Proceedings of the 30th International Conference on Knowledge-Based and Intelligent Information & Engineering Systems (KES 2026)
机构
*
School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen, Guangdong, China(哈尔滨工业大学(深圳)计算机科学与技术学院)
;
School of Artificial Intelligence and Computer Science, Jiangnan University, Wuxi, China(江南大学人工智能与计算机学院)
;
Guangdong Provincial Key Laboratory of Intelligent Information Processing(广东省智能信息处理重点实验室)
;
Pengcheng Laboratory(鹏城实验室)
;
Chinese People’s Liberation Army General Hospital, Beijing, China(中国人民解放军总医院)
BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts
BenSyc: 孟加拉语上下文中大语言模型对话谄媚与人类对齐的基准测试
Kazi Noshin, Sajib Acharjee Dip, Ranat Das Prangon, Fardin Hassan Tamim, Syed Ishtiaque Ahmed, Liqing Zhang, Sharifa Sultana
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Virginia Tech(弗吉尼亚理工大学)
;
Bangladesh University of Engineering and Technology(孟加拉工程与技术大学)
;
BRAC University(BRAC大学)
;
University of Toronto(多伦多大学)
PreAct-Bench: Benchmarking Predictive Monitoring in LLMs
PreAct-Bench:大语言模型中的预测性监控基准
Hainiu Xu, Italo Luis da Silva, Jiangnan Ye, Yuhao Wang, Wei Liu, Linyi Yang, Jonathan Richard Schwarz, Nicola Paoletti, Yulan He, Hanqi Yan
机构
*
King’s College London(伦敦国王学院)
;
National University of Singapore(新加坡国立大学)
;
Southern University of Science and Technology(南方科技大学)
;
Thomson Reuters Foundational Research(汤姆森路透基础研究)
;
Imperial College London(伦敦帝国学院)
;
The Alan Turing Institute(艾伦·图灵研究所)
CommentsICAHS, \c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works
机构
*
Department of Computer Science and Information Engineering, National Taiwan University(国立台湾大学计算机科学与资讯工程系)
;
National Taiwan University AI Center of Research Excellence(国立台湾大学人工智能研究中心)
Conditional Vendi Score: Prompt-Aware Diversity Evaluation for Generative AI Models and LLMs
条件 Vendi 分数:生成式 AI 模型和 LLM 的提示感知多样性评估
Mohammad Jalali, Azim Ospanov, Amin Gohari, Farzan Farnia
机构
*
Department of Computer Science and Engineering, The Chinese University of Hong Kong(计算机科学与工程系,香港中文大学)
;
Department of Information Engineering, The Chinese University of Hong Kong(信息工程系,香港中文大学)
专题命中
安全评测
:alignment(abstract);分类 cs.AI、cs.LG
AI总结
针对文本提示引导的生成模型,提出条件 Vendi 和条件 RKE 分数,通过条件熵分离模型自身多样性,并证明收敛性及在多个任务中恢复真实多样性排序。
Earth-OneVision: Extending Remote Sensing Multimodal Large Language Models to More Sensor Modalities and Tasks
Earth-OneVision:将遥感多模态大语言模型扩展到更多传感器模态和任务
Miaoxin Cai, Guanqun Wang, Wei Zhang, Guangyao Zhou, Yin Zhuang, Tong Zhang, Hao Wang, He Chen, Jun Li
机构
*
National Key Laboratory of Science and Technology on Space-Born Intelligent Information Processing (SBIIP), Beijing Institute of Technology(北京理工大学空间智能信息处理国家重点实验室)
;
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院空天信息创新研究院)
;
Key Laboratory of Technology in Geo-Spatial Information Processing and Application System, Chinese Academy of Sciences(中国科学院地理空间信息处理与应用系统技术重点实验室)
;
Advanced Research Institute of Multidisciplinary Sciences, Beijing Institute of Technology(北京理工大学前沿交叉科学研究院)
;
School of Mechatronical Engineering, Beijing Institute of Technology(北京理工大学机电学院)
;
School of Earth and Space Sciences, Peking University(北京大学地球与空间科学学院)
;
School of Electronics, Peking University(北京大学电子学院)
;
School of Computer Science and Hubei Key Laboratory of Intelligent Geo-Information Processing(华中科技大学计算机科学与技术学院&湖北省智能地理信息处理重点实验室)
Sim2Schedule: A Simulator-Guided LLM Framework for Autonomous Open-Pit Mine Scheduling
Sim2Schedule: 一种模拟器引导的LLM框架用于自主露天矿调度
Mustavi Ibne Masum, Thiago Eustaquio Alves de Oliveira, Mahzabeen Emu
机构
*
Department of Computer Science, Lakehead University(湖头大学计算机科学系)
;
Quantum Communications and Computing Research Center and Department of Electrical and Computer Engineering, Memorial University of Newfoundland(新斯科舍纪念大学量子通信与计算研究中心及电气与计算机工程系)
;
Department of Electrical and Computer Engineering, Memorial University of Newfoundland(新斯科舍纪念大学电气与计算机工程系)
Mitigating hallucinations in healthcare LLMs with granular fact-checking and domain-specific adaptation
通过细粒度事实核查和领域特定适应减轻医疗保健大语言模型中的幻觉
Musarrat Zeba, Abdullah Al Mamun, Kishoar Jahan Tithee, Debopom Sutradhar, Mohaimenul Azam Khan Raiaan, Saddam Mukta, Reem E. Mohamed, Md Rafiqul Islam, Yakub Sebastian, Mukhtar Hussain, Sami Azam
机构
*
Applied Artificial Intelligence and Intelligent Systems (AAIINS) Laboratory(应用人工智能与智能系统实验室)
;
Department of Computer Science and Engineering(计算机科学与工程系)
;
Department of Data Science and Artificial Intelligence(数据科学与人工智能系)
;
Department of Software Engineering(软件工程系)
;
Faculty of Science and Information Technology(科学与信息技术学院)
;
Faculty of Science and Technology(科学与技术学院)
Integrating Virtual Reality and Large Language Models for Team-Based Non-Technical Skills Training and Evaluation in the Operating Room
将虚拟现实与大型语言模型结合用于手术室基于团队的非技术技能训练与评估
Jacob Barker, Doga Demirel, Cullen Jackson, Anna Johansson, Robbin Miraglia, Darian Hoagland, Stephanie B. Jones, John Mitchell, Daniel B. Jones, Suvranu De
机构
*
Beth Israel Deaconess Medical Center Center(贝希斯尔德医疗中心中心)
;
Department of Surgery, Northwell Health(外科,北well健康)
;
College of Engineering, Florida Agricultural and Mechanical University and Florida State University(工程学院,佛罗里达农业与机械大学和佛罗里达州立大学)
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
重新审视印度语言机器翻译和摘要细粒度评估的度量可靠性
Amir Hossein Yari, Kalmit Kulkarni, Ahmad Raza Khan, Fajri Koto
机构
*
Sharif University of Technology(谢里夫理工学院)
;
Vellore Institute of Technology(韦洛雷理工学院)
;
IIT Kharagpur(印度理工学院达卡分校)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)