AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
AgentMisalignment:衡量基于LLM的代理中失调行为的倾向性
Akshat Naik, Emma Gouné, Patrick Quinn, Guillermo Bosch, Francisco Javier Campos Zabala, Jason Ross Brown, Edward James Young
机构
*
Department of Computer Science(计算机科学系)
;
University of Oxford(牛津大学)
;
Institute of Intelligent Systems and Robotics(智能系统与机器人研究所)
;
Sorbonne Université(索邦大学)
;
The Leverhulme Centre for the Future of Intelligence(未来智能中心)
;
University of Cambridge(剑桥大学)
;
Independent Researcher(独立研究者)
;
Department of Computer Science and Technology(计算机科学与技术系)
;
Department of Engineering(工程系)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
机构
*
Division of Radiology and Biomedical Engineering, Graduate School of Medicine, The University of Tokyo(放射医学与生物医学工程系,东京大学医学研究生院)
;
Department of Computational Diagnostic Radiology and Preventive Medicine, The University of Tokyo Hospital(计算诊断放射学与预防医学系,东京大学医院)
;
Department of Radiology, Kanazawa University Hospital(金泽大学医院放射科)
;
Faculty of Medicine, The University of Tokyo(东京大学医学系)
;
Department of Radiology, School of Medicine, Jichi Medical University(立命馆大学医学系放射科)
;
Department of Radiology, The University of Tokyo Hospital(东京大学医院放射科)
专题命中
评测与基准
:LLM(summary_cn,abstract);large language model(title,abstract);language model(title,abstract);分类 cs.AI
LLM-ReSum: A Framework for LLM Reflective Summarization through Self-Evaluation
LLM-ReSum: 一个通过自我评估实现LLM反思性总结的框架
Huyen Nguyen, Haoxuan Zhang, Yang Zhang, Haihua Chen, Junhua Ding
机构
*
dept. of Information Science University of North Texas Denton, Texas, USA(信息科学系 俄克拉荷马州立大学 丹顿 俄克拉荷马州 美国)
;
dept. of Data Science University of North Texas Denton, Texas, USA(数据科学系 俄克拉荷马州立大学 丹顿 俄克拉荷马州 美国)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI
CommentsThis paper has been accepted as an invited paper for publication in Proceedings of The 12th IEEE International Conference on Big Data Computing Service and Machine Learning Applications. This is the accepted manuscript. The final authenticated version will be available via IEEE Xplore
机构
*
School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算与信息系统学院)
;
Human-Computer Interaction Institute, Carnegie Mellon University(卡内基梅隆大学人机交互研究所)
专题命中
评测与基准
:LLM(title,title_cn);large language model(abstract);language model(abstract);分类 cs.CL
SegTME-UNI2: A Foundation Model-Based Framework for Generalisable Multiclass Cell Segmentation and LLM-Driven Tumour Microenvironment Characterisation in Histopathology
Wan Siti Halimatul Munirah Wan Ahmad, Faris Syahmi Samidi, Mohammad Badal Ahmmed, Vimal Angela Thiviyanathan, Selvam Thavaraj, Anwar P. P. Abdul Majeed
机构
*
Department of Data Science and Artificial Intelligence, School of Computing and Artificial Intelligence, Faculty of Engineering and Technology, Sunway University(双威大学工程与技术学院计算与人工智能学院数据科学与人工智能系)
;
Faculty of Dentistry, Universiti Malaya(马来亚大学牙科学院)
From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents
从知道到行动:基准测试LLM代理的自我意识能力
Yifan Li, Shengbin Yue, Boyu Feng, Jinhu Qi, Bo Ke, Zixing Song, Hongru Wang, Zhongyu Wei, Irwin King
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Fudan University(复旦大学)
;
University of Edinburgh(爱丁堡大学)
;
Tencent(腾讯)
;
University of Bristol(布里斯托大学)
BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems
BELLS-O:评估LLM监督系统的运营权衡
Leonhard Waibl, Felix Michalak, Hadrien Mariaccia
机构
*
University of Graz, Graz, Austria(格拉茨大学)
;
Supervised Program for Alignment Research (SPAR)(对齐研究监督计划 (SPAR))
;
Centre pour la Sécurité de l'IA (CeSIA), Paris, France(人工智能安全研究中心 (CeSIA),巴黎,法国)
Implicit Identity Technologies for LLMs: Fingerprinting and Watermarking across Datasets, Models, and Generated Content
LLM的隐式身份技术:跨数据集、模型和生成内容的指纹识别与水印
Bing Liu, Shunping Wang, Yufan Zhu, Xinyi Yu, Jing Huang, Linkang Du, Hongbin Pei, Wei Luo
机构
*
School of Cyber Science and Engineering, Xi’an Jiaotong University, Xi’an, China(西安交通大学计算机科学与工程学院)
;
State Grid Henan Marketing Service Center, Henan, China(国网河南营销服务中心)
;
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China(中国科学院大学网络安全学院)
;
School of Information Technology, Deakin University, Geelong, Australia(迪金大学信息技术学院)
专题命中
评测与基准
:LLM(title_cn,summary_cn);large language model(abstract);language model(abstract);分类 cs.CL、cs.LG
CommentsAccepted by IJCAI-ECAI 2026. 11 pages, 1 figure. Survey and taxonomy of LLM fingerprinting and watermarking for identity, provenance, generated-content attribution, and asset protection
Safety-Aware Evaluation of LLM-Generated Driver Intervention Messages through Multi-Task Risk Fusion
通过多任务风险融合的LLM生成驾驶员干预信息的安全感知评估
Keito Inoshita
机构
*
Faculty of Business and Commerce, Kansai University(关西大学商学部)
;
Data Science and AI Innovation Research Promotion Center, Shiga University(滋贺大学数据科学与人工智能创新研究推进中心)
RouteJudge: An Open Platform for Reproducible and Preference-Aware LLM Routing
RouteJudge: 一个可复现且偏好感知的LLM路由开放平台
Guannan Lai, Haoran Hu, Han-Jia Ye
机构
*
School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
;
National Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)
;
SinapisAI