CommentsAccepted to Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), July 20--24, 2026, Melbourne, VIC, Australia
Toward User Comprehension Supports for LLM Agent Skill Specifications
向LLM代理技能规范提供用户理解支持
Zikai Alex Wen
机构
*
University of Washington, Tacoma School of Engineering \& Technology Tacoma, Washington, USA
;
University of Washington, Tacoma School of Engineering \& Technology
CommentsDataset: github.com/Murrough-Foley/web-content-extraction-benchmark, doi.org/10.5281/zenodo.19316874. Leaderboard: webcontentextraction.org. Preprint also deposited at doi.org/10.5281/zenodo.19664685
MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
MCP-Atlas:一个大规模的工具使用能力基准测试,使用真实的MCP服务器
Chaithanya Bandi, Razvan-Gabriel Dumitru, Ben Hertzberg, Divyansh Agarwal, Geobio Boo, Tejas Polakam, Sami Hassaan, Jeff Da, HiJae Kim, Vipul Gupta, Manasi Sharma, Andrew Park, Martin Dimakis, Ernesto Gabriel Hernandez Montoya, Dan Rambado, Ivan Salazar, Rafael Cruz, MohammadHossein Rezaei, Chetan Rane, Ben Levin, Daniel Yue Zhang, Brad Kenstler, Bing Liu
机构
*
National University of Singapore(新加坡国立大学)
;
Scale AI
专题命中
评测与基准
:LLM(abstract,abstract_cn);large language model(abstract);language model(abstract);分类 cs.AI
GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety
GrandGuard:面向老年人与聊天机器人交互安全的分类、基准及防护措施
Changxuan Fan, Xi Yang, Yueyuan Zheng, Bin Zhou, Yuanping Wang, Wenbin Hu, Huihao Jing, Ki Sen Hung, Dazhao Du, Haoran Li, Janet Hui-wen Hsiao, Yangqiu Song
机构
*
The Hong Kong University of Science and Technology(香港科技大学)
Transcription and Recognition of Italian Parliamentary Speeches Using Vision-Language Models
使用视觉-语言模型进行意大利议会演讲的转录与识别
Luigi Curini, Alfio Ferrara, Giovanni Pagano, Sergio Picascia
机构
*
Università degli Studi di Milano(米兰大学)
;
Department of Social and Political Sciences(社会科学系)
;
Department of Literary Studies, Philology and Linguistics(文学研究、语言学与语言学系)
;
Department of Computer Science(计算机科学系)
Commentsto be published in: ParlaCLARIN V: Interoperability, Multilinguality, and Multimodality in Parliamentary Corpora, organized within the 15th Language Resource and Evaluation Conference (2026)
机构
*
Multimodal Language Department(多模态语言部门)
;
Max Planck Institute for Psycholinguistics(马克斯·普朗克心理语言学研究所)
;
Department of Linguistics(语言学系)
;
Boğaziçi University(博多伊奇大学)
;
Donders Institute for Brain Cognition and Behaviour(多纳尔斯脑认知与行为研究所)
;
Radboud University(拉德堡德大学)
;
Department of Linguistics and Communication(语言学与沟通系)
;
University of Birmingham(伯明翰大学)
Are Vision-Language Models Ready for Dietary Assessment? Exploring the Next Frontier in AI-Powered Food Image Recognition
视觉-语言模型是否准备好进行饮食评估?探索AI驱动的食品图像识别的下一个前沿
Sergio Romero-Tapiador, Ruben Tolosana, Blanca Lacruz-Pleguezuelos, Laura Judith Marcos Zambrano, Guadalupe X. Bazán, Isabel Espinosa-Salinas, Julian Fierrez, Javier Ortega-Garcia, Enrique Carrillo de Santa Pau, Aythami Morales
机构
*
Biometrics and Data Pattern Analytics Lab, Universidad Autonoma de Madrid(生物度量与数据模式分析实验室,马德里自治大学)
;
IMDEA Food, CEI UAM+CSIC(IMDEA食品,CEI UAM+CSIC)