PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models
PlanBench-V: 面向视觉语言模型的空间规划地图基准
Minxin Chen, He Zhu, Junyou Su, Wen Wang, Yijie Deng, Wenjia Zhang
机构
*
Behavioral and Spatial AI Lab(行为与空间人工智能实验室)
;
Tongji University(同济大学)
;
Peking University(北京大学)
;
College of Architecture and Urban Planning(建筑与城市规划学院)
CommentsWithdrawn by the authors due to pending intellectual property considerations. The authors have determined that the current version contains material that should not have been publicly disseminated at this stage
GenTI: Benchmarking LLMs for Autonomous IDPS Rule Generation for Unseen Attacks
GenTI: 针对未知攻击的自主IDPS规则生成的LLM基准测试
Hassan Jalil Hadi, Rehana Yasmin, Ali Shoker
机构
*
Cyber Security and Resilience Technology (CyberSaR), King Abdullah University of Science and Technology (KAUST)(网络安全与韧性技术(CyberSaR),国王阿卜杜勒·阿齐兹大学科学与技术学院(KAUST))
DocHop-QA: Towards Multi-Hop Reasoning over Multimodal Document Collections
DocHop-QA: 向多跳推理多模态文档集合迈进
Jiwon Park, Seohyun Pyeon, Jinwoo Kim, Rina Carines Cabal, Zhenyuan He, Yihao Ding, Soyeon Caren Han
机构
*
Pohang University of Science and Technology(釜山科学技术大学)
;
The University of Sydney(悉尼大学)
;
The University of Western Australia(西澳大学)
;
The University of Melbourne(墨尔本大学)
UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning
UltraVR:面向证据推理的诊断性超分辨率图像VQA基准
Gexin Huang, Yanting Yang, Myeongkyun Kang, Beidi Zhao, Jun Zhou, Chen Zhou, Gang Wang, Zu-hua Gao, Xiaoxiao Li
机构
*
University of British Columbia(不列颠哥伦比亚大学)
;
Vector Institute(向量研究所)
;
BC Cancer Agency(不列颠哥伦比亚癌症中心)
;
The Hong Kong Polytechnic University(香港理工大学)
机构
*
Shanghai Institute of AI for Education(上海人工智能教育研究院)
;
School of Computer Science(计算机科学学院)
;
East China Normal University(东华大学)
;
Tencent Inc.(腾讯公司)
;
Shanghai Innovation Institute(上海创新研究院)
Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models
Almieyar-Oryx-BloomBench:一个用于视觉语言模型认知知情评估的双语多模态基准
Mohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor, Marzia Nouri, Doratossadat Dastgheib, Mohamed Hefeeda, Ehsaneddin Asgari
机构
*
University of British Columbia(不列颠哥伦比亚大学)
;
Zuse School(Zuse学校)
;
Qatar Computing Research Institute (QCRI)(卡塔尔计算研究所)
;
Hamad Bin Khalifa University(哈马德·本·哈利法大学)
Harnessing Structural Context for Entity Alignment Foundation Models
利用结构上下文进行实体对齐基础模型
Xingyu Chen, Yuanning Cui, Zequn Sun, Wei Hu
机构
*
State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China(南京大学新型软件技术国家重点实验室)
;
Nanjing University of Information Science and Technology, Nanjing, China(南京信息科学技术大学)
;
National Institute of Healthcare Data Science, Nanjing University, Nanjing, China(南京大学健康数据科学国家研究院)
Rollout-Level Advantage-Prioritized Experience Replay for GRPO
基于轨迹级别优势优先经验回放的GRPO
Gyeongtae Yoo, Sanghyeok Park, Soohyuk Jang, Ik-hwan Kim, Sungroh Yoon
机构
*
Department of Electrical and Computer Engineering, Seoul National University(首尔国立大学电子与计算机工程系)
;
Interdisciplinary Program in AI, Seoul National University(首尔国立大学人工智能跨学科项目)
;
AIIS, ASRI, INMC, and ISRC, Seoul National University(首尔国立大学人工智能研究所、人工智能研究机构、智能网络与计算中心及人工智能科学研究中心)
机构
*
MMLab, The Chinese University of Hong Kong(中大香港实验室)
;
CPII under InnoHK(创新香港 CPII)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Shandong University(山东大学)
;
Huawei Technologies(华为技术)
Agent Memory: Characterization and System Implications of Stateful Long-Horizon Workloads
Agent记忆:有状态长时任务工作负载的表征与系统影响
Yasmine Omri, Ziyu Gan, Zachary Broveak, Robin Geens, Zexue He, Alex Pentland, Marian Verhelst, Tsachy Weissman, Thierry Tambe
机构
*
Massachusetts Institute of Technology(麻省理工学院)
;
Stanford University(斯坦福大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
MIT Media Lab(麻省理工学院媒体实验室)
CASS-RTL: Correctness-Aware Subspace Steering for RTL Generation with LLMs
CASS-RTL:面向LLM的RTL生成的正确性感知子空间引导
Mohammad Akyash, Nowfel Mashnoor, Kimia Azar, Hadi Kamali
机构
*
Department of Electrical and Computer Engineering (ECE), University of Central Florida, Orlando, FL 32816, USA(电子与计算机工程系,中央佛罗里达大学,奥兰多,佛罗里达州32816,美国)
Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows
更多智能体有帮助吗?LLM智能体工作流的受控与协议对齐评估
Yuhang Fu, Ruishan Fang, Jiaqi Shao, Huiyu Zheng, Zhengtao Zhu, Bing Luo, Tao Lin
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Westlake University(西湖大学)
;
Zhejiang University(浙江大学)
;
Duke Kunshan University(杜克大学昆山分校)
;
Hong Kong University of Science and Technology(香港科技大学)
;
Zhejiang University of Technology(浙江工业大学)
Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation
面向长时域任务的安全具身AI:机器人操作跨层分析
Dabin Kim, Daemin Park, Sangyub Lee, Jinsik Kim, Yeongtak Oh, Jongho Shin, Sungroh Yoon
机构
*
UNIST InnoCORE AI-Space Solar Initiative(UNIST创新核心人工智能空间太阳能计划)
;
Ulsan National Institute of Science and Technology (UNIST)(乌山国立科学技术研究院)
;
Automation and Systems Research Institute(自动化与系统研究所)
;
Department of Electrical and Computer Engineering(电气与计算机工程系)
;
Interdisciplinary Program in Artificial Intelligence(人工智能跨学科项目)
;
LG Electronics(LG电子)
Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing
面向沉浸式视频角色扮演的奖励分解强化学习
Miao Wang, Yuling Shi, Yijiang Li, Yeheng Chen, Xiaodong Gu, Bin Li, Bo Gao, Jun Wang, Zengxin Han, Jingtong Wu, Yaduan Ruan
机构
*
Nanjing University(南京大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
University of California, San Diego(加州大学圣地亚哥分校)
;
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
;
School of Information Engineering, Beijing Institute of Graphic Communication(北京印刷学院信息工程学院)
;
Ant International, Ant Group(蚂蚁集团国际部)
;
Independent Researcher(独立研究者)
Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code Translation
超越代码对:基于对话的数据生成用于LLM代码翻译
Le Chen, Nuo Xu, Winson Chen, Bin Lei, Pei-Hung Lin, Dunzhi Zhou, Rajeev Thakur, Caiwen Ding, Ali Jannesari, Chunhua Liao
机构
*
Argonne National Laboratory(阿贡国家实验室)
;
University of Minnesota(明尼苏达大学)
;
Iowa State University(爱荷华州立大学)
;
Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)