ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
ERGeoBench:多模态大语言模型中具身推理与地理定位的综合基准
Kaiwen Xue, Tao Wei, Guoxin Zhang, Zhonghong Ou, Kaoyan Lu, Yu Feng, Yifan Zhu, Haoran Luo
机构
*
Beijing University of Posts and Telecommunications(北京邮电大学)
;
State Key Laboratory of Networking and Switching Technology(网络与交换技术国家重点实验室)
;
School of Materials Science and Engineering(材料科学与工程学院)
;
China Mobile Research Institute(中国移动研究院)
;
College of Computing and Data Science(计算与数据科学学院)
Steins;Gate Drive: Semantic Safety Arbitration over Structured Futures for Latency-Decoupled LLM Planning
Steins;Gate Drive: 基于结构化未来语义安全仲裁的延迟解耦LLM规划
Anjie Qiu, Hans D. Schotten
机构
*
Institute for Wireless Communication and Navigation(无线通信与导航研究所)
;
RPTU University Kaiserslautern-Landau(凯撒斯劳滕-兰道大学)
;
German Research Center for Artificial Intelligence(德国人工智能研究中心)
SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?
SpatiaLab: 视觉-语言模型能否在真实环境中进行空间推理?
Azmine Toushik Wasi, Wahid Faisal, Abdur Rahman, Mahfuz Ahmed Anik, Munem Shahriar, Mohsin Mahmud Topu, Sadia Tasnim Meem, Rahatun Nesa Priti, Sabrina Afroz Mitu, Md. Iqramul Hoque, Shahriyar Zaman Ridoy, Mohammed Eunus Ali, Majd Hawasly, Mohammad Raza, Md Rizwan Parvez
机构
*
Computational Intelligence and Operations Laboratory(计算智能与运筹实验室)
;
Shahjalal University of Science and Technology(沙赫jalal科技大学)
;
BRAC University(BRAC大学)
;
North South University(北南大学)
;
Monash University(墨尔本大学)
;
Qatar Computing Research Institute(卡塔尔计算研究院)
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
超越高斯瓶颈:基于拓扑对齐的视觉Transformer特征空间编码
Andrew Bond, Ilkin Umut Melanlioglu, Erkut Erdem, Aykut Erdem
机构
*
Department of Computer Engineering, Koç University, Istanbul, Turkey(科克大学计算机工程系,伊斯坦布尔,土耳其)
;
Department of Computer Engineering, Hacettepe University, Ankara, Turkey(哈恰塔佩大学计算机工程系,安卡拉,土耳其)
;
KUIS AI Research Center, Istanbul, Turkey(KUIS人工智能研究中心,伊斯坦布尔,土耳其)
;
Department of Electrical and Electronics Engineering, Koç University, Istanbul, Turkey(科克大学电气与电子工程系,伊斯坦布尔,土耳其)
机构
*
State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Peking University(北京大学)
;
Beijing Institute of Technology(北京理工大学)
;
Tsinghua University(清华大学)
CommentsThis preprint is withdrawn due to significant errors in the emergent geometric isomorphism results that necessitate full rewriting, coupled with unresolved author disagreement on authorship. A corrected and revised manuscript will be released separately
Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models
超越识别:评估视觉语言模型中的视觉视角理解
Gracjan Góral, Alicja Ziarko, Piotr Miłoś, Michał Nauman, Maciej Wołczyk, Michał Kosiński
机构
*
University of Warsaw(华沙大学)
;
Polish Academy of Sciences(波兰科学院)
;
Stanford University(斯坦福大学)
;
University of California, Berkeley(加州大学伯克利分校)
;
IDEAS NCBR
;
Lute
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Chongqing Chang’an Technology Co., Ltd(重庆长安科技有限公司)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
College of AI, Tsinghua University(清华大学人工智能学院)
;
Zhongguancun Academy(中关村学院)
BLINK: Behavioral Latent Modeling of NK Cell Cytotoxicity
BLINK:自然杀伤细胞杀伤行为的隐式建模
Iman Nematollahi, Jose Francisco Villena-Ossa, Alina Moter, Kiana Farhadyar, Gabriel Kalweit, Abhinav Valada, Toni Cathomen, Evelyn Ullrich, Maria Kalweit