Nikos Theodoridis, Tim Brophy, Reenu Mohandas, Ganesh Sistu, Fiachra Collins, Anthony Scanlan, Ciaran Eising
机构
*
Department of Electronic and Computer Engineering, University of Limerick(利默尼克大学电子与计算机工程系)
;
Data Driven Computer Engineering Research Centre, University of Limerick(利默尼克大学数据驱动计算机工程研究中心)
;
Lero, The Irish Software Research Centre, University of Limerick(利默尼克大学Lero爱尔兰软件研究中心)
;
Valeo Vision Systems(瓦莱奥视觉系统)
Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
过程链:用于过程问答的层次视觉语言推理
Guanhua Chen, Yutong Yao, Shenghe Sun, Ci-Jun Gao, Shudong Liu, Lidia S. Chao, Feng Wan, Derek F. Wong
机构
*
NLP
;
CT Lab, Department of Computer and Information Science, University of Macau(计算机与信息科学系,澳门大学CT实验室)
;
Department of Electrical and Computer Engineering, University of Macau(电气与计算机工程系,澳门大学)
;
Centre for Cognitive and Brain Sciences, University of Macau(认知与脑科学中心,澳门大学)
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
视觉思维混合:探索上下文自适应推理模式选择用于通用视觉推理
Zejun Li, Yingxiu Zhao, Jiwen Zhang, Siyuan Wang, Yang Yao, Runzhou Zhao, Jun Song, Bo Zheng, Zhongyu Wei
机构
*
Fudan University(复旦大学)
;
Alibaba Group Holding Limited(阿里巴巴集团控股有限公司)
;
Future Living Lab of Alibaba(阿里巴巴未来生活实验室)
;
University of Southern California(南加州大学)
;
Shanghai Innovation Institute(上海创新研究院)
DermAgent: A Self-Reflective Agentic System for Dermatological Image Analysis with Multi-Tool Reasoning and Traceable Decision-Making
DermAgent: 一种用于皮肤病图像分析的自反思代理系统,具备多工具推理和可追溯决策制定
Yize Liu, Siyuan Yan, Ming Hu, Lie Ju, Xieji Li, Feilong Tang, Wei Feng, Zongyuan Ge
机构
*
AIM for Health Lab, Faculty of Information Technology, Monash University, Melbourne, Australia(健康人工智能实验室,信息科技学院,墨尔本大学,澳大利亚)
;
Faculty of Information Technology, Monash University, Melbourne, Australia(信息科技学院,墨尔本大学,澳大利亚)
;
University College London, Institute of Ophthalmology, London, United Kingdom(伦敦大学学院,眼科研究所,英国)
专题命中
视觉推理
:grounding(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
机构
*
School of Computer Science, National Engineering Research Center for Multimedia Software and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University, China(计算机学院,国家多媒体软件工程技术研究中心和湖北多媒体与网络通信工程重点实验室,武汉大学,中国)
;
Alibaba Group, Hangzhou, China(阿里巴巴集团,杭州,中国)
;
Independent Researcher(独立研究者)
;
Department of Machine Learning, Mohamed bin Zayed University of Artificial Intelligence, United Arab Emirates(机器学习系,Mohamed bin Zayed人工智能大学,阿拉伯联合酋长国)
专题命中
视觉推理
:grounding(abstract);multimodal large language model(abstract);分类 cs.CV
From Table to Cell: Attention for Better Reasoning with TABALIGN
从表格到单元格:用于基于TABALIGN的更好推理的注意力
Tung Sum Thomas Kwok, Zeyong Zhang, Xinyu Wang, Chunhe Wang, Xiaofeng Lin, Hanwei Wu, Lei Ding, Guang Cheng, Zhijiang Guo
机构
*
University of California, Los Angeles(加州大学洛杉矶分校)
;
New Jersey Institute of Technology(新泽西理工学院)
;
McGill University(麦吉尔大学)
;
Université de Montréal(蒙特利尔大学)
;
University of Manitoba(曼尼托巴大学)
;
SimpleWay
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
Junho Kim, Eun Sun Lee, Gwangtak Bae, Seunggu Kang, Young Min Kim
机构
*
Dept. of Electrical and Computer Engineering, Seoul National University(电子与计算机工程系,首尔国立大学)
;
Interdisciplinary Program in Artificial Intelligence and INMC, Seoul National University(人工智能交叉计划和INMC,首尔国立大学)
GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines
GeoLaux:一个评估多模态大语言模型在需要辅助线的长步骤问题上几何性能的基准
Yumeng Fu, Jiayin Zhu, Lingling Zhang, Wenjun Wu, Bo Zhao, Shaoxuan Ma, Yushun Zhang, Jun Liu
机构
*
School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)
;
Ministry of Education Key Laboratory of Intelligent Networks and Network Security, China(教育部智能网络与网络安全重点实验室)
;
Shaanxi Province Key Laboratory of Big Data Knowledge Engineering, China(陕西省大数据知识工程重点实验室)
;
School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.AI
From Street View to Visual Network: Mapping the Visibility of Urban Landmarks with Vision-Language Models
从街景到视觉网络:利用视觉语言模型映射城市地标可见性
Zicheng Fan, Kunihiko Fujiwara, Pengyuan Liu, Fan Zhang, Filip Biljecki
机构
*
organization= Department of Architecture, National University of Singapore , country= Singapore
;
organization= Research \& Development Institute, Takenaka Corporation , country= Japan
;
organization= Urban Analytics Subject Group, Urban Studies \& Social Policy Division, University of Glasgow , country= United Kingdom
;
organization= Institute of Remote Sensing
;
GIS, Peking University , country= China
;
organization= Department of Real Estate, National University of Singapore , country= Singapore
专题命中
视觉定位与Grounding
:vision-language model(title);VLM(abstract,abstract_cn);vision language model(abstract);分类 cs.CV
机构
*
City University of Hong Kong(香港城市大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
CAIR, HKISI, Chinese Academy of Sciences(中国科学院计算智能研究所)
;
UESTC(电子科技大学)
;
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
EARL:一种统一的分析引导强化学习框架,用于第一人称交互推理与像素定位
Yuejiao Su, Xinshen Zhang, Zhen Ye, Lei Yao, Lap-Pui Chau, Yi Wang
机构
*
Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong SAR(香港理工大学电子与电气工程系)
;
Division of Emerging Interdisciplinary Areas (EMIA), The Hong Kong University of Science and Technology, Hong Kong SAR(香港理工大学新兴跨学科领域研究中心)
专题命中
视觉定位与Grounding
:grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV
SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding
SVAG-Bench:多实例时空视频动作定位的大规模基准
Tanveer Hannan, Shuaicong Wu, Mark Weber, Suprosanna Shit, Jindong Gu, Rajat Koner, Aljoša Ošep, Laura Leal-Taixé, Thomas Seidl
机构
*
LMU Munich(慕尼黑大学)
;
MCML
;
Technical University of Munich(慕尼黑技术大学)
;
University of Zurich(苏黎世大学)
;
University of Oxford(牛津大学)
;
Amazon(亚马逊)
;
NVIDIA(英伟达)
机构
*
Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
;
Zhongguancun Academy(中关村学院)
;
Huazhong University of Science and Technology(华中科技大学)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
South China University of Technology(华南理工大学)
;
University of Oxford(牛津大学)
;
Peking University(北京大学)