机构
*
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
OPPO AI Center, OPPO Inc.(OPPO公司OPPO人工智能中心)
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
视觉语言动作模型所言即所指?论忠实性在具身推理中的作用
Matthew Foutter, Matteo Cercola, Lena Wild, Yunshan Wang, Michelle Li, Daniele Gammelli, Marco Pavone
机构
*
Stanford University(斯坦福大学)
;
Politecnico di Milano(米兰理工大学)
;
KTH Royal Institute of Technology(皇家理工学院)
;
Italian Institute of Artificial Intelligence (AI4I)(意大利人工智能研究所)
;
NVIDIA Research(英伟达研究院)
ACE: Agentic Control for Embodied Manipulation via Zero-shot Workflow Reasoning
ACE:通过零样本工作流推理实现具身操纵的智能体控制
Iok Tong Lei, QianZhi Li, Ying Jie Yap, Yujie Zhang, Rui Zhong, Haichao Gui, Xiaolong Liu, Zhidong Deng
机构
*
Department of Computer Science, Tsinghua University(清华大学计算机科学系)
;
National College for Excellent Engineers, Beihang University(北京航空航天大学卓越工程师学院)
;
Wuxi Dexteroushands Robotic Technology Co.(无锡灵犀机器人技术有限公司)
机构
*
University of Science and Technology of China(中国科学技术大学)
;
King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
;
Vast Intelligence Lab(旷视智能实验室)
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
VKnowU:评估多模态语言模型中的视觉知识理解
Tianxiang Jiang, Sheng Xia, Yicheng Xu, Linquan Wu, Xiangyu Zeng, Limin Wang, Yu Qiao, Yi Wang
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Nanjing University(南京大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
City University of Hong Kong(香港城市大学)
专题命中
视觉推理
:multimodal large language model(abstract);分类 cs.CV
Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs
机构
*
The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心)
;
Eisai Inc.(卫材株式会社)
;
IISc, Bangalore(印度科学研究所班加罗尔分校)
;
Cohere Labs Community(Cohere实验室社区)
专题命中
视觉定位与Grounding
:vision language model(title,abstract);grounding(title,abstract);分类 cs.CV
GeoSearcher: Anchor-Guided Progressive Reasoning for Remote Sensing Visual Grounding with Process Supervision
GeoSearcher: 基于锚点引导的渐进推理遥感视觉定位与过程监督
Dianyu Wang, Peirong Zhang, Xuyang Li, Xiaoxuan Liu, Lei Wang
机构
*
Key Laboratory of Target Cognition and Application Technology (TCAT), Chinese Academy of Sciences(中国科学院目标认知与应用技术重点实验室)
;
School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院)
专题命中
视觉定位与Grounding
:grounding(title,abstract);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV
Detecting Hallucinations in Retrieval-Augmented Generation through Grounding-Aware Sensitivity by Perturbation (GASP)
通过扰动的基于接地感知敏感性检测检索增强生成中的幻觉(GASP)
Mohamed Aly Bouke
机构
*
Centre for Intelligent Cloud Computing, CoE for Advanced Cloud, Faculty of Information Science and Technology, Multimedia University(智能云计算中心、高级云研究中心、信息科学与技术学院、多媒体大学)
机构
*
Australian Institute for Machine Learning, Adelaide University(澳大利亚机器学习研究所,阿德莱德大学)
;
Ruyi Dynamics(如意动力)
;
Fudan University(复旦大学)
;
Swiss Federal Institute of Technology Lausanne (EPFL)(瑞士洛桑联邦理工学院)
Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models
检索头能看见图像吗?长上下文视觉语言模型中的多模态检索头
Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao, Yiming Du, Haobo Li, Xiyu Ren, Ginny Wong, Simon See, Lishu Luo, Haodong Duan, Pasquale Minervini, Yangqiu Song
机构
*
HKUST(香港科技大学)
;
University of Edinburgh(爱丁堡大学)
;
CUHK(香港中文大学)
;
NVAITC, NVIDIA, Santa Clara, USA(NVIDIA Santa Clara 分公司)
;
Tsinghua University(清华大学)
VLMGuard: Bootstrapping Malicious Prompt Detectors from Unlabeled Vision-Language Prompts in the Wild
VLMGuard:从野生未标记视觉语言提示中引导恶意提示检测器
Junlin Fang, Wenyu Chen, Reshmi Ghosh, Robert Sim, Ahmed Salem, Vitor R. Carvalho, Emily Lawton, Sharon Li, Jack W. Stokes, Sean Du
机构
*
College of Computing and Data Science(计算与数据科学学院)
;
Nanyang Technological University(南洋理工大学)
;
School of Physical and Mathematical Sciences(物理与数学科学学院)
;
Microsoft Corp.(微软公司)
;
Department of Computer Sciences(计算机科学系)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
Language-Guided Grasping under Partial Observation for Mobile Manipulation in Field Inspection and Maintenance
用于现场检查和维护中移动操作的部分观察下语言引导抓取
Dilermando Almeida, Juliano Negri, Guilherme Lazzarini, Thiago H. Segreto, Ranulfo Bezerra, Gustavo J. G. Lahr, Ricardo V. Godoy, Marcelo Becker
机构
*
Department of Mechanical Engineering, Federal University of Uberlândia(联邦大学伯南迪利亚机械工程系)
;
Department of Mechanical Engineering, University of São Paulo(圣保罗大学机械工程系)
;
Graduate School of Information Sciences, Tohoku University(东北大学信息科学研究生院)
;
Faculdade Israelita de Ensino e Pesquisa Albert Einstein, Hospital Israelita Albert Einstein(艾伯特·爱因斯坦以色列教学与研究学院,艾伯特·爱因斯坦医院)
GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video
GuideMe:流视频中的多域任务指导与干预
Fang Liu, Jinpeng Chen, Ke Xu, Yuhao Liu, Huankang Guan, Xudong Lu, Bo Yang, Gerhard Hancke, Rui Liu, Rynson W. H. Lau
机构
*
City University of Hong Kong(香港城市大学)
;
Huawei Research(华为研究院)
;
University of Science and Technology of China(中国科学技术大学)
;
Chinese University of Hong Kong(香港中文大学)
;
City University of Hong Kong (Dongguan)(香港城市大学(东莞))
专题命中
视觉定位与Grounding
:multimodal large language model(abstract);MLLM(abstract);分类 cs.CV