arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7370 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7370 篇

2503.01616 2025-03-04 cs.RO 78%

RoboDexVLM: Visual Language Model-Enabled Task Planning and Motion Control for Dexterous Robot Manipulation

Haichao Liu, Sikai Guo, Pengfei Mai, Jiahang Cao, Haoang Li, Jun Ma

专题命中 视觉定位与Grounding :visual language model(title);vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.11144 2025-02-19 cs.HC 78%

CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding

Xingyu "Bruce" Liu, Ruolin Wang, Dingzeyu Li, Xiang 'Anthony' Chen, Amy Pavel

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03888 2025-02-18 cs.CL 78%

Multi3Hate: Multimodal, Multilingual, and Multicultural Hate Speech Detection with Vision-Language Models

Minh Duc Bui, Katharina von der Wense, Anne Lauscher

专题命中 视觉定位与Grounding :vision-language model(title,abstract)

Comments Accepted to NAACL 2025 Main (Camera-Ready Version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.02243 2025-02-18 cs.CL q-bio.NC 78%

Language Writ Large: LLMs, ChatGPT, Grounding, Meaning and Understanding

Stevan Harnad

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 54 pages, 29 references

Journal ref Frontiers in Artificial Intelligence 7: 1490698 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12812 2025-02-10 cs.CL 78%

Grounding Fallacies Misrepresenting Scientific Publications in Evidence

Max Glockner, Yufang Hou, Preslav Nakov, Iryna Gurevych

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Accepted to NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.05753 2025-02-10 cs.LG cs.AI cs.CV 78%

Grounding Continuous Representations in Geometry: Equivariant Neural Fields

David R Wessels, David M Knigge, Samuele Papa, Riccardo Valperga, Sharvaree Vadgama, Efstratios Gavves, Erik J Bekkers

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10491 2025-01-22 math.LO 78%

Denotational semantics for languages of epistemic grounding based on Prawitz's theory of grounds

Antonio Piccolomini d'Aragona

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10490 2025-01-22 math.LO 78%

Calculi of epistemic grounding based on Prawitz's theory of grounds

Antonio Piccolomini d'Aragona

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03200 2025-01-07 cs.CL 78%

The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Alon Jacovi, Andrew Wang, Chris Alberti, Connie Tao, Jon Lipovetz, Kate Olszewska, Lukas Haas, Michelle Liu, Nate Keating, Adam Bloniarz, Carl Saroufim, Corey Fry, Dror Marcus, Doron Kukliansky, Gaurav Singh Tomar, James Swirhun, Jinwei Xing, Lily Wang, Madhu Gurumurthy, Michael Aaron, Moran Ambar, Rachana Fellinger, Rui Wang, Zizhao Zhang, Sasha Goldshtein, Dipanjan Das

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06821 2024-12-11 cs.HC 78%

FinFlier: Automating Graphical Overlays for Financial Visualizations with Knowledge-Grounding Large Language Model

Jianing Hao, Manling Yang, Qing Shi, Yuzhe Jiang, Guang Zhang, Wei Zeng

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 17 pages, 13 figures, this paper is published on IEEE Transactions on Visualization and Computer Graphics

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12960 2024-11-21 cs.RO 78%

I Can Tell What I am Doing: Toward Real-World Natural Language Grounding of Robot Experiences

Zihan Wang, Brian Liang, Varad Dhat, Zander Brumbaugh, Nick Walker, Ranjay Krishna, Maya Cakmak

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11323 2024-11-19 cs.RO 78%

SayComply: Grounding Field Robotic Tasks in Operational Compliance through Retrieval-Based Language Models

Muhammad Fadhil Ginting, Dong-Ki Kim, Sung-Kyun Kim, Bandi Jai Krishna, Mykel J. Kochenderfer, Shayegan Omidshafiei, Ali-akbar Agha-mohammadi

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03359 2024-11-07 cs.CV cs.AI cs.LG 78%

Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection

Geng Yu, Jianing Zhu, Jiangchao Yao, Bo Han

专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI、cs.LG

Comments accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.02118 2024-11-05 cs.HC cs.CL 78%

Grounding Emotional Descriptions to Electrovibration Haptic Signals

Guimin Hu, Zirui Zhao, Lukas Heilmann, Yasemin Vardar, Hasti Seifi

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16472 2024-10-23 cs.CL 78%

DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding

Manan Suri, Puneet Mathur, Franck Dernoncourt, Rajiv Jain, Vlad I Morariu, Ramit Sawhney, Preslav Nakov, Dinesh Manocha

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments EMNLP 2024 (Main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09106 2024-10-21 cs.CL 78%

"We Demand Justice!": Towards Social Context Grounding of Political Texts

Rajkumar Pujari, Chengfei Wu, Dan Goldwasser

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Accepted as an oral at EMNLP 2024 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.05663 2024-10-10 cs.RO 78%

Abstract Hardware Grounding towards the Automated Design of Automation Systems

Yu-Zhe Shi, Qiao Xu, Fanxu Meng, Lecheng Ruan, Qining Wang

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments In International Conference on Intelligent Robotics and Applications (ICIRA'24)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.04819 2024-10-08 cs.CL 78%

MINER: Mining the Underlying Pattern of Modality-Specific Neurons in Multimodal Large Language Models

Kaichen Huang, Jiahao Huo, Yibo Yan, Kun Wang, Yutao Yue, Xuming Hu

专题命中 视觉定位与Grounding :multimodal large language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16671 2024-10-08 cs.CL 78%

StructLM: Towards Building Generalist Models for Structured Knowledge Grounding

Alex Zhuang, Ge Zhang, Tianyu Zheng, Xinrun Du, Junjie Wang, Weiming Ren, Stephen W. Huang, Jie Fu, Xiang Yue, Wenhu Chen

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.00518 2024-09-30 cs.RO 78%

When Robots Get Chatty: Grounding Multimodal Human-Robot Conversation and Collaboration

Philipp Allgeuer, Hassan Ali, Stefan Wermter

专题命中 视觉定位与Grounding :grounding(title,abstract)

Journal ref International Conference on Artificial Neural Networks, Sep 2024 (pp. 306-321)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.07980 2024-08-16 cs.LO 78%

Efficiently grounding FOL using bit vectors

Lucas Van Laer, Simon Vandevelde, Joost Vennekens

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Short version published at LPNMR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01139 2024-08-09 cs.CL 78%

It Couldn't Help But Overhear: On the Limits of Modelling Meta-Communicative Grounding Acts with Supervised Learning

Brielen Madureira, David Schlangen

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Accepted to SIGdial 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14321 2024-07-22 cs.CL cs.IR cs.MM 78%

Multimodal Misinformation Detection using Large Vision-Language Models

Sahar Tahmasebi, Eric Müller-Budack, Ralph Ewerth

专题命中 视觉定位与Grounding :vision-language model(title,abstract)

Comments Accepted for publication in: Conference on Information and Knowledge Management (CIKM) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12858 2024-07-19 cs.CL cs.AI cs.CV cs.LG 78%

Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)

Krishnaram Kenthapadi, Mehrnoosh Sameki, Ankur Taly

专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI、cs.LG

Comments Survey Article for the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD 2024) Tutorial

Journal ref Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.02584 2024-07-18 cs.SD eess.AS 78%

Towards Weakly Supervised Text-to-Audio Grounding

Xuenan Xu, Ziyang Ma, Mengyue Wu, Kai Yu

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.00209 2024-07-09 cs.CL 78%

EventGround: Narrative Reasoning by Grounding to Eventuality-centric Knowledge Graphs

Cheng Jiayang, Lin Qiu, Chunkit Chan, Xin Liu, Yangqiu Song, Zheng Zhang

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01824 2024-07-03 cs.HC cs.CL cs.RO 78%

Empathic Grounding: Explorations using Multimodal Interaction and Large Language Models with Conversational Agents

Mehdi Arjmand, Farnaz Nouraei, Ian Steenstra, Timothy Bickmore

专题命中 视觉定位与Grounding :grounding(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01749 2024-06-05 cs.CL 78%

Towards Harnessing Large Language Models for Comprehension of Conversational Grounding

Kristiina Jokinen, Phillip Schneider, Taiga Mori

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments Accepted to IWSDS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.07468 2024-04-16 cs.RO 78%

SGGNet$^2$: Speech-Scene Graph Grounding Network for Speech-guided Navigation

Dohyun Kim, Yeseung Kim, Jaehwi Jang, Minjae Song, Woojin Choi, Daehyung Park

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 7 pages, 6 figures, Published at 2023 IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), [Dohyun Kim, Yeseung Kim, Jaehwi Jang, and Minjae Song] contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.03092 2024-04-05 cs.CL cs.RO 78%

Unsupervised, Bottom-up Category Discovery for Symbol Grounding with a Curious Robot

Catherine Henry, Casey Kennington

专题命中 视觉定位与Grounding :grounding(title,abstract)

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏