Segment Anything Model for automated image data annotation: empirical studies using text prompts from Grounding DINO
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accept to ACL 2024; 19 Pages, 15 Figures, 6 Tables
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Medical Imaging with Deep Learning (MIDL) 2024 (Oral)
专题命中 视觉定位与Grounding :VLM(title,abstract);visual language model(abstract)
专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract)
专题命中 视觉定位与Grounding :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments In EMNLP 2023 main conference proceedings (to appear)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Project website: https://chat-with-nerf.github.io/
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Extended version for ICLRW-2020 (BeTR-RL) paper
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Conference on Robot Learning (CoRL) 2021
专题命中 视觉定位与Grounding :grounding(title,abstract);visual question answering(abstract)
Comments Accepted at LANTERN2021
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted at the International Symposium on Experimental Robotics (ISER) 2020. Videos at http://speechrobot.cs.uni-freiburg.de
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments 4th Conference on Robot Learning (CoRL 2020), Cambridge MA, USA
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted at WACV 2019. Also at NeurIPS 2017 workshop on Visually-Grounded Interaction and Language (ViGIL)
合成对象组合用于检测、分割和接地任务中的可扩展和准确学习
机构 * University of Washington(华盛顿大学) ; Allen Institute for Artificial Intelligence(人工智能研究院)
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.AI
AI总结 合成对象组合(SOC)通过新的以对象为中心的组合策略,实现了准确且可扩展的数据合成流程,提升了检测、分割和接地任务的性能,尤其在低数据和封闭词汇场景中表现突出。
Comments Project website: https://github.com/weikaih04/Synthetic-Detection-Segmentation-Grounding-Data
专题命中 视觉定位与Grounding :grounding(title,abstract);分类 cs.CV、cs.LG
Comments The second and the third authors contributed equally to this paper, and are listed in alphabetical order. Project Page: http://visual-physics-grounding.csail.mit.edu/
SEED: 用于可解释文本伪造检测的简单ViT与演化框架
机构 * State Key Laboratory of Internet of Things for Smart City, Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系智慧城市物联网国家重点实验室) ; School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)
专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);分类 cs.CV
AI总结 提出SEED系统,结合相似性引导数据增强、单一ViT联合检测与定位、以及基于MLLM的演化报告生成,在ACM MM 2026文本伪造挑战赛中获得第三名。
EditFlow3D:保留轨迹的3D资产自动化局部编辑
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);分类 cs.CV
AI总结 EditFlow3D是一种无需训练的3D局部编辑框架,通过VLM驱动工作流生成掩码与视觉引导,结合掩码引导差分流和轨迹保留技术,在新基准EditFlow-Bench上验证了其编辑准确性与非目标区域保留能力优于现有方法。
LDU-Bench:面向不同布局电路背景下光刻缺陷理解的多模态大语言模型评估基准
专题命中 视觉定位与Grounding :multimodal large language model(abstract,abstract_cn);MLLM(summary_cn);分类 cs.CV
AI总结 本文提出LDU-Bench这一多模态基准,将光刻审查工作流程分解为四项任务,评估发现现有多模态大语言模型的缺陷分类能力无法稳定迁移至下游审查阶段,为工业MLLM评估提供了统一平台。
Comments 12 pages, 3 figures, and 5 tables, including appendices
OVEarth-Bench:面向开放词汇地球观测的类别广度与查询多样性评估
专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract);分类 cs.CV
AI总结 本文提出OVEarth-Bench基准,从类别广度与查询多样性两方面扩展开放词汇地球观测评估,经实验发现MLLM类方法性能最优,为该领域未来研究提供指导。
TongueReenact:几何锚定的人脸重演舌部合成
机构 * University of Technology Sydney(悉尼科技大学) ; Shandong University of Science and Technology(山东科技大学)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract);分类 cs.CV
AI总结 本文提出首个跨身份舌部动态迁移框架,结合 foundation model 辅助的自举流水线与空间约束潜在掩码扩散模型,在舌部指标上较基线提升超两倍,且 VLM 评估证实其感知优势。
证据见证组合:针对闭集多模态答案的单预填充风险检测
专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 该研究提出WEP方法,利用MLLM的白盒预填充路径,无需额外操作即可检测闭集视觉答案的推理风险,在3个MLLM和4个基准上提升了平均错误AP。
Comments 22 pages, 6 figures; includes supplementary material. Code: https://github.com/SouthWinter/WEP
RAD:面向机器人观测的真实异常检测数据集与基准
机构 * Massachusetts Institute of Technology(麻省理工学院) ; Peking University(北京大学) ; Carnegie Mellon University(卡内基梅隆大学) ; Great Bay University(大湾大学) ; Harvard University(哈佛大学) ; Tsinghua University(清华大学) ; Nanjing University(南京大学) ; Deakin University(德克萨斯大学)
专题命中 视觉定位与Grounding :VLM(summary_cn,abstract_cn);vision-language model(abstract);分类 cs.CV
AI总结 提出RAD数据集,包含13类日常物体和4种缺陷,从60多个机器人视角在非受控光照下采集,用于评估2D特征、3D重建和视觉语言模型在姿态无关的异常检测中的表现,发现2D方法优于3D和VLM方法。
循环代码取证:面向图像伪造检测的代理工具使用
机构 * University of Science and Technology of China(中国科学技术大学) ; Shanghai Innovation Institute(上海创新研究院) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract_cn);multimodal large language model(abstract);分类 cs.AI
AI总结 本文提出ForenAgent框架,通过多轮交互使MLLM自主生成并迭代优化Python工具,提升图像伪造检测的灵活性和可解释性,构建FABench数据集进行系统训练和评估。
Comments 18 pages, 7 figures
VG3S:基于视觉几何的高斯散射用于语义占用预测
机构 * Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(电子与计算机工程系,香港科学与技术大学)
专题命中 视觉定位与Grounding :grounding(summary_cn,abstract);分类 cs.CV
AI总结 VG3S 通过引入视觉基础模型的几何 grounding 能力,提升语义占用预测的准确性与泛化能力。
Comments Accepted by IROS 2026
Open-KNEAD:通过智能分解实现基于知识的营养估计
机构 * Purdue University(普渡大学)
专题命中 视觉定位与Grounding :MLLM(abstract,abstract_cn);grounding(abstract);multimodal large language model(abstract);分类 cs.CV
AI总结 研究探讨多模态大语言模型用于膳食评估时检索增强基础方法是否仍有价值。提出无需训练、可本地部署的Open-KNEAD框架,通过营养感知检索关联食物项与FNDDS代码,提高份量估计,还能恢复非美国烹饪风格估计偏差,优势显著并开源相关框架与知识库。
Comments 10 pages main paper, 5 pages supplementary
EvoGuard: 一种基于代理强化学习的可扩展框架,用于实际和演化的AI生成图像检测
机构 * The University of Tokyo(东京大学) ; National Institute of Informatics(国家信息研究所)
专题命中 视觉定位与Grounding :MLLM(summary_cn,abstract_cn);multimodal large language model(abstract);分类 cs.CV
AI总结 本文提出EvoGuard框架,利用多模态大语言模型和非MLLM检测器,通过能力感知的动态编排机制实现高效AIGI检测,优化后无需细粒度标注,实验表明其在准确性和可扩展性上均达最优。
Comments Template changed