arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-04-10 至 2026-04-10 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 3 篇

2604.08050 2026-04-10 cs.CV 83%

ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning

ABMAMBA: 多模态大语言模型中的对齐层次双向扫描用于高效的视频描述

Daichi Yashima, Shuhei Kurita, Yusuke Oda, Shuntaro Suzuki, Seitaro Otsuki, Komei Sugiura

机构 * Keio University(庆应义塾大学) National Institute of Informatics(国立信息学研究所) National Institute of Informatics Research and Development Center for Large Language Models(国立信息学研究所大型语言模型研发中心)

专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV

AI总结 本文提出ABMamba,一种具有线性计算复杂度的多模态大语言模型,通过替代二次注意力机制,实现视频序列的高效处理,在视频描述任务中表现出色。

Comments Accepted to ICPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08063 2026-04-10 cs.CV 57%

EEG2Vision: A Multimodal EEG-Based Framework for 2D Visual Reconstruction in Cognitive Neuroscience

EEG2Vision:一种基于EEG的多模态框架,用于认知神经科学中的2D视觉重建

Emanuele Balloni, Emanuele Frontoni, Chiara Matti, Marina Paolanti, Roberto Pierdicca, Emiliano Santarnecchi

机构 * Università Politecnica delle Marche(马尔凯理工大学) University of Macerata(马切拉塔大学) Horizon Intelligence Labs(地平线智能实验室) Harvard Medical School(哈佛医学院) Massachusetts General Hospital(麻省总医院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

AI总结 EEG2Vision通过多模态大语言模型和图像扩散模型提升低密度电极配置下的视觉重建质量,验证了低分辨率EEG设备在实时脑-图像应用中的可行性。

Comments 17 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14611 2026-04-10 cs.HC 50%

Exploring MLLMs Perception of Network Visualization Principles

探索 MLLMs 对网络可视化原则的感知

Jacob Miller, Markus Wallinger, Ludwig Felder, Timo Brand, Henry Förster, Johannes Zink, Chunyang Chen, Stephen Kobourov

专题命中 其他VLM :multimodal large language model(abstract)

AI总结 研究测试多模态大语言模型是否能在网络布局属性感知任务中匹配人类表现,发现MLLMs在同等条件下表现优于非专家,且提示工程可提升性能,表明其可能依赖视觉代理而非计算实际值。

详情

展开后加载摘要…

URL PDF HTML 收藏