arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视频大模型

视频理解、视频生成、视频语言模型和时序视觉推理。

2026-08-19 至 2026-08-19 共收录 3 信号源:cs.CV, eess.IV, cs.MM

1. 视频理解 3 篇

2608.17402 2026-08-19 cs.CV 新提交 74%

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

MoE-ViE:用于高效图像与视频理解的混合专家视觉编码器

Bonan Zhang, Shiyu Dong, Quan Hung Tran, Katharina Gschwind, Shuqi Yang, Sijia Chen, Adel Ahmadyan, Seungwhan Moon, Lu Zhang, Ahmed Kirmani, Babak Damavandi, Anuj Kumar

机构 * Meta

专题命中 视频理解 :video understanding(title);分类 cs.CV

AI总结 本研究提出MoE-ViE,通过细粒度MoE拓扑、无辅助损失平衡变体、专用MoE内核及帧级蒸馏等设计,实现高效图像与视频理解,性能优于对应密集模型及更大规模SOTA编码器。

Comments Accepted to ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17935 2026-08-19 cs.CV 新提交 61%

Beyond Instrument Motion: Recognizing Tissue Tension Toward Surgical Skill Assessment

超越器械运动:面向手术技能评估的组织张力识别

Marko Haralovi, Zhiqi Miao, Alexander Machiel Bont, Jiapan Guo, Frans van Workum, Estefania Talavera

机构 * University of Zagreb(萨格勒布大学) University of Twente(特文特大学) University of Groningen(格罗宁根大学) Radboud University Medical Center(拉德堡德大学医学中心) Canisius-Wilhelmina Hospital(卡尼修斯-威廉明娜医院)

专题命中 视频理解 :video understanding(abstract,comments);分类 cs.CV

AI总结 针对现有手术视频理解方法未捕捉组织张力的问题,本文构建SurgTension数据集,提出TensionTRAC框架,实现了腹腔镜及机器人辅助直肠癌手术的组织张力识别,为手术技能评估提供客观基准。

Comments The paper is accepted by ECCV 2026 Workshop On Medical Video Understanding and submitted the camera-ready version to the ECCV organization

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01460 2026-08-19 cs.CV 版本更新 57%

Reinforcing Consistency in Video MLLMs with Structured Rewards

通过结构化奖励强化视频MLLMs的一致性

Yihao Quan, Zeru Shi, Jinman Zhao, Ruixiang Tang

机构 * Rutgers University(罗格斯大学) University of Toronto(多伦多大学)

专题命中 视频理解 :video understanding(abstract);分类 cs.CV

AI总结 研究通过结构化奖励提升视频MLLMs的一致性,发现传统监督不足,提出结合事实和时间单元的奖励机制,提升视频理解准确性。

Comments Accepted by COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏