arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4721 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4721 篇

2509.10887 2026-03-20 cs.CV 79%

AutoOEP -- A Multi-modal Framework for Online Exam Proctoring

AutoOEP -- 一种用于在线考试监考的多模态框架

Aryan Kashyap Naveen

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出AutoOEP,通过计算机视觉和机器学习实现高效的自动监考,采用双摄像头和LSTM网络检测可疑行为,准确率达90.7%。

Comments arXiv admin comment: This version has been removed by arXiv administrators as the submitter did not have the rights to agree to the license at the time of submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17541 2026-03-19 cs.CV 79%

Temporal Gains, Spatial Costs: Revisiting Video Fine-Tuning in Multimodal Large Language Models

时间收益,空间代价:重新审视多模态大语言模型中的视频微调

Linghao Zhang, Jungang Li, Yonghua Hei, Sicheng Tao, Song Dai, Yibo Yan, Zihao Dongfang, Weiting Liu, Chenxi Qin, Hanqian Li, Xin Zou, Jiahao Zhang, Shuhang Xun, Haiyun Jiang, Xuming Hu

机构 * SJTU(上海交通大学) HKUST(GZ)(香港科技大学(珠海)) HKUST(香港科技大学) CityU(城市大学) FDU(复旦大学) HIT(哈尔滨工业大学) TJU(天津大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文研究了视频微调对多模态大语言模型视觉能力的影响,发现视频性能提升的同时,静态图像表现可能受限,提出混合帧策略以缓解空间与时间理解的平衡问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15392 2026-03-17 cs.MM cs.HC 79%

Multimodal Cyber-physical Interaction in XR: Hybrid Doctoral Thesis Defense

XR中的多模态物理交互:混合博士论文答辩

Ahmad Alhilal, Kit Yung Lam, Lik-Hang Lee, Xuetong Wang, Sijia Li, Matti Siekkinen, Tristan Braud, Pan Hui

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.MM

AI总结 本文提出一种多模态框架,突破传统答辩模式,支持现场、虚拟现实和浏览器等多种参与方式,通过全身体姿追踪实现自然交互,验证了其在混合活动中的有效性。

Comments 10 pages, 3 figures, magazine paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07463 2026-03-17 cs.CV cs.LG eess.IV 79%

MSEG-VCUQ: Multimodal SEGmentation with Enhanced Vision Foundation Models, Convolutional Neural Networks, and Uncertainty Quantification for High-Speed Video Phase Detection Data

MSEG-VCUQ: 多模态分割与增强视觉基础模型、卷积神经网络和不确定性量化用于高速视频相位检测数据

Chika Maduabuchi, Ericmoore Jossou, Matteo Bucci

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出MSEG-VCUQ,结合U-Net和SAM提升高速视频相位检测分割精度,引入不确定性量化和首个开源多模态数据集,实现更可靠的相位分割。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13759 2026-03-17 cs.CV 79%

QTrack: Query-Driven Reasoning for Multi-modal MOT

QTrack: 基于查询的多模态 MOT 推理

Tajamul Ashraf, Tavaheed Tariq, Sonia Yadav, Abrar Ul Riyaz, Wasif Tak, Moloud Abdar, Janibul Bashir

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出基于查询的多模态 MOT 推理方法,通过自然语言查询定位和跟踪特定目标,构建了 RMOT26 基准并提出 QTrack 模型,结合多模态推理与跟踪定位,提升语言引导的跟踪性能。

Comments Project Page: https://gaashlab.github.io/QTrack/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13367 2026-03-17 cs.CV cs.LG 79%

Multimodal Deep Learning for Dynamic and Static Neuroimaging: Integrating MRI and fMRI for Alzheimer Disease Analysis

多模态深度学习用于动态和静态神经影像:整合MRI和fMRI进行阿尔茨海默病分析

Anima Kujur, Zahra Monfared

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出一种多模态深度学习框架,整合MRI和fMRI对阿尔茨海默病、轻度认知障碍和正常认知状态进行多类分类,通过3D卷积神经网络提取结构特征,用递归架构学习fMRI的时间特征,并融合这些表示进行联合空间-时间学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12746 2026-03-16 cs.CV 79%

Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World

动态思维:多模态大语言模型如何感知、跟踪和推理物理4D世界的动态

Yuzhi Huang, Kairun Wen, Rongxin Gao, Dongxuan Liu, Yibin Lou, Jie Wu, Jing Xu, Jian Zhang, Zheng Yang, Yunlong Lin, Chenxin Li, Panwang Pan, Junbin Lu, Jingyan Jiang, Xinghao Ding, Yue Huang, Zhi Wang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出Dyn-Bench基准测试,评估多模态大语言模型在动态场景中的时空推理和动态物体感知能力,发现现有模型难以同时保持强性能,而结构化整合方法显著提升其动态感知能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09287 2026-03-11 cs.CV 79%

Exploring Modality-Aware Fusion and Decoupled Temporal Propagation for Multi-Modal Object Tracking

探索多模态感知融合与解耦时间传播用于多模态目标跟踪

Shilei Wang, Pujian Lai, Dong Gao, Jifeng Ning, Gong Cheng

专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 MDTrack通过模态感知融合和解耦时间传播提升多模态目标跟踪性能,实现统一的模态处理和独立的时间信息捕捉。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11954 2026-03-10 cs.AI 79%

UniCast: A Unified Framework for Instance-Conditioned Multimodal Time-Series Forecasting

UniCast: 一个实例条件多模态时间序列预测的统一框架

Sehyuk Park, Soyeon Caren Han, Eduard Hovy

机构 * The Universito of Melbourne(墨尔本大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 UniCast通过实例条件提示和动态模态路由,实现多模态时间序列预测的统一框架,有效提升预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05581 2026-03-09 cs.LG cs.AI eess.SP 79%

Spatiotemporal Heterogeneity of AI-Driven Traffic Flow Patterns and Land Use Interaction: A GeoAI-Based Analysis of Multimodal Urban Mobility

人工智能驱动交通流模式与土地利用相互作用的时空异质性:基于GeoAI的多模式城市交通分析

Olaf Yunus Laitinen Imanov

机构 * Department of Applied Mathematics(应用数学系) Computer Science (DTU Compute), Technical University of Denmark, Kongens Lyngby, Denmark(计算机科学(DTU Compute),丹麦技术大学,Kongens Lyngby)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本研究提出GeoAI混合框架,通过整合MGWR、RF和ST-GCN,分析多模式城市交通的时空异质性及土地利用互动,提升交通管理与政策设计的科学性。

Comments 13 pages, 7 figures, 9 tables. Submitted to Computers, Environment and Urban Systems (Elsevier)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04509 2026-03-06 cs.CV 79%

Recognition of Daily Activities through Multi-Modal Deep Learning: A Video, Pose, and Object-Aware Approach for Ambient Assisted Living

通过多模态深度学习识别日常活动:一种面向环境辅助养老的视频、姿态和物体感知方法

Kooshan Hashemifard, Pau Climent-Pérez, Francisco Florez-Revuelta

机构 * Alicante Institute for Health and Biomedical Research (ISABIAL)(阿利坎特健康与生物医学研究所(ISABIAL)) ValgrAI - Valencian Graduate School and Research Network of Artificial Intelligence(瓦尔格AI——瓦伦西亚人工智能研究生院与研究网络)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出一种多模态深度学习方法,结合视频、姿态和物体信息,用于老年人日常活动识别,提升环境辅助养老系统的准确性与实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02481 2026-03-04 cs.CV 79%

ModalPatch: A Plug-and-Play Module for Robust Multi-Modal 3D Object Detection under Modality Drop

ModalPatch: 一种插拔式模块,用于在模态缺失情况下鲁棒的多模态3D目标检测

Shuangzhi Li, Lei Ma, Xingyu Li

机构 * University of Alberta, Canada(阿尔伯塔大学) The University of Tokyo, Japan(东京大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 ModalPatch是一种插拔式模块,通过利用传感器数据的时间特性实现感知连续性,提升多模态3D目标检测在模态缺失情况下的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01701 2026-03-03 cs.DB cs.AI 79%

Beyond Single-Modal Analytics: A Framework for Integrating Heterogeneous LLM-Based Query Systems for Multi-Modal Data

超越单一模态分析:一种整合异构LLM查询系统的框架用于多模态数据

Ruyu Li, Tinghui Zhang, Haodi Ma, Daisy Zhe Wang, Yifan Wang

机构 * University of Hawaii at Manoa(夏威夷大学曼纳oa分校) University of Florida(佛罗里达大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出Meta Engine,一种整合异构LLM查询系统的框架,通过统一的语义查询引擎提升多模态数据处理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00405 2026-03-03 cs.LG cs.AI 79%

UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings

UME-R1: 探索基于推理的生成多模态嵌入

Zhibin Lan, Liqiang Niu, Fandong Meng, Jie Zhou, Jinsong Su

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) WeChat AI, Tencent Inc, China(腾讯公司微信AI部门) Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室) Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 UME-R1通过生成嵌入和强化学习提升多模态嵌入性能,实现判别与生成嵌入的互补,展现生成嵌入在推理和下游任务中的优势。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03905 2026-03-03 cs.CV 79%

EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation

EchoMimicV3:13亿参数足矣实现统一的多模态和多任务人类动画

Rang Meng, Yan Wang, Weipeng Wu, Ruobing Zheng, Yuming Li, Chenguang Ma

机构 * Terminal Technology Department, Alipay, Ant Group(蚂蚁集团支付宝终端技术部)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 EchoMimicV3通过统一多任务和多模态人类动画,以13亿参数实现高效且稳定的多任务和多模态动画生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23937 2026-03-02 cs.RO cs.CV 79%

Enhancing Vision-Language Navigation with Multimodal Event Knowledge from Real-World Indoor Tour Videos

通过真实世界室内游览视频的多模态事件知识增强视觉语言导航

Haoxuan Xu, Tianfu Li, Wenbo Chen, Yi Liu, Xingxing Zuo, Yaoxian Song, Haoang Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tsinghua University(清华大学) Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)(马尔代夫 bin Zayed 大学人工智能学院) Hangzhou City University(杭州城市学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于多模态事件知识的视觉语言导航增强方法,通过构建大规模时空知识图谱并结合层次检索机制,提升长视界推理和粗粒度指令处理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08578 2026-03-02 cs.CV 79%

Multimodal Knowledge Distillation for Egocentric Action Recognition Robust to Missing Modalities

多模态知识蒸馏用于抗缺失模态的自体视觉动作识别

Maria Santos-Villafranca, Dustin Carrión-Ojeda, Alejandro Perez-Yus, Jesus Bermudez-Cameo, Jose J. Guerrero, Simone Schaub-Meyer

机构 * I3A – University of Zaragoza(萨拉戈萨大学I3A研究中心) Technical University of Darmstadt, Department of Computer Science(达姆施塔特技术大学计算机科学系)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 KARMMA通过多模态知识蒸馏实现抗缺失模态的自体视觉动作识别,以轻量模型提升机器人部署效率。

Comments Project Page: https://visinf.github.io/KARMMA

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22920 2026-02-27 cs.CV 79%

OSDaR-AR: Enhancing Railway Perception Datasets via Multi-modal Augmented Reality

OSDaR-AR: 通过多模态增强现实提升铁路感知数据集

Federico Nesti, Gianluca D'Amico, Mauro Marinoni, Giorgio Buttazzo

机构 * Department of Excellence in Robotics & AI, Scuola Superiore Sant’Anna(机器人与人工智能卓越部门,圣安娜高等学院) Simulatrix MV srl(Simulatrix MV公司)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出OSDaR-AR数据集,通过多模态增强现实技术整合逼真虚拟对象,提升铁路感知任务的数据质量与真实性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11910 2026-02-26 cs.CV 79%

Seeing the Forest and the Trees: Query-Aware Tokenizer for Long-Video Multimodal Language Models

看清森林与树木:面向长视频多模态语言模型的查询感知分词器

Siyou Li, Huanan Wu, Juexi Shao, Yinghao Ma, Yujian Gan, Yihao Luo, Yuwei Wang, Dong Nie, Lu Wang, Wenqing Wu, Le Zhang, Massimo Poesio, Juntao Yu

机构 * Queen Mary University of London(伦敦女王学院) University of Sheffield(谢菲尔德大学) Imperial College London(伦敦帝国学院) Pengcheng Laboratory(鹏城实验室) Meta Inc(Meta公司) Meituan Inc(美团公司) Nanjing University of Science(南京理工大学) University of Birmingham(伯明翰大学) Utrecht University(乌得勒支大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 QTSplus是一种轻量高效的视觉token选择模块,通过动态选择重要视觉证据提升长视频多模态语言模型的性能,显著降低计算成本并提高处理效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19063 2026-02-24 cs.CV 79%

Direction-aware 3D Large Multimodal Models

具有方向感知的3D大多模态模型

Quan Liu, Weihao Xuan, Junjue Wang, Naoto Yokoya, Ling Shao, Shijian Lu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出方向感知的3D大多模态模型,通过补充自身姿态并改进点云数据对齐,提升3D多模态模型的性能和通用性。

Comments In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.17252 2026-02-20 cs.CV cs.SY eess.IV eess.SY 79%

A Multi-modal Detection System for Infrastructure-based Freight Signal Priority

基于基础设施的货运信号优先的多模态检测系统

Ziyan Zhang, Chuheng Wei, Xuanpeng Zhao, Siyan Li, Will Snyder, Mike Stas, Peng Hao, Kanok Boriboonsomsin, Guoyuan Wu

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于激光雷达和摄像头的多模态货运车辆检测系统,通过混合传感架构和卡尔曼滤波实现稳定实时性能,用于支持基于基础设施的货运信号优先应用。

Comments 12 pages, 15 figures. Accepted at ICTD 2026. Final version to appear in ASCE Proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10902 2026-02-18 cs.CL 79%

Multimodal Peer Review Simulation with Actionable To-Do Recommendations for Community-Aware Manuscript Revisions

多模态同行评审模拟:面向社区意识的论文修订可操作待办推荐系统

Mengze Hong, Di Jiang, Weiwei Zhao, Yawen Li, Yihang Wang, Xinyuan Luo, Yanjie Sun, Chen Jason Zhang

机构 * Hong Kong Polytechnic University(香港理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Independent Researcher(独立研究者)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本研究提出一个多模态同行评审模拟系统,通过整合文本和视觉信息,利用检索增强生成技术提升评审质量,并生成可操作的待办列表,以提高论文修订的效率和准确性。

Comments Accepted by TheWebConf 2026 Demo Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21842 2026-02-17 cs.CV cs.CR 79%

Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?

模态失语:统一多模态模型能否从记忆中描述图像?

Michael Aerni, Joshua Swanson, Kristina Nikolić, Florian Tramèr

机构 * Michael Aerni(独立研究者) Joshua Swanson(独立研究者) Kristina Nikolić(独立研究者) Florian Tramèr(独立研究者)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 研究发现统一多模态模型在视觉记忆与文本表达间存在系统性缺陷,导致安全框架可能因单一模态防护而失效。

Comments Accepted to ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12593 2026-02-16 cs.IR cs.AI 79%

RQ-GMM: Residual Quantized Gaussian Mixture Model for Multimodal Semantic Discretization in CTR Prediction

RQ-GMM:用于CTR预测的多模态语义离散化残差量化高斯混合模型

Ziye Tong, Jiahao Liu, Weimin Zhang, Hongji Ruan, Derick Tang, Zhanpeng Zeng, Qinsong Zeng, Peng Zhang, Tun Lu, Ning Gu

机构 * Tencent(腾讯) Fudan University(复旦大学) Beijing Jiaotong University(北京交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 RQ-GMM通过残差量化高斯混合模型提升CTR预测中多模态语义离散化效果,实现代码本利用和重建准确性的显著提升。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09638 2026-02-11 cs.CV 79%

VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model

VideoAfford: 通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding

Hanqing Wang, Mingyu Liu, Xiaoyu Chen, Chengwei MA, Yiming Zhong, Wenti Yin, Yuhao Liu, Zhiqing Cui, Jiahao Yuan, Lu Dai, Zhiyuan Ma, Hui Xiong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 VideoAfford通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding,结合动态交互先验和空间感知损失函数,提升机器人操作的可操作区域识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08861 2026-02-10 cs.CV 79%

TiFRe: Text-guided Video Frame Reduction for Efficient Video Multi-modal Large Language Models

TiFRe: 基于文本的视频帧减少用于高效视频多模态大语言模型

Xiangtian Zheng, Zishuo Wang, Yuxin Peng

机构 * Wangxuan Institute of Computer Technology, Peking University(王轩计算机技术研究所,北京大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 TiFRe通过文本引导的帧采样和帧匹配机制,有效减少视频输入帧数,同时保留关键信息,提升视频多模态大语言模型的效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13564 2026-02-10 cs.CV 79%

State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models

具有门控注意力和可学习采样的状态空间分层压缩用于大多模态模型中的小时级视频理解

Geewook Kim, Minjoon Seo

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种基于状态空间模型和门控注意力的高效压缩方法,用于减少大模型中小时级视频的token消耗,同时保持性能。

Comments AAAI 2026 (Oral). Project page: https://github.com/naver-ai/mambamia

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13928 2026-02-05 cs.CV cs.IR 79%

LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts

LoVR:一种多模态背景下长视频检索的基准

Qifeng Cai, Hao Liang, Zhaoyang Han, Hejun Dong, Meiyi Qiang, Ruichuan An, Quanqing Xu, Bin Cui, Wentao Zhang

机构 * East China Normal University Shanghai China Peking University \& Zhongguancun Academy Beijing China Huazhong University of Science Beihang University Beijing China Peking University Beijing China East China Normal University Peking University \& Zhongguancun Academy Beihang University Peking University

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 LoVR是一个针对长视频检索的多模态基准,通过高质量标注和细粒度数据提升视频理解挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02453 2026-02-04 cs.AI 79%

Thinking with Comics: Enhancing Multimodal Reasoning through Structured Visual Storytelling

用漫画思考:通过结构化视觉叙事增强多模态推理

Andong Chen, Wenxin Zhu, Qiuyu Ding, Yuchen Song, Muyun Yang, Tiejun Zhao

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出用漫画作为中间视觉表示,通过结构化视觉叙事提升多模态推理效率和性能。

Comments Working paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01125 2026-02-03 cs.CL cs.LG 79%

Long-range Modeling and Processing of Multimodal Event Sequences

多模态事件序列的长程建模与处理

Jichu Li, Yilun Zhong, Zhiting Li, Feng Zhou, Quyu Kong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出了一种基于LLM的多模态时间点过程框架,通过自适应序列压缩解决长上下文问题,提升多模态事件序列的建模与生成能力。

详情

展开后加载摘要…

URL PDF HTML 收藏