arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4703 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 视频多模态 4703 篇

2602.14589 2026-02-17 cs.AI cs.CL cs.LG 81%

MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMs

MATEO:一种多模态基准,用于LVLMs中的时间推理和规划

Gabriel Roccabruna, Olha Khomyn, Giuseppe Riccardi

机构 * Signals and Interactive Systems Lab, University of Trento, Italy(特伦托大学信号与交互系统实验室) University of Trento(特伦托大学) Amazon(亚马逊)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 MATEO是一个多模态基准,用于评估和提升大型视觉语言模型在时间推理和规划方面的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10549 2026-02-12 cs.CV cs.AI 81%

Enhancing Weakly Supervised Multimodal Video Anomaly Detection through Text Guidance

通过文本引导增强弱监督多模态视频异常检测

Shengyang Sun, Jiashen Hua, Junyi Feng, Xiaojin Gong

机构 * School of Computer Science and Technology, Hangzhou Dianzi University(杭州电子科技大学计算机科学与技术学院) Alibaba Cloud(阿里云) College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出文本引导框架,通过多阶段文本增强和多尺度瓶颈Transformer融合模块,提升弱监督多模态视频异常检测的性能。

Comments Accepted by IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18862 2026-02-11 cs.CV cs.AI 81%

TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models

TAMMs: 卫星图像时间序列中基于时序感知的多模态模型用于变化理解和预测

Zhongbin Guo, Yuhao Wang, Ping Jian, Chengzhi Li, Xinyue Chen, Zhen Yang, Ertai E

机构 * School of Computer Science & Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) School of Computing, National University of Singapore(计算学院,新加坡国立大学)

专题命中 视频多模态 :multimodal(title);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 TAMMs通过时序感知多模态模型统一执行卫星图像时间序列中的变化理解和未来预测,提升长程时间动态建模能力。

Comments Published as a conference paper at The Fourteenth International Conference on Learning Representations (ICLR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08057 2026-02-10 cs.CV cs.AI 81%

Weak to Strong: VLM-Based Pseudo-Labeling as a Weakly Supervised Training Strategy in Multimodal Video-based Hidden Emotion Understanding Tasks

弱到强:基于VLM的伪标签作为多模态视频隐藏情绪理解任务中的弱监督训练策略

Yufei Wang, Haixu Liu, Tianxiang Xu, Chuancheng Shi, Hongsheng Xing

机构 * The University of New South Wales, Sydney, New South Wales, Australia(新南威尔士大学) The University of Sydney, Sydney, New South Wales, Australia(悉尼大学) School of Software and Microelectronics, Peking University, Zibo, Shandong, China(北京大学软件与微电子学院) Shandong University of Technology, Beijing, China(山东理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出基于VLM的伪标签弱监督方法,用于多模态视频隐藏情绪识别,提升准确率至0.69并建立新基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00559 2026-02-03 cs.CV cs.AI 81%

Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models

在视频多模态大语言模型中学习对抗组合性幻觉

Wenbin Xing, Quanxing Zha, Lizheng Zu, Mengran Li, Ming Li, Junchi Yan

机构 * Sun Yat-sen University(中山大学) Huaqiao University(华侨大学) Shenzhen University(深圳大学) Guangming Laboratory(光明实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 针对视频多模态大语言模型中的组合性幻觉问题,提出TriCD框架,通过对比解码和三路径校准机制提升模型在对抗幻觉时的准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18192 2026-01-28 cs.CV cs.HC cs.MM 81%

MindCine: Multimodal EEG-to-Video Reconstruction with Large-Scale Pretrained Models

MindCine: 基于大规模预训练模型的多模态EEG到视频重建

Tian-Yi Zhou, Xuan-Hao Liu, Bao-Liang Lu, Wei-Long Zheng

机构 * School of Computer Science, Shanghai Jiao Tong University, 800 Dongchuan Road, Shanghai, China(计算机科学学院,上海交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

AI总结 MindCine通过多模态联合学习和预训练大型EEG模型,实现有限数据下的高保真EEG到视频重建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08512 2026-01-28 cs.CV cs.AI 81%

MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding

MLVTG: 基于Mamba的特征对齐与LLM驱动的多模态视频时间定位

Zhiyi Zhu, Xiaoyu Wu, Zihao Liu, Linlin Yang

机构 * State Key Laboratory of Media Convergence and Communication(媒体融合与传播国家重点实验室)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 MLVTG通过结合MambaAligner和LLMRefiner,实现了多模态视频时间定位的高精度定位与语义净化。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17466 2026-01-27 cs.HC cs.AI cs.CL 81%

Autiverse: Eliciting Autistic Adolescents' Daily Narratives through AI-guided Multimodal Journaling

Autiverse: 通过AI引导的多模态日记记录 eliciting 自闭症青少年的日常叙述

Migyeong Yang, Kyungah Lee, Jinyoung Han, SoHyun Park, Young-Ho Kim

机构 * NAVER AI Lab(NAVER AI实验室) Dodakim Child Development Center(Dodakim儿童发展中心) Sungkyunkwan University(松均大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 Autiverse通过AI引导的多模态日记记录帮助自闭症青少年组织日常经历,提升叙述能力并增强自主感。

Comments 19 pages excluding reference. Conditionally accepted to ACM CHI 2026

Journal ref In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26), April 13-17, 2026, Barcelona, Spain. ACM, New York, NY, USA, 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17041 2026-01-27 cs.CV cs.AI 81%

Arabic Sign Language Recognition using Multimodal Approach

阿拉伯手语识别的多模态方法

Ghadeer Alanazi, Abir Benabid

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出结合Leap Motion和RGB摄像头的多模态方法,用于提高阿拉伯手语识别的准确率,实验结果显示78%的识别准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12432 2026-01-21 cs.CV cs.MM 81%

SkeFi: Cross-Modal Knowledge Transfer for Wireless Skeleton-Based Action Recognition

SkeFi: 无线传感器跨模态知识转移用于基于骨骼的动作识别

Shunyu Huang, Yunjiao Zhou, Jianfei Yang

机构 * School of Electrical and Electronics Engineering, Nanyang Technological University, Singapore(南洋理工大学电子与电气工程学院)

专题命中 视频多模态 :cross-modal(title,abstract);分类 cs.CV、cs.MM

AI总结 SkeFi通过跨模态知识转移方法,利用无线传感器提升基于骨骼的动作识别性能,实现毫米波和LiDAR上的先进表现。

Comments Published in IEEE Internet of Things Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10228 2026-01-16 cs.CV cs.MM eess.IV 81%

Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge

优化多模态大语言模型以实现视角视频理解:HD-EPIC VQA挑战的解决方案

Sicheng Yang, Yukai Huang, Shitong Sun, Weitong Cai, Jiankang Deng, Jifei Song, Zhensong Zhang

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

AI总结 本文提出一种优化多模态大语言模型的方法,通过预处理、微调和时间链式思考提示,在HD-EPIC VQA挑战中实现了41.6%的准确率。

Comments 4 pages, 1 figure, CVPR 2025 EgoVis Workshop, 2nd Place in HD-EPIC Challenge

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06573 2026-01-13 cs.AI cs.MM 81%

QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models

QMAVIS:利用大多模态模型融合实现长视频音频理解

Zixing Lin, Jiale Wang, Gee Wah Ng, Lee Onn Mak, Chan Zhi Yang Jeriel, Jun Yang Lee, Yaohao Li

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University, Singapore(南洋理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM

AI总结 QMAVIS通过融合大型多模态模型、大型语言模型和语音识别模型,实现了长视频音频理解,展示了在VideoMME数据集上38.75%的性能提升,并在其他数据集上表现出竞争力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.06843 2026-01-13 cs.CV cs.CL 81%

Speak While Watching: Unleashing TRUE Real-Time Video Understanding Capability of Multimodal Large Language Models

边看边说:解锁多模态大语言模型的实时视频理解能力

Junyan Lin, Junlong Tong, Hao Wu, Jialiang Zhang, Jinming Liu, Xin Jin, Xiaoyu Shen

机构 * Department of Computing, The Hong Kong Polytechnic University(香港理工大学计算机系) Ningbo Key Laboratory of Spatial Intelligence and Digital Derivative, Institute of Digital Twin, EIT(宁波空间智能与数字衍生关键实验室,数字孪生研究院,EIT) Shanghai Jiao Tong University(上海交通大学) Ocean University of China(中国海洋大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 本文提出并行流式框架,通过三种设计解决多模态大语言模型在实时视频理解中的位置连续性约束问题,实现边看边说的实时交互。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05495 2026-01-12 cs.CV cs.CL 81%

MMViR: A Multi-Modal and Multi-Granularity Representation for Long-range Video Understanding

MMViR:一种多模态和多粒度表示用于长视频理解

Zizhong Li, Haopeng Zhang, Jiawei Zhang

机构 * IFM Lab, University of California, Davis(加州大学戴维斯分校信息融合实验室) ALOHA Lab, University of Hawaii at Mānoa(夏威夷大学马诺阿分校ALOHA实验室)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.CL

AI总结 MMViR通过多模态和多粒度表示提升长视频理解,实现更高效的检索和更优的性能表现。

Comments 13 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21641 2025-12-29 cs.CV cs.AI 81%

TrackTeller: Temporal Multimodal 3D Grounding for Behavior-Dependent Object References

TrackTeller: 基于时间的多模态3D定位用于行为依赖的对象引用

Jiahong Yu, Ziqi Wang, Hailiang Zhao, Wei Zhai, Xueqiang Yan, Shuiguang Deng

机构 * Zhejiang University(浙江大学) Fudan University(复旦大学) Huawei Technologies Ltd.(华为技术有限公司)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 TrackTeller通过融合LiDAR图像、语言条件解码和时间推理,提升动态3D场景中基于语言的物体定位与跟踪性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.13460 2025-12-29 cs.CV cs.AI 81%

Chain-of-Evidence Multimodal Reasoning for Few-shot Temporal Action Localization

基于证据链的多模态推理用于少样本时序动作定位

Mengshi Qi, Hongwei Ji, Wulian Yun, Xianlin Zhang, Huadong Ma

机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出基于证据链的多模态推理方法,通过结合文本和视觉信息提升少样本时序动作定位的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14058 2025-12-17 cs.CV cs.AI 81%

Real-time prediction of workplane illuminance distribution for daylight-linked controls using non-intrusive multimodal deep learning

基于非侵入式多模态深度学习的实时工作平面照度分布预测用于日光联动控制

Zulin Zhuang, Yu Bian

机构 * School of Architecture(建筑学院) State Key Laboratory of Subtropical Building(亚热带建筑科学国家重点实验室) South China University of Technology(华南理工大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究提出了一种基于非侵入式多模态深度学习的实时工作平面照度分布预测方法,用于提高日光联动控制的能效

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20268 2025-12-16 cs.CV cs.MM 81%

GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection

GMFVAD:利用细粒度多模态特征提升视频异常检测

Guangyu Dai, Dong Chen, Siliang Tang, Yueting Zhuang

机构 * Zhejiang University(浙江大学) Zhengzhou University(郑州大学)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.MM

AI总结 GMFVAD通过细粒度多模态特征减少冗余信息,提升视频异常检测性能。

Comments Accepted for publication in the Proceedings of the ICONIP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06190 2025-12-09 cs.CV cs.AI cs.LG eess.IV 81%

Multi-Modal Zero-Shot Prediction of Color Trajectories in Food Drying

多模态零样本预测食品干燥中的颜色轨迹

Shichen Li, Ahmadreza Eslaminia, Chenhui Shao

机构 * Department of Mechanical Science and Engineering, University of Illinois at Urbana-Champaign, Urbana, IL, USA(机械科学与工程系,伊利诺伊大学厄巴纳-香槟分校) Department of Mechanical Engineering, University of Michigan, Ann Arbor, MI, USA(机械工程系,密歇根大学安娜堡分校)

专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种多模态零样本方法,通过整合高维时间颜色信息和干燥参数,实现对食品干燥中颜色轨迹的高效预测,显著提升预测精度和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07755 2025-12-08 cs.CV cs.AI cs.GR cs.RO 81%

SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models

SAT:多模态语言模型的动态空间能力训练

Arijit Ray, Jiafei Duan, Ellis Brown, Reuben Tan, Dina Bashkirova, Rose Hendrix, Kiana Ehsani, Aniruddha Kembhavi, Bryan A. Plummer, Ranjay Krishna, Kuo-Hao Zeng, Kate Saenko

机构 * Boston University(波士顿大学) University of Washington(华盛顿大学) Allen Institute for AI(人工智能研究院) Microsoft Research(微软研究院) New York University(纽约大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 SAT通过模拟数据提升多模态语言模型在动态空间推理中的能力,实验表明其在多个基准测试中优于现有方法。

Comments Accepted to COLM 2025. Project webpage: https://arijitray.com/SAT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00714 2025-12-02 cs.CV cs.AI 81%

Deep Learning-Based Computer Vision Models for Early Cancer Detection Using Multimodal Medical Imaging and Radiogenomic Integration Frameworks

基于深度学习的计算机视觉模型用于多模态医学影像和放射基因组整合框架中的早期癌症检测

Emmanuella Avwerosuoghene Oghenekaro

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出基于深度学习的计算机视觉模型,结合多模态医学影像和放射基因组学,用于早期癌症检测,提升肿瘤基因型预测和治疗抗性分析能力。

Journal ref International Journal of Computer Applications Technology and Research, vol. 14, no. 11, pp. 1-14, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12263 2025-12-02 cs.CV cs.AI 81%

CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models

CrossVid: 一个用于评估多模态大语言模型在跨视频推理中的综合基准

Jingyao Li, Jingyun Wang, Molin Tan, Haochen Wang, Cilin Yan, Likun Shi, Jiayin Cai, Xiaolong Jiang, Yao Hu

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 CrossVid是首个用于评估多模态大语言模型跨视频推理能力的综合基准,通过多样化的任务和数据集验证了模型在复杂视频推理任务中的表现。

Comments Accepted to AAAI 2026 (main track). For code and data, see https://github.com/chuntianli666/CrossVid

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06859 2025-11-25 cs.AI cs.CV 81%

MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction

MeteorPred:一种用于恶劣天气事件预测的气象多模态大模型和数据集

Shuo Tang, Jian Xu, Jiadong Zhang, Yi Chen, Qizhao Jin, Lingdong Shen, Chenglin Liu, Shiming Xiang

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Zhongguancun Academy, Beijing(中关村学院) China Meteorological Administration(中国气象局)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 MeteorPred提出一种多模态大模型和数据集,用于提升恶劣天气事件预测的准确性和自动化水平。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07290 2025-11-11 eess.IV cs.CV cs.MM 81%

CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video

Xinyi Wang, Angeliki Katsenou, Junxiao Shen, David Bull

机构 * School of Computer Science, University of Bristol(布里斯托大学计算机科学学院)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 14 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14809 2025-11-05 cs.CV cs.MM cs.RO 81%

Light Future: Multimodal Action Frame Prediction via InstructPix2Pix

Zesen Zhong, Duomin Zhang, Yijia Li

机构 * School of Data Science, The Chinese University of Hong Kong, Shenzhen(数据科学学院,香港中文大学(深圳))

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.MM

Comments 9 pages including appendix, 4 tables, 8 figures, to be submitted to WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17394 2025-10-28 cs.CV cs.AI 81%

HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs

Zhaolin Cai, Fan Li, Ziwei Zheng, Yanjun Qin

机构 * Xinjiang University(新疆大学) Xi'an Jiaotong University(西安交通大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21761 2025-10-28 cs.RO cs.AI cs.CV 81%

J-ORA: A Framework and Multimodal Dataset for Japanese Object Identification, Reference, Action Prediction in Robot Perception

Jesse Atuhurra, Hidetaka Kamigaito, Taro Watanabe, Koichiro Yoshino

机构 * Division of Information Science, NAIST(NAIST信息科学系) Guardian Robot Project, RIKEN(RIKEN守护机器人项目)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

Comments Accepted to IROS2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.17038 2025-10-21 cs.RO cs.AI cs.CV 81%

DINO-CVA: A Multimodal Goal-Conditioned Vision-to-Action Model for Autonomous Catheter Navigation

Pedram Fekri, Majid Roshanfar, Samuel Barbeau, Seyedfarzad Famouri, Thomas Looi, Dale Podolsky, Mehrdad Zadeh, Javad Dargahi

机构 * Gina Cody School of Engineering and Computer Science, Concordia University(甘娜·柯迪工程与计算机科学学院,康科迪亚大学) The Wilfred and Joyce Posluns Centre for Image Guided Innovation & Therapeutic Intervention (PCIGITI) at the Hospital for Sick Children (SickKids)(威廉与乔伊斯·波斯卢斯影像引导创新与治疗干预中心(PCIGITI)(SickKids医院)) Electrical and Computer Engineering Department, Kettering University(电气与计算机工程系,凯特林大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.08559 2025-10-10 cs.CV cs.AI 81%

SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models

Andong Deng, Taojiannan Yang, Shoubin Yu, Lincoln Spencer, Mohit Bansal, Chen Chen, Serena Yeung-Levy, Xiaohan Wang

机构 * University of Central Florida(中央佛罗里达大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Stanford University(斯坦福大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02790 2025-10-06 cs.CV cs.CL 81%

From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding

Xiangfeng Wang, Xiao Li, Yadong Wei, Xueyu Song, Yang Song, Xiaoqiang Xia, Fangrui Zeng, Zaiyi Chen, Liu Liu, Gu Xu, Tong Xu

机构 * University of Science and Technology of China(中国科学技术大学) ByteDance China(字节跳动中国)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL

Comments Accepted by EMNLP 2025 Industry Track

详情

展开后加载摘要…

URL PDF HTML 收藏