arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4878 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4878 篇

2608.04154 2026-08-06 cs.CV cs.AI 新提交 76%

TRNet: Topography-Guided Frequency Rectification and Structure-Aware Decoding for Multimodal Paddy Rice Segmentation

TRNet:用于多模态水稻分割的地形引导频率校正与结构感知解码

Kaiwen Xiao, Chunlong Fu, Liping Zheng, Yanfeng Su

机构 * School of Computer Science, Sichuan University Jinjiang College(四川大学锦江学院计算机学院)

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.AI

AI总结 TRNet模型针对山区丘陵水稻分割难题,采用双编码器与地形引导解码,在A、B区域水稻IoU较双编码器U-Net显著提升,验证了地形作为上下文先验的有效性。

Comments 16 pages, 9 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01238 2026-08-04 cs.CL cs.MM 新提交 76%

Evaluating VLMs on Multimodal Aristotelian Persuasion Tasks

评估视觉语言模型(VLMs)在亚里士多德式多模态说服任务上的表现

Khondoker Ittehadul Islam

机构 * Saarland University(萨尔大学)

专题命中 其他多模态 :multimodal(title);分类 cs.CL、cs.MM

AI总结 该研究采用ImageArg数据集评估VLMs在亚里士多德式多模态说服任务的表现,发现Qwen系列模型在相关检测任务上性能提升并发布代码。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22784 2026-07-13 cs.CV cs.AI cs.RO 版本更新 76%

Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching

基于监督跨模态特征匹配的单帧点像素配准

Yu Han, Zhiwei Huang, Yanting Zhang, Fangjun Ding, Shen Cai, Xiaoyu Tang, Yanchao Dong, Rui Fan

机构 * School of Information and Intelligent Science, Donghua University(信息与智能科学学院,东华大学)

专题命中 其他多模态 :cross-modal(title);分类 cs.CV、cs.AI

AI总结 研究激光雷达点云和相机图像的点像素配准难题,提出基于投影的无检测器框架及重复性评分机制,通过实验证明该方法在单帧激光雷达下能达先进性能,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03956 2026-03-05 cs.CV cs.AI 76%

Towards Generalized Multimodal Homography Estimation

面向通用多模态透视图估计

Jinkun You, Jiaxin Cheng, Jie Zhang, Yicong Zhou

机构 * Department of Computer and Information Science, University of Macau(计算机与信息科学系,澳门大学)

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.AI

AI总结 本文提出了一种训练数据合成方法和改进网络,用于提升多模态透视图估计的泛化能力和精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22623 2026-02-27 cs.LG cs.AI cs.CL 76%

ContextRL: Enhancing MLLM's Knowledge Discovery Efficiency with Context-Augmented RL

ContextRL: 通过上下文增强强化学习提升大语言模型的知识发现效率

Xingyu Lu, Jinpeng Wang, YiFan Zhang, Shijie Ma, Xiao Hu, Tianke Zhang, Haonan fan, Kaiyu Jiang, Changyi Liu, Kaiyu Tang, Bin Wen, Fan Yang, Tingting Gao, Han Li, Chun Yuan

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Chinese Academy of Sciences(中国科学院) Tsinghua University(清华大学)

专题命中 其他多模态 :MLLM(title);分类 cs.CL、cs.AI

AI总结 ContextRL通过上下文增强强化学习提升大语言模型的知识发现效率,有效缓解奖励黑客问题并提升性能。

Comments 14 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00597 2026-02-03 cs.CL cs.AI 76%

Hermes the Polyglot: A Unified Framework to Enhance Expressiveness for Multimodal Interlingual Subtitling

赫мес:一种增强多模态跨语言字幕表达力的统一框架

Chaoqun Cui, Shijing Wang, Liangbin Huang, Qingqing Gu, Zhaolong Huang, Xiao Zeng, Wenji Mao

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences Beijing China School of AI, University of\ Academy of Sciences Beijing China Beijing Jiaotong University Beijing China Geely AI lab Ningbo Zhejiang China MAIS, Institute of Automation, Chinese Academy of Sciences School of AI, University of\ Academy of Sciences Beijing Jiaotong University Geely AI lab

专题命中 其他多模态 :multimodal(title);分类 cs.CL、cs.AI

AI总结 赫мес通过整合说话人分离、术语识别和表达力增强模块,提升了多模态跨语言字幕的表达力和连贯性,实现了最先进的字幕生成性能。

Comments Accepted to The Web Conference (WWW) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22501 2025-12-23 cs.CV cs.AI 76%

How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT?

多模态遥感数据集如何通过SpatialNet-ViT实现分类转变?

Gautam Siddharth Kashyap, Manaswi Kulahara, Nipun Joshi, Usman Naseem

机构 * Macquarie University(麦考瑞大学) TERI School Of Advanced Studies(TERI高级研究学院) Cornell University(康奈尔大学)

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.AI

AI总结 本文提出SpatialNet-ViT模型,结合视觉转换器和多任务学习,以提升遥感数据分类的准确性和泛化能力。

Comments Accepted in the 2025 IEEE International Geoscience and Remote Sensing Symposium (IGARSS 2025), scheduled for 3 - 8 August 2025 in Brisbane, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.08889 2025-12-10 cs.CV cs.AI 76%

No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers

无标签,无问题:利用多模态验证器训练视觉推理器

Damiano Marsili, Georgia Gkioxari

机构 * California Institute of Technology(加州理工学院)

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.AI

AI总结 本文提出无需标注的视觉推理训练框架,结合AI驱动的验证器提升推理与定位能力,超越现有开源和专有模型。

Comments Project webpage: https://glab-caltech.github.io/valor/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.05361 2025-11-10 cs.CL cs.AI 76%

A multimodal multiplex of the mental lexicon for multilingual individuals

Maria Huynh, Wilder C. Rodrigues

专题命中 其他多模态 :multimodal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22751 2025-10-28 cs.AI cs.CL 76%

Multi-Modal Fact-Verification Framework for Reducing Hallucinations in Large Language Models

Piyushkumar Patel

机构 * Microsoft(微软)

专题命中 其他多模态 :multi-modal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07885 2025-08-12 cs.RO cs.AI cs.CV cs.SY eess.SY 76%

Autonomous Navigation of Cloud-Controlled Quadcopters in Confined Spaces Using Multi-Modal Perception and LLM-Driven High Semantic Reasoning

Shoaib Ahmmad, Zubayer Ahmed Aditto, Md Mehrab Hossain, Noushin Yeasmin, Shorower Hossain

机构 * Department of Mechanical Engineering(机械工程系) Rajshahi University of Engineering and Technology(拉贾沙希工程与技术大学) Department of Industrial and Production Engineering(工业与生产工程系) Shahjalal University of Science and Technology(沙赫jalal科学与技术大学) Bangladesh University of Engineering and Technology(孟加拉工程与技术大学) Department of Urban and Regional Planning(城市与区域规划系) Department of Computer Science Engineering(计算机科学与工程系) United International University(联合国际大学)

专题命中 其他多模态 :multi-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19195 2025-05-27 cs.AI cs.CV 76%

CardioCoT: Hierarchical Reasoning for Multimodal Survival Analysis

Shaohao Rui, Haoyang Su, Jinyi Xiang, Lian-Ming Wu, Xiaosong Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) Fudan University(复旦大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07190 2025-03-11 cs.CV cs.CL 76%

Multi-Modal 3D Mesh Reconstruction from Images and Text

Melvin Reka, Tessa Pulli, Markus Vincze

专题命中 其他多模态 :multi-modal(title);分类 cs.CV、cs.CL

Comments under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06828 2025-03-11 eess.IV cs.AI cs.CV 76%

Towards a Multimodal MRI-Based Foundation Model for Multi-Level Feature Exploration in Segmentation, Molecular Subtyping, and Grading of Glioma

Somayeh Farahani, Marjaneh Hejazi, Antonio Di Ieva, Emad Fatemizadeh, Sidong Liu

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07016 2025-01-14 eess.IV cs.AI cs.CV 76%

A Multi-Modal Deep Learning Framework for Pan-Cancer Prognosis

Binyu Zhang, Shichao Li, Junpeng Jian, Zhu Meng, Limei Guo, Zhicheng Zhao

专题命中 其他多模态 :multi-modal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.11585 2024-03-26 cs.AI cs.CL 76%

Causal Intersectionality and Dual Form of Gradient Descent for Multimodal Analysis: a Case Study on Hateful Memes

Yosuke Miyanishi, Minh Le Nguyen

专题命中 其他多模态 :multimodal(title);分类 cs.CL、cs.AI

Comments Accepted to LREC-COLING 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.08915 2024-03-15 cs.CV cs.AI 76%

Cross-Modal Learning of Housing Quality in Amsterdam

Alex Levering, Diego Marcos, Devis Tuia

专题命中 其他多模态 :cross-modal(title);分类 cs.CV、cs.AI

Comments Presented at SIGSpatial GeoAI workshop '21

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17793 2024-02-29 cs.AI cs.CL cs.LG 76%

A Surprising Failure? Multimodal LLMs and the NLVR Challenge

Anne Wu, Kianté Brantley, Yoav Artzi

专题命中 其他多模态 :multimodal(title);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.12645 2023-07-27 cs.CL cs.AI 76%

Exploring Multi-Modal Representations for Ambiguity Detection & Coreference Resolution in the SIMMC 2.0 Challenge

Javier Chiyah-Garcia, Alessandro Suglia, José Lopes, Arash Eshghi, Helen Hastie

专题命中 其他多模态 :multi-modal(title);分类 cs.CL、cs.AI

Comments Accepted to AAAI 2022 DSTC10 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.10355 2022-07-22 cs.IR cs.AI cs.LG cs.MM 76%

Unimodal vs. Multimodal Siamese Networks for Outfit Completion

Mariya Hendriksen, Viggo Overes

专题命中 其他多模态 :multimodal(title);分类 cs.AI、cs.MM

Comments 3 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.10403 2022-05-19 cs.AI cs.MM 76%

Towards Integrative Multi-Modal Personal Health Navigation Systems: Framework and Application

Nitish Nag, Hyungik Oh, Mengfan Tang, Mingshu Shi, Ramesh Jain

专题命中 其他多模态 :multi-modal(title);分类 cs.AI、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.00811 2021-07-05 cs.RO cs.CL cs.CV 76%

Target-dependent UNITER: A Transformer-Based Multimodal Language Comprehension Model for Domestic Service Robots

Shintaro Ishikawa, Komei Sugiura

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.CL

Comments Accepted for presentation at IROS2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.04430 2021-06-29 cs.CV cs.AI 76%

TransBTS: Multimodal Brain Tumor Segmentation Using Transformer

Wenxuan Wang, Chen Chen, Meng Ding, Jiangyun Li, Hong Yu, Sen Zha

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.00240 2021-06-02 cs.CL cs.CV 76%

Volta at SemEval-2021 Task 6: Towards Detecting Persuasive Texts and Images using Textual and Multimodal Ensemble

Kshitij Gupta, Devansh Gautam, Radhika Mamidi

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.CL

Comments 7 pages, accepted at SemEval-2021 co-located with ACL-IJCNLP 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2012.08290 2020-12-16 cs.CL cs.CV 76%

Enhance Multimodal Transformer With External Label And In-Domain Pretrain: Hateful Meme Challenge Winning Solution

Ron Zhu

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2010.00802 2020-10-05 cs.CV cs.AI cs.MA 76%

PrognoseNet: A Generative Probabilistic Framework for Multimodal Position Prediction given Context Information

Thomas Kurbiel, Akash Sachdeva, Kun Zhao, Markus Buehren

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.AI

Comments 8 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2001.02372 2020-01-09 cs.CV cs.AI cs.LG 76%

Multimodal Semantic Transfer from Text to Image. Fine-Grained Image Classification by Distributional Semantics

Simon Donig, Maria Christoforaki, Bernhard Bermeitinger, Siegfried Handschuh

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.AI

Comments 19 pages, second half in German as published in DHd2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1903.12520 2019-04-01 cs.CV cs.CL 76%

Multimodal Emotion Classification

Anurag Illendula, Amit Sheth

专题命中 其他多模态 :multimodal(title);分类 cs.CV、cs.CL

Comments Accepted at the 2nd Emoji Workshop co-located with The Web Conference 2019

Journal ref Companion Proceedings of the 2019 World Wide Web Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.07941 2016-11-24 cs.CV cs.AI 76%

Multi-Modal Mean-Fields via Cardinality-Based Clamping

Pierre Baqué, François Fleuret, Pascal Fua

专题命中 其他多模态 :multi-modal(title);分类 cs.CV、cs.AI

Comments Submitted for review to CVPR 2017

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.09239 2016-06-30 cs.CL cs.CV cs.LG 76%

Learning Concept Taxonomies from Multi-modal Data

Hao Zhang, Zhiting Hu, Yuntian Deng, Mrinmaya Sachan, Zhicheng Yan, Eric P. Xing

专题命中 其他多模态 :multi-modal(title);分类 cs.CV、cs.CL

Comments To appear in ACL 2016

详情

展开后加载摘要…

URL PDF HTML 收藏