arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 45985 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4643 篇

2603.01816 2026-03-03 cs.MM 79%

Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding

声音、面孔与情感:多模态情绪-认知描述用于心理健康理解

Zhiyuan Zhou, Yanrong Guo, Shijie Hao

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.MM

AI总结 ECMC通过多模态数据生成情绪-认知描述,提升心理健康评估的准确性和可解释性。

Comments Accepted at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01055 2026-03-03 cs.AI 79%

MMCOMET: A Large-Scale Multimodal Commonsense Knowledge Graph for Contextual Reasoning

MMCOMET:一种大规模多模态常识知识图谱用于上下文推理

Eileen Wang, Hiba Arnaout, Dhita Pratama, Shuo Yang, Dangyang Liu, Jie Yang, Josiah Poon, Jeff Pan, Caren Han

机构 * University of Sydney(悉尼大学) University of Melbourne(墨尔本大学) University of Edinburgh(爱丁堡大学) The University of Sydney(悉尼大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 MMCOMET是一种大规模多模态常识知识图谱,通过整合视觉、物理和社会知识,提升了复杂推理任务如图像描述和叙事生成的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00694 2026-03-03 cs.RO cs.AI 79%

Wild-Drive: Off-Road Scene Captioning and Path Planning via Robust Multi-modal Routing and Efficient Large Language Model

Wild-Drive: 通过鲁棒多模态路由和高效大语言模型实现越野场景描述与路径规划

Zihang Wang, Xu Li, Benwu Wang, Wenkai Zhu, Xieyuanli Chen, Dong Kong, Kailin Lyu, Yinan Du, Yiming Peng, Haoyang Che

机构 * School of Instrument Science and Engineering, Southeast University(东南大学仪器科学与工程学院) Southeast University Nanjing Jiangbei New Area Innovation Research Institute(东南大学南京江滨新区创新研究院) National Key Laboratory of Equipment State Sensing and Smart Support, National University of Defense Technology(国防科技大学装备状态感知与智能支撑国家重点实验室) School of Transportation, Shandong University of Science and Technology(山东科技大学交通学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 图文多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

AI总结 Wild-Drive通过鲁棒多模态路由和高效大语言模型实现越野场景描述与路径规划,提升复杂环境下的可解释性和稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.06223 2026-03-03 cs.LG cs.CV stat.ML 79%

Beyond DAGs: A Latent Partial Causal Model for Multimodal Learning

超越DAGs:一种用于多模态学习的潜在部分因果模型

Yuhang Liu, Zhen Zhang, Dong Gong, Erdun Gao, Biwei Huang, Mingming Gong, Anton van den Hengel, Kun Zhang, Javen Qinfeng Shi

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种用于多模态学习的潜在部分因果模型,通过解耦表示提升模型在少量样本学习和领域泛化中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23229 2026-02-27 cs.CV 79%

Large Multimodal Models as General In-Context Classifiers

大多模态模型作为通用上下文分类器

Marco Garosi, Matteo Farina, Alessandro Conti, Massimiliano Mancini, Elisa Ricci

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出CIRCLE方法,通过伪标签迭代优化,使LMM在开放世界分类中超越VLM,展示LMM作为统一分类器的潜力。

Comments CVPR Findings 2026. Project website at https://circle-lmm.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22734 2026-02-27 cs.CV 79%

Asymmetric Idiosyncrasies in Multimodal Models

多模态模型中的非对称个性化特征

Muzi Tao, Chufan Shi, Huijuan Wang, Shengbang Tong, Xuezhe Ma

机构 * University of Southern California(南加州大学) New York University(纽约大学)

专题命中 图文多模态 :multimodal(title);cross-modal(abstract);分类 cs.CV

AI总结 本文研究了多模态模型中描述模型的个性化特征及其对文本到图像模型的影响,发现生成图像丢失了描述中的关键变化,提出了一种新的量化方法。

Comments Project page: https://muzi-tao.github.io/asymmetric-idiosyncrasies/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22644 2026-02-27 cs.CV 79%

Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models

插件、即插即用并加固:一种低成本模块用于鲁棒多模态图像理解模型

Siqi Lu, Wanying Xu, Yongbin Zheng, Wenting Luan, Peng Sun, Jianhang Yao

机构 * National University of Defense Technology(国防科技大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种低成本模块,通过频域分析解决多模态模型中缺失模态导致的性能问题,提升模型鲁棒性和整体学习效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11221 2026-02-27 cs.CL 79%

The Automatic Verification of Image-Text Claims (AVerImaTeC) Shared Task

图像-文本主张的自动验证(AVerImaTeC)共享任务

Rui Cao, Zhenyun Deng, Yulong Chen, Michael Schlichtkrull, Andreas Vlachos

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CL

AI总结 本文提出图像-文本主张自动验证共享任务,通过检索证据和验证真实主张,评估系统性能并展示最佳结果。

Comments Shared Task Overview and Summary for the Ninth FEVER Workshop, Co-located at EACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18811 2026-02-24 cs.CV 79%

Learning Multi-Modal Prototypes for Cross-Domain Few-Shot Object Detection

跨域少样本目标检测中的多模态原型学习

Wanqi Wang, Jingcai Guo, Yuxiang Cai, Zhi Chen

机构 * University of Chinese Academy of Sciences(中国科学院大学) The Hong Kong Polytechnic University(香港理工大学) Zhejiang University(浙江大学) The University of Southern Queensland(昆士兰大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出LMP方法,通过结合文本和视觉信息,提升跨域少样本目标检测的精度和性能。

Comments Accepted to CVPR 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06530 2026-02-20 cs.CV cs.CR 79%

Universal Anti-forensics Attack against Image Forgery Detection via Multi-modal Guidance

面向图像伪造检测的通用反取证攻击框架 via 多模态指导

Haipeng Li, Rongxuan Peng, Anwei Luo, Shunquan Tan, Changsheng Chen, Anastasia Antsiferova

机构 * Shenzhen University, China(深圳大学) Nanyang Technological University, Singapore(南洋理工大学) Shenzhen MSU-BIT University, China(深圳MSU-BIT大学) Lomonosov Moscow State University's Institute for Artificial Intelligence, Russia(罗蒙诺索夫莫斯科国立大学人工智能研究所)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出ForgeryEraser框架,通过多模态指导损失实现对图像伪造检测器的通用反取证攻击,揭示了VLMs依赖性带来的对抗性漏洞,并展示了对先进AIGC检测器的性能影响。

Comments 17 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15903 2026-02-19 cs.CV 79%

Detecting Deepfakes with Multivariate Soft Blending and CLIP-based Image-Text Alignment

利用多变量软融合和基于CLIP的图像-文本对齐检测深度伪造

Jingwei Li, Jiaxin Tong, Pengfei Wu

机构 * Zhejiang Gongshang University(浙江工商大学)

专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出MSBA-CLIP框架,通过多变量软融合和CLIP引导的伪造强度估计,提升深度伪造检测的准确性和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22914 2026-02-18 cs.CV cs.LG 79%

cadrille: Multi-modal CAD Reconstruction with Reinforcement Learning

cadrille: 多模态CAD重建与强化学习

Maksim Kolodiazhnyi, Denis Tarasov, Dmitrii Zhemchuzhnikov, Alexander Nikulin, Ilya Zisman, Anna Vorontsova, Anton Konushin, Vladislav Kurenkov, Danila Rukhovich

机构 * Lomonosov Moscow State University(罗蒙诺索夫莫斯科国立大学) AXXX ETH Zurich(苏黎世联邦理工学院) Innopolis University(因诺波利斯大学) Institute of Mechanics, Armenia(亚美尼亚力学研究所)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 cadrille通过强化学习实现多模态CAD重建,首次在CAD任务中应用RL微调,提升重建性能并设定新基准。

Comments ICLR 2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15190 2026-02-18 cs.CL 79%

AIC CTU@AVerImaTeC: dual-retriever RAG for image-text fact checking

AIC CTU@AVerImaTeC:双检索器RAG用于图像-文本事实核查

Herbert Ullrich, Jan Drchal

专题命中 图文多模态 :image-text(title);multimodal(abstract);分类 cs.CL

AI总结 AIC CTU@AVerImaTeC提出双检索器RAG方法,通过结合文本和图像检索模块,实现高效的图像-文本事实核查,具有低运行成本和易复现性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15124 2026-02-18 cs.CV 79%

Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition

基于多模态大语言模型的零样本人-物交互检测

Shiyu Xuan, Dongkai Wang, Zechao Li, Jinhui Tang

机构 * School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院) School of Computing and Artificial Intelligence, Southwestern University of Finance and Economics(西南财经大学计算机与人工智能学院) Nanjing Forestry University(南京林业大学)

专题命中 图文多模态 :MLLM(title);multi-modal(abstract);分类 cs.CV

AI总结 本文提出基于多模态大语言模型的零样本人-物交互检测框架,通过解耦检测与识别任务,结合确定性生成方法和空间感知模块,实现高效准确的零样本交互识别。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14889 2026-02-17 cs.LG cs.CV cs.ET cs.HC cs.NE 79%

Web-Scale Multimodal Summarization using CLIP-Based Semantic Alignment

基于CLIP的语义对齐的网络级多模态摘要

Mounvik K, N Harshit

机构 * School of Computer Science Engineering(计算机科学与工程学院) VIT-AP University(VIT-AP大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出基于CLIP的语义对齐网络级多模态摘要框架,通过结合网络文本和图像数据生成摘要,实现高准确率的多模态对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20110 2026-02-17 cs.CV 79%

Cross-Modal Mapping: Mitigating the Modality Gap for Few-Shot Image Classification

跨模态映射:缓解模态差距以实现少样本图像分类

Xi Yang, Pai Peng, Wulin Xie, Xiaohuan Lu, Jie Wen

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文提出跨模态映射方法,通过全局对齐和三元组损失优化,缓解模态差距,提升少样本图像分类性能。

Comments The authors request withdrawal of this article. This version was submitted in error. Compared to the intended final version, it contains inaccuracies and fails to accurately reflect the authors' work and conclusions

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09080 2026-02-11 cs.LG cs.AI 79%

Looping Back to Move Forward: Recursive Transformers for Efficient and Flexible Large Multimodal Models

循环回溯以前进:递归变换器用于高效灵活的大型多模态模型

Ruihan Xu, Yuting Gao, Lan Wang, Jianing Li, Weihao Chen, Qingpei Guo, Ming Yang, Shiliang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) AntGroup(蚂蚁集团)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 RecursiveVLM通过递归细化机制提升多模态模型效率,实现参数复用和性能提升。

Comments This is a primary contribution in the Recursive Vision-Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06218 2026-02-11 cs.CV cs.LG 79%

Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings

跨模态冗余与视觉-语言嵌入的几何学

Grégoire Dhimoïla, Thomas Fel, Victor Boutin, Agustin Picard

机构 * Brown University(布朗大学) ENS Paris Saclay(巴黎萨克雷大学) IRT Saint Exupéry(IRT圣埃克苏佩里) Kempner Institute, Harvard University(哈佛大学凯姆纳研究所)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文通过等能假设和对齐稀疏自编码器,揭示了视觉-语言模型中跨模态对齐的几何结构,发现稀疏双模态原子承载了跨模态对齐信号,单模态原子解释了模态差距,去除单模态原子可消除差距而不影响性能。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02197 2026-02-03 cs.LG cs.AI 79%

Hierarchical Adaptive Eviction for KV Cache Management in Multimodal Language Models

分层自适应淘汰用于多模态语言模型的KV缓存管理

Xindian Ma, Yidi Lu, Peng Zhang, Jing Zhang

机构 * College of Intelligence and Computing, Tianjin University(智能与计算学院,天津大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出HAE框架,通过双注意力修剪和动态解码淘汰策略优化多模态语言模型的KV缓存管理,减少内存使用并提升推理效率。

Comments 10 oages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09981 2026-02-03 cs.CV 79%

DR$^2$Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models

DR$^2$Seg: 分解式两阶段 rollout 用于多模态大语言模型中的高效推理分割

Yulin He, Wei Chen, Zhikang Jian, Tianhang Guo, Wenjuan Zhou, Minglong Li, Shaowu Yang, Wenjing Yang

机构 * School of Computer, National University of Defense Technology, Changsha, China(计算机学院,国防科技大学,中国长沙)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 DR$^2$Seg通过分解式两阶段rollout策略提升多模态大语言模型中的推理效率和分割准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.17493 2026-02-03 cs.CV 79%

Meme Similarity and Emotion Detection using Multimodal Analysis

基于多模态分析的膜拜相似性与情感检测

Aidos Konyspay, Pakizar Shamoi, Malika Ziyada, Zhusup Smambayev

机构 * School of Information Technology and Engineering(信息科技与工程学院) Kazakh-British Technical University(哈萨克-英国技术大学) Faculty of Information Technology(信息技术学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本研究通过多模态分析方法,利用CLIP模型和DistilBERT模型对迷因的相似性和情感进行检测,发现愤怒和快乐是主要情绪,为在线视觉交流和内容管理提供新思路。

Comments 2025 International Conference on Activity and Behavior Computing (ABC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22492 2026-02-02 cs.CV 79%

PromptMAD: Cross-Modal Prompting for Multi-Class Visual Anomaly Localization

PromptMAD: 多类视觉异常定位的跨模态提示

Duncan McCain, Hossein Kashiani, Fatemeh Afghah

机构 * Holcombe Department of Electrical and Computer Engineering(电气与计算机工程系) Computer Engineering Clemson University(计算机工程学系 哥伦比亚大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 PromptMAD通过跨模态提示和Focal损失函数,在多类视觉异常检测中实现高精度像素级定位与高效性能。

Comments Accepted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20433 2026-02-02 cs.CV 79%

MARE: Multimodal Alignment and Reinforcement for Explainable Deepfake Detection via Vision-Language Models

MARE: 多模态对齐与强化学习用于通过视觉-语言模型的可解释深度伪造检测

Wenbo Xu, Wei Lu, Xiangyang Luo, Jiantao Zhou

机构 * School of Computer Science and Engineering, MoE Key Laboratory of Information Technology, Guangdong Province Key Laboratory of Information Security Technology, Sun Yat-sen University, Guangzhou 510006, China(计算机科学与工程学院,信息技术MOE实验室,广东省信息安全技术重点实验室,中山大学,广州510006,中国) State Key Laboratory of Mathematical Engineering and Advanced Computing(数学工程与先进计算国家重点实验室) Department of Computer and Information Science, University of Macau.(计算机与信息科学系,澳门大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 MARE通过多模态对齐和强化学习提升视觉-语言模型在深度伪造检测中的准确性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18190 2026-01-27 cs.CV 79%

Multi-Perspective Subimage CLIP with Keyword Guidance for Remote Sensing Image-Text Retrieval

多视角子图像CLIP与关键词引导的遥感图像-文本检索

Yifan Li, Shiying Wang, Jianqiang Huang

机构 * School of Computer Technology and Applications(计算机技术与应用学院) Qinghai University(青海大学) Qinghai Provincial Laboratory for Intelligent Computing and Application(青海省智能计算与应用实验室)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

AI总结 MPS-CLIP通过关键词引导的多视角细粒度对齐提升遥感图像-文本检索性能,实现35.18%和48.40%的mR成绩。

Comments 7 pages, 3 figures. Code: https://github.com/Lcrucial1f/MPS-CLIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17197 2026-01-27 cs.CL cs.LG 79%

Reasoning Beyond Literal: Cross-style Multimodal Reasoning for Figurative Language Understanding

超越字面:跨风格多模态推理用于隐喻语言理解

Seyyed Saeid Cheshmi, Hahnemann Ortiz, James Mooney, Dongyeop Kang

机构 * University of Minnesota(明尼苏达大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出一种三步框架,通过跨风格多模态推理提升隐喻语言理解能力,实验显示推理轨迹和跨风格训练能显著提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24330 2026-01-27 cs.CV 79%

SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning

SenseNova-MARS: 通过强化学习赋能多模态代理推理与搜索

Yong Xien Chng, Tao Hu, Wenwen Tong, Xueheng Li, Jiandong Chen, Haojia Yu, Jiefan Lu, Hewei Guo, Hanming Deng, Chengjun Xie, Gao Huang, Dahua Lin, Lewei Lu

机构 * SenseTime Research(商汤科技研究院) Tsinghua University(清华大学) University of Science and Technology of China(中国科学技术大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 SenseNova-MARS通过强化学习赋能多模态代理推理与搜索,提升视觉-语言模型在复杂视觉任务中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14757 2026-01-22 cs.CV 79%

ReinPath: A Multimodal Reinforcement Learning Approach for Pathology

ReinPath:一种用于病理学的多模态强化学习方法

Kangcheng Zhou, Jun Jiang, Qing Zhang, Shuang Zheng, Qingli Li, Shugong Xu

机构 * East China Normal University(华东师范大学) Shanghai University(上海大学) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 ReinPath提出了一种多模态强化学习方法,通过构建高质量病理学VQA数据集,结合语义奖励策略和群体相对策略优化,提升了病理图像和文本的多模态推理能力,并在零样本分类任务中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14052 2026-01-21 cs.CV 79%

Vision Also You Need: Navigating Out-of-Distribution Detection with Multimodal Large Language Model

视觉也需要:利用多模态大语言模型进行分布外检测导航

Haoran Xu, Yanlin Liu, Zizhao Tong, Jiaze Li, Kexue Fu, Yuyang Zhang, Longxiang Gao, Shuaiguang Li, Xingyu Li, Yanran Xu, Changwei Wang

机构 * Zhejiang University(浙江大学) Tsinghua University(清华大学) University of Chinese Academy of Sciences(中国科学院大学) Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(教育部计算电力网络与信息安全重点实验室,山东计算机科学中心(国家超算中心济南中心),齐鲁工业大学(山东科学院)) Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science(山东省计算电力互联网与服务计算重点实验室,山东省计算机科学基础研究中心) University of Electronic Science and Technology of China(电子科技大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) RWTH Aachen University(亚琛工业大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出MM-OOD方法,利用多模态大语言模型的推理能力,通过多轮对话增强分布外检测,提升近远OOD任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11243 2026-01-19 cs.CV 79%

Image-Text Knowledge Modeling for Unsupervised Multi-Scenario Person Re-Identification

图像-文本知识建模用于无监督多场景人物重识别

Zhiqi Pang, Lingling Zhao, Yang Liu, Chunyu Wang, Gaurav Sharma

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

AI总结 本文提出图像-文本知识建模用于无监督多场景人物重识别,通过三阶段框架提升跨场景识别性能。

Comments 12 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06475 2026-01-15 cs.CL 79%

HapticLLaMA: A Multimodal Sensory Language Model for Haptic Captioning

HapticLLaMA:一种多模态感官语言模型用于触觉描述

Guimin Hu, Daniel Hershcovich, Hasti Seifi

机构 * University of Copenhagen(哥本哈根大学) Arizona State University(亚利桑那州立大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 HapticLLaMA是一种多模态感官语言模型,通过触觉信号生成描述,利用两种分词器和强化学习提升触觉感知描述能力。

详情

展开后加载摘要…

URL PDF HTML 收藏