arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 6897 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态训练与对齐 6897 篇

2503.14377 2025-03-19 eess.IV cs.CV cs.LG 70%

Advancing Medical Representation Learning Through High-Quality Data

Negin Baghbanzadeh, Adibvafa Fallahpour, Yasaman Parhizkar, Franklin Ogidi, Shuvendu Roy, Sajad Ashkezari, Vahid Reza Khazaie, Michael Colacci, Ali Etemad, Arash Afkanpour, Elham Dolatabadi

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02304 2025-03-18 cs.CV 70%

A Token-level Text Image Foundation Model for Document Understanding

Tongkun Guan, Zining Wang, Pei Fu, Zhengtao Guo, Wei Shen, Kai Zhou, Tiezhu Yue, Chen Duan, Hao Sun, Qianyi Jiang, Junfeng Luo, Xiaokang Yang

专题命中 多模态训练与对齐 :multi-modal(abstract);MLLM(abstract);分类 cs.CV

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12902 2025-03-11 cs.CV 70%

DoPTA: Improving Document Layout Analysis using Patch-Text Alignment

Nikitha SR, Tarun Ram Menta, Mausoom Sarkar

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22489 2025-02-27 cs.CV 70%

Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation

Zhaochong An, Guolei Sun, Yun Liu, Runjia Li, Min Wu, Ming-Ming Cheng, Ender Konukoglu, Serge Belongie

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Published at ICLR 2025 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.17273 2025-02-25 cs.CV 70%

X Modality Assisting RGBT Object Tracking

Zhaisheng Ding, Haiyan Li, Ruichao Hou, Yanyu Liu, Shidong Xie

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11560 2025-02-18 cs.AI cs.LG 70%

A Survey of Automatic Prompt Engineering: An Optimization Perspective

Wenwu Li, Xiangfeng Wang, Wenhao Li, Bo Jin

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

Comments 19 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10423 2025-02-18 cs.NE cs.CV cs.LG eess.IV 70%

Spiking Neural Network Feature Discrimination Boosts Modality Fusion

Katerina Maria Oikonomou, Ioannis Kansizoglou, Antonios Gasteratos

专题命中 多模态训练与对齐 :multi-modal(abstract);audio-visual(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07905 2025-02-13 cs.CV cs.LG 70%

DeepSeek on a Trip: Inducing Targeted Visual Hallucinations via Representation Vulnerabilities

Chashi Mahiul Islam, Samuel Jacob Chacko, Preston Horne, Xiuwen Liu

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments 19 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00339 2025-02-04 cs.CL cs.CY 70%

Challenges and Innovations in LLM-Powered Fake News Detection: A Synthesis of Approaches and Future Directions

Jingyuan Yi, Zeqiu Xu, Tianyi Huang, Peiyang Yu

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.11276 2025-01-22 eess.IV cs.CV 70%

ITCFN: Incomplete Triple-Modal Co-Attention Fusion Network for Mild Cognitive Impairment Conversion Prediction

Xiangyang Hu, Xiangyu Shen, Yifei Sun, Xuhao Shan, Wenwen Min, Liyilei Su, Xiaomao Fan, Ahmed Elazab, Ruiquan Ge, Changmiao Wang, Xiaopeng Fan

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments 5 pages, 1 figure, accepted by IEEE ISBI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06215 2025-01-14 cs.CV cs.CL cs.LG cs.MM eess.AS 70%

Fitting Different Interactive Information: Joint Classification of Emotion and Intention

Xinger Li, Zhiqiang Zhong, Bo Huang, Yang Yang

专题命中 多模态训练与对齐 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01416 2025-01-03 cs.CV 70%

Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension

Yaxian Wang, Henghui Ding, Shuting He, Xudong Jiang, Bifan Wei, Jun Liu

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments AAAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.20646 2024-12-31 cs.CV 70%

Enhancing Visual Representation for Text-based Person Searching

Wei Shen, Ming Fang, Yuxia Wang, Jiafeng Xiao, Diping Li, Huangqun Chen, Ling Xu, Weifeng Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.16444 2024-12-24 cs.CL cs.LG 70%

Effective Context Modeling Framework for Emotion Recognition in Conversations

Cuong Tran Van, Thanh V. T. Tran, Van Nguyen, Truong Son Hy

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.00667 2024-12-18 cs.CL 70%

Does Vision Accelerate Hierarchical Generalization in Neural Language Learners?

Tatsuki Kuribayashi, Timothy Baldwin

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CL

Comments COLING 2025; 15 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.05149 2024-12-09 cs.CL 70%

Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora

Michael Y. Hu, Aaron Mueller, Candace Ross, Adina Williams, Tal Linzen, Chengxu Zhuang, Ryan Cotterell, Leshem Choshen, Alex Warstadt, Ethan Gotlieb Wilcox

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09401 2024-12-06 cs.CV 70%

Unsupervised Modality-Transferable Video Highlight Detection with Representation Activation Sequence Learning

Tingtian Li, Zixun Sun, Xinyu Xiao

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted by IEEE Transactions on Image Processing, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10669 2024-11-19 cs.CV 70%

Awaker2.5-VL: Stably Scaling MLLMs with Parameter-Efficient Mixture of Experts

Jinqiang Long, Yanqi Dai, Guoxing Yang, Hongpeng Lin, Nanyi Fei, Yizhao Gao, Zhiwu Lu

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10480 2024-11-19 cs.CV cs.AI cs.CL cs.LG cs.MM 70%

Hateful Meme Detection through Context-Sensitive Prompting and Fine-Grained Labeling

Rongxin Ouyang, Kokil Jaidka, Subhayan Mukerjee, Guangyu Cui

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments AAAI-25 Student Abstract, Oral Presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07750 2024-11-13 eess.IV cs.CV 70%

LapGSR: Laplacian Reconstructive Network for Guided Thermal Super-Resolution

Aditya Kasliwal, Ishaan Gakhar, Aryan Kamani, Pratinav Seth, Ujjwal Verma

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.02034 2024-10-29 cs.CV 70%

Mini-Monkey: Alleviating the Semantic Sawtooth Effect for Lightweight MLLMs via Complementary Image Pyramid

Mingxin Huang, Yuliang Liu, Dingkang Liang, Lianwen Jin, Xiang Bai

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11499 2024-10-16 q-bio.GN cs.AI cs.LG 70%

BSM: Small but Powerful Biological Sequence Model for Genes and Proteins

Weixi Xiang, Xueting Han, Xiujuan Chai, Jing Bai

专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.12081 2024-10-15 cs.AI 70%

Robust Graph Matching Using An Unbalanced Hierarchical Optimal Transport Framework

Haoran Cheng, Dixin Luo, Hongteng Xu

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.07113 2024-10-10 cs.CV 70%

Personalized Visual Instruction Tuning

Renjie Pi, Jianshu Zhang, Tianyang Han, Jipeng Zhang, Rui Pan, Tong Zhang

专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16538 2024-10-01 cs.CV 70%

OpenDlign: Open-World Point Cloud Understanding with Depth-Aligned Images

Ye Mao, Junpeng Jing, Krystian Mikolajczyk

专题命中 多模态训练与对齐 :multimodal(abstract);image-text(abstract);分类 cs.CV

Comments 17 pages (Accepted by NeurIPS 2024)

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16199 2024-09-20 cs.CV cs.AI cs.CL cs.LG cs.MM 70%

LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Renrui Zhang, Jiaming Han, Chris Liu, Peng Gao, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Yu Qiao

专题命中 多模态训练与对齐 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted by ICLR 2024. Code is available at https://github.com/OpenGVLab/LLaMA-Adapter

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.12150 2024-09-19 cs.IR cs.AI cs.LG 70%

Decoding Style: Efficient Fine-Tuning of LLMs for Image-Guided Outfit Recommendation with Preference

Najmeh Forouzandehmehr, Nima Farrokhsiar, Ramin Giahi, Evren Korpeoglu, Kannan Achan

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.AI

Comments CIKM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09942 2024-09-17 cs.CV 70%

Knowledge-enhanced Visual-Language Pretraining for Computational Pathology

Xiao Zhou, Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Weidi Xie, Yanfeng Wang

专题命中 多模态训练与对齐 :cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments ECCV2024(Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00845 2024-09-04 cs.CV 70%

Image-to-Lidar Relational Distillation for Autonomous Driving Data

Anas Mahmoud, Ali Harakeh, Steven Waslander

专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted at ECCV 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.11424 2024-08-22 cs.CV 70%

EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning

Bohao Xing, Zitong Yu, Xin Liu, Kaishen Yuan, Qilang Ye, Weicheng Xie, Huanjing Yue, Jingyu Yang, Heikki Kälviäinen

专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏