arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4878 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4878 篇

2506.11737 2025-06-16 cs.CV cs.CL cs.MM 67%

Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model

Dinh Viet Cuong, Hoang-Bao Le, An Pham Ngoc Nguyen, Liting Zhou, Cathal Gurrin

机构 * School of Computing, Dublin City University(都柏林城市大学计算机学院) ADAPT Centre(ADAPT中心)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13852 2025-05-22 cs.CL cs.AI cs.CV cs.LG 67%

Retrospective Learning from Interactions

Zizhao Chen, Mustafa Omer Gul, Yiwei Chen, Gloria Geng, Anne Wu, Yoav Artzi

机构 * Department of Computer Science and Cornell Tech, Cornell University(计算机科学系和康奈尔科技学院,康奈尔大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.10154 2025-05-13 cs.RO 67%

A Clinical Tuning Framework for Continuous Kinematic and Impedance Control of a Powered Knee-Ankle Prosthesis

Emma Reznick, T. Kevin Best, Robert Gregg

机构 * IEEE

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

Comments Published in IEEE JTEHM. IEEE Journal of Translational Engineering in Health and Medicine (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09829 2025-04-25 cs.RO cs.LG cs.SY eess.SY 67%

SE(3)-Equivariant Robot Learning and Control: A Tutorial Survey

Joohwan Seo, Soochul Yoo, Junwoo Chang, Hyunseok An, Hyunwoo Ryu, Soomi Lee, Arvind Kruthiventy, Jongeun Choi, Roberto Horowitz

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

Comments Accepted to International Journcal of Control, Automation and Systems (IJCAS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12316 2025-04-18 cs.CL cs.AI cs.CV 67%

Data Metabolism: An Efficient Data Design Schema For Vision Language Model

Jingyuan Zhang, Hongzhi Zhang, Zhou Haonan, Chenxi Sun, Xingguang ji, Jiakang Wang, Fanheng Kong, Yahui Liu, Qi Wang, Fuzheng Zhang

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments To be presented at ICLR 2025, First Workshop on Open Science for Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.01759 2025-03-12 cs.SI cs.CL cs.CV cs.MM 67%

VGA: Vision and Graph Fused Attention Network for Rumor Detection

Lin Bai, Caiyan Jia, Ziying Song, Chaoqun Cui

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04353 2025-02-10 cs.CL cs.AI cs.CV 67%

CognArtive: Large Language Models for Automating Art Analysis and Decoding Aesthetic Elements

Afshin Khadangi, Amir Sartipi, Igor Tchappi, Gilbert Fridgen

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.12735 2024-12-18 cs.CV cs.AI cs.CL 67%

GIRAFFE: Design Choices for Extending the Context Length of Visual Language Models

Mukai Li, Lei Li, Shansan Gong, Qi Liu

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18908 2024-12-04 cs.HC 67%

DuetML: Human-LLM Collaborative Machine Learning Framework for Non-Expert Users

Wataru Kawabe, Yusuke Sugano

专题命中 其他多模态 :multimodal(abstract);MLLM(abstract)

Comments 22 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23883 2024-11-01 cs.CL cs.AI cs.LG cs.MM 67%

'No' Matters: Out-of-Distribution Detection in Multimodality Long Dialogue

Rena Gao, Xuetong Wu, Siwen Luo, Caren Han, Feng Liu

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments 16 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.08424 2024-10-28 cs.RO cs.LG 67%

Conditional Neural Expert Processes for Learning Movement Primitives from Demonstration

Yigit Yildirim, Emre Ugur

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

Comments This work has been submitted to the IEEE RA-L for possible publication. Submitted to Robotics and Automation Letters on July 5, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18931 2024-09-30 cs.SI cs.CY 67%

Social Media Bot Policies: Evaluating Passive and Active Enforcement

Kristina Radivojevic, Christopher McAleer, Catrell Conley, Cormac Kennedy, Paul Brenner

专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.16567 2024-08-30 cs.RO cs.LG 67%

Identifying Terrain Physical Parameters from Vision -- Towards Physical-Parameter-Aware Locomotion and Navigation

Jiaqi Chen, Jonas Frey, Ruyi Zhou, Takahiro Miki, Georg Martius, Marco Hutter

专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10993 2024-05-21 q-bio.QM 67%

No winners: Performance of lung cancer prediction models depends on screening-detected, incidental, and biopsied pulmonary nodule use cases

Thomas Z. Li, Kaiwen Xu, Aravind Krishnan, Riqiang Gao, Michael N. Kammer, Sanja Antic, David Xiao, Michael Knight, Yency Martinez, Rafael Paez, Robert J. Lentz, Stephen Deppen, Eric L. Grogan, Thomas A. Lasko, Kim L. Sandler, Fabien Maldonado, Bennett A. Landman

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

Comments Submitted to Radiology: AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17971 2024-04-03 cs.CV cs.AI cs.CL 67%

All in an Aggregated Image for In-Image Learning

Lei Wang, Wanyu Xu, Zhiqiang Hu, Yihuai Lan, Shan Dong, Hao Wang, Roy Ka-Wei Lee, Ee-Peng Lim

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.12383 2023-12-20 cs.AI cs.CL cs.CV 67%

Visual AI and Linguistic Intelligence Through Steerability and Composability

David Noever, Samantha Elizabeth Miller Noever

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.11441 2023-11-07 cs.CV cs.AI cs.CL cs.HC 67%

Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, Jianfeng Gao

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.18842 2023-05-31 cs.CL cs.AI cs.CV 67%

Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge

Xingyu Fu, Sheng Zhang, Gukyeong Kwon, Pramuditha Perera, Henghui Zhu, Yuhao Zhang, Alexander Hanbo Li, William Yang Wang, Zhiguo Wang, Vittorio Castelli, Patrick Ng, Dan Roth, Bing Xiang

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ACL 2023 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.03701 2022-10-10 cs.RO 67%

VIRDO++: Real-World, Visuo-tactile Dynamics and Perception of Deformable Objects

Youngsun Wi, Andy Zeng, Pete Florence, Nima Fazeli

专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.09559 2022-05-31 stat.ME stat.CO stat.ML 67%

Continuously-Tempered PDMP Samplers

Matthew Sutton, Robert Salomone, Augustin Chevallier, Paul Fearnhead

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01295 2022-02-16 cs.CV cs.AI cs.CL 67%

Information Symmetry Matters: A Modal-Alternating Propagation Network for Few-Shot Learning

Zhong Ji, Zhishen Hou, Xiyao Liu, Yanwei Pang, Jungong Han

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.11403 2021-10-25 cs.CV cs.AI cs.CL cs.LG 67%

SCENIC: A JAX Library for Computer Vision Research and Beyond

Mostafa Dehghani, Alexey Gritsenko, Anurag Arnab, Matthias Minderer, Yi Tay

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.08013 2021-09-17 cs.CV cs.CL cs.LG cs.MM 67%

Detecting Propaganda Techniques in Memes

Dimitar Dimitrov, Bishr Bin Ali, Shaden Shaar, Firoj Alam, Fabrizio Silvestri, Hamed Firooz, Preslav Nakov, Giovanni Da San Martino

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments propaganda, disinformation, fake news, memes, multimodality. arXiv admin note: text overlap with arXiv:2105.09284

Journal ref ACL-2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2104.08614 2021-04-20 cs.SD cs.AI cs.CL cs.LG cs.RO eess.AS 67%

Cetacean Translation Initiative: a roadmap to deciphering the communication of sperm whales

Jacob Andreas, Gašper Beguš, Michael M. Bronstein, Roee Diamant, Denley Delaney, Shane Gero, Shafi Goldwasser, David F. Gruber, Sarah de Haas, Peter Malkin, Roger Payne, Giovanni Petri, Daniela Rus, Pratyusha Sharma, Dan Tchernov, Pernille Tønnesen, Antonio Torralba, Daniel Vogt, Robert J. Wood

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.10370 2020-09-23 cs.CV cs.AI cs.HC cs.MM 67%

Visual Methods for Sign Language Recognition: A Modality-Based Review

Bassem Seddik, Najoua Essoukri Ben Amara

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI、cs.MM

Comments This survey paper is accepted as Springer book chapter, currently under edition

详情

展开后加载摘要…

URL PDF HTML 收藏
1909.09070 2019-09-20 cs.AI cs.CL cs.CV 67%

Look, Read and Enrich. Learning from Scientific Figures and their Captions

Jose Manuel Gomez-Perez, Raul Ortega

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted in the 10th International Conference on Knowledge capture (K-CAP 2019)

详情

展开后加载摘要…

URL PDF HTML 收藏
1908.10285 2019-08-28 cs.CL cs.AI cs.CV 67%

Is the Red Square Big? MALeViC: Modeling Adjectives Leveraging Visual Contexts

Sandro Pezzelle, Raquel Fernández

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted at EMNLP-IJCNLP 2019

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.00377 2018-06-05 cs.CV cs.AI cs.CL cs.LG 67%

Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering

Aishwarya Agrawal, Dhruv Batra, Devi Parikh, Aniruddha Kembhavi

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments 15 pages, 10 figures. To appear in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1611.09490 2017-05-30 cs.RO 67%

Generalized Shared Control versus Classical Shared Control: Illustrative Examples

Pete Trautman

专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1506.08126 2015-06-29 cs.CL cs.AI cs.MM stat.ML 67%

Humor in Collective Discourse: Unsupervised Funniness Detection in the New Yorker Cartoon Caption Contest

Dragomir Radev, Amanda Stent, Joel Tetreault, Aasish Pappu, Aikaterini Iliakopoulou, Agustin Chanfreau, Paloma de Juan, Jordi Vallmitjana, Alejandro Jaimes, Rahul Jha, Bob Mankoff

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL、cs.AI、cs.MM

Comments 10 pages, in submission

详情

展开后加载摘要…

URL PDF HTML 收藏