arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4549 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 音频语音多模态 4549 篇

2606.05763 2026-06-08 eess.AS cs.SD 版本更新 85%

M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition

M2S-AVSR:面向鲁棒视听语音识别的模态感知多视角自监督表示

Fei Su, Cancan Li, Ming Li, Juan Liu

机构 * School of Artificial Intelligence and the School of Computer Science, Wuhan University, China(人工智能学院和计算机科学学院,武汉大学,中国) School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(人工智能学院,香港中文大学(深圳),中国) School of Artificial Intelligence, Wuhan University, China(人工智能学院,武汉大学,中国)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 eess.AS

AI总结 提出一种模态感知多视角自监督表示框架(M2S-AVSR),通过多视角编码学习视角不变视觉语音表示,并利用模态感知模块进行细粒度融合,以应对视角变化、音频失真和视觉遮挡等挑战,在多个基准上取得最优性能。

Comments submitted to IEEE Transactions on Audio, Speech, and Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14569 2026-06-03 cs.CL cs.LG 85%

Social Caption: Evaluating Social Understanding in Multimodal Models

Social Caption: 评估多模态模型的社会理解能力

Leena Mathur, Bhaavanaa Thumu, Youssouf Kebe, Louis-Philippe Morency

机构 * School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CL

AI总结 提出基于交互理论的SOCIAL CAPTION框架,从社会推理、整体社会分析和定向社会分析三个维度评估多模态大语言模型的社会理解能力,并分析影响性能的因素。

Comments 25 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.23113 2026-05-25 cs.CV 85%

Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake Localization

不一致感知多模态薛定谔桥用于深度伪造定位

Jiayu Xiong, Jing Wang, Qi Zhang, Wanlong Wang, Jun Xue

机构 * Department of Computer Science and Techonology, Huaqiao University(华侨大学计算机科学与技术系) Xiamen Key Laboratory of Computer Vision and Pattern Recognition, Huaqiao University(厦门计算机视觉与模式识别重点实验室) Tongji University(同济大学) School of Cyber Science and Engineering, Wuhan University(武汉大学网络空间安全学院)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 提出IaMSB框架,利用薛定谔桥统一一致性估计、跨模态信息选择和桥步调度,实现高精度区间级深度伪造定位。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07154 2026-05-11 cs.CV 85%

PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition

PRIMED: 通过偏见竞争实现适应性模态抑制的指引用视听分割

Yuchen He, Jing Zhang

机构 * School of Information Science and Engineering, East China University of Science and Technology(信息科学与工程学院,东华大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 PRIMED通过偏见竞争理论,实现适应性模态抑制,提升指引用视听分割的准确性,通过模态先验解码器和token蒸馏器增强多模态融合,实验表明其在Ref-AVS基准上达到最优性能。

Comments 11 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.14129 2026-04-16 cs.CV 85%

Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models

别让视频说话:面向音频-视觉语言模型的音频对比偏好优化

Ami Baid, Zihui Xue, Kristen Grauman

机构 * University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出ACPO方法,通过引入输出对比和输入对比目标,解决音频-视觉语言模型中视频驱动的音频幻觉问题,提升音频真实性和多模态能力。

Comments Project page: https://vision.cs.utexas.edu/projects/acpo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11964 2026-04-15 cs.HC cs.MM 85%

When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs

当绘画不够时:探索通过草图进行意图对齐的自发性言语在多模态大语言模型中的应用

Weiyan Shi, Dorien Herremans, Kenny Tsu Wei Choo

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract,abstract_cn);分类 cs.MM

AI总结 本文探讨了在多模态大语言模型中,通过草图与自发性言语的结合来提升设计初期意图对齐的效果,通过实验表明加入自发性言语能显著提高生成图像的意图匹配度。

Comments Accepted at DIS 2026 PWiP

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08209 2026-04-10 cs.CV 85%

OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering

OmniJigsaw:通过模态协调重排序增强多模态推理

Yiduo Jia, Muzhi Zhu, Hao Zhong, Mingyu Liu, Yuling Xi, Hao Chen, Bin Qin, Yongjie Yang, Zhenbo Luo, Chunhua Shen

机构 * Zhejiang University(浙江大学) Xiaomi Inc.(小米公司)

专题命中 音频语音多模态 :omni-modal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 OmniJigsaw通过模态协调重排序任务提升多模态推理能力,采用三重策略实现跨模态整合,并通过两级数据过滤提升模型适应性,验证了其在多模态自监督学习中的有效性。

Comments Project page: https://aim-uofa.github.io/OmniJigsaw/

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08147 2026-04-10 cs.SD cs.CV 85%

Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning

通过教师引导的双路径实现语义噪声削减

Linge Wang, Yingying Chen, Bingke Zhu, Lu Zhou, Jinqiao Wang

机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所基础模型研究中心) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Objecteye Inc.(北京眼神科技有限公司)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文提出TG-DP框架,通过分离重建与对齐路径,减少语义噪声,提升音频视频表示学习效果,实现零样本检索性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22732 2026-03-25 cs.CV 85%

SOUPLE: Enhancing Audio-Visual Localization and Segmentation with Learnable Prompt Contexts

SOUPLE: 通过可学习的提示上下文增强音频-视觉定位与分割

Khanh Binh Nguyen, Chae Jung Park

机构 * Deakin University(德克萨斯大学) National Cancer Center(国立癌症中心)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);image-text(abstract);分类 cs.CV

AI总结 SOUPLE通过可学习的上下文标记提升音频-视觉定位与分割性能,利用视觉特征生成条件上下文以连接音频与视觉输入。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17181 2026-03-25 cs.CV cs.LG cs.SD 85%

Investigating self-supervised representations for audio-visual deepfake detection

探究用于音频-视觉深度伪造检测的自监督表示

Dragos-Alexandru Boldisor, Stefan Smeu, Dan Oneata, Elisabeta Oneata

机构 * Bitdefender Politehnica Bucharest(巴蒂亚努大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 本文系统评估了自监督表示在多模态和多领域中的检测效果、信息可解释性和跨模态互补性,发现音频引导的表示在深度伪造检测中表现最佳。

Comments Accepted to the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21900 2026-03-10 cs.SD eess.AS 85%

EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs

EmoOmni:弥合情感理解与表达在多模态大语言模型中的鸿沟

Wenjie Tian, Zhixian Zhao, Jingbin Hu, Huakang Chen, Haohe Liu, Binshen Mu, Lei Xie

机构 * Northwestern Polytechnical University, Xi'an, China(西北工业大学,西安,中国) University of Surrey, Guildford, United Kingdom(萨里大学,Guildford,英国)

专题命中 音频语音多模态 :omni-modal(title,abstract);multimodal(abstract);audio-visual(abstract);分类 eess.AS

AI总结 EmoOmni通过引入情感链式思维机制,提升多模态情感对话中的理解与表达能力,构建了评估基准并验证了其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21646 2026-03-04 cs.CL 85%

Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion

可扩展的多语言多模态机器翻译与语音-文本融合

Yexing Du, Youcheng Pan, Zekun Wang, Zheng Chu, Yichong Huang, Kaiyuan Liu, Bo Yang, Yang Xiang, Ming Liu, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学) Pengcheng Laboratory(鹏城实验室)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);image-text(abstract);分类 cs.CL

AI总结 本文提出一种语音引导的机器翻译框架,通过融合语音和文本提升多语言翻译性能,并在多个基准上取得新突破。

Comments Accepted in ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17645 2026-01-27 cs.SD cs.CL cs.CV cs.MM eess.AS 85%

AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

AVMeme Exam: 一个多模态、多语言、多文化的基准测试,用于评估LLM的上下文和文化知识与思维

Xilin Jiang, Qiaolin Wang, Junkai Wu, Xiaomin He, Zhongweiyang Xu, Yinghao Ma, Minshuo Piao, Kaiyi Yang, Xiuwen Zheng, Riki Shimizu, Yicong Chen, Arsalan Firoozi, Gavin Mischler, Sukru Samet Dindar, Richard Antonello, Linyang He, Tsun-An Hsieh, Xulin Fan, Yulun Wu, Yuesheng Ma, Chaitanya Amballa, Weixiong Chen, Jiarui Hai, Ruisi Li, Vishal Choudhari, Cong Han, Yinghao Aaron Li, Adeen Flinker, Mounya Elhilali, Emmanouil Benetos, Mark Hasegawa-Johnson, Romit Roy Choudhury, Nima Mesgarani

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 AVMeme Exam 通过多模态、多语言、多文化的基准测试,评估LLM在上下文和文化理解方面的局限性,揭示模型在无文本音乐和文化思维上的不足。

Comments avmemeexam.github.io/public

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18552 2026-01-21 cs.MM cs.CL cs.CV cs.LG cs.SD eess.AS 85%

Multimodal Emotion Recognition using Audio-Video Transformer Fusion with Cross Attention

基于音频-视频Transformer融合与交叉注意力的多模态情感识别

Joe Dhanith P R, Shravan Venkatraman, Vigya Sharma, Santhosh Malarvannan

机构 * School of Computer Science and Engineering, Vellore Institute of Technology (VIT) University(计算机科学与工程学院,维洛雷理工学院(VIT大学))

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 本文提出AVT-CA模型,通过音频-视频Transformer融合与交叉注意力机制,提升多模态情感识别的准确率和F1分数。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05191 2025-12-30 cs.CV 85%

MokA: Multimodal Low-Rank Adaptation for MLLMs

MokA:多模态低秩适应用于大规模语言模型

Yake Wei, Yu Miao, Dongzhan Zhou, Di Hu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大模型与智能治理研究重点实验室) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程研究中心) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 MokA通过多模态低秩适应提升大规模语言模型的多模态微调效率和效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.02733 2025-11-24 cs.CV cs.AI cs.LG cs.MM cs.SD eess.AS 85%

AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos

AV-Lip-Sync+: 借助AV-HuBERT利用多模态不一致性的深度伪造检测方法

Sahibzada Adil Shahzad, Ammarah Hashmi, Yan-Tsung Peng, Yu Tsao, Hsin-Min Wang

机构 * Social Networks and Human-Centered Computing Program, Taiwan International Graduate Program, Academia Sinica(学术.sinica.edu.tw 社会网络与人本计算计划,台湾国际研究生计划,台湾.sinica.edu.tw) Department of Computer Science, National Chengchi University(计算机科学系,国立政治大学) Institute of Information Systems and Applications, National Tsing Hua University(信息系统与应用研究所,国立清华大学) Research Center for Information Technology Innovation, Academia Sinica(资讯技术创新研究中心,学术.sinica.edu.tw) Institute of Information Science, Academia Sinica(资讯研究所,学术.sinica.edu.tw)

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(abstract);分类 cs.CV、cs.AI、cs.MM

AI总结 本文提出AV-Lip-Sync+方法,利用AV-HuBERT和多模态不一致性检测深度伪造,实现对正面人脸视频的高精度检测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11124 2025-11-17 cs.CL cs.AI cs.CV cs.MM cs.SD 85%

AV-Dialog: Spoken Dialogue Models with Audio-Visual Input

Tuochao Chen, Bandhav Veluri, Hongyu Gong, Shyamnath Gollakota

机构 * Paul G. Allen School of Computer Science & Engineering, University of Washington(华盛顿大学保罗·G·阿伦计算机科学与工程学院) Meta AI Research(Meta AI研究院)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10059 2025-11-14 cs.CV 85%

When Eyes and Ears Disagree: Can MLLMs Discern Audio-Visual Confusion?

Qilang Ye, Wei Zeng, Meng Liu, Jie Zhang, Yupeng Hu, Zitong Yu, Yu Zhou

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13630 2025-11-12 cs.CV 85%

AVAR-Net: A Lightweight Audio-Visual Anomaly Recognition Framework with a Benchmark Dataset

Amjid Ali, Zulfiqar Ahmad Khan, Altaf Hussain, Muhammad Munsif, Adnan Hussain, Sung Wook Baik

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments I would like to request the withdrawal of my paper . The reason for this request is that I am currently working on additional experiments and analyses, which will lead to updates in the results section. Once these updates are complete, I will resubmit the revised version. Thank you for your understanding

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01299 2025-11-04 eess.AS 85%

Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking

Siyin Wang, Zengrui Jin, Changli Tang, Qiujia Li, Bo Li, Chen Chen, Yuchen Hu, Wenyi Yu, Yixuan Li, Jimin Zhuang, Yudong Yang, Mingqiu Wang, Michael Han, Yifan Ding, Junwen Bai, Tom Ouyang, Shuo-yiin Chang, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Lu Lu, Guangzhi Sun, Zhehuai Chen, Ji Wu, Bowen Zhou, Yuxuan Wang, Tara Sainath, Yonghui Wu, Chao Zhang

专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract);audio-visual(abstract);分类 eess.AS

Comments 22 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10051 2025-10-14 cs.CV 85%

Complementary and Contrastive Learning for Audio-Visual Segmentation

Sitong Gong, Yunzhi Zhuge, Lu Zhang, Pingping Zhang, Huchuan Lu

机构 * School of Information and Communication Engineering, Dalian University of Technology(信息与通信工程学院,大连理工大学) School of Future Technology and the School of Artificial Intelligence, Dalian University of Technology(未来技术学院和人工智能学院,大连理工大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

Comments Accepted to IEEE Transactions on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20862 2025-10-01 cs.CV 85%

AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding

Chaeyoung Jung, Youngjoon Jang, Joon Son Chung

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24214 2025-09-30 cs.CV 85%

Scalable Audio-Visual Masked Autoencoders for Efficient Affective Video Facial Analysis

Xuecheng Wu, Junxiao Xue, Xinyi Yin, Yunyun Shi, Liangyu Fu, Danlei Huang, Yifan Wang, Jia Zhang, Jiayu Nie, Jun Wang

机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) School of Cyber Science and Engineering, Zhengzhou University(郑州大学网络科学与工程学院) School of Software, Northwestern Polytechnical University(西北工业大学软件学院) Institute of Advanced Technology, University of Science and Technology of China(中国科学技术大学先进技术研究院) Inspur Group(Inspur集团) Research Center for Space Computing System, Zhejiang Lab(浙江实验室空间计算系统研究中心)

专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07173 2025-09-30 cs.CL 85%

Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models

Leyi Pan, Zheyu Fu, Yunpeng Zhai, Shuchang Tao, Sheng Guan, Shiyu Huang, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Felix Henry, Aiwei Liu, Lijie Wen

机构 * Tsinghua University(清华大学) Tongyi Lab(通义实验室) Peking University(北京大学) OpenRL Lab(OpenRL实验室)

专题命中 音频语音多模态 :audio-visual(title,abstract);cross-modal(abstract);omni-modal(abstract);分类 cs.CL

Comments 22 pages, 10 figures, 12 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18912 2025-09-24 cs.CV 85%

Frequency-Domain Decomposition and Recomposition for Robust Audio-Visual Segmentation

Yunzhe Shen, Kai Peng, Leiye Liu, Wei Ji, Jingjing Li, Miao Zhang, Yongri Piao, Huchuan Lu

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09595 2025-09-18 cs.CV 85%

Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis

Yikang Ding, Jiwen Liu, Wenyuan Zhang, Zekun Wang, Wentao Hu, Liyuan Cui, Mingming Lao, Yingchao Shao, Hui Liu, Xiaohan Li, Ming Chen, Xiaoqiang Liu, Yu-Shen Liu, Pengfei Wan

机构 * Kling Team, Kuaishou Technology(快手科技 Kling 团队)

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);audio-visual(abstract);分类 cs.CV

Comments Technical Report. Project Page: https://klingavatar.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06375 2025-08-29 cs.MM cs.AI cs.CL cs.CV cs.LG cs.SD 85%

OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset

Jeongkyun Park, Jung-Wook Hwang, Kwanghee Choi, Seung-Hyun Lee, Jun Hwan Ahn, Rae-Hong Park, Hyung-Min Park

机构 * 1 Department of Artificial Intelligence, Sogang University, Seoul 04107, Republic of Korea 2 Department of Electronic Engineering, Sogang University, Seoul 04107, Republic of Korea 3 Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA 15213, USA 4 Mindslab Inc., Gyeonggi-do 13493, Republic of Korea 5 ICT Convergence Disaster/Safety Research Institute, Sogang University, Seoul 04107, Republic of Korea

专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI

Comments Accepted to ICASSP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16188 2025-08-29 cs.CL cs.CV cs.MM cs.SD eess.AS 85%

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

Weiting Tan, Jiachen Lian, Hirofumi Inaguma, Paden Tomasello, Philipp Koehn, Xutai Ma

机构 * Johns Hopkins University(约翰霍普金斯大学) Meta AI Research(Meta AI 研究)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

Comments EMNLP 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22886 2025-08-01 cs.CV 85%

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

Kaining Ying, Henghui Ding, Guangquan Jie, Yu-Gang Jiang

机构 * Fudan University, China(复旦大学)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);MLLM(abstract);分类 cs.CV

Comments ICCV 2025, Project Page: https://henghuiding.com/OmniAVS/

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22781 2025-07-31 cs.CV 85%

HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training

Xuecheng Wu, Danlei Huang, Heli Sun, Xinyi Yin, Yifan Wang, Hao Wang, Jia Zhang, Fei Wang, Peihao Guo, Suyu Xing, Junxiao Xue, Liang He

机构 * Xi'an Jiaotong University(西安交通大学) Zhengzhou University(郑州大学) University of Science and Technology of China(中国科学技术大学) Dalian University of Technology(大连理工大学) Zhejiang Lab(浙江实验室)

专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);cross-modal(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏