arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-04-27 至 2026-04-27 共收录 9 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 跨模态检索 9 篇

2210.05513 2026-04-27 cs.CV 70%

ViFiCon: Vision and Wireless Association Via Self-Supervised Contrastive Learning

ViFiCon:通过自监督对比学习实现视觉与无线关联

Nicholas Meegan, Hansi Liu, Bryan Bo Cao, Abrar Alali, Kristin Dana, Marco Gruteser, Shubham Jain, Ashwin Ashok

机构 * Rutgers University(罗切斯特大学) Stony Brook University(石溪大学) Saudi Electronic University(沙特电子大学) Georgia State University(佐治亚州立大学)

专题命中 跨模态检索 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 ViFiCon通过自监督对比学习实现视觉与无线模态的关联,利用RGB-D相机和WiFi测距数据,减少隐私和能耗,无需标注数据即可实现高精度关联。

Comments 8 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22292 2026-04-27 cs.CL cs.AI 62%

ReLeVAnT: Relevance Lexical Vectors for Accurate Legal Text Classification

ReLeVAnT:用于准确法律文本分类的相关词向量

Ishaan Gakhar, Harsh Nandwani

机构 * Perssonify

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出ReLeVAnT框架,通过n-gram处理、对比分数匹配和浅层神经网络实现法律文档二分类,准确率达99.3%。

Comments 9 Pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22061 2026-04-27 cs.CL cs.AI cs.LG 62%

Lightweight Retrieval-Augmented Generation and Large Language Model-Based Modeling for Scalable Patient-Trial Matching

轻量级检索增强生成与基于大语言模型的建模用于可扩展的患者试验匹配

Xiaodi Li, Yang Xiao, Munhwan Lee, Konstantinos Leventakos, Young J. Juhn, David Jones, Terence T. Sio, Wei Liu, Maria Vassilaki, Nansu Zong

机构 * Department of Artificial Intelligence and Informatics, Mayo Clinic(人工智能与信息学系,梅奥诊所) Computer Science Department, University of Tulsa(图兰大学计算机科学系) Mayo Clinic Comprehensive Cancer Center, Mayo Clinic(梅奥诊所综合癌症中心,梅奥诊所) Division of Community Pediatric and Adolescent Medicine, Department of Pediatrics, Mayo Clinic(社区儿科与青少年医学分会,儿科部,梅奥诊所) Department of Neurology, Mayo Clinic(神经病学部,梅奥诊所) Department of Radiation Oncology, Mayo Clinic(放射肿瘤学部,梅奥诊所) Department of Quantitative Health Sciences, Mayo Clinic(定量健康科学部,梅奥诊所)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出轻量级框架,结合检索增强生成和大语言模型建模,解决患者试验匹配中长异构电子健康记录和复杂资格标准的可扩展性、泛化性和计算效率问题。

Comments 31 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22678 2026-04-27 cs.CL 57%

BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering

BERAG:基于贝叶斯集成的检索增强生成用于基于知识的视觉问答

Jinghong Chen, Jingbiao Mei, Guangyu Yang, Bill Byrne

机构 * Department of Engineering(工程系)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CL

AI总结 BERAG通过将语言模型条件于单个检索文档而非单一上下文,解决检索增强生成中文档贡献不明和长上下文中的信息丢失问题,提升视觉问答任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22374 2026-04-27 cs.CL 57%

Selective Contrastive Learning For Gloss Free Sign Language Translation

选择性对比学习用于无 gloss 手语翻译

Changhao Lai, Rui Zhao, Xuewen Zhong, Jinsong Su, Yidong Chen

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) Key Lab of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian-Taiwan (XMU), Ministry of Culture and Tourism, China(福建省 Fujian-Taiwan 非物质文化遗产数字保护与智能处理重点实验室,文化部,中国) National Language Resources Monitoring and Research Center for Education and Teaching Media, Xiamen University, China(教育与教学媒体语言资源监测与研究中心,厦门大学,中国)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL

AI总结 本文提出选择性对比学习方法,通过动态选择负样本提升手语翻译的跨模态对齐效果,减少噪声干扰。

Comments Accepted by ACL 2026 as the main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22180 2026-04-27 cs.IR cs.AI 57%

ResRank: Unifying Retrieval and Listwise Reranking via End-to-End Joint Training with Residual Passage Compression

ResRank:通过端到端联合训练与残差段落压缩统一检索与列表级重排序

Xiaojie Ke, Shuai Zhang, Liansheng Sun, Yongjin Wang, Hengjun Jiang, Xiangkun Liu, Cunxin Gu, Jian Xu, Guanjun Jiang

机构 * Qwen Applications Business Group of Alibaba(阿里巴巴Qwen应用业务组)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.AI

AI总结 ResRank通过端到端联合训练与残差段落压缩技术,解决传统重排序方法中输入长度增加导致的排名质量下降和推理延迟问题,实现高效且有效的检索与重排序统一。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.22174 2026-04-27 cs.CV 57%

Unlocking Optical Prior: Spectrum-Guided Knowledge Transfer for SAR Generalized Category Discovery

解锁光学先验:基于频谱引导的知识迁移用于SAR广义类别发现

Jingyuan Xia, Ruikang Hu, Ye Li, Zhixiong Yang, Xu Lan, Zhejun Lu

机构 * College of Electronic Science and Technology, National University of Defense Technology(电子科学与技术学院,国防科技大学) State Key Laboratory of Complex System Simulation and Modeling Technology(复杂系统仿真与建模技术国家重点实验室)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出基于频谱引导的知识迁移框架,通过频谱差异建模提升SAR图像中光学先验的适应性,实现广义类别发现任务的高效性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21806 2026-04-27 cs.CV 57%

TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval

TEMA: 图像锚定,文本引导的多修改复合图像检索

Zixu Li, Yupeng Hu, Zhiheng Fu, Zhiwei Chen, Yongqi Li, Liqiang Nie

机构 * School of Software, Shandong University(山东大学软件学院) Department of Computing, Hong Kong Polytechnic University(香港理工大学计算机系) School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)计算机科学与技术学院)

专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV

AI总结 TEMA通过构建多修改数据集和提出文本导向实体映射架构,解决复合图像检索中的实体覆盖不足和句法-实体对齐问题,提升检索准确性和效率。

Comments Accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13211 2026-04-27 cs.CV 57%

3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at Scale

3DAlign-DAER:动态注意力策略与高效检索策略用于大规模细粒度3D文本对齐

Yijia Fan, Jusheng Zhang, Kaitong Cai, Jing Yang, Jian Wang, Keze Wang

机构 * Uni3D-g

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出3DAlign-DAER框架,通过动态注意力策略和高效检索策略实现文本与3D几何的细粒度对齐,解决大规模3D数据库下的对齐性能下降问题。

详情

展开后加载摘要…

URL PDF HTML 收藏