Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language Models
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.LG
Comments 9 pages, 7 figures
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV、cs.LG
Comments 9 pages, 7 figures
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI
Comments ICML 2024 (Oral)
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.AI、cs.LG
Comments Accepted to workshop on ReGenAI@CVPR 2024
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract,comments);分类 cs.CV
Comments ICLR'24 (Spotlight) ; Project at https://mllm-ie.github.io ; Code at https://github.com/tsujuifu/pytorch_mgie
RecoReward:用于推荐的推荐器引导多模态描述生成
专题命中 其他VLM :MLLM(summary_cn,abstract);multimodal large language model(abstract)
AI总结 研究针对多模态推荐中传统方法不足,提出RecoReward,训练时用行为衍生奖励,保留仅内容推理。在直播推荐中利用用户历史等估计亲和力,通过推荐器亲和力分数提供反馈,实验表明该方法能提升MLLM性能,利于下游推荐且保留内容服务。
Comments 16 pages, 4 figures
个性化多模态大语言模型中的群体偏好崩溃
机构 * Computer Vision Center, Universitat Autònoma de Barcelona(计算机视觉中心,巴塞罗那自治大学)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI
AI总结 研究个性化多模态大语言模型中群体偏好崩溃问题,提出PrefMoE框架,它分离配置文件信息与偏好表示,分解偏好,保留个性化残差并通过单独路径路由因素,实验表明该框架能改善偏好敏感个性化并减少偏好崩溃。
EmoBench-M:多模态大语言模型情感智能评估基准
机构 * Shenzhen University(深圳大学) ; Guangdong Laboratory of Artificial Intelligence(广东人工智能与数字经济实验室) ; Auckland University of Technology(奥克兰理工大学) ; Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; SKL-IOTSC, CIS, University of Macau(澳门科学技术大学SKL-IOTSC、CIS)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI
AI总结 本文提出EmoBench-M,通过建立心理理论基础,评估多模态大语言模型在13种场景中的情感识别、理解及社会复杂情感分析能力,揭示模型在情感智能方面的性能差距。
基于脉冲的多模态大语言模型:通过模态特定的时间尺度和时间压缩
机构 * 1 Institute of Automation, Chinese Academy of Sciences 2 University of Chinese Academy of Sciences 3 Beijing Academy of Artificial Intelligence 4 Zhongguancun Academy 5 Peking University 6 Key Laboratory of Brain Cognition ; Brain-inspired Intelligence Technology 7 Spiking Intelligence Lab, Tianqiao \& Chrissy Chen Institute [0.5em] Equal contribution Corresponding authors
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract_cn);分类 cs.AI
AI总结 本文提出SpikeMLLM,首个基于脉冲的多模态大语言模型框架,通过模态特定时间尺度和时间压缩技术,在减少时间步数的同时保持高性能,实验显示其在多个基准上表现优异。
ABMAMBA: 多模态大语言模型中的对齐层次双向扫描用于高效的视频描述
机构 * Keio University(庆应义塾大学) ; National Institute of Informatics(国立信息学研究所) ; National Institute of Informatics Research and Development Center for Large Language Models(国立信息学研究所大型语言模型研发中心)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
AI总结 本文提出ABMamba,一种具有线性计算复杂度的多模态大语言模型,通过替代二次注意力机制,实现视频序列的高效处理,在视频描述任务中表现出色。
Comments Accepted to ICPR 2026
基于多模态大语言模型的可扩展且可解释的学习者-视频交互预测
机构 * EPFL(瑞士联邦理工学院洛桑)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI
AI总结 本文提出利用多模态大语言模型预测学习者观看、暂停、跳过和回放行为,以评估视频内容的认知负荷,通过7700万次视频控制事件验证了模型的可扩展性和解释性。
Comments Accepted as long paper to the 27th International Conference on Artificial Intelligence in Education (AIED 2026)
当视觉隐私保护与多模态大语言模型相遇
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
AI总结 本文研究多模态大语言模型在提供便利性的同时如何保护视觉隐私,提出基于黑盒模型的优化框架,通过帕累托最优和关键历史增强优化实现隐私与性能的平衡。
Journal ref Int J Comput Vis (IJCV) 134, 167 (2026)
OrchMLLM: 通过批量后平衡 orchestrate 多模态数据以加速多模态大语言模型训练
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI
AI总结 OrchMLLM通过批量后平衡调度器和全局协调器解决多模态数据训练中的模态不一致问题,提升训练效率和可扩展性。
利用多模态大语言模型生成合成缺陷图像用于电力线路绝缘子检测
机构 * Department of Electrical and Computer Engineering, Wayne State University(电气与计算机工程系,韦恩州立大学)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
AI总结 利用多模态大语言模型生成合成缺陷图像,提升电力线路绝缘子缺陷识别的准确率和数据效率。
Comments Submitted to Engineering Applications of Artificial Intelligence, Feb. 16, 2026
多模态大语言模型用于低资源语言:巴斯克语的案例研究
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI
AI总结 本文通过开发巴斯克语多模态大语言模型,证明低比例多模态数据即可获得良好性能,且无需巴斯克语指导的LLM即可实现强模型。
基于大语言模型的引导式目标发现指导:用于基于事实的多模态大语言模型医疗报告生成
机构 * Zhejiang University(浙江大学) ; The Second Affiliated Hospital Zhejiang University School of Medicine(浙江大学医学院附属第二医院) ; Hangzhou Pu Jian Medical Technology Co., Ltd.(杭州普健医疗科技有限公司)
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV
AI总结 本文提出Fact-Flow框架,通过引导大语言模型生成事实准确的医疗报告,利用LLM自动生成标注数据集,提升医疗报告生成的事实准确性。
Comments 10 pages, 1 figure
多模态大语言模型在图像标签任务中是好的标注者吗?
机构 * RIKEN Center for Advanced Intelligence Project, Japan(日本先进人工智能研究中心) ; The University of Tokyo, Japan(东京大学) ; Southeast University, China(东南大学) ; China University of Mining(中国矿业大学)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
AI总结 本文提出TagLLM框架,通过候选生成和标签歧义消除方法,有效缩小多模态大语言模型生成标注与人工标注之间的差距,提升图像标签任务的效率和性能。
多模态大语言模型如何支持视障人士获取视觉信息:一项与盲人和低视力人士的日记研究
机构 * Cornell University(康奈尔大学)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI
AI总结 本研究探讨了多模态大语言模型如何通过视觉助手技能支持视障人士获取视觉信息,并发现其在实际应用中的表现及改进方向。
Comments 24 pages, 17 figures, 7 tables, appendix section, to appear main track CHI 2026
Artic: 面向AI的实时通信用于多模态大语言模型视频助手
机构 * Peking University(北京大学)
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.AI
AI总结 Artic提出了一种面向AI的实时通信框架,通过自适应比特率、上下文感知流式传输和退化视频理解基准提升多模态大语言模型视频助手的准确性和降低延迟。
InternSVG:通过多模态大语言模型实现统一的SVG任务
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Nanjing University(南京大学) ; Donghua University(东华大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
AI总结 InternSVG通过多模态大语言模型实现统一的SVG任务建模,结合大规模数据集和两阶段训练策略,提升SVG理解、编辑和生成性能。
MLLM-VADStory: 基于领域知识的多模态大语言模型用于视频广告剧情洞察
机构 * Meta
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV
AI总结 MLLM-VADStory通过领域知识引导多模态大语言模型,系统量化和生成视频广告剧情洞察,提升视频广告创意效果。
Exo2Ego:基于外部知识引导的多模态大语言模型用于第一人称视频理解
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV
AI总结 Exo2Ego通过迁移学习提升内向视频理解能力,利用外向知识增强模型性能。
Comments This paper is accepted by AAAI 2026
你可能可以自由发言:通过答案提取提升多模态大语言模型的细粒度视觉识别能力
机构 * University of Massachusetts, Amherst(马萨诸塞大学阿姆赫斯特分校) ; Brown University(布朗大学)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
AI总结 本文提出nlg2choice方法,通过两阶段策略提升多模态大语言模型在细粒度视觉识别任务中的表现,通过开放性问题和受限解码提高检索效率。
Comments Accepted to WACV26. 12 pages, 8 tables, 5 figures
当隐私与恢复相遇: surrogate驱动的隐私保护在MLLM编辑中的被忽视的一半
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV
AI总结 本文提出SPPE数据集和统一方法,通过引导生成任务实现隐私恢复,平衡隐私保护与MLLM可用性。
Comments 9 pages,7figures
机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) ; School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) ; Tencent Youtu Lab(腾讯优图实验室) ; Xiamen University(厦门大学) ; CASIA(中国科学院自动化研究所)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
Comments NeurIPS DB 2025 Spotlight, Project Page: https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Evaluation
机构 * Stanford University(斯坦福大学) ; Institute for Computational and Mathematical Engineering (ICME)(计算与数学工程研究所)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.CV
Comments Proceedings of the 10th Machine Learning for Healthcare Conference, PMLR 298, 2025
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI
机构 * State Key Laboratory for Novel Software Technology(新型软件技术国家重点实验室) ; Nanjing University(南京大学) ; School of Computer Science(计算机学院)
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.LG
Comments 17 pages, 13 figures, the second version
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.CV
机构 * National Taiwan Normal University(台湾国立正常大学)
专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract);分类 cs.AI
Comments Accepted at IEEE ASRU 2025
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ; Beijing Future Brain Education Technology Co., Ltd.(北京未来脑教育科技有限公司) ; The Hong Kong University of Science and Technology(香港科技大学) ; Tsinghua University(清华大学) ; Guangxi Zhuang Autonomous Region Big Data Research Institute(广西壮族自治区大数据研究院)
专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract);分类 cs.AI
Comments Accepted by ACL Findings 2025