Reconstructing Nonlinear Dynamical Systems from Multi-Modal Time Series
专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(abstract)
Comments 19 pages, 8 figures
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 其他多模态 :multi-modal(title,abstract);cross-modal(abstract)
Comments 19 pages, 8 figures
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted for publication at SemEval-2022 Workshop, Task 5: MAMI - Multimedia Automatic Misogyny Identification co-located with NAACL 2022
专题命中 其他多模态 :multi-modal(title,abstract);multimodal(abstract)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM
Comments 22 pages, 11 figures. Under review at Digital Scholarship in the Humanities
Journal ref Digital Scholarship in the Humanities, 2021 (ahead of press)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Updated acknowledgements
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM
Comments 13 pages, 2 figures
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Published in IEEE Access
Journal ref IEEE Access, 2021
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract)
Comments Accepted by IJCAI 2020. SOLE copyright holder is IJCAI (international Joint Conferences on Artificial Intelligence)
专题命中 其他多模态 :cross-modal(title,abstract);multi-modal(abstract)
Comments 8 pages, in Proceedings of the 2019 International Joint Conference on Neural Networks (IJCNN)
专题命中 其他多模态 :multi-modal(title,abstract);multimodal(abstract)
专题命中 其他多模态 :multimodal(title,abstract);multi-modal(abstract)
专题命中 其他多模态 :cross-modal(title,abstract);multimodal(abstract)
专题命中 其他多模态 :multimodal(title,abstract);cross-modal(abstract)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 20 pages, 16 figures. Accepted to CVPR 2024 Workshop on Prompting in Vision. Project page: https://folbaeni.gitlab.io/multimodal-icl
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 11 pages, 6 figures, Multimodal learning, NSCLC, Survival analysis, Transformer
EvoGuard: 一种基于代理强化学习的可扩展框架,用于实际和演化的AI生成图像检测
机构 * The University of Tokyo(东京大学) ; National Institute of Informatics(国家信息研究所)
专题命中 其他多模态 :MLLM(summary_cn,abstract_cn);multimodal(abstract);分类 cs.CV
AI总结 本文提出EvoGuard框架,利用多模态大语言模型和非MLLM检测器,通过能力感知的动态编排机制实现高效AIGI检测,优化后无需细粒度标注,实验表明其在准确性和可扩展性上均达最优。
Comments Template changed
构造性失真:通过注意力引导的图像扭曲改进MLLMs
机构 * University of Illinois Urbana–Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Skan AI ; Texas A&M University(德克萨斯大学阿姆斯特朗分校) ; University of California, Irvine(加州大学 Irvine 分校)
专题命中 其他多模态 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 本文提出AttWarp方法,通过注意力引导的图像扭曲增强MLLMs对细节和空间关系的感知能力,提升准确性和推理能力。
Comments Accepted at ICLR 2026
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);cross-modal(abstract);multimodal foundation model(abstract)
Comments ICML2024
基于模拟响应概率的多模态题目参数估计
机构 * College Board(大学理事会)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI
AI总结 本研究基于Qwen3.5微调多模态大语言模型,利用含图像与文本的多项选择题训练语料,通过学习学生能力相关的选择概率,实现了题目参数的准确估计。
Comments Submitted and Accepted for AIME-Con 2026
分析多模态大语言模型代理在城市感知中生成解释的角色效应
机构 * Universidade Tecnologica Federal do Parana(巴西南里奥格兰德联邦技术大学) ; University of Toronto(多伦多大学)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
AI总结 通过对比不同角色提示和无角色设置下多模态大语言模型生成的文本,发现标题描述趋同,但理由描述随社会经济和政治属性系统变化,感知标签无显著差异。
Comments 17 pages, 9 figures
iStructTab:面向图像与表格数据多模态学习的结构化特征排序
机构 * West Virginia University(西弗吉尼亚大学) ; The University of Utah(犹他大学) ; Scientific Computing and Imaging Institute(科学计算与成像研究所)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
AI总结 iStructTab基于图增强描述子排序算法(GEDS),将其融入感知顺序的高效Transformer框架,在多模态基准上有效减少特征分散,提升了图像与表格数据多模态学习的预测性能及鲁棒性。
Comments This paper has been accepted for presentation at the 28th International Conference on Pattern Recognition (ICPR 2026) in Lyon, France Code: https://github.com/zadid6pretam/iStructTab PyPI: pip install istructtab
Journal ref International Conference on Pattern Recognition (ICPR 2026)
受损多模态语言模型再现失语症患者的图片命名模式
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI
AI总结 研究未针对临床模拟设计的通用语言模型能否再现失语症患者图片命名错误模式,通过对多模态语言模型进行扰动配置,建立定量框架再现个体失语症错误模式,表明语言模型或可成为中风后失语症患者数字替身。
Comments 15 pages, 8 figures; supplementary materials (18 pages, 6 sections) included
MultiFair:具有双级梯度调制的多模态平衡公平感知医学分类
机构 * School of Computing and Informatics, University of Louisiana at Lafayette(路易斯安那州立大学拉法叶分校计算机与信息学学院) ; Louisiana Center for Health Innovation and College of Nursing & Health Sciences, University of Louisiana at Lafayette(路易斯安那州立大学拉法叶分校健康创新中心及护理与健康科学学院) ; Massachusetts Eye and Ear, Harvard Medical School(哈佛医学院马萨诸塞眼耳医院) ; Department of Computer Science, University of Central Florida(佛罗里达州立大学计算机科学系) ; Department of Electrical Engineering and Computer Science, Florida Atlantic University(佛罗里达Atlantic大学电子工程与计算机科学系)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
AI总结 针对多模态医学分类中数据模态学习不均衡和模型对特定群体不公平的问题,提出MultiFair方法,通过双级梯度调制过程动态调整训练梯度,在多数据集上评估,有效解决了上述挑战。
Comments This work has been accepted for publication in IEEE Transactions on Medical Imaging
回答前分割:用于多模态大语言模型视觉推理的像素定位
专题命中 其他多模态 :MLLM(title);multimodal(abstract);分类 cs.CV、cs.AI
AI总结 本文针对多模态大语言模型视觉推理,提出SegAnswer方法,将放大单元从边界框转为像素级分割掩码,能精准定位感兴趣区域,过滤冗余信息,与视觉令牌构建方式更契合,经多基准测试验证了其在像素定位及分割任务上的性能。
残差引导的专家特化用于不完整多模态学习
机构 * Pohang University of Science and Technology (POSTECH)(釜山科学技术大学)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
AI总结 提出MARS框架,利用缺失引起的表示偏差引导专家特化,通过残差信号路由样本到对应专家,并引入差异感知噪声正则化缓解训练-测试路由差距,在缺失场景下提升多模态分类与分割性能。
Comments ECCV 2026
GeoLocator:一种集成位置的大型多模态模型,用于推断地理隐私
机构 * Spatial Sciences Institute, University of Southern California(南加州大学空间科学研究所) ; Evolutionary Assets(进化资产) ; Viterbi school of engineering, University of Southern California(南加州大学维特比工程学院)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
AI总结 本文提出GeoLocator模型,用于推断输入影像和社会媒体中的地理位置信息,揭示了模型在地理隐私保护中的潜在风险。
Comments 16pages, 2 figures
Journal ref Appl. Sci. 14(16), 7091 (2024)
通过奖励设计解耦多模态大语言模型中的感知与推理
机构 * Department of Computer Science, Dartmouth College(达特茅斯大学计算机科学系) ; Department of Computer Science, University of Central Florida(中央佛罗里达大学计算机科学系) ; Independent Researcher(独立研究者)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
AI总结 研究多模态大模型中感知与推理的瓶颈,发现感知是主要约束,并通过奖励设计提升视觉基础推理,平均提升5.56分。
Comments 24 pages, 15 Figures, 10 Tables
双不确定性引导的多模态推理策略学习
机构 * Tencent Hunyuan(腾讯文汇) ; University of Maryland(马里兰大学) ; University of North Carolina(北卡罗来纳大学)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL、cs.AI
AI总结 提出DUPL方法,通过量化感知不确定性和输出不确定性来引导策略更新,在多个多模态推理基准上显著提升模型准确率,优于现有方法。
双调优:面向多模态大语言模型训练的推理效能驱动的数据筛选
机构 * Ant Group(蚂蚁集团)
专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL
AI总结 本文提出双调优框架,通过评估训练数据对推理训练的收益及推理训练的效能,指导多模态大语言模型的数据筛选与训练策略匹配。
Comments Project Page: https://digital-avatar.github.io/ai/ThinkingBoundary/