Beyond Images: An Integrative Multi-modal Approach to Chest X-Ray Report Generation
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM
Comments NeurIPS 2023 spotlight
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Project page & demo: https://aka.ms/biomedjourney
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to EMNLP 2023 main conference
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
Comments ICCV 2023
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.MM、eess.AS
专题命中 多模态生成 :any-to-any(title);multimodal(abstract);分类 cs.CV、cs.CL、eess.AS
Comments Project Page: https://codi-gen.github.io
专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI、cs.MM
Comments IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted by the 30th ACM International Conference on Multimedia (ACM MM 2022)
专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)
Comments Published in Advanced Robotics
Journal ref Advanced Robotics, 36:5-6, 261-278, 2022
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CV、cs.AI、cs.MM
Comments arXiv admin note: text overlap with arXiv:2105.14211
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments CVPR Camera ready
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM
Comments ACM MM 2021 (Industrial Track). Code: https://github.com/researchmm/generate-it
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.MM
Comments ACL Fingdings 2021
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments visit https://fairytailor.org/ and https://github.com/EdenBD/MultiModalStory-demo for web demo and source code
专题命中 多模态生成 :multi-modal(title,abstract);cross-modal(abstract)
Comments 5 pages, 2 figures
专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract)
Journal ref IEEE Transactions on Biomedical Engineering, 2019
专题命中 多模态生成 :multi-modal(title,abstract);multimodal(abstract)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments EMNLP 2018 (accepted)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Comments arXiv admin note: text overlap with arXiv:1612.04757
DGSSM:基于扩散引导的状态空间模型的多模态显著目标检测
机构 * Dept. of Computer Science and Engineering, Indian Institute of Technology, Guwahati(计算机科学与工程系,印度理工学院,果阿提)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI
AI总结 本文提出DGSSM,一种结合扩散模型结构先验和多尺度状态空间编码的多模态显著目标检测框架,通过迭代Mamba扩散细化机制提升边界精度,实验表明其在多个评估指标上优于现有方法。
Comments Accepted at ICPR 2026. Diffusion-guided Mamba framework for multimodal salient object detection. Evaluated on 13 benchmarks (RGB, RGB-D, RGB-T)
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments This article is an earlier version of my work arXiv:2410.01820 "PixelBytes: Catching Unified Representation for Multimodal Generation."
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.CL、cs.AI
Comments Multi-modal, Large Language Models, Tokenizer, Understanding and Generation
SEED: 用于可解释文本伪造检测的简单ViT与演化框架
机构 * State Key Laboratory of Internet of Things for Smart City, Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系智慧城市物联网国家重点实验室) ; School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)
专题命中 多模态生成 :MLLM(summary_cn,abstract);分类 cs.CV
AI总结 提出SEED系统,结合相似性引导数据增强、单一ViT联合检测与定位、以及基于MLLM的演化报告生成,在ACM MM 2026文本伪造挑战赛中获得第三名。
UniGen-AR:通过自回归建模统一视觉生成
机构 * Carnegie Mellon University(卡内基梅隆大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳 - 香槟分校) ; Toyota Research Institute(丰田研究院)
专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);multi-modal(abstract);分类 cs.CV
AI总结 研究统一视觉生成问题,提出UniGen-AR框架,将通用多模态语言模型与视觉自回归解码器配对,能为多任务生成图像值输出,相比基于扩散的基线,推理延迟低达19倍,确立视觉自回归建模为统一视觉生成的高效主干。
GeoSearcher: 基于锚点引导的渐进推理遥感视觉定位与过程监督
机构 * Key Laboratory of Target Cognition and Application Technology (TCAT), Chinese Academy of Sciences(中国科学院目标认知与应用技术重点实验室) ; School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences(中国科学院大学电子电气与通信工程学院)
专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 提出GeoSearcher,通过锚点引导的渐进推理和过程监督,将遥感视觉定位转化为两阶段过程,解决小目标定位和复杂查询的挑战。
Comments 14 pages, 11 figures, 7 tables
模态解耦的在线递归编辑
机构 * Harbin Institute of Technology, Shenzhen, China.(哈尔滨工业大学(深圳)) ; Peng Cheng Laboratory, China.(鹏城实验室) ; Huazhong University of Science and Technology, China(华中科技大学)
专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);cross-modal(abstract);分类 cs.AI
AI总结 本文提出M-ORE,一种用于持续多模态大语言模型适应的模态解耦在线递归编辑器,通过统一的近端投影公式和Sherman-Morrison递归实现常数级的每编辑开销,从而在保持模块局部统计信息和固定正交低秩编辑子空间的同时,减少长周期干扰,提升可靠性、通用性和局部性。