arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-08-04 至 2026-08-04 共收录 19 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 19 篇

2606.04015 2026-08-04 eess.SP 版本更新 88%

GenED-SC: Generative Editing Semantic Communication with Integrated Multi-Modal LLMs

GenED-SC:集成多模态大模型的生成式编辑语义通信

Shuoyao Wang, Suzhi Bi, Mingze Gong, Zhanpeng Wang, Li Ping Qian, Qiang Ye

专题命中 多模态生成 :MLLM(summary_cn,abstract);multi-modal(title);multimodal(abstract)

AI总结 提出一种两阶段语义图像传输框架,结合JSCC判别传输与MLLM生成编辑,在低信噪比下提升语义保真度和感知质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01113 2026-08-04 cs.CV 新提交 84%

CoT-Edit: Let CoT Guide Instruction Video Editing

CoT-Edit:让思维链(CoT)指导指令视频编辑

Sen Liang, Fengbin Guan, Youliang Zhang, Xin Li, Zhibo Chen

机构 * University of Science and Technology of China(中国科学技术大学) Zhongguancun Academy(中关村学院) Tsinghua University(清华大学)

专题命中 多模态生成 :MLLM(summary_cn,abstract);multimodal(abstract);分类 cs.CV

AI总结 本文提出CoT-Edit的plan--guide--edit框架,以CoT增强的MLLM为规划器生成空间先验,结合扩散编辑器实现高保真指令视频编辑,性能优于多个基准方法

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00410 2026-08-04 cs.AI cs.CL cs.CV 新提交 82%

Where did the ambiguity go? Examining how multimodal models interpret polysemous words

歧义去了哪里?探究多模态模型如何解释多义词

Jasin Cekinmez, Addison J. Wu, Raja Marjieh, Thomas L. Griffiths

机构 * Princeton University(普林斯顿大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 该研究对比17个文本到图像模型和15个文本生成模型,发现多模态模型生成图像的词义多样性低于文本,揭示了基础模型在不同模态间意义表达的迁移 gap。

Comments Oral Presentation, Sci-FM Workshop @ COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16898 2026-08-04 cs.CV cs.CL cs.CR 版本更新 81%

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing

跨分支冲突作为一种保护手段:在统一多模态图像编辑中保护面部身份

Weiwei Tan, Junxian Li, Rui Wang, Zhenhua Xu, Yanjun Zhang, Yu Leo Zhang

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV、cs.CL

AI总结 研究统一多模态模型中个人肖像编辑的安全问题,提出CCS统一对抗保护框架,通过联合驱动ViT和VAE表示及破坏跨分支兼容性,防止模型恢复身份信息,在抑制身份保留编辑上表现优于现有方法。

Comments 14 pages, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01942 2026-08-04 cs.CV cs.CL cs.MM 新提交 80%

CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

CultureVidBench:文本到视频生成中的文化理解基准测试

Xianjing Han, Yuhan Su, Yang Deng, Dong Ma, Wee Peng Tay, Bin Zhu

机构 * Nanyang Technological University(南洋理工大学) Centrale Supélec(中央高等电力学院) Singapore Management University(新加坡管理大学) University of Cambridge(剑桥大学)

专题命中 多模态生成 :MLLM(abstract,abstract_cn);multimodal(abstract);分类 cs.CV、cs.CL、cs.MM

AI总结 本研究推出CultureVidBench基准,评估7款T2V模型的文化理解能力,发现现有模型难捕捉细粒度文化细节,尤其针对代表性不足的文化区域与线索。

Comments Project page:https://hanxjing.github.io/CultureVidBench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10767 2026-08-04 cond-mat.mes-hall cond-mat.mtrl-sci quant-ph 版本更新 71%

Lambert W Function Framework for Graphene Nanoribbon Quantum Sensing: Theory, Verification, and Multi-Modal Applications

基于拉姆伯特W函数框架的石墨烯纳米带量子传感:理论、验证与多模应用

F. A. Chishtie, K. Roberts, N. Jisrawi, S. R. Valluri, A. Soni, P. C. Deshmukh

专题命中 多模态生成 :multi-modal(title)

AI总结 基于拉姆伯特W函数的石墨烯纳米带量子传感框架,通过理论验证和多模应用,实现了灵敏度增强和性能预测。

Comments 21 pages, 10 figures, published version at Results in Engineering journal

Journal ref Results in Engineering, Volume 32, 2026, 112238

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26817 2026-08-04 cs.RO cs.CV 版本更新 70%

From Uncertainty to Determinism: Coarse-to-Fine Visual Floorplan Localization without Ray Matching

从不确定性到确定性:无射线匹配的由粗到细视觉楼层平面图定位

Shiyong Meng, Bolei Chen, Ping Zhong, Yang Wan, Rongzhi Wang, Jiazhi Xia, Jianxin Wang

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 该研究针对视觉楼层平面图定位的多模态姿态分布问题,提出无射线匹配的由粗到细框架,通过图像条件姿态扩散模型与局部细化器实现定位,在S3D和ZInD基准上达到最优精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02817 2026-08-04 cs.CV 版本更新 70%

MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling

MMPhysVideo: 通过联合多模态建模提升视频生成的物理合理性

Shubo Lin, Xuanyang Zhang, Wei Cheng, Weiming Hu, Gang Yu, Jin Gao

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, CASIA(中国科学院自动化研究所多模态人工智能系统国家重点实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) StepFun(阶跃星辰) Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information(多模态信息超级智能安全北京市重点实验室) School of Information Science and Technology, Shanghai Tech University(上海科技大学信息科学与技术学院)

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 MMPhysVideo通过联合多模态建模提升视频生成的物理合理性,将感知线索统一为伪RGB格式,提出双向控制教师架构以减少跨模态干扰,并通过MMPhysPipe构建物理丰富的多模态数据集,提升物理合理性和视觉质量。

Comments Project Page: https://shubolin028.github.io/MMPhysVideo-Page

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02477 2026-08-04 cs.IR 新提交 67%

Unpaired Modality-Agnostic Generative Recommendation

非配对模态无关生成式推荐

Weihao Shen, Wei Chen, Fuwei Zhang, Meng Yuan, Yuqin Lan, Guojun Liu, Qingsong Hua, Wei Lin, Fuzhen Zhuang

专题命中 多模态生成 :multimodal(abstract);cross-modal(abstract)

AI总结 提出UnpairGR模型,从多类观测中学习统一语义ID空间,在三类基准数据集上验证其可提升完全与不完全观测下的推荐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02505 2026-08-04 cs.AI cs.CV cs.IR 新提交 62%

Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation

无实体的溯因推理?科学假说生成的表征奠基与溯因循环

Michael Farmer

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出无需实体的科学溯因推理,构建含弃权机制的溯因循环架构,以惯例空间解决跨领域检索难题,通过案例抽象架构并提出DAB-30基准作为评估方案。

Comments 20 pages, 4 figures. DAB-30 execution reported in companion paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01298 2026-08-04 cs.CV cs.AI cs.LG 新提交 62%

UDT: Reconciling U-Nets and Diffusion Transformers with Data-Adaptive Token Reduction

UDT:通过数据自适应令牌缩减协调U-Net与扩散Transformer

Junno Yun, Yaşar Utku Alçalar, Mehmet Akçakaya

机构 * University of Minnesota(明尼苏达大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出UDT架构,通过数据自适应令牌合并协调U-Net与DiTs,在ImageNet图像生成任务中收敛速度更快、FID指标更优,为DiTs提供了新的高效骨干网络。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00279 2026-08-04 eess.IV cs.CV 新提交 57%

Learning to See Locally and Align Clinically with Pathology Semantics for Radiology Report Generation

学习局部感知并与病理学语义临床对齐的放射学报告生成方法

Xuan Cuong Ngo

专题命中 多模态生成 :image-text(abstract);分类 cs.CV

AI总结 针对放射学视觉-语言模型的图像-文本对齐缺陷,提出病理学感知对齐框架PALM,结合共享病理学原型与掩码证据建模,在多数据集上提升报告生成与异常鲁棒性

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02218 2026-08-04 cs.AI 新提交 57%

PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs

PosterMELD:面向可编辑打印就绪输出的可控设计多样性多智能体论文转海报生成

Haojie Hu, Chenhao Dang, Yaojia Liu, Hengrui Kang, Conghui He, Weijia Li

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 PosterMELD是模板条件化多智能体海报生成流水线,可导出可编辑PPTX和PNG,在621篇论文实验中PRR达81.3%,成本仅为Codex+Skill的3.5%,性能优于P2P、PosterGen等方法。

Comments 9 pages, 5 figures, and 4 tables. Code and resources are available at https://github.com/Shannon4Science/PosterMELD

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01896 2026-08-04 cs.CV 新提交 57%

GeoCore-9B: Towards Geo-Aware Generative Foundation Models in Earth Observation

GeoCore-9B:面向地球观测的地理感知生成基础模型

Jeonghyeok Do, Munchurl Kim

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

AI总结 本文针对现有EO生成模型的地理空间约束冲突问题,推出90亿参数的GeoCore-9B,采用Flow Matching的DiT与地理空间元数据条件生成,结合地理空间语义对齐损失,在多项EO任务上实现最优性能。

Comments Please visit our project page at https://kaist-viclab.github.io/GeoCore-9B_site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01314 2026-08-04 cs.CV 新提交 57%

Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning

Remember-R1:通过强化学习缓解长上下文视觉遗忘

Jianmin Chen, Jiaqi Tang, Wei Wei, Xiaogang Xu, Jiafei Wu, Zhe Liu, Qianzhou Wang, Yingying Yan, Botong Geng, Yuyang Xia, Lei Zhang, Qifeng Chen

机构 * Northwestern Polytechnical University(西北工业大学) Hong Kong University of Science and Technology(香港科技大学) Zhejiang University(浙江大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 该研究针对多模态大语言模型的长上下文视觉遗忘问题,提出强化学习框架Remember-R1,通过过程级监督优化视觉依赖与注意力,提升了多模态推理性能。

Comments Accepted by ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25498 2026-08-04 cs.SD cs.AI 版本更新 57%

SymphonyGen: 3D Hierarchical Orchestral Generation with Controllable Harmony Skeleton

SymphonyGen:具有可控和声骨架的3D分层交响乐生成

Xuzheng He, Nan Nan, Zhilin Wang, Ziyue Kang, Zhuoru Mo, Ao Li, Yu Pan, Xiaobing Li, Feng Yu, Xiaohong Guan

机构 * Department of AI Music and Music Information Technology, Central Conservatory of Music(中央音乐学院人工智能音乐与音乐信息科技系) Frontier Institute of Science and Technology, and Interdisciplinary Research Center of Frontier Science and Technology, Xi’an Jiaotong University(西安交通大学前沿科学与技术研究院及交叉科学与技术 interdisciplinary research center) University of Science and Technology of China(中国科学技术大学) Shenzhen University(深圳大学)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI

AI总结 SymphonyGen通过3D分层框架实现当代电影交响乐生成,采用分层解码器架构提升效率,引入短谱条件和GRPO优化,增强和声纯净度与音乐表现力。

Comments Accepted at ISMIR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00919 2026-08-04 cs.CV cs.RO 版本更新 57%

DriveCode: Domain Specific Numerical Encoding for LLM-Based Autonomous Driving

DriveCode: 针对基于LLM的自动驾驶的领域特定数值编码

Zhiye Wang, Yanbo Jiang, Rui Zhou, Bo Zhang, Fang Zhang, Zhenhua Xu, Yaqin Zhang, Jianqiang Wang

机构 * School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院) The School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动性学院) The Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) DiDi, Beijing, China(滴滴出行) State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University(清华大学智能绿色车辆与移动性国家重点实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出DriveCode,一种将数字表示为专用嵌入而非离散文本标记的新型数值编码方法,提升LLM在自动驾驶中的数值推理与解码效率。

Comments The project page is available at https://shiftwilliam.github.io/DriveCode

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06423 2026-08-04 cs.CL 版本更新 57%

On the Wings of Imagination: Conflicting Script-based Multi-role Framework for Humor Caption Generation

在想象之翼上:用于幽默标题生成的冲突脚本多角色框架

Wenbo Shang, Yuxi Sun, Jing Ma, Xin Huang

机构 * Hong Kong Baptist University(香港 Baptist 大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL

AI总结 本文提出基于幽默理论GTVH的HOMER框架,通过多角色LLM协作生成幽默标题,实验表明其在多模态幽默标题生成中优于现有方法。

Comments Paper published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00272 2026-08-04 cs.RO 新提交 50%

Localization in Spatiotemporal Fields via Environmental PDEs

基于环境偏微分方程的时空场定位方法

Jose Fuentes, Abdullah Al Redwan Newaz, Ana Cavalcanti, Leonardo Bobadilla

专题命中 多模态生成 :multimodal(abstract)

AI总结 该研究提出基于PDE控制的环境时空场的定位框架,采用Rao-Blackwellized粒子滤波器分解车辆状态,仿真与实验表明其定位精度优于标准粒子滤波器,验证了环境场用于实际定位的可行性。

Comments Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏