arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4959 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4959 篇

2602.02437 2026-02-23 cs.CV cs.AI 62%

UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing

UniReason 1.0: 一个统一的推理框架,用于世界知识对齐的图像生成与编辑

Dianyi Wang, Chaofan Ma, Feng Han, Size Wu, Wei Song, Yibin Wang, Zhixiong Zhang, Tianhang Wang, Siyuan Wang, Zhongyu Wei, Jiaqi Wang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学) University of Southern California(美国南加州大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 UniReason 1.0通过统一推理框架整合图像生成与编辑,利用世界知识增强文本推理和视觉优化,提升复杂合成任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11096 2026-02-12 cs.CL cs.AI 62%

Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away

推理模型中的安全恢复仅需几步早期引导步骤

Soumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Furong Huang, Dinesh Manocha, Amrit Singh Bedi

机构 * University of Maryland, College Park(马里兰大学哥伦比亚学院) IIT, Bombay(孟买印度理工学院) University of Central Florida(佛罗里达中央大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 SafeThink通过在推理早期干预减少攻击成功率,提升推理模型的安全性

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21963 2026-02-12 cs.CY cs.AI cs.CL cs.SI 62%

Industrialized Deception: The Collateral Effects of LLM-Generated Misinformation on Digital Ecosystems

工业化的欺骗:大语言模型生成的虚假信息对数字生态系统的影响

Alexander Loth, Martin Kappes, Marc-Oliver Pahl

机构 * Frankfurt University of Applied Sciences(法兰克福应用科学大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出JudgeGPT和RogueGPT工具,研究人类对AI生成虚假信息的感知与检测,探讨生成与检测之间的竞争及缓解策略。

Comments Accepted at ACM TheWebConf '26 Companion

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07983 2026-02-10 cs.AI cs.CL 62%

Accelerating Social Science Research via Agentic Hypothesization and Experimentation

通过代理假设和实验加速社会科学研究

Jishu Sen Gupta, Harini SI, Somesh Kumar Singh, Syed Mohamad Tawseeq, Yaman Kumar Singla, David Doermann, Rajiv Ratn Shah, Balaji Krishnamurthy

机构 * Adobe Media and Data Science Research (MDSR)(Adobe媒体与数据科学研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 EXPERIGEN通过代理框架实现端到端的科学发现,发现更多显著且预测性强的假设,并通过A/B测试验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08785 2026-02-10 cs.CV cs.AI 62%

DeltaSpace: A Semantic-aligned Feature Space for Flexible Text-guided Image Editing

DeltaSpace: 一种语义对齐的特征空间用于灵活的文本引导图像编辑

Yueming Lyu, Kang Zhao, Bo Peng, Huafeng Chen, Yue Jiang, Yingya Zhang, Jing Dong, Caifeng Shan

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

AI总结 DeltaSpace通过语义对齐的特征空间实现文本引导图像编辑的灵活训练和推理,支持零样本推理和无需文本的训练。

Comments 18 pages. arXiv admin note: text overlap with arXiv:2303.06285

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07176 2026-02-10 cs.CL cs.AI cs.ET cs.HC 62%

Open TutorAI: An Open-source Platform for Personalized and Immersive Learning with Generative AI

Open TutorAI: 一个基于生成AI的开源平台,用于个性化和沉浸式学习

Mohamed El Hajji, Tarek Ait Baha, Aicha Dakir, Hammou Fadili, Youssef Es-Saady

机构 * IRF-SIC Laboratory, Ibnou Zohr University(IRF-SIC实验室,伊本·扎赫尔大学) Regional Center for Education(教育与培训专业地区中心) Polydisciplinary Faculty of Taroudant, Ibnou Zohr University(塔鲁旦多学科学院,伊本·扎赫尔大学) Higher School of Technology of Guelmim, Ibnou Zohr University(盖尔米姆技术高等学校,伊本·扎赫尔大学) Paragraphe laboratory, Paris 8 and CY Cergy Paris Universities(Paragraphe实验室,巴黎8大学和CY塞克巴黎大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 Open TutorAI 是一个基于生成AI的开源平台,通过个性化和沉浸式学习体验提升教育效果。

Comments 19 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02276 2026-02-06 cs.LG cs.AI cs.CL q-bio.QM 62%

CellForge: Agentic Design of Virtual Cell Models

CellForge: 虚拟细胞模型的代理设计

Xiangru Tang, Zhuoyun Yu, Jiapeng Chen, Yan Cui, Daniel Shao, Weixu Wang, Fang Wu, Yuchen Zhuang, Wenqi Shi, Zhi Huang, Arman Cohan, Xihong Lin, Fabian Theis, Smita Krishnaswamy, Mark Gerstein

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 CellForge通过多代理协作自主设计虚拟细胞模型,生成高竞争力的计算方法,推动计算生物学的自主科学方法开发。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02820 2026-02-04 cs.LG cs.AI cs.CV 62%

From Tokens to Numbers: Continuous Number Modeling for SVG Generation

从标记到数字:用于SVG生成的连续数字建模

Michael Ogezi, Martin Bell, Freda Shi, Ethan Smith

机构 * Cheriton School of Computer Science, University of Waterloo, Waterloo, ON, Canada(滑铁卢大学计算机科学学院) Vector Institute(向量研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出连续数字建模(CNM)方法,通过直接建模连续数值提升SVG生成的效率与质量,实现训练速度提升30%及更高的视觉保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01193 2026-02-03 cs.CL cs.CV 62%

Bridging Lexical Ambiguity and Vision: A Mini Review on Visual Word Sense Disambiguation

弥合词汇歧义与视觉:关于视觉词义消歧的简要综述

Shashini Nilukshi, Deshan Sumanathilaka

机构 * School of Computing Informatics(计算与信息学学院) Institute of Technology(技术研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文综述了视觉词义消歧的发展,探讨了对比模型和LLM在解决词汇歧义中的作用,并指出未来发展方向。

Comments 2 figures, 2 Tables, Accepted at IEEE TIC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20911 2026-01-30 cs.CV cs.AI 62%

Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs

非马尔可夫多轮对话图像生成与历史条件化大语言模型

Haochen Zhang, Animesh Sinha, Felix Juefei-Xu, Haoyu Ma, Kunpeng Li, Zhipeng Fan, Meng Dong, Xiaoliang Dai, Tingbo Hou, Peizhao Zhang, Zecheng He

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出非马尔可夫多轮对话图像生成方法,通过历史条件化框架和数据构建策略提升多轮一致性与指令遵循性,同时保持单轮编辑能力。

Comments 19 pages, 19 figures, plan for TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03147 2026-01-30 cs.HC cs.CL cs.LG cs.SD eess.AS 62%

A conversational gesture synthesis system based on emotions and semantics

基于情感和语义的对话手势合成系统

Thanh Hoang-Minh

机构 * Department of Information Technology, VNUHCM -- University of Science(越南科学大学信息科技系)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、eess.AS

AI总结 本文提出DeepGesture,一种基于扩散的 gesture 合成系统,通过多模态信号生成具有情感和语义条件的手势,提升数字人类的自然表达能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19613 2026-01-28 cs.CL cs.AI 62%

Up to 36x Speedup: Mask-based Parallel Inference Paradigm for Key Information Extraction in MLLMs

最高36倍提速:面向MLLMs关键信息提取的基于掩码的并行推断范式

Xinzhong Wang, Ya Guo, Jing Li, Huan Chen, Yi Tu, Yijie Hong, Gongshen Liu, Huijia Zhu

机构 * Shanghai Jiao Tong University(上海交通大学) Ant Info Security Lab, Ant Group(蚂蚁集团信息安全部实验室) Inner Mongolia Research Institute, Shanghai Jiao Tong University(内蒙古研究院,上海交通大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出基于掩码的并行推断范式,通过并行生成提升MLLMs关键信息提取的效率,实现最高36倍的提速。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12529 2026-01-28 cs.HC cs.AI cs.CL 62%

Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing

接受修改后:人工智能辅助科学写作的价值

Sanchaita Hazra, Doeun Lee, Bodhisattwa Prasad Majumder, Sachin Kumar

机构 * The University of Utah(犹他大学) The Ohio State University(俄亥俄州立大学) Allen Institute for AI(人工智能研究院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了AI辅助科学写作的有效性,发现AI生成摘要在披露来源信息后可达到与人工摘要相当的可接受性,且作者编辑行为受对AI作者身份的感知驱动。

Comments Published in ACM IUI 2026 (Paphos, Cyprus)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18234 2026-01-27 cs.CY cs.AI cs.CL 62%

Generative AI in Saudi Arabia: A National Survey of Adoption, Risks, and Public Perceptions

生成式人工智能在沙特阿拉伯:国家调查:采用、风险和公众认知

Abdulaziz AlDakheel, Ali Alshehre, Esraa Alamoudi, Moslim AlKhabbaz, Ahmed Aljohani, Raed Alharbi

机构 * College of Computing and Informatics, Saudi Electronic University(计算机与信息学院,沙特电子大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过全国调查分析沙特阿拉伯GenAI的采用情况、风险和公众认知,发现93%的受访者积极使用GenAI进行文本任务,但整体意识和理解不均衡,需加强AI素养和伦理培训。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17673 2026-01-27 cs.CV cs.AI 62%

Uni-RS: A Spatially Faithful Unified Understanding and Generation Model for Remote Sensing

Uni-RS: 一种用于遥感的具有空间忠实性的统一理解和生成模型

Weiyu Zhang, Yuan Hu, Yong Li, Yu Liu

机构 * Institute of Remote Sensing and Geographic Information System, School of Earth and Space Sciences, Peking University(遥感与地理信息系统研究所,地球与空间科学学院,北京大学) Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology(土木与环境工程系,香港科学与技术大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Uni-RS通过空间布局规划、空间感知查询监督和图像描述空间布局变化,提升遥感文本到图像生成的空间忠实性,同时保持多模态理解任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17096 2026-01-27 cs.CY cs.AI cs.CL 62%

Beyond Instrumental and Substitutive Paradigms: Introducing Machine Culture as an Emergent Phenomenon in Large Language Models

超越工具性和替代性范式:引入机器文化作为大型语言模型中的涌现现象

Yueqing Hu, Xinyang Peng, Yukun Zhao, Lin Qiu, Ka-lai Hung, Kaiping Peng

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出机器文化作为大型语言模型中的一种新兴现象,挑战传统工具性和替代性范式,揭示模型在文化表现上的独特特性。

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17027 2026-01-27 cs.CV cs.AI 62%

Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility

科学图像合成:基准测试、方法论与下游应用

Honglin Lin, Chonghan Qin, Zheng Liu, Qizhi Pei, Yu Li, Zhanping Zhong, Xin Gao, Yanfeng Wang, Conghui He, Lijun Wu

机构 * Shanghai Jiao Tong University(上海交通大学) OpenDataLab, Shanghai Artificial Intelligence Laboratory(OpenDataLab,上海人工智能实验室) The University of Hong Kong(香港大学) Peking University(北京大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出ImgCoder框架和SciGenBench基准,通过逻辑驱动方法提升科学图像生成的结构精度,并展示微调LMMs在科学图像上的效果,验证了高保真合成在多模态推理中的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16007 2026-01-23 cs.CV cs.AI 62%

PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models

PhysicsMind: 为基础多模态大语言模型和世界模型中的物理推理和预测进行仿真与现实力学基准测试

Chak-Wing Mak, Guanyu Zhu, Boyi Zhang, Hongji Li, Xiaowei Chi, Kevin Zhang, Yichen Wu, Yangfan He, Chun-Kai Fan, Wentao Lu, Kuangzhi Ge, Xinyu Fang, Hongyang He, Kuan Lu, Tianxiang Xu, Li Zhang, Yongxin Ni, Youhua Li, Shanghang Zhang

机构 * Peking University(北京大学) Mohamed bin Zayed University of Artificial Intelligence(莫扎伊德大学人工智能学院) National University of Singapore(新加坡国立大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Science and Technology of China(中国科学技术大学) Cornell University(康奈尔大学) Hong Kong Polytechnic University(香港理工大学) City University of Hong Kong(香港城市大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 PhysicsMind是一个结合现实和仿真环境的统一基准,用于评估基础多模态大语言模型和世界模型在物理推理和预测中的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15664 2026-01-23 cs.CV cs.AI 62%

Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling

Skywork UniPic 3.0:通过序列建模实现统一的多图像合成

Hongyang Wei, Hongbo Liu, Zidong Wang, Yi Peng, Baixin Xu, Size Wu, Xuying Zhang, Xianglong He, Zexiang Liu, Peiyu Wang, Xuchen Song, Yangguang Li, Yang Liu, Yahui Zhou

机构 * Skywork

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Skywork UniPic 3.0通过序列建模实现统一的多图像合成,采用新颖的训练范式和高效的数据流程,在单图像编辑和多图像合成任务中均取得优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05665 2026-01-22 cs.CL cs.CV 62%

Interleaved Latent Visual Reasoning with Selective Perceptual Modeling

交错的潜在视觉推理与选择性感知建模

Shuai Dong, Siyuan Wang, Xingyu Liu, Chenglin Li, Haowen Hou, Zhongyu Wei

机构 * China University of Geosciences, Wuhan(中国地质大学(武汉)) Shanghai Innovation Institute(上海创新研究院) University of Southern California(南加州大学) Fudan University(复旦大学) Zhejiang University(浙江大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 ILVR通过交错潜在视觉表示与文本生成,实现动态状态演变与精确感知建模的统一,提升多模态推理性能。

Comments 18 pages, 11 figures. Code available at https://github.com/XD111ds/ILVR

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14041 2026-01-21 cs.CL cs.AI 62%

Top 10 Open Challenges Steering the Future of Diffusion Language Model and Its Variants

推动扩散语言模型及其变体未来发展的十大开放挑战

Yunhe Wang, Kai Han, Huiling Zhen, Yuchuan Tian, Hanting Chen, Yongbing Huang, Yufei Cui, Yingte Shu, Shan Gao, Ismail Elezi, Roy Vaughan Miles, Songcen Xu, Feng Wen, Chao Xu, Sinan Zeng, Dacheng Tao

机构 * Huawei Noah’s Ark Lab(华为诺亚实验室) Peking University(北京大学) Huawei Technologies(华为技术有限公司) Nanyang Technological University(南洋理工大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出推动扩散语言模型及其变体发展的十大挑战,并提出以多尺度标记化、主动遮罩和潜在思考为核心的四支柱战略路线图,旨在突破因果瓶颈,实现复杂结构推理和多模态整合。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08623 2026-01-14 cs.CV cs.AI cs.CR cs.LG 62%

SafeRedir: Prompt Embedding Redirection for Robust Unlearning in Image Generation Models

SafeRedir: 基于提示嵌入重定向的图像生成模型鲁棒去学习

Renyang Liu, Kangjie Chen, Han Qiu, Jie Zhang, Kwok-Yan Lam, Tianwei Zhang, See-Kiong Ng

机构 * National University of Singapore(国立新加坡大学) Nanyang Technological University(南洋理工大学) Tsinghua University(清华大学) CFAR and IHPC, A*STAR(CFAR和IHPC,A*STAR)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 SafeRedir通过提示嵌入重定向实现图像生成模型的鲁棒去学习,无需修改模型即可有效消除不安全内容并提升对抗性攻击抵抗力。

Comments Code at https://github.com/ryliu68/SafeRedir

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17068 2026-01-13 cs.CV cs.AI 62%

ReBrain: Brain MRI Reconstruction from Sparse CT Slice via Retrieval-Augmented Diffusion

ReBrain: 通过检索增强扩散模型从稀疏CT切片重建脑部MRI

Junming Liu, Yifei Sun, Weihua Cheng, Yujin Kang, Yirong Chen, Ding Wang, Guosun Zeng

机构 * Tongji University(同济大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 ReBrain通过检索增强扩散模型从稀疏CT切片重建脑部MRI,利用BBDM和ControlNet实现结构连续性,提升稀疏条件下的跨模态重建性能。

Comments 16 pages, 12 figures, 7 tables; Accepted by WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05254 2026-01-12 cs.CV cs.AI cs.LG 62%

Explaining Low Perception Model Competency with High-Competency Counterfactuals

用高能力反事实解释低感知模型能力

Sara Pohland, Claire Tomlin

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出五种生成高能力反事实图像的方法,用于解释模型预测不确定性的原因,并通过实验验证了其在生成语言解释中的有效性。

Journal ref Explainable Artificial Intelligence. xAI 2025. Communications in Computer and Information Science, vol 2580

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10963 2026-01-07 cs.CV cs.CL 62%

MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning

MMMG:大规模、跨学科、多层级的文本到图像推理生成基准

Yuxuan Luo, Yuhui Yuan, Junwen Chen, Haonan Cai, Ziyi Yue, Yuwei Yang, Fatima Zohra Daha, Ji Li, Zhouhui Lian

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出MMMG基准,用于评估文本到图像生成模型的推理能力,揭示现有模型在知识图像生成中的不足,并发布FLUX-Reason作为开放基线。

Comments 85 pages, 70 figures, code: https://github.com/MMMGBench/MMMG, project page: https://mmmgbench.github.io/

Journal ref Advances in Neural Information Processing Systems 38 (NeurIPS 2025) Datasets and Benchmarks Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17103 2026-01-07 cs.CL cs.AI 62%

Forging Time Series with Language: A Large Language Model Approach to Synthetic Data Generation

用语言锻造时间序列:一种基于大语言模型的合成数据生成方法

Cécile Rousseau, Tobia Boschi, Giandomenico Cornacchia, Dhaval Salwala, Alessandra Pascale, Juan Bernabe Moreno

机构 * IBM Research Europe(IBM欧洲研究院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 SDForger通过大语言模型生成高质量多变量时间序列,利用文本条件化实现高效合成数据生成及多模态建模。

Journal ref NeurIPS 2025, https://openreview.net/forum?id=A2pmvkqOgp

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23185 2025-12-30 eess.IV cs.AI cs.CV 62%

EIR: Enhanced Image Representations for Medical Report Generation

EIR: 增强的医学报告生成图像表示

Qiang Sun, Zongcheng Ji, Yinlong Xiao, Peng Chang, Jun Yu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 EIR通过跨模态Transformer融合元数据与图像表示,结合医学领域预训练模型,有效解决信息不对称和领域差距问题,提升胸部X光报告生成的准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22288 2025-12-30 cs.LG cs.AI cs.CV 62%

Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model

Co-GRPO:针对掩码扩散模型的协同优化组相对策略优化

Renping Zhou, Zanlin Ni, Tianyi Chen, Zeyu Liu, Yang Yue, Yulin Wang, Yuxuan Wang, Jingshu Liu, Gao Huang

机构 * Leap Lab, Tsinghua University(清华大学 leap 实验室) Anyverse Dynamics

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 Co-GRPO通过协同优化模型和调度参数提升掩码扩散模型的生成质量。

Comments 17 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22237 2025-12-30 cs.CV cs.AI 62%

Meta-information Guided Cross-domain Synergistic Diffusion Model for Low-dose PET Reconstruction

元信息引导的跨领域协同扩散模型用于低剂量PET重建

Mengxiao Geng, Ran Hong, Xiaoling Xu, Bingxuan Li, Qiegen Liu

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出元信息引导的跨领域协同扩散模型,通过整合跨模态先验知识提升低剂量PET重建质量与生理细节保留能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19609 2025-12-23 cs.CV cs.AI 62%

MapTrace: Scalable Data Generation for Route Tracing on Maps

MapTrace: 用于地图路线追踪的可扩展数据生成

Artemis Panagopoulou, Aveek Purohit, Achin Kulshrestha, Soroosh Yazdani, Mohit Goyal

机构 * Google XR(谷歌XR) University of Pennsylvania(宾夕法尼亚大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 MapTrace通过合成数据生成提升地图路线追踪性能,微调模型使成功率提升6.4个百分点,减少路径追踪误差。

详情

展开后加载摘要…

URL PDF HTML 收藏