An Efficient and Multi-Modal Navigation System with One-Step World Model
一种高效且多模态的导航系统与一步世界模型
机构 * Tsinghua University(清华大学) ; Xiaomi (China)(小米(中国))
专题命中 多模态生成 :multi-modal(title,abstract)
AI总结 本文提出了一种高效多模态导航系统,通过一步世界模型和3D U-Net骨干网络,提升导航效率和鲁棒性。
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
一种高效且多模态的导航系统与一步世界模型
机构 * Tsinghua University(清华大学) ; Xiaomi (China)(小米(中国))
专题命中 多模态生成 :multi-modal(title,abstract)
AI总结 本文提出了一种高效多模态导航系统,通过一步世界模型和3D U-Net骨干网络,提升导航效率和鲁棒性。
探索通用型自主多模态代理用于病理报告生成
专题命中 多模态生成 :multimodal(title,abstract)
AI总结 研究探索通用型自主多模态代理在病理报告生成中的应用,发现其在无额外信息时诊断准确率较低,但提供形态学描述时表现更佳,但仍不及人类专家。
Comments 6 pages, 1 figure, accepted paper for BVM 2026
Journal ref BVM 2026, https://bvm-conf.org
推动辅助机器人:多模态导航与生物物理监测用于下一代轮椅
专题命中 多模态生成 :multi-modal(title,abstract)
AI总结 本文提出了一种多模态轮椅控制系统,结合多种输入接口与生物物理监测,提升患者独立性并实现护理人员实时监督。
多子空间多模态建模用于扩散模型:估计、收敛与专家混合
机构 * Shanghai Jiao Tong University(上海交通大学) ; East China Normal University(华东师范大学)
专题命中 多模态生成 :multi-modal(title,abstract)
AI总结 本文提出基于多子空间多模态建模的扩散模型,通过引入低秩混合高斯混合结构,有效捕捉多模态信息并提升生成性能。
多模态任务感知语义通信的分布式信息瓶颈理论
机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(电子信息与通信学院,华中科技大学) ; Peng Cheng Laboratory(鹏城实验室) ; Pazhou Laboratory (Huangpu)(琶洲实验室(黄埔)) ; School of Mechanical Engineering and Electronic Information, China University of Geosciences(机械工程与电子信息学院,中国地质大学) ; State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学)
专题命中 多模态生成 :multi-modal(title,abstract)
AI总结 本文提出了一种多模态任务感知语义通信的分布式信息瓶颈框架,通过量化模态贡献来优化资源利用,提升通信效率与任务性能。
IGDMRec: 基于行为的物品图扩散用于多模态推荐
专题命中 多模态生成 :multimodal(title,abstract)
AI总结 IGDMRec通过行为条件图扩散和条件去噪网络,利用用户行为信息去噪多模态推荐的语义物品图,提升推荐性能。
Comments 12 pages, 6 figures. This paper has been accepted for publication in IEEE Transactions on Multimedia. The final published version will be available via IEEE Xplore
在稀缺多模态数据下实现无线网络高效领域泛化
专题命中 多模态生成 :multi-modal(title,abstract)
AI总结 本文提出了一种两阶段学习框架,通过基于物理的损失函数和协作域适应方法,在稀缺多模态数据下提升无线网络的领域泛化性能。
Comments Submitted to IEEE TWC
TranSimHub:一个多模态感知与决策的空地协同仿真平台
机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Shanghai AI Laboratory(上海人工智能实验室) ; Stanford University(斯坦福大学) ; Nanyang Technological University(南洋理工大学) ; SenseTime Group Ltd(商汤科技有限公司) ; Beihang University(北京航空航天大学) ; Nokia Bell Labs(诺基亚贝尔实验室)
专题命中 多模态生成 :multi-modal(title,abstract)
AI总结 TranSimHub是一个用于空地协同智能的统一仿真平台,支持多模态感知与决策研究,提供同步渲染和可控场景编辑功能。
Comments 9 pages, 4 figures
MAGE-ID:一种多模态生成框架用于入侵检测系统
机构 * Department of Computer Science, Hunter College and The Graduate Center, City University of New York(计算机科学系,亨特学院和研究生中心,纽约市立大学)
专题命中 多模态生成 :multimodal(title,abstract)
AI总结 MAGE-ID通过多模态生成框架提升入侵检测系统的数据增强效果,实现更平衡和一致的多模态合成,显著提升检测性能。
MolEdit: 多模态分子语言模型的知识编辑
机构 * University of Virginia(弗吉尼亚大学) ; Florida State University(佛罗里达州立大学)
专题命中 多模态生成 :multimodal(title,abstract)
AI总结 MolEdit通过多专家知识适配器和专家意识编辑切换器,提升多模态分子语言模型的编辑可靠性与局部性,实现分子与描述词的高效互转。
专题命中 多模态生成 :multimodal(title,abstract)
机构 * Hong Kong University of Science and Technology(香港理工大学)
专题命中 多模态生成 :multimodal(title,abstract)
Comments NeurIPS 2025
机构 * National University of Singapore(新加坡国立大学) ; University of Maryland, College Park(马里兰大学 College Park 分校) ; University of California, Los Angeles(加州大学洛杉矶分校)
专题命中 多模态生成 :any-to-any(title);multimodal(abstract)
Comments 44 pages, 9 figures, 13 tables, paper accepted by NeurIPS 2025
机构 * Seoul National University(首尔国立大学) ; Hanbat National University(翰baum国立大学) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 多模态生成 :multimodal(title,abstract)
机构 * EECS, MIT(麻省理工学院电子工程与计算机科学系) ; IAIFI, MIT(麻省理工学院天文研究所) ; CfA, Harvard(哈佛大学天文台)
专题命中 多模态生成 :multimodal(title,abstract)
机构 * CREST, ENSAE Institut Polytechnique de Paris(CREST,ENSAE 巴黎高等理工学院)
专题命中 多模态生成 :multi-modal(title,abstract)
专题命中 多模态生成 :multi-modal(title);multimodal(abstract)
Comments 8 pages, 6 figures, accepted in IEEE International Workshop on Computer-Aided Modeling and Design of Communication Links and Networks (CAMAD)
专题命中 多模态生成 :multi-modal(title,abstract)
Comments 31 pages, 9 figures
机构 * École Polytechnique, IP Paris(巴黎理工学院,IP巴黎)
专题命中 多模态生成 :multimodal(title,abstract)
机构 * State Key Laboratory of CAD&CG(计算机辅助设计与图形学国家重点实验室) ; National Key Laboratory for Novel Software Technology(新型软件技术国家实验室)
专题命中 多模态生成 :multimodal(title,abstract)
专题命中 多模态生成 :multi-modal(title);multimodal(abstract)
机构 * Brunel University of London(伦敦布鲁内尔大学) ; SnT, Université du Luxembourg(卢森堡大学SnT分校) ; University of Hull(霍尔姆斯大学)
专题命中 多模态生成 :multi-modal(title);multimodal(abstract)
机构 * Peng Cheng Laboratory(鹏城实验室)
专题命中 多模态生成 :multimodal(title,abstract)
Comments arXiv admin note: text overlap with arXiv:2505.16138
专题命中 多模态生成 :multimodal(title,abstract)
Comments 17 pages; 1 table; 6 figures; extended version of accepted version, published at the 2025 Winter Simulation Conference (WSC '25)
专题命中 多模态生成 :multi-modal(title,abstract)
专题命中 多模态生成 :multi-modal(title,abstract)
专题命中 多模态生成 :multi-modal(title,abstract)
Comments Accepted at IEEE International Conference on Distributed Computing Systems (ICDCS 2025)
机构 * School of Astronautics, Harbin Institute of Technology(哈尔滨工业大学航天学院) ; Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) ; Institute of Systems Engineering and Collaborative Laboratory for Intelligent Science and Systems, Macau University of Science and Technology(澳门科学大学系统工程研究所) ; School of Vehicle and Mobility, Tsinghua University(清华大学车辆与移动系统学院)
专题命中 多模态生成 :multimodal(title,abstract)
机构 * Mila -- Quebec AI Institute(魁北克人工智能研究所) ; McGill University(麦吉尔大学) ; Allen Institute for AI (Ai2)(人工智能研究所) ; University of Idaho(爱达荷大学)
专题命中 多模态生成 :multimodal(title,abstract)
Comments 10 pages, ICML 2025 (TerraBytes)
专题命中 多模态生成 :multi-modal(title,abstract)
Comments 12 pages, 12 figures, 2 tables