From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI
专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments This paper has been accepted by ICML2024
用于LEED合规的神经符号人工智能:以文档为中心的基准测试、确定性数值检查以及多模态何时有害
机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)
专题命中 其他多模态 :multimodal(title);分类 cs.AI
AI总结 研究小型本地部署语言模型对LEED文档筛选及符号组件作用,引入神经符号管道,通过对齐PDF、检索证据、语言模型验证和数值检查器检查来验证合规性,实验表明特定模型在任务中表现较好,管道及基线提供了参考。
DeepInterestGR: 利用多模态大语言模型挖掘深度多兴趣用于生成式推荐
机构 * Southeast University(东南大学)
专题命中 其他多模态 :multi-modal(title);分类 cs.CV
AI总结 提出DeepInterestGR框架,通过多LLM兴趣挖掘、奖励标记深度兴趣和兴趣增强物品离散化,解决生成式推荐中的浅层兴趣问题,在三个Amazon数据集上显著提升推荐性能。
用生成式多模态模型模拟临床干预
机构 * Department of Computer Science and Applied Mathematics, Weizmann Institute of Science(魏茨曼科学研究所计算机科学与应用数学系) ; Department of Molecular Cell Biology, Weizmann Institute of Science(魏茨曼科学研究所分子细胞生物学系) ; NVIDIA ; Novo Nordisk Foundation Center for Basic Metabolic Research, University of Copenhagen(诺沃维克基金会基础代谢研究中心,哥本哈根大学) ; Faculty of Medical and Health Sciences, Tel Aviv University(特拉维夫大学医学与健康科学学院) ; The Jesse Z and Sara Lea Shafer Institute for Endocrinology and Diabetes, National Center for Childhood Diabetes, Schneider Children’s Medical Center of Israel(杰西Z和索菲亚·李·沙弗内分泌学与糖尿病研究所,以色列儿童糖尿病国家中心,施耐德儿童医学中心) ; Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
专题命中 其他多模态 :multimodal(title);分类 cs.AI
AI总结 本文提出HealthFormer模型,通过训练人类表型项目数据,生成人类生理轨迹,实现对个体生理变化的预测和干预模拟,提升临床风险评分和疾病预测能力。
由多模态大语言模型视觉先验引导的3D烟雾场景重建
机构 * Hefei University of Technology(合肥工业大学) ; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院) ; Anhui University(安徽大学) ; United Arab Emirates University(阿联酋大学) ; Nanyang Technological University(南洋理工大学)
专题命中 其他多模态 :multimodal(title);分类 cs.CV
AI总结 本文提出整合视觉先验与高效3D场景建模的框架,通过增强烟雾退化图像和开发Smoke-GS框架,提升烟雾场景重建与视图合成的鲁棒性与一致性。
多模态自动驾驶的混合渲染:融合神经网络与基于物理的模拟
机构 * aiMotive
专题命中 其他多模态 :multimodal(title);分类 cs.CV
AI总结 本文提出混合渲染方法,结合神经网络与基于物理的模拟,提升自动驾驶模拟中新颖视角合成质量及实时渲染效率。
通过多模态多实例学习对自身免疫性疾病进行外周血T细胞受体谱系的分类
机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China(清华大学深圳国际研究生院,清华大学,深圳,中国)
专题命中 其他多模态 :multimodal(title);分类 cs.AI
AI总结 EAMil通过多模态多实例学习方法,利用TCR测序数据高精度诊断SLE和RA,并有效识别疾病相关基因。
Comments 4 figures, 3 tabels, 8 pages
改进多模态蒸馏以应对激光雷达语义分割中的域移位
机构 * CNRS, IRISA, Univ. Bretagne Sud(CNRS、IRISA、布列塔尼大学) ; LIGM, Ecole des Ponts, Univ Gustave Eiffel, CNRS(LIGM、巴黎理工学院、古斯塔夫·埃菲尔大学、CNRS)
专题命中 其他多模态 :multimodal(title);分类 cs.CV
AI总结 本研究提出改进多模态蒸馏方法,通过冻结预训练主干网络并训练MLP头,提升激光雷达语义分割在域移位下的性能。
Comments Accepted at BMVC 2025
专题命中 其他多模态 :multi-modal(title);分类 cs.AI
机构 * Shanghai Jiao Tong University, Shanghai, China(上海交通大学) ; Xi'an Jiaotong-Liverpool University, Suzhou, China(西安交通大学-利物浦大学)
专题命中 其他多模态 :multi-modal(title);分类 cs.CL
机构 * Deakin University(德金大学) ; Monash University(莫纳什大学)
专题命中 其他多模态 :multimodal(title);分类 cs.AI
Comments 15 pages, 1 figure, Accepted to MAI-XAI@ECAI2025
机构 * University of California San Diego(加州大学圣迭戈分校) ; Georgia Institute of Technology(佐治亚理工学院)
专题命中 其他多模态 :multimodal(title);分类 cs.AI
Comments To appear in IMWUT'25. Code is available at: https://github.com/Orienfish/SensorChat
机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) ; University of Southern California(南加州大学)
专题命中 其他多模态 :multimodal(title);分类 cs.CV
机构 * University of Science and Technology of China(中国科学技术大学) ; Suzhou Institute for Advanced Research, USTC(先进研究所) ; Hunan University(湖南大学) ; Monash University(墨尔本大学) ; State Key Laboratory of Precision and Intelligent Chemistry, USTC(精密与智能化学国家重点实验室)
专题命中 其他多模态 :multimodal(title);分类 cs.AI
专题命中 其他多模态 :multimodal(title);分类 cs.AI
Comments 24 pages
机构 * Department of Computer Science and Operational Research(计算机科学与运筹学系) ; Université de Montréal - Mila(蒙特利尔大学 - Mila)
专题命中 其他多模态 :multi-modal(title);分类 cs.AI
专题命中 其他多模态 :multi-modal(title);分类 cs.CV
专题命中 其他多模态 :multi-modal(title);分类 cs.CV
专题命中 其他多模态 :multimodal(title);分类 cs.CL
Comments Accepted in Proceedings of the 27th International Conference on. Human-Computer Interaction, 2025
专题命中 其他多模态 :multi-modal(title);分类 cs.CV
Comments NeurIPS 2024
专题命中 其他多模态 :multi-modal(title);分类 cs.CV
Comments Accepted in Proceedings of the 3rd International Conference on Computing Advancements, 2024
专题命中 其他多模态 :multimodal(title);分类 cs.AI
Comments Paper accepted to the 2025 CHI Conference on Human Factors in Computing Systems (CHI 2025)
专题命中 其他多模态 :multi-modal(title);分类 cs.CV
专题命中 其他多模态 :multimodal(title);分类 cs.CV
Comments 11th Italian Workshop on Artificial Intelligence and Robotics (AIRO 2024), Published in CEUR Workshop Proceedings AI*IA Series
专题命中 其他多模态 :multi-modal(title);分类 cs.AI
专题命中 其他多模态 :multimodal(title);分类 cs.CV
专题命中 其他多模态 :multimodal(title);分类 cs.AI
Comments Siggraph Asia 2024 Art Paper
专题命中 其他多模态 :multimodal(title);分类 cs.CL
专题命中 其他多模态 :multi-modal(title);分类 cs.CV