arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4868 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4868 篇

2603.01106 2026-03-03 cs.AI 79%

DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage

DIVA-GRPO:通过难度自适应变体优势增强多模态推理

Haowen Gao, Zhenyu Zhang, Liang Pang, Fangda Guo, Hongjian Dou, Guannan Lv, Shaoguo Liu, Tingting Gao, Huawei Shen, Xueqi Cheng

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, CAS, Beijing, China(人工智能安全国家重点实验室,计算技术研究所,中国科学院,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国) Kuaishou Technology, Beijing, China(快手科技,北京,中国)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 DIVA-GRPO通过难度自适应变体优势方法提升多模态推理能力,解决GRPO在困难问题上的奖励稀疏性和优势消失问题,提升训练稳定性与推理性能。

Comments Accepted to ICLR 2026. Code and models are available at https://github.com/Siaaaaaa1/DIVA-GRPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00289 2026-03-03 cs.CV 79%

Seeking Necessary and Sufficient Information from Multimodal Medical Data

从多模态医学数据中寻求必要和充分的信息

Boyu Chen, Weiye Bao, Junjie Liu, Michael Shen, Bo Peng, Paul Taylor, Zhu Li, Mengyue Yang

机构 * University College London, London, UK(伦敦大学学院) Imperial College London, London, UK(伦敦帝国学院) Mingdu Tech, China(明都科技) University of Bristol, Bristol, UK(布里斯托大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出通过概率必要性和充分性学习多模态医学数据中的必要和充分特征,以提升模型性能和鲁棒性。

Comments 11 pages, 1 figure. Submitted to MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.27492 2026-03-03 cs.CV 79%

ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning

ThinkMorph:多模态交错链式推理中的涌现特性

Jiawei Gu, Yunzhuo Hao, Huichen Will Wang, Linjie Li, Michael Qizhe Shieh, Yejin Choi, Ranjay Krishna, Yu Cheng

机构 * National University of Singapore(新加坡国立大学) Zhejiang University(浙江大学) University of Washington(华盛顿大学) Stanford University(斯坦福大学) absolute AI The Chinese University of Hong Kong(香港中文大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 ThinkMorph通过统一模型提升多模态推理性能,展现视觉操控与模式切换等新兴能力。

Comments project page: https://thinkmorph.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02339 2026-03-03 cs.CL 79%

AStar: Boosting Multimodal Reasoning with Automated Structured Thinking

AStar: 通过自动化结构化思维提升多模态推理

Jinyang Wu, Mingkuan Feng, Guocheng Zhai, Shuai Zhang, Zheng Lian, Fangrui Lv, Pengpeng Shao, Ruihan Jin, Zhengqi Wen, Jianhua Tao

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 AStar通过自动化结构化思维提升多模态推理效率,实现更高准确率和更强迁移能力。

Comments Accepted by AAAI 2026 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21743 2026-02-27 cs.CV 79%

Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization

通过难度感知分组归一化增强多模态大语言模型推理

Jinghan Li, Junfeng Fang, Jinda Lu, Yuan Wang, Xiaoyan Guo, Tianyu Zhang, Xiang Wang, Xiangnan He

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学)

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);分类 cs.CV

AI总结 本文提出难度感知分组归一化方法,通过感知复杂度和推理不确定性表征样本,提升多模态大语言模型的推理稳定性与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23597 2026-02-19 cs.CV cs.IR 79%

Scalable Residual Feature Aggregation Framework with Hybrid Metaheuristic Optimization for Robust Early Pancreatic Neoplasm Detection in Multimodal CT Imaging

可扩展的残差特征聚合框架与混合元启发式优化用于多模CT影像中胰腺肿瘤的稳健早期检测

Janani Annur Thiruvengadam, Kiran Mayee Nabigaru, Anusha Kovi

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种结合残差特征聚合与混合元启发式优化的框架,用于多模CT影像中胰腺肿瘤的稳健早期检测,实现了高准确率和泛化能力。

Comments Accepted at 11th International Conference on Big Data Analytics (ICBDA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.15769 2026-02-18 cs.CL 79%

ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attribution

ViTaB-A:在视觉表格归因上评估多模态大语言模型

Yahia Alqurnawi, Preetom Biswas, Anmol Rao, Tejas Anvekar, Chitta Baral, Vivek Gupta

机构 * School of Computing and Augmented Intelligence(计算与增强智能学院)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 ViTaB-A研究了多模态大语言模型在视觉表格归因中的表现,发现其在证据归因方面存在显著缺陷,影响透明性和可追溯性应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00287 2026-02-18 cs.AI cs.CY 79%

SIGMUS: Semantic Integration for Knowledge Graphs in Multimodal Urban Spaces

SIGMUS:多模态城市空间中知识图谱的语义整合

Brian Wang, Mani Srivastava

机构 * University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 SIGMUS通过大型语言模型实现多模态城市数据的自动语义整合,构建知识图谱以连接不同数据源与事件关系。

Comments 9 pages, accepted at UrbComp 2025 KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10143 2026-02-12 cs.CV 79%

MPA: Multimodal Prototype Augmentation for Few-Shot Learning

MPA: 多模态原型增强用于少样本学习

Liwen Wu, Wei Wang, Lei Zhao, Zhan Gao, Qika Lin, Shaowen Yao, Zuozhu Liu, Bin Pu

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 MPA通过多模态原型增强方法,在少样本学习中实现更优性能,通过语义增强、多视图增强和不确定类吸收器提升模型表现。

Comments This paper has been accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.00437 2026-02-12 cs.CV 79%

ADGaussian: Generalizable Gaussian Splatting for Autonomous Driving via Multi-modal Joint Learning

ADGaussian:通过多模态联合学习实现自动驾驶的通用高斯点云重建

Qi Song, Chenghong Li, Haotong Lin, Sida Peng, Rui Huang

机构 * School of Science and Engineering, The Chinese University of Hong Kong (Shenzhen)(中国香港中文大学(深圳)科学与工程学院) Zhejiang University(浙江大学)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 ADGaussian通过多模态联合学习实现自动驾驶场景的高斯点云重建,提升零样本泛化能力。

Comments The paper is accepted by ICRA 2026 and the project page can be found at https://maggiesong7.github.io/research/ADGaussian/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08524 2026-02-10 cs.CV 79%

GeoFocus: Blending Efficient Global-to-Local Perception for Multimodal Geometry Problem-Solving

GeoFocus:融合高效全局到局部感知的多模态几何问题求解

Linger Deng, Yuliang Liu, Wenwen Yu, Zujia Zhang, Jianzhong Ju, Zhenbo Luo, Xiang Bai

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) MiLM Plus, Xiaomi Inc(小米公司)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 GeoFocus通过关键局部感知器和顶点语言提升多模态几何问题求解的准确性和效率

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07207 2026-02-10 cs.IR cs.AI 79%

Multimodal Enhancement of Sequential Recommendation

多模态序列推荐增强

Bucher Sahyouni, Matthew Vowels, Liqun Chen, Simon Hadfield

机构 * University of Surrey(萨里大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 MuSTRec通过结合多模态和序列推荐范式,利用物品-物品图和频率自注意力模块提升推荐性能,实验表明其在多个数据集上取得显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07024 2026-02-10 cs.RO cs.CV 79%

A Distributed Multi-Modal Sensing Approach for Human Activity Recognition in Real-Time Human-Robot Collaboration

一种用于实时人机协作中人类活动识别的分布式多模态传感方法

Valerio Belcamino, Nhat Minh Dinh Le, Quan Khanh Luu, Alessandro Carfì, Van Anh Ho, Fulvio Mastrogiovanni

机构 * University of Genoa(热那亚大学) The University of Danang–University of Science and Technology(丹江-科学技术大学) Japan Advanced Institute of Science and Technology(日本先进科学研究院)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 本文提出一种结合数据手套和视觉触觉传感器的多模态方法,用于实时人机协作中的人类活动识别,实验显示该方法在不同场景下均表现出高精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06197 2026-02-09 cs.HC cs.AI 79%

Personagram: Bridging Personas and Product Design for Creative Ideation with Multimodal LLMs

Personagram: 通过多模态大语言模型连接人设与产品设计以促进创意构思

Taewook Kim, Matthew K. Hong, Yan-Ying Chen, Jonathan Q. Li, Monica P Van, Shabnam Hakimi, Matthew Kay, Matthew Klenk

机构 * Toyota Research Institute(丰田研究院) Northwestern University(西北大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 Personagram利用多模态大语言模型帮助设计师探索基于人口普查的人设,提取并重新组合产品特征,提升创意构思的可行性和参与度。

Comments 22 pages, 10 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02736 2026-02-04 cs.CL 79%

Time-Critical Multimodal Medical Transportation: Organs, Patients, and Medical Supplies

时间敏感的多模式医疗运输:器官、患者和医疗物资

Elaheh Sabziyan Varnousfaderani, Syed A. M. Shihab, Mohammad Taghizadeh

机构 * College of Aeronautics and Engineering(航空航天工程学院) Kent State University(肯特州立大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本研究提出了一种多模式医疗运输调度算法,通过整合地面和空中车辆,优化运输效率并降低运营成本。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01884 2026-02-03 cs.AI cs.LG 79%

Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models

熵引导的数据高效训练用于多模态推理奖励模型

Shidong Yang, Tongwen Huang, Hao Wen, Yong Wang, Li Chen, Xiangxiang Chu

机构 * School of Software, Tsinghua University(清华大学软件学院) AMAP, Alibaba Group(阿里巴巴集团AMAP)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出熵引导训练方法,通过熵指导数据筛选和训练策略提升多模态推理奖励模型的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00590 2026-02-03 cond-mat.mtrl-sci cond-mat.soft cs.AI cs.LG physics.data-an 79%

Multimodal Machine Learning for Integrating Heterogeneous Analytical Systems

多模态机器学习用于整合异质分析系统

Shun Muroga, Hideaki Nakajima, Taiyo Shimizu, Kazufumi Kobashi, Kenji Hata

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出多模态机器学习框架,通过整合多种分析数据,提升复杂材料的表征精度与解释性。

Comments 12 pages, 4 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00067 2026-02-03 cs.LG cs.AI 79%

Modality as Heterogeneity: Node Splitting and Graph Rewiring for Multimodal Graph Learning

模态作为异质性:用于多模态图学习的节点分裂与图重 wiring

Yihan Zhang, Ercan E. Kuruoglu

机构 * Institute of Data and Information, Shenzhen International Graduate School, Tsinghua University(数据与信息研究所,深圳国际研究生院,清华大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 NSG-MoE通过节点分裂与图重 wiring机制,结合结构化MoE架构,有效解决多模态图学习中的模态混淆问题,提升模型的结构信息保留与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13292 2026-01-14 q-bio.QM cs.AI eess.IV 79%

An interpretable generative multimodal neuroimaging-genomics framework for decoding Alzheimer's disease

可解释的生成多模态神经影像-基因组框架用于解码阿尔茨海默病

Giorgio Dolci, Federica Cruciani, Md Abdur Rahaman, Anees Abrol, Jiayu Chen, Zening Fu, Ilaria Boscolo Galazzo, Gloria Menegaz, Vince D. Calhoun

机构 * Department of Computer Science, University of Verona(威尼斯大学计算机科学系) Department of Engineering for Innovation Medicine, University of Verona(威尼斯大学创新医学工程系) Tri-Institutional Center for Translational Research in Neuroimaging and Data Science (TReNDS), Georgia State University, Georgia Institute of Technology, Emory University(神经影像与数据科学转化研究三机构中心(TReNDS),佐治亚州立大学,佐治亚理工学院,埃默里大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出了一种可解释的生成多模态神经影像-基因组框架,用于解码阿尔茨海默病,通过多模态数据和单核苷酸多态性实现AD检测和MCI预测,并揭示了与疾病相关的生物学机制。

Comments 33 pages, 8 figures (main text + supplementary materials), submitted to a journal

Journal ref J. Neural Eng. 22 056021 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03579 2026-01-08 cs.CV 79%

SpatiaLoc: Leveraging Multi-Level Spatial Enhanced Descriptors for Cross-Modal Localization

SpatiaLoc: 借助多级空间增强描述符进行跨模态定位

Tianyi Shang, Pengjie Xu, Zhaojun Deng, Zhenyu Li, Zhicong Chen, Lijun Wu

机构 * Fuzhou University(福州大学) Shandong Academy of Sciences(山东省科学院) Qingdao University(青岛大学) Tongji University(同济大学)

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 SpatiaLoc通过多级空间增强描述符提升跨模态定位性能,采用粗到细策略结合贝塞尔曲线和频率域建模,实现更精确的机器人定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14435 2026-01-08 cs.CV cs.LG 79%

MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models

MoTE:混合三元专家用于内存高效的大型多模态模型

Hongyu Wang, Jiayu Xu, Ruiping Wang, Yan Feng, Yitao Zhai, Peng Pei, Xunliang Cai, Xilin Chen

机构 * Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全重点实验室,计算技术研究所) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 MoTE通过训练更多低精度三元专家,实现内存高效的大规模多模态模型训练,提升端任务性能并降低内存需求。

Comments Work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.18376 2026-01-07 cs.LG cs.CV 79%

Empowering Source-Free Domain Adaptation via MLLM-Guided Reliability-Based Curriculum Learning

通过MLLM引导的可靠性基于课程学习增强源无关领域适应

Dongjie Chen, Kartik Patwari, Zhengfeng Lai, Xiaoguang Zhu, Sen-ching Cheung, Chen-Nee Chuah

机构 * University of California, Davis(加州大学戴维斯分校) University of Kentucky(肯塔基大学)

专题命中 其他多模态 :MLLM(title);multimodal(abstract);分类 cs.CV

AI总结 通过MLLM引导的可靠性基于课程学习增强源无关领域适应,提出一种新的框架,利用多个冻结的MLLMs的稳健监督蒸馏到目标模型,实现稳定且噪声感知的训练。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22603 2025-12-30 cs.CL 79%

Structured Prompting and LLM Ensembling for Multimodal Conversational Aspect-based Sentiment Analysis

结构化提示与大语言模型集成用于多模态对话基于方面的情感分析

Zhiqiang Gao, Shihao Gao, Zixing Zhang, Yihao Guo, Hongyu Chen, Jing Han

机构 * Hunan University(湖南大学) University of Cambridge(剑桥大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CL

AI总结 本文提出结构化提示与大语言模型集成方法,用于多模态对话基于方面的情感分析,有效提升情感识别与翻转检测的准确性。

Journal ref ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11467 2025-12-25 cs.CV eess.IV 79%

A Multicore and Edge TPU-Accelerated Multimodal TinyML System for Livestock Behavior Recognition

一种多核和边缘TPU加速的多模态TinyML系统用于牲畜行为识别

Qianxue Zhang, Eiman Kanjo

机构 * Medical AI Lab, Hebei Provincial Engineering Research Center for AI-Based Cancer Treatment Decision-Making, The First Hospital of Hebei Medical University(医学人工智能实验室,河北省人工智能辅助癌症治疗决策工程研究中心,河北省医科大学第一医院) Computing Department, Imperial College London(计算部门,帝国理工学院伦敦分校) Professor Pervasive Sensing & TinyML and the Head of the Smart Sensing Lab at Nottingham Trent University(感知与TinyML教授及智能感知实验室主任,诺丁汉特伦特大学) Provost’s Visiting Professor in tinyML at Imperial College London(帝国理工学院伦敦分校副校长兼任TinyML客座教授)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种多核和边缘TPU加速的多模态TinyML系统,用于高效识别牲畜行为,实现高模型压缩和低延迟的实时推理。

Comments 12 pages, 10 figures

Journal ref IEEE Internet of Things Journal, vol. 13, no. 1, pp. 666-677, 1 Jan.1, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18279 2025-12-24 cs.CV 79%

UniMPR: A Unified Framework for Multimodal Place Recognition with Heterogeneous Sensor Configurations

UniMPR: 一种用于异构传感器配置的多模态地点识别统一框架

Zhangshuo Qi, Jingyi Xu, Luqi Cheng, Shichen Wen, Yiming Ma, Guangming Xiong

机构 * Beijing Institute of Technology(北京理工大学) Shanghai Jiao Tong University(上海交通大学) The University of New South Wales(新南威尔士大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 UniMPR提出了一种统一框架,能够适应多种异构传感器配置,通过极坐标BEV特征空间和多分支网络实现多模态地点识别的高效与鲁棒性。

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03695 2025-12-23 cs.LG cs.AI 79%

Hierarchical Federated Foundation Models over Wireless Networks for Multi-Modal Multi-Task Intelligence: Integration of Edge Learning with D2D/P2P-Enabled Fog Learning Architectures

无线网络上的分层联邦基础模型用于多模态多任务智能:整合边缘学习与D2D/P2P-enabled雾学习架构

Payam Abdisarabshali, Fardis Nadimi, Kasra Borazjani, Naji Khosravan, Minghui Liwang, Wei Ni, Dusit Niyato, Michael Langberg, Seyyedali Hosseinalipour

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.AI

AI总结 本文提出分层联邦基础模型,整合边缘学习与D2D/P2P-enabled雾学习架构,解决多模态多任务智能中的异构性问题。

Comments 7 pages, 2 figures, 1 table

Journal ref IEEE Communications Magazine, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12461 2025-12-16 cs.LG cs.AI q-bio.NC 79%

Cross-Modal Representational Knowledge Distillation for Enhanced Spike-Informed LFP Modeling

跨模态表征知识蒸馏用于增强基于尖峰的LFP建模

Eray Erturk, Saba Hashemi, Maryam M. Shanechi

机构 * Ming Hsieh Department of Electrical and Computer Engineering(明希斯电气与计算机工程系) Thomas Lord Department of Computer Science(托马斯·劳德计算机科学系) Alfred E. Mann Department of Biomedical Engineering(阿尔弗雷德·E·曼生物医学工程系) Neuroscience Graduate Program University of Southern California(神经科学研究生项目美国南加州大学)

专题命中 其他多模态 :cross-modal(title,abstract);分类 cs.AI

AI总结 本文提出跨模态知识蒸馏框架,通过将预训练的尖峰模型知识转移至LFP模型,提升LFP建模的准确性和泛化能力。

Comments Published at the 39th Annual Conference on Neural Information Processing Systems 2025. Code is available at https://github.com/ShanechiLab/CrossModalDistillation

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.11067 2025-12-15 cs.DB cs.AI 79%

KathDB: Explainable Multimodal Database Management System with Human-AI Collaboration

KathDB:具有人机协作的可解释多模数据库管理系统

Guorui Xiao, Enhao Zhang, Nicole Sullivan, Will Hansen, Magdalena Balazinska

机构 * University of Washington(华盛顿大学)

专题命中 其他多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 KathDB是一种结合关系语义和基础模型推理能力的多模数据库管理系统,通过人机协作实现查询解析、执行和结果解释的可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05258 2025-12-08 cs.CV cs.LG cs.RO 79%

Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving

多模态数据高效3D场景理解用于自动驾驶

Lingdong Kong, Xiang Xu, Jiawei Ren, Wenwei Zhang, Liang Pan, Kai Chen, Wei Tsang Ooi, Ziwei Liu

机构 * WorldBench Team Project Lead(WorldBench团队项目负责人)

专题命中 其他多模态 :multi-modal(title,abstract);分类 cs.CV

AI总结 LaserMix++通过多模态方法提升自动驾驶中LiDAR数据高效3D场景理解,以更少标注实现更高精度。

Comments IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10044 2025-12-02 cs.CR cs.AI 79%

Large Language Models for Power System Security: A Novel Multi-Modal Approach for Anomaly Detection in Energy Management Systems

用于电力系统安全的大型语言模型:一种用于能源管理系统异常检测的新型多模态方法

Aydin Zaboli, Junho Hong, Alexandru Stefanov, Chen-Ching Liu, Chul-Sang Hwang

机构 * Department of Electrical and Computer Engineering, University of Michigan -- Dearborn, MI, 48128 USA.(电气与计算机工程系,密歇根大学迪尔伯恩分校) Department of Electrical Sustainable Energy, Technische Universiteit Delft, 2628 CD Delft, Netherlands.(可持续能源系,代尔夫特理工大学) Bradley Department of Electrical and Computer Engineering, Virginia Polytechnic Institute and State University, Blacksburg, VA 24061, USA.(布雷德利电气与计算机工程系,弗吉尼亚理工学院和州立大学) Smart Grid Research Division System Reliability Research Team, Korea Electrotechnology Research Institute (KERI), Gwangju-si, 61751, South Korea.(智能电网研究分会系统可靠性研究团队,韩国电力技术研究所(KERI))

专题命中 其他多模态 :multi-modal(title);multimodal(abstract);分类 cs.AI

AI总结 本文提出了一种基于大型语言模型的多模态方法,用于电力系统中能源管理系统的安全防护和异常检测。

Comments 10 Figures; 6 Tables; Accepted, IEEE ACCESS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏