arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Huazhong University of Science and Technology(华中科技大学)

共收录 702
2603.05384 2026-03-06 cs.CV

ORMOT: A Dataset and Framework for Omnidirectional Referring Multi-Object Tracking

ORMOT: 一个用于全方位参照多目标跟踪的数据集和框架

Sijia Chen, Zihan Zhou, Yanqiu Yu, En Yu, Wenbing Tao

机构 * State Key Laboratory of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(多谱信息智能处理技术国家重点实验室,人工智能与自动化学院,华中科技大学)

AI总结 提出ORMOT任务,构建ORSet数据集和ORTrack框架,解决传统数据集视野限制问题,提升模型对长周期语言描述的理解能力。

Comments https://github.com/chen-si-jia/ORMOT

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05235 2026-03-06 cs.AI

Reclaiming Lost Text Layers for Source-Free Cross-Domain Few-Shot Learning

恢复丢失的文本层以实现无源跨域少样本学习

Zhenyu Zhang, Guangyao Chen, Yixiong Zou, Yuhua Li, Ruixuan Li

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) National Key Laboratory for Multimedia Information Processing, Peking University(北京大学多媒体信息处理国家重点实验室)

AI总结 本文提出了一种方法,通过重新利用文本编码器中被忽视的层信息,提升源无关跨域少样本学习的性能。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04338 2026-03-05 cs.CV

ArtHOI: Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors

ArtHOI: 通过视频先验进行4D重建的可变形人-物交互合成

Zihao Huang, Tianqi Liu, Zhaoxi Chen, Shaocong Xu, Saining Zhang, Lixing Xiao, Zhiguo Cao, Wei Li, Hao Zhao, Ziwei Liu

机构 * School of AIA, Huazhong University of Science and Technology(华中科技大学人工智能学院) the S-Lab, Nanyang Technological University (NTU)(南洋理工大学S实验室) the Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院) Zhejiang University (ZJU)(浙江大学) the Institute for AI Industry Research (AIR), Tsinghua University (THU)(清华大学人工智能产业研究院)

AI总结 ArtHOI通过4D重建从视频先验生成可变形人-物交互,提升接触精度和物理合理性。

Comments Project Page: https://arthoi.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03806 2026-03-05 cs.CV cs.AI

Separators in Enhancing Autoregressive Pretraining for Vision Mamba

分离器在增强视觉Mamba自回归预训练中的应用

Hanpeng Liu, Zidan Wang, Shuoxi Zhang, Kaiyuan Gao, Kun He

机构 * Huazhong University of Science and Technology(华中科技大学)

AI总结 本文提出STAR方法,通过引入分隔符技术扩展视觉Mamba的输入序列长度,提升其在ImageNet-1k上的准确率至83.5%

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03714 2026-03-05 cs.CL cs.AI cs.CV cs.MM

Order Is Not Layout: Order-to-Space Bias in Image Generation

秩序并非布局:图像生成中的秩序到空间偏差

Yongkang Zhang, Zonglin Zhao, Yuechen Zhang, Fei Ding, Pei Li, Wenxuan Wang

机构 * Renmin University of China, China(中国人民大学) Huazhong Agricultural University, China(华中农业大学) Huazhong University of Science and Technology, China(华中科技大学) Jiangnan University, China(江南大学) Nanchang University, China(南昌大学)

AI总结 本文研究了图像生成中因文本实体顺序导致的空间布局偏差问题,提出OTS-Bench进行量化评估,并通过微调和干预策略减少该偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02635 2026-03-04 cs.LG

SaFeR-ToolKit: Structured Reasoning via Virtual Tool Calling for Multimodal Safety

SaFeR-ToolKit: 通过虚拟工具调用实现多模态安全的结构化推理

Zixuan Xu, Tiancheng He, Huahui Yi, Kun Wang, Xi Chen, Gongli Xi, Qiankun Li, Kang Li, Yang Liu, Zhigang Zeng

机构 * Huazhong University of Science and Technology(华中科技大学) Beijing University of Posts and Telecommunications(北京邮电大学) West China Hospital, Sichuan University(四川大学华西医院) Nanyang Technological University(南洋理工大学)

AI总结 SaFeR-ToolKit通过虚拟工具调用实现多模态安全的结构化推理,提升安全性、帮助性和推理严谨性,同时保持通用能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02556 2026-03-04 cs.CV cs.AI cs.CL cs.LG

Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs

通过对比的视角:VLMs中的自改进视觉推理

Zhiyu Pan, Yizheng Wu, Jiashen Hua, Junyi Feng, Shaotian Yan, Bing Deng, Zhiguo Cao, Jieping Ye

机构 * Huazhong University of Science and Technology(华中科技大学) Alibaba Cloud(阿里云)

AI总结 通过视觉对比提升VLMs的推理能力,提出VC-STaR框架,有效减少推理幻觉并提升多种VLMs的视觉推理性能。

Comments 19 pages, 9 figures, accepted to ICLR 2026 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02329 2026-03-04 cs.CV

HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding

HAMMER: 通过跨模态整合利用大语言模型进行意图驱动的3D affordance grounding

Lei Yao, Yong Chen, Yuejiao Su, Yi Wang, Moyun Liu, Lap-Pui Chau

机构 * The Hong Kong Polytechnic University(香港理工大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 HAMMER通过跨模态整合多模态大语言模型,实现意图驱动的3D affordance grounding,提升3D表示的准确性和鲁棒性。

Comments Accepted by CVPR 2026. Project Page: https://rayyoh.github.io/Hammer

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01650 2026-03-04 cs.CV

PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts

PromptStereo: 通过结构和运动提示实现零样本立体匹配

Xianqi Wang, Hao Yang, Hangtian Wang, Junda Cheng, Gangwei Xu, Min Lin, Xin Yang

机构 * Huazhong University of Science and Technology(华中科技大学) Optics Valley Laboratory(光谷实验室)

AI总结 PromptStereo通过结构和运动提示提升零样本立体匹配的泛化性能,提出PRU模块增强单目深度模型的潜在表示。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05612 2026-03-04 cs.LG cs.AI

Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle

Shuffle-R1: 通过数据导向的动态洗牌提升多模态大语言模型的强化学习框架

Linghao Zhu, Yiran Guan, Dingkang Liang, Jianzhong Ju, Zhenbo Luo, Bin Qin, Jian Luan, Yuliang Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) MiLM Plus, Xiaomi Inc.(MiLM Plus,小米公司)

AI总结 Shuffle-R1通过动态洗牌和轨迹采样提升多模态大语言模型的强化学习效率,实现更高效的训练效果。

Comments This paper has been accepted by ICLR 2026. Conference link: https://iclr.cc/virtual/2026/poster/10007559 OpenReview link: https://openreview.net/forum?id=mYP33u1QBK Project page at: https://xenozlh.github.io/Shuffle-R1/

Journal ref The Fourteenth International Conference on Learning Representations (ICLR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01755 2026-03-03 cs.NI cs.AI

Federated Agentic AI for Wireless Networks: Fundamentals, Approaches, and Applications

联邦代理AI用于无线网络:基础、方法与应用

Lingyi Cai, Yu Zhang, Ruichen Zhang, Yinqiu Liu, Tao Jiang, Dusit Niyato, Wei Ni, Abbas Jamalipour

机构 * Research Center of 6G Mobile Communications, School of Cyber Science and Engineering, Huazhong University of Science and Technology(6G移动通信研究中心,信息科学与工程学院,华中科技大学) College of Computing and Data Science, Nanyang Technological University(计算与数据科学学院,南洋理工大学) School of Engineering, Edith Cowan University(工程学院,埃迪斯·科文大学) School of Computer Science and Engineering, University of New South Wales (UNSW)(计算机科学与工程学院,新南威尔士大学) School of Electrical and Computer Engineering, University of Sydney(电气与计算机工程学院,悉尼大学) Graduate School of Information Sciences, Tohoku University(信息科学研究生院,东北大学)

AI总结 本文提出适用于无线网络的联邦代理AI方法,通过协作本地学习和参数共享提升无线网络的自主性和自我改进能力。

Comments 7 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20999 2026-03-03 cs.CV

VII: Visual Instruction Injection for Jailbreaking Image-to-Video Generation Models

VII: 视觉指令注入用于劫持图像到视频生成模型

Bowen Zheng, Yongli Xiang, Ziming Hong, Zerong Lin, Chaojian Yu, Tongliang Liu, Xinge You

机构 * National Anti-Counterfeit Engineering Research Center, Huazhong University of Science and Technology(华中科技大学反伪工程研究中心) School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) Sydney AI Centre, The University of Sydney(悉尼大学AI中心)

AI总结 VII通过将恶意意图伪装为良性视觉指令,有效劫持图像到视频生成模型,实现高攻击成功率并降低拒绝率。

Comments Project page: https://Zbwwwwwwww.github.io/VII

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01026 2026-03-03 cs.CV

RaUF: Learning the Spatial Uncertainty Field of Radar

RaUF: 学习雷达的时空不确定性场

Shengpeng Wang, Kuangyu Wang, Wei Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Wuhan University(武汉大学)

AI总结 RaUF通过学习雷达测量的各向异性特性,解决方位模糊和虚假回波问题,提升空间检测的可靠性与不确定性校准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00971 2026-03-03 cs.CV

Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning

揭示认知罗盘:基于理论of-Mind的多模态情感推理

Meng Luo, Bobo Li, Shanqing Xu, Shize Zhang, Qiuchan Chen, Menglu Han, Wenhao Chen, Yanxiang Huang, Hao Fei, Mong-Li Lee, Wynne Hsu

机构 * National University of Singapore(新加坡国立大学) Huazhong University of Science and Technology(华中科技大学) The Hong Kong Polytechnic University(香港理工大学)

AI总结 本文提出HitEmotion基准和TMPO方法,通过理论of-Mind引导多模态情感推理,提升模型情感理解和认知能力。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23147 2026-03-03 cs.CV

GeoTeacher: Geometry-Guided Semi-Supervised 3D Object Detection

GeoTeacher: 基于几何的半监督3D物体检测

Jingyu Li, Xiaolong Zhao, Zhe Liu, Wenxiao Wu, Li Zhang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Tongji University(同济大学) Hong Kong University(香港大学) Huazhong University of Science and technology(华中科技大学)

AI总结 GeoTeacher通过几何关系监督模块和体素级数据增强策略,提升半监督3D物体检测的几何感知能力。

Comments Accepted for publication in 2026 IEEE International Conference on Robotics and Automation (ICRA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15888 2026-03-03 cs.CL cs.AI

Distribution-Aligned Decoding for Efficient LLM Task Adaptation

面向高效大语言模型任务适应的分布对齐解码

Senkang Hu, Xudong Han, Jinqi Jiang, Yihang Tao, Zihan Fang, Yong Dai, Sam Tak Wu Kwong, Yuguang Fang

机构 * Hong Kong JC STEM Lab of Smart City(香港JC智能城市STEM实验室) City University of Hong Kong(香港城市大学) University of Sussex(苏塞克斯大学) Huazhong University of Science and Technology(华中科技大学) Fudan University(复旦大学) Lingnan University(岭大大学)

AI总结 SVDecode通过引导输出分布对齐提升大语言模型任务适应性能,理论证明其与全微调等价,实验显示在多个任务上准确率提升显著。

Comments Accepted by NeurIPS'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23832 2026-03-02 cs.RO

OmniTrack: General Motion Tracking via Physics-Consistent Reference

OmniTrack: 通过物理一致的参考进行通用运动追踪

Yuhan Li, Peiyuan Zhi, Yunshen Wang, Tengyu Liu, Sixu Yan, Wenyu Liu, Xinggang Wang, Baoxiong Jia, Siyuan Huang

机构 * Huazhong University of Science and Technology(华中科技大学) State Key Lab of General AI, Beijing Institute for General Artificial Intelligence (BIGAI)(通用人工智能国家重点实验室,北京通用人工智能研究院) Shanghai Jiao Tong University(上海交通大学)

AI总结 OmniTrack通过分离物理可行性与通用运动追踪,提升了机器人运动跟踪的准确性和稳定性,支持复杂动作和动态远程操作。

Comments website: https://omnitrack-humanoid.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20903 2026-02-27 cs.CV

TextPecker: Rewarding Structural Anomaly Quantification for Enhancing Visual Text Rendering

TextPecker: 通过奖励结构异常量化提升视觉文本渲染

Hanshen Zhu, Yuliang Liu, Xuecheng Wu, An-Lan Wang, Hao Feng, Dingkang Yang, Chao Feng, Can Huang, Jingqun Tang, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) ByteDance(字节跳动)

AI总结 TextPecker通过结构异常感知强化学习策略提升视觉文本渲染的结构忠实度和语义对齐度。

Comments Accepted by CVPR 2026; Code: https://github.com/CIawevy/TextPecker

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13587 2026-02-27 cs.CV

UniFuture: A 4D Driving World Model for Future Generation and Perception

UniFuture: 一个面向未来生成与感知的4D驾驶世界模型

Dingkang Liang, Dingyuan Zhang, Xin Zhou, Sifan Tu, Tianrui Feng, Xiaofan Li, Yumeng Zhang, Mingyang Du, Xiao Tan, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) Baidu Inc.(百度公司)

AI总结 UniFuture通过统一4D建模提升自动驾驶中的未来生成与几何感知能力。

Comments Accepted by ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.11564 2026-02-27 cs.CV

SGIFormer: Semantic-guided and Geometric-enhanced Interleaving Transformer for 3D Instance Segmentation

SGIFormer: 基于语义引导和几何增强的交错变换器用于3D实例分割

Lei Yao, Yi Wang, Moyun Liu, Lap-Pui Chau

机构 * Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(电子工程系,香港理工大学) School of Mechanical Science and Engineering, Huazhong University of Science and Technology(机械科学与工程学院,华中科技大学)

AI总结 SGIFormer通过语义引导和几何增强的交错变换器,提升3D实例分割的准确性和效率。

Journal ref IEEE Transactions on Circuits and Systems for Video Technology, vol. 35, no. 3, pp. 2276-2288, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22033 2026-02-26 cs.CV

RT-RMOT: A Dataset and Framework for RGB-Thermal Referring Multi-Object Tracking

RT-RMOT:一种用于RGB-热成像参照多目标跟踪的数据集和框架

Yanqiu Yu, Zhifan Jin, Sijia Chen, Tongfei Chu, En Yu, Liman Liu, Wenbing Tao

机构 * Huazhong University of Science and Technology(华中科技大学) South-Central Minzu University(西南民族大学)

AI总结 RT-RMOT提出一种融合RGB和热成像特征的多目标跟踪框架,通过改进的RL策略优化提升全天候跟踪性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21983 2026-02-26 cs.RO

Humanizing Robot Gaze Shifts: A Framework for Natural Gaze Shifts in Humanoid Robots

使机器人目光转移人性化:一种在人形机器人中实现自然目光转移的框架

Jingchao Wei, Jingkai Qin, Yuxiao Cao, Jingcheng Huang, Xiangrui Zeng, Min Li, Zhouping Yin

机构 * School of Mechanical Science and Engineering, Huazhong University of Science and Technology(机械科学与工程学院,华中科技大学)

AI总结 本文提出RGS框架,通过结合认知注意力机制与生物模仿运动生成,实现人形机器人自然且上下文合适的人类化目光转移。

Comments submitted to AIM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02515 2026-02-24 cs.LG

FedSDAF: Leveraging Source Domain Awareness for Enhanced Federated Domain Generalization

FedSDAF: 利用源域意识提升联邦域泛化

Hongze Li, Zesheng Zhou, Zhenbiao Cao, Xinhui Li, Wei Chen, Xiaojin Zhang

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) School of Computer Science and Technology, Tiangong University(天津工业大学计算机科学与技术学院)

AI总结 FedSDAF通过利用源域意识特征提升联邦域泛化的性能,采用双适配器架构解耦本地专长与全局共识,引入双向知识蒸馏机制实现高效知识交流。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.16719 2026-02-20 cs.DB cs.AI

GPU-Accelerated Algorithms for Graph Vector Search: Taxonomy, Empirical Study, and Research Directions

基于GPU的图向量搜索算法:分类、实证研究与研究方向

Yaowen Liu, Xuejia Chen, Anxin Tian, Haoyang Li, Qinbin Li, Xin Zhang, Alexander Zhou, Chen Jason Zhang, Qing Li, Lei Chen

机构 * Hong Kong Polytechnic University, Hong Kong SAR(香港理工大学) The Hong Kong University of Science and Technology, Hong Kong SAR(香港理工大学) Huazhong University of Science and Technology, China(华中科技大学)

AI总结 本文研究了基于GPU的图向量搜索算法,通过实证分析揭示了距离计算和数据传输对性能的影响,并提出了系统设计中的关键权衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14577 2026-02-17 cs.CV

DriveFine: Refining-Augmented Masked Diffusion VLA for Precise and Robust Driving

DriveFine: 通过细化增强的掩码扩散VLA实现精确且鲁棒的驾驶

Chenxu Dang, Sining Ang, Yongkang Li, Haochen Tian, Jie Wang, Guang Li, Hangjun Ye, Jie Ma, Long Chen, Yan Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Xiaomi EV(小米电动车) Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学)

AI总结 DriveFine提出一种结合灵活解码与自我纠正能力的掩码扩散VLA模型,通过插件式块MoE和混合强化学习策略提升自动驾驶规划的精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14518 2026-02-17 cs.AI

Diagnosing Knowledge Conflict in Multimodal Long-Chain Reasoning

多模态长链推理中的知识冲突诊断

Jing Tang, Kun Wang, Haolang Lu, Hongjin Chen, KaiTao Chen, Zhongxiang Sun, Qiankun Li, Lingjuan Lyu, Guoshun Nan, Zhigang Zeng

机构 * Huazhong University of Science and Technology(华中科技大学) Nanyang Technological University(南洋理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Renmin University of China(中国人民大学)

AI总结 本研究提出了一种统一的知识冲突概念,揭示了多模态长链推理中不同冲突类型的特征和处理机制,为诊断和控制推理失败提供了原理性方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.14301 2026-02-17 cs.LG cs.AI cs.MA

DeepFusion: Accelerating MoE Training via Federated Knowledge Distillation from Heterogeneous Edge Devices

DeepFusion: 通过异构边缘设备的联邦知识蒸馏加速MoE训练

Songyuan Li, Jia Hu, Ahmed M. Abdelmoniem, Geyong Min, Haojun Huang, Jiwei Huang

机构 * School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦女王玛丽大学电子工程与计算机科学学院) Department of Computer Science, Faculty of Environment, Science and Economy, University of Exeter(埃克塞特大学环境、科学与经济学院计算机科学系) School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) Beijing Key Laboratory of Petroleum Data Mining, China University of Petroleum(北京石油数据挖掘重点实验室,中国石油大学)

AI总结 DeepFusion通过联邦知识蒸馏融合异构边缘设备的LLM知识,实现高效MoE训练,减少通信成本并提升模型性能。

Comments Index Terms: Large language models, Mixture-of-experts, Federated knowledge distillation, Edge device heterogeneity

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11636 2026-02-13 cs.CV cs.AI

ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning

ScalSelect: 可扩展的无训练多模态数据选择用于高效的视觉指令微调

Changti Wu, Jiahuai Mao, Yuzhuo Miao, Shijie Lian, Bin Yu, Xiaopeng Lin, Cong Huang, Lei Zhang, Kai Chen

机构 * East China Normal University(华东师范大学) Zhongguancun Academy(中关村学院) The Hong Kong Polytechnic University(香港理工大学) Harbin Institute of Technology(哈尔滨工业大学) Huazhong University of Science and Technology(华中科技大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)

AI总结 ScalSelect提出一种无训练、可扩展的多模态数据选择方法,通过线性时间复杂度实现高效视觉指令微调,实验显示其性能接近甚至超越全数据训练。

Comments The code is available at \href{https://github.com/ChangtiWu/ScalSelect}{ScalSelect}

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11007 2026-02-12 cs.CV

LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance Segmentation

LaSSM:通过局部聚合和状态空间模型实现高效的语义-空间查询解码用于3D实例分割

Lei Yao, Yi Wang, Yawen Cui, Moyun Liu, Lap-Pui Chau

机构 * Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(电子工程系,香港理工大学) School of Mechanical Science and Engineering, Huazhong University of Science and Technology(机械科学与工程学院,华中科技大学)

AI总结 LaSSM通过局部聚合和状态空间模型实现高效语义-空间查询解码,提升3D实例分割性能。

Comments Accepted at IEEE-TCSVT

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02172 2026-02-12 cs.CV cs.AI cs.MM

GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting

GaussianCross: 通过高斯点撒技术实现跨模态自监督3D表示学习

Lei Yao, Yi Wang, Yi Zhang, Moyun Liu, Lap-Pui Chau

机构 * Hong Kong Polytechnic University(香港理工大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 GaussianCross通过高斯点撒技术实现跨模态自监督3D表示学习,提升3D点云的表示质量和泛化能力。

Comments 14 pages, 8 figures, accepted by MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏