arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Huazhong University of Science and Technology(华中科技大学)

2026-04-22 至 2026-04-22 共收录 5
2604.19432 2026-04-22 cs.CV

DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval

DINO Eats CLIP:超越已知领域进行开放集3D物体检索

Xinwei He, Yansong Zheng, Qianru Han, Zhichuan Wang, Yuxuan Cai, Yang Zhou, Jingbo Xia, Yulong Wang, Jinhai Xiang, Xiang Bai

机构 * Huazhong Agricultural University(华中农业大学) Huazhong University of Science and Technology(华中科技大学) Shenzhen University(深圳大学)

AI总结 本文提出DEC框架,通过动态多视图整合和虚拟特征合成模块,提升开放集3D物体检索的性能,解决已知类别过拟合问题。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19349 2026-04-22 cs.CV

RAFT-MSF++: Temporal Geometry-Motion Feature Fusion for Self-Supervised Monocular Scene Flow

RAFT-MSF++: 基于时间几何-运动特征融合的自监督单目场景流估计

Xunpei Sun, Zuoxun Hou, Yi Chang, Gang Chen, Wei-Shi Zheng

机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) Beijing Institute of Space Mechanics and Electricity(北京空间机械与电子研究所) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)

AI总结 RAFT-MSF++提出一种自监督多帧框架,通过递归融合时间特征联合估计深度和场景流,引入几何-运动特征模块和遮挡正则化模块提升鲁棒性,实验显示在KITTI数据集上取得优异性能。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21954 2026-04-22 cs.SE cs.AI

Fine-Tuning Code Language Models to Detect Cross-Language Bugs

对代码语言模型进行微调以检测跨语言错误

Zengyang Li, Yimeng Li, Binbin Huang, Peng Liang, Ran Mo, Hui Liu, Yutao Ma

机构 * School of Computer Science, Central China Normal University(中央师范大学计算机科学学院) State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Computer Science, Wuhan University(武汉大学计算机学院) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)

AI总结 本文研究预训练代码语言模型在跨语言错误检测中的潜力,开发了CLCFinder工具并构建了包含三种编程语言组合的跨语言错误数据集,通过微调13种CodeLMs发现小模型表现优于大模型,且数据集规模和代码注释对性能影响各异。

Comments Preprint accepted for publication in ACM Transactions on Software Engineering and Methodology (TOSEM), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16161 2026-04-22 cs.CV cs.CL

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

OmniParser V2:结构化思维点用于统一的视觉文本解析及其在多模态大语言模型中的通用性

Wenwen Yu, Zhibo Yang, Jianqiang Wan, Sibo Song, Jun Tang, Wenqing Cheng, Yuliang Liu, Xiang Bai

机构 * School of Information Science and Engineering, East China University of Science and Technology(东华大学信息科学与工程学院) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) Alibaba Group(阿里巴巴集团)

AI总结 本文提出OmniParser V2,通过结构化思维点提示方案统一视觉文本解析任务,简化流程并提升性能,在多个数据集上取得最佳结果,并验证其在多模态大语言模型中的通用性。

Comments Accepted by IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18275 2026-04-22 cs.CV

Visual Adversarial Attack on Vision-Language Models for Autonomous Driving

面向自动驾驶的视觉对抗攻击研究

Tianyuan Zhang, Lu Wang, Xinwei Zhang, Yitong Zhang, Boyi Jia, Siyuan Liang, Shengshan Hu, Qiang Fu, Aishan Liu, Xianglong Liu

机构 * Beihang University(北航) National University of Singapore(国立新加坡大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 本文提出ADvLM框架,针对自动驾驶中视觉语言模型的特殊需求,解决文本指令变异性和视觉场景时间序列性问题,实现高效对抗攻击。

Comments Accepted by Machine Intelligence Research

详情

展开后加载摘要…

URL PDF HTML 收藏