arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Science and Technology of China(中国科学技术大学)

2026-07-15 至 2026-07-15 共收录 5
2607.12959 2026-07-15 cs.CV 新提交

ViCo3D: Empowering LiDAR-based Collaborative 3D Object Detection with Vision Foundation Models

ViCo3D:利用视觉基础模型增强基于激光雷达的协作式3D目标检测

Haojie Ren, Songrui Luo, Lingfeng Wang, Yan Xia, Yao Li, Jing Li, Lu Zhang, Jiajun Deng, Yanyong Zhang

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 研究针对V2X系统中基于激光雷达协作式3D感知的不足,提出ViCo3D框架。通过将点云投影为图像让VFM提取特征,引入融合模块及跨智能体融合策略,实现了先进的3D检测性能,提升了协作增益。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12645 2026-07-15 cs.LG 新提交

AdaPCLA: Adaptive Prior-Calibrated Logit Adjustment for Long-Tailed Longitudinal EHR Generation

AdaPCLA:用于长尾纵向电子健康记录生成的自适应先验校准逻辑调整

Shuai Cui, Chen Wenxuan, Wenjie Du, Jian Lou, Dan Li, Wenjie Feng

机构 * University of Science and Technology of China(中国科学技术大学) Sun Yat-sen University(中山大学)

AI总结 研究针对纵向电子健康记录生成中标准模型的不足,提出AdaPCLA框架,通过数据分布感知训练策略、模拟退火训练及零样本分布控制实现自适应拟合与生成,理论分析刻画相关界限,实验验证其在多方面有提升。

Comments 40 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.12576 2026-07-15 cs.SD 新提交

UD-ASD: A Unified Diffusion Model for Anomalous Sound Detection

UD-ASD:一种用于异常声音检测的统一扩散模型

Pengxiang Gao, Yu Qiu, Yanzhi Song

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 研究针对异常声音检测,提出含轻量级模块的统一扩散模型,先将音频转对数梅尔频谱图,通过嵌入机器ID引导模型为特定机器重建数据,经高斯混合模型拟合误差分布,实验验证该模型在DCASE2022任务2中相比基线有显著提升。

Comments 5 pages, 3 figures, Interspeech 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.12783 2026-07-15 cs.IR cs.AI 版本更新

SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise

SQuTR:一种在语音噪声下 spoken query 到文本检索的鲁棒性基准

Yuejie Li, Ke Yang, Yueying Hua, Berlin Chen, Jianhao Nie, Yueping He, Caixin Kang

机构 * Huazhong University of Science and Technology(华中科技大学) The University of Hong Kong(香港大学) Soochow University(苏州大学) University of Science and Technology of China(中国科学技术大学) Wuhan University(武汉大学) Tsinghua University(清华大学) The University of Tokyo(东京大学)

AI总结 SQuTR通过大规模数据集和统一评估协议,评估语音检索系统在复杂噪声环境下的鲁棒性,揭示了极端噪声下检索性能显著下降的问题。

Comments Accepted by SIGIR 2026

Journal ref Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '26), July 20--24, 2026, Melbourne, VIC, Australia

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.19632 2026-07-15 cs.CV 版本更新

CreatiParser: Generative Image Parsing of Raster Graphic Designs into Editable Layers

CreatiParser: 从位图图形设计生成可编辑的图层

Weidong Chen, Dexiang Hong, Zhendong Mao, Yutao Cheng, Xinyan Liu, Lei Zhang, Yongdong Zhang

机构 * School of Information Science and Technology, University of Science and Technology of China(科学技术大学信息科学与技术学院) ByteDance Intelligent Creation(字节跳动智能创作) School of Computer Science and Technology, Harbin Institute of Technology (Weihai)(哈尔滨工业大学(威海)计算机科学与技术学院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合国家科学中心人工智能研究院)

AI总结 本文提出CreatiParser框架,将位图图形设计分解为可编辑的文本、背景和贴纸图层,结合视觉语言模型和多分支扩散架构,提升生成质量与编辑灵活性,实验显示在Parser-40K和Crello数据集上性能优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏