arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Huazhong University of Science and Technology(华中科技大学)

共收录 702
2603.19678 2026-03-23 cs.CV

Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identification

视觉-语言属性解耦与强化用于终身人物重识别

Kunlun Xu, Haotong Cheng, Jiangmeng Li, Xu Zou, Jiahuan Zhou

机构 * Wangxuan Institute of Computer Technology, Peking University, Beijing, China(北京大学王轩计算机技术研究院,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan, China(华中科技大学人工智能与自动化学院,武汉,中国)

AI总结 本文提出VLADR方法,通过视觉-语言属性解耦与强化提升跨域知识迁移,增强抗遗忘与泛化能力,实验表明在抗遗忘和泛化能力上优于现有方法。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13032 2026-03-20 cs.CV

Multimodal OCR: Parse Anything from Documents

多模态OCR:从文档中解析一切

Handong Zheng, Yumeng Li, Kaile Zhang, Liang Xin, Guangwei Zhao, Hao Liu, Jiayu Chen, Jie Lou, Qi Fu, Rui Yang, Shuo Jiang, Weijian Luo, Weijie Su, Weijun Zhang, Xingyu Zhu, Yabin Li, Yiwei ma, Yu Chen, Yuqiu Ji, Zhaohui Yu, Guang Yang, Colin Zhang, Lei Zhang, Yuliang Liu, Xiang Bai

机构 * Huazhong University of Science and Technology(华中科技大学) hi lab, Xiaohongshu Inc(小红书实验室,小红书公司)

AI总结 本文提出多模态OCR(MOCR),通过联合解析文本和图形,生成统一的文本表示。方法将图表、表格等视觉元素作为解析目标,提升文档重建的准确性,并通过端到端训练实现跨模态监督,最终在文档和结构化图形解析任务中取得优异表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18541 2026-03-20 cs.CV

Remedying Target-Domain Astigmatism for Cross-Domain Few-Shot Object Detection

矫正目标域色散以提升跨域少样本目标检测

Yongwei Jiang, Yixiong Zou, Yuhua Li, Ruixuan Li

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)

AI总结 本文针对跨域少样本目标检测中目标域色散问题,提出生物启发的注意力细化框架,通过正负模式细化和文本语义对齐模块提升检测精度,实验证明在六个基准上取得新突破。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.18035 2026-03-20 cs.LG

Taming Epilepsy: Mean Field Control of Whole-Brain Dynamics

癫痫控制:整体脑动力学的均场控制

Ming Li, Ting Gao, Jingqiao Dua

机构 * School of Mathematics and Information Science(数学与信息科学学院) Guangzhou University(广州大学) School of Sciences(科学学院) Great Bay University(大亚湾大学) Guangdong Provincial Key Laboratory of Mathematical and Neural Dynamical Systems(广东省数学与神经动力系统重点实验室) School of Mathematics and Statistics(数学与统计学学院) Huazhong University of Science and Technology(华中科技大学) Center for Mathematical Science(数学科学中心)

AI总结 本文提出基于图正则化的Koopman均场游戏框架,通过Reservoir Computing和Alternating Population and Agent Control Network实现脑神经动力学的非线性控制,有效抑制癫痫发作并保持脑功能拓扑结构。

Comments 22 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17851 2026-03-19 cs.RO

DexViTac: Collecting Human Visuo-Tactile-Kinematic Demonstrations for Contact-Rich Dexterous Manipulation

DexViTac:收集人类视觉-触觉-运动示范以实现富接触的灵巧操作

Xitong Chen, Yifeng Pan, Min Li, Xiaotian Ding

机构 * State Key Laboratory of Intelligent Manufacturing Equipment and Technology(智能制造装备与技术国家重点实验室) Huazhong University of Science and Technology(华中科技大学) Wuhan Huaweike Intelligent Technology Co., Ltd.(武汉华为凯科技有限公司)

AI总结 DexViTac通过高保真采集第一人称视觉、高密度触觉感知、末端执行器姿态和手部运动学,构建了超过2400个视觉-触觉-运动学示范的多模态数据集,提升了接触丰富灵巧操作的学习效率和成功率。

Comments 9 pages, 9 figures.Project page: https://xitong-c.github.io/DexViTac/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17412 2026-03-19 cs.CV cs.LG

Mutually Causal Semantic Distillation Network for Zero-Shot Learning

互因果语义蒸馏网络用于零样本学习

Shiming Chen, Shuhuang Chen, Guo-Sen Xie, Xinge You

机构 * Huazhong University of Science and Technology(华中科技大学) National Anti-Counterfeit Engineering Research Center(国家防伪工程技术研究中心) Nanjing University of Science and Technology(南京理工大学)

AI总结 本文提出互因果语义蒸馏网络MSDN++,通过双向因果注意力子网络学习视觉与属性间的内在语义表示,提升零样本学习性能。

Comments Accepted to IJCV. arXiv admin note: text overlap with arXiv:2203.03137

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03674 2026-03-18 cs.LG

Out-of-Distribution Graph Models Merging

分布外图模型融合

Yidi Wang, Ziyue Qiao, Jiawei Gu, Xubin Zheng, Pengyang Wang, Xiaobing Pei, Xiao Luo

机构 * School of Computing and Information Technology, Great Bay University(大湾大学计算与信息学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) Department of Computer Science, University of California, Los Angeles(加州大学洛杉矶分校计算机科学系) Department of Computer and Information Science, University of Macau(澳门大学计算机与信息科学系)

AI总结 本文提出一种图生成策略,通过MoE模块和掩码机制融合多个领域预训练图模型,实现跨领域适应,无需源目标域数据。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06140 2026-03-18 cs.CV

Boosting the Local Invariance for Better Adversarial Transferability

提升局部不变性以增强对抗迁移性

Bohan Liu, Xiaosen Wang

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)

AI总结 本文提出LI-Boost方法,通过提升对抗扰动的局部不变性来增强模型间对抗迁移性,实验表明其在多种攻击类型上均有效。

Comments Code is available at https://github.com/Trustworthy-AI-Group/TransferAttack

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13365 2026-03-17 cs.CV cs.AI

WaveComm: Lightweight Communication for Collaborative Perception via Wavelet Feature Distillation

WaveComm: 通过小波特征蒸馏实现轻量级协作感知通信

Erdemt Bao, Jin Yang

机构 * School of Mechanical Science and Engineering, Huazhong University of Science and Technology(华中科技大学机械科学与工程学院)

AI总结 WaveComm通过小波变换减少通信开销,保持低带宽环境下感知性能,实验显示在通信量减少至原值86.3%和87.0%时仍保持先进性能。

Comments Accepted by ICRA 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13341 2026-03-17 cs.CV cs.AI

Mind the Discriminability Trap in Source-Free Cross-domain Few-shot Learning

注意源无关跨域少样本学习中的判别性陷阱

Zhenyu Zhang, Yixiong Zou, Yuhua Li, Ruixuan Li, Guangyao Chen

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) National Key Laboratory for Multimedia Information Processing, Peking University(北京大学多媒体信息处理国家重点实验室)

AI总结 本文研究源无关跨域少样本学习中增强视觉模态判别性会抑制VLM性能的现象,提出通过扰动视觉学习以引导跨模态对齐的方法,实验显示在多种设置和数据集上取得新状态-of-the-art结果。

Comments CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13325 2026-03-17 cs.MA cs.AI

Auditing Cascading Risks in Multi-Agent Systems via Semantic-Geometric Co-evolution

通过语义-几何共演化审计多智能体系统中的级联风险

Zixun Luo, Yuhang Fan, Hengyu Lin, Yufei Li, Youzhi Zhang

机构 * Huazhong University of Science and Technology(华中科技大学) Lingnan University(岭大大学) Tsinghua University(清华大学) Centre for Artificial Intelligence and Robotics(人工智能与机器人中心) Hong Kong Institute of Science and Innovation(香港创新科技研究院) Chinese Academy of Sciences(中国科学院)

AI总结 本文提出基于语义-几何共演化的框架,通过动态图模型和Ollivier-Ricci曲率检测多智能体系统中的级联风险,提前预警并定位根本原因。

Comments This work has been accepted to ICLR 2026 Workshop: Principled Design for Trustworthy AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12903 2026-03-16 cs.CV

Spectral-Geometric Neural Fields for Pose-Free LiDAR View Synthesis

光谱-几何神经场用于无姿态激光雷达视角合成

Yinuo Jiang, Jun Cheng, Yiran Wang, Cheng Cheng

机构 * School of AIA, Huazhong University of Science and Technology(华中科技大学人工智能学院)

AI总结 本文提出SG-NLF框架,通过融合光谱信息与几何一致性实现无姿态激光雷达视角合成,提升重建质量和姿态精度。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12624 2026-03-16 cs.CV eess.IV

Prompt-Driven Lightweight Foundation Model for Instance Segmentation-Based Fault Detection in Freight Trains

基于提示的轻量级基础模型用于基于实例分割的货运列车故障检测

Guodong Sun, Qihang Liang, Xingyu Pan, Moyun Liu, Yang Zhang

机构 * School of Mechanical Engineering, Hubei University of Technology(湖北工业大学机械工程学院) Hubei Key Laboratory of Modern Manufacturing Quality Engineering, Hubei University of Technology(湖北工业大学现代制造质量工程重点实验室) School of Mechanical Science and Engineering, Huazhong University of Science and Technology(华中科技大学机械科学与工程学院)

AI总结 本文提出一种轻量级自提示实例分割框架,用于货运列车故障检测,通过自提示生成模块实现基础模型到领域任务的知识迁移,并采用轻量级视觉Transformer backbone,实现低计算成本的高效部署。

Comments 14 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10469 2026-03-12 cs.RO

DepthCache: Depth-Guided Training-Free Visual Token Merging for Vision-Language-Action Model Inference

DepthCache: 一种基于深度的无训练视觉标记融合方法,用于视觉-语言-动作模型推理

Yuquan Li, Lianjie Ma, Han Ding, Lijun Zhu

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) School of Mechanical Science and Engineering, Huazhong University of Science and Technology(华中科技大学机械科学与工程学院)

AI总结 DepthCache通过利用深度作为结构先验,实现无训练的视觉标记融合,提升视觉-语言-动作模型推理速度,同时保持较高的任务成功率。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20325 2026-03-12 cs.CV

AD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models

AD-R1: 闭环强化学习用于端到端自动驾驶的中立世界模型

Tianyi Yan, Tao Tang, Xingtai Gui, Yongkang Li, Jiasen Zhesng, Weiyao Huang, Lingdong Kong, Wencheng Han, Xia Zhou, Xueyang Zhang, Yifei Zhan, Kun Zhan, Cheng-zhong Xu, Jianbing Shen

机构 * SKL-IOTSC, University of Macau(SKL-IOTSC,澳门大学) Li Auto Inc. Sun Yat-sen University(中山大学) Huazhong University of Science and Technology(华中科技大学) Northwestern University(西北大学) National University of Singapore(新加坡国立大学)

AI总结 AD-R1通过引入中立世界模型和反事实合成技术,提升自动驾驶系统在危险预测和安全控制方面的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09718 2026-03-11 cs.CV

GSStream: 3D Gaussian Splatting based Volumetric Scene Streaming System

GSStream: 基于3D高斯点云的体积分流系统

Zhiye Tang, Qiudan Zhang, Lei Zhang, Junhui Hou, You Yang, Xu Wang

机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院)

AI总结 GSStream通过协作视口预测和深度强化学习比特率适应,实现高效的3D高斯点云体积分流交付。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09703 2026-03-11 cs.CV

ProGS: Towards Progressive Coding for 3D Gaussian Splatting

ProGS:迈向3D高斯散射的渐进编码

Zhiye Tang, Lingzhuo Liu, Shengjie Jiao, Qiudan Zhang, Junhui Hou, You Yang, Xu Wang

机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院)

AI总结 ProGS通过八叉树结构实现3D高斯散射的高效渐进编码,显著提升压缩效率和视觉保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08113 2026-03-10 cs.CV

SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving

SAMoE-VLA:一种面向自动驾驶的场景自适应混合专家视觉-语言-动作模型

Zihan You, Hongwei Liu, Chenxu Dang, Zhe Wang, Sining Ang, Aoqi Wang, Yan Wang

机构 * Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究院(AIR),清华大学) School of Instrument Science and Engineering, Southeast University(仪器科学与工程学院,东南大学) Zhili College, Tsinghua University(紫荆学院,清华大学) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) Department of Automation, University of Science and Technology of China(自动化学院,中国科学技术大学) Department of Automation, University of Science and Technology Beijing(自动化学院,北京科技大学)

AI总结 SAMoE-VLA通过场景自适应混合专家机制提升自动驾驶中的视觉-语言-动作推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07928 2026-03-10 cs.RO

Omnidirectional Humanoid Locomotion on Stairs via Unsafe Stepping Penalty and Sparse LiDAR Elevation Mapping

全方位台阶行走的人形机器人:通过不安全踏步惩罚与稀疏Li DAR高度映射

Yuzhi Jiang, Yujun Liang, Junhao Li, Han Ding, Lijun Zhu

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) State Key Laboratory of Intelligent Manufacturing Equipment and Technology, Huazhong University of Science and Technology(华中科技大学智能制造装备技术国家重点实验室) School of Mechanical Science and Engineering, Huazhong University of Science and Technology(华中科技大学机械科学与工程学院)

AI总结 本文提出一种单阶段训练框架,结合密集不安全踏步惩罚和稀疏LiDAR高度映射,实现人形机器人在台阶上的安全全方位行走。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07630 2026-03-10 cs.CV

Real-Time Glottis Detection Framework via Spatial-decoupled Feature Learning for Nasal Transnasal Intubation

通过空间解耦特征学习的实时声带检测框架用于鼻内气管插管

Jinyu Liu, Gaoyang Zhang, Yang Zhou, Ruoyi Hao, Yang Zhang, Hongliang Ren

机构 * Hubei Key Laboratory of Modern Manufacturing Quality Engineering, Hubei University of Technology(湖北现代制造质量工程重点实验室,湖北工业大学) School of Mechanical Science and Engineering, Huazhong University of Science and Technology(华中科技大学机械科学与工程学院) Department of Electronic Engineering, The Chinese University of Hong Kong(香港中文大学电子工程系) Key Laboratory of Symbolic Computation and Knowledge Engineering, Ministry of Education(教育部符号计算与知识工程重点实验室) National Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家实验室)

AI总结 本文提出Mobile GlottisNet,通过空间解耦特征学习实现实时声带检测,适用于嵌入式和边缘设备,提升鼻内气管插管的急救应用效率。

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07590 2026-03-10 cs.CV cs.LG

Models as Lego Builders: Assembling Malice from Benign Blocks via Semantic Blueprints

模型作为乐高积木:通过语义蓝图从良性积木中组装恶意内容

Chenxi Li, Xianggan Liu, Dake Shen, Yaosong Du, Zhibo Yao, Hao Jiang, Linyi Jiang, Chengwei Cao, Jingzhe Zhang, RanYi Peng, Peiling Bai, Xiande Huang

机构 * DAIL Tech(DAIL科技) NLP & KG Lab, Huazhong University of Science and Technology(自然语言处理与知识图谱实验室,华中科技大学)

AI总结 本文提出StructAttack,通过语义蓝图将看似无害的槽类型组合成恶意内容,揭示LVLMs在视觉模态整合中的安全漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02767 2026-03-10 cs.CV cs.AI

ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion

通过协同多模态对齐和训练时融合实现图像与文本一体化:ITO

Hanpeng Liu, Yaqian Li, Zidan Wang, Shuoxi Zhang, Zonglin Zhao, Zihao Bo, Rinyoichi Takezoe, Kaiwen Long, Kun He

机构 * School of Computer Science(计算机科学学院) Huazhong University of Science and Technology(华中科技大学) Li Auto Inc.(力汽车公司) Institute of AI for Industries, Chinese Academy of Sciences(产业人工智能研究院,中国科学院)

AI总结 ITO通过协同多模态对齐与训练时融合机制,提升图像与文本表示的一致性,有效解决模态间结构化交互问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15163 2026-03-10 cs.LG stat.ML

The Exploration of Error Bounds in Classification with Noisy Labels

在噪声标签下分类中误差界限的探索

Haixia Liu, Boxiao Li, Can Yang, Yang Wang

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) School of Mathematics and Statistics(数学与统计学学院) Huazhong University of Science and Technology(华中科技大学) The University of Hong Kong(香港大学)

AI总结 本文研究了噪声标签下深度学习分类中的误差界限,通过分解统计误差和近似误差,提出理论分析方法以提高分类性能。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07023 2026-03-10 cs.CL cs.AI

Hit-RAG: Learning to Reason with Long Contexts via Preference Alignment

通过偏好对齐学习长上下文的推理

Junming Liu, Yuqi Li, Shiping Wen, Zhigang Zeng, Tingwen Huang

机构 * Tongji University(同济大学) The City University of New York(纽约城市大学) University of Technology Sydney(悉尼大学) Huazhong University of Science and Technology(华中科技大学) Shenzhen University of Advanced Technology(深圳先进技术大学)

AI总结 Hit-RAG通过多阶段偏好对齐框架解决长上下文推理中的注意力稀释和幻觉问题,提升模型在长上下文场景下的推理能力。

Comments 21 pages, 2 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06669 2026-03-10 cs.NI cs.AI

Hybrid Orchestration of Edge AI and Microservices via Graph-based Self-Imitation Learning

基于图的自我模仿学习的边缘AI与微服务混合编排

Chen Yang, Jin Zheng, Yang Zhuolin, Lai Pan, Zhang Xiao, Hu Menglan, Yin Haiyan

机构 * School of Computer Science, South-Central Minzu University(中央民族大学南中央学院) Key Laboratory of Cyber-Physical Fusion Intelligent Computing, State Ethnic Affairs Commission(国家民族事务委员会网络物理融合智能计算重点实验室) Hubei Key Laboratory of Smart Internet Technology, School of Electronic Information and Communication, Huazhong University of Science and Technology(湖北智能互联网技术重点实验室,华中科技大学电子信息与通信学院) Centre for Frontier AI Research, Agency for Science, Technology and Research (A*STAR)(前沿人工智能研究中心,科技研究局(A*STAR))

AI总结 本文提出SIL-GPO框架,通过图注意力网络和自我模仿学习优化边缘AI与微服务的混合编排,显著降低延迟并提升资源利用率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19195 2026-03-10 cs.CV cs.AI

Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks

重新思考驾驶世界模型作为感知任务的合成数据生成器

Kai Zeng, Zhanqian Wu, Kaixin Xiong, Xiaobao Wei, Xiangyu Guo, Zhenxin Zhu, Kalok Ho, Lijun Zhou, Bohan Zeng, Ming Lu, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Wentao Zhang

机构 * Peking University(北京大学) Xiaomi EV(小米电动车) Huazhong University of Science and Technology(华中科技大学) Beijing Key Laboratory of Data Intelligence and Security (Peking University)(北京数据智能与安全重点实验室(北京大学)) Zhongguancun Academy(中关村学院)

AI总结 Dream4Drive通过生成高质量的合成数据提升自动驾驶感知任务性能

Journal ref ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13687 2026-03-09 cs.CV

Towards Scalable Pre-training of Visual Tokenizers for Generation

面向生成任务的视觉分词器可扩展预训练

Jingfeng Yao, Yuda Song, Yucong Zhou, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学) MiniMax

AI总结 VTP通过联合优化图像-文本对比、自监督和重建损失,提升视觉分词器的生成性能和扩展性。

Comments Our pre-trained models are available at https://github.com/MiniMax-AI/VTP

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06357 2026-03-09 cs.CV

LATO: 3D Mesh Flow Matching with Structured TOpology Preserving LAtents

LATO:基于结构拓扑保持的3D网格流匹配

Tianhao Zhao, Youjia Zhang, Hang Long, Jinshen Zhang, Wenbing Li, Yang Yang, Gongbo Zhang, Jozef Hladký, Matthias Nießner, Wei Yang

机构 * Huazhong University of Science and Technology(华中科技大学) Technical University of Munich(慕尼黑技术大学) Independent Researcher(独立研究者) Peking University(北京大学)

AI总结 LATO通过拓扑保持的潜在表示实现高效3D网格生成,结合流匹配技术生成复杂几何且拓扑结构良好的网格。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06200 2026-03-09 cs.CV

Adaptive Language-Aware Image Reflection Removal Network

自适应语言感知图像反光去除网络

Siyan Fang, Yuntao Wang, Jinpu Zhang, Ziwen Li, Yuehuan Wang

机构 * Huazhong University of Science and Technology(华中科技大学) National University of Defense Technology(国防科技大学)

AI总结 ALANet通过自适应语言感知策略有效去除复杂反光,提升语言引导下的反光去除性能。

Comments IJCAI 2025

Journal ref Proceedings of the 34th International Joint Conference on Artificial Intelligence (IJCAI-25), pages 973-981, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.05845 2026-03-09 cs.CV

Cog2Gen3D: Sculpturing 3D Semantic-Geometric Cognition for 3D Generation

Cog2Gen3D: 三维语义-几何认知雕刻用于三维生成

Haonan Wang, Hanyu Zhou, Haoyue Liu, Tao Gu, Luxin Yan

机构 * School of Artificial and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) School of Computing, National University of Singapore(新加坡国立大学计算机学院) School of Computing, Macquarie University(麦考瑞大学计算机学院)

AI总结 Cog2Gen3D通过结合语义和绝对几何信息,提出了一种三维认知引导的扩散框架,以实现可控的三维生成。

详情

展开后加载摘要…

URL PDF HTML 收藏