arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Huazhong University of Science and Technology(华中科技大学)

共收录 94
2606.26711 2026-07-01 cs.CV 新提交

Mask to Concept: Auto-Promptable SAM3 via Efficient Test-Time Concept Embedding Search for Few-Shot Annotation

掩码到概念:通过高效测试时概念嵌入搜索实现自动可提示的SAM3用于少样本标注

Quan Zhou, Shaoqing Zhai, Qiang Hu, Jia Chen, Qiang Li, Zhiwei Wang

机构 * Wuhan University of Technology(武汉理工大学) Huazhong University of Science and Technology(华中科技大学) Changzhou United lmaging Healthcare Surgical Technology Co. Ltd.(常州联影医疗手术技术有限公司)

AI总结 提出Mask to Concept (M2C)框架,无需外部模块或重训练,仅用少量标注图像,通过可学习概念嵌入搜索和混合不确定性估计,实现SAM3在医学图像中的自动少样本标注,达到SOTA性能。

Comments Accepted by MICCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27504 2026-06-29 cs.CV 新提交

ReWorld: Learning Better Representations for World Action Models

ReWorld:为世界动作模型学习更好的表示

Tianze Xia, Lijun Zhou, Kaixin Xiong, Jingfeng Yao, Yu Zhu, Zhenxin Zhu, Bing Wang, Guang Chen, Hangjun Ye, Wenyu Liu, Haiyang Sun, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Xiaomi EV(小米汽车)

AI总结 提出ReWorld框架,通过优化中间表示(未来预测监督、跨模态对齐、难负样本监督)提升自动驾驶世界动作模型的规划性能,在nuScenes和NAVSIM上取得显著提升。

Comments 19 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27187 2026-06-26 cs.CV cs.CL 新提交

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

HarmVideoBench:大型多模态模型中有害视频理解的基准测试

Jiajun Wu, Haoyu Kang, Yining Sun, Jiacheng Hou, Heng Zhang, Danyang Zhang, Zhenjun Zhao, Haochi Zhang, Leixin Sun, Eric Hanchen Jiang, Yushan Li, Ruiyu Li, Mengkai Huang, Yan Gao, Xu Zhang, Guancheng Wan

机构 * Central South University(中南大学) Tsinghua University(清华大学) South China Normal University(华南师范大学) ByteDance Inc(字节跳动公司) University of Zaragoza(阿拉维达大学) CosmosMind Wuhan University(武汉大学) University of California, Los Angeles(加州大学洛杉矶分校) Southeast University(东南大学) Tencent(腾讯公司) Nankai University(南开大学) Supermicro Computer Inc(Supermicro计算机公司) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出HarmVideoBench,一个包含1379个视频和4137道选择题的多层次诊断基准,从三个维度评估模型对有害视频的深层理解,并引入BCR方法将宏平均准确率从61.7%提升至84.4%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26938 2026-06-26 cs.CV 新提交

Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE

关注重要之处:面向扩散MoE的显著性引导精确路由

Haoyou Deng, Keyu Yan, Chaojie Mao, Xiang Wang, Yu Liu, Changxin Gao, Nong Sang

机构 * Key Laboratory of Image Processing and Intelligent Control, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院图像处理与智能控制重点实验室) Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室)

AI总结 针对扩散MoE中路由无法准确分配计算资源给显著令牌的问题,提出SharpMoE后训练框架,利用干净潜在特征作为无噪声引导信号,并引入轨迹路由损失,实现精确资源分配,在视觉生成中达到最优性能。

Comments ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26171 2026-06-26 cs.CV cs.AI 新提交

LCG: Long-Context Consistent Image Generation with Sparse Relational Attention

LCG: 基于稀疏关系注意力的一致长上下文图像生成

Zihao Wang, Yijia Xu, Haoze Zheng, Xuran Ma, Haokun Gui, Harry Yang

机构 * Huazhong University of Science and Technology(华中科技大学) Peking University(北京大学) Hong Kong University of Science and Technology(香港科技大学)

AI总结 提出LCG框架,利用稀疏关系注意力(SRA)和路由一致性约束(RCC),在长序列图像生成中保持语义和外观一致性,并构建了大规模数据集LCCD进行训练和评估。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25754 2026-06-25 cs.RO 新提交

Stage-Aware and Roughness-Constrained Diffusion Policy for Multi-Stage Robotic Polishing

阶段感知与粗糙度约束的扩散策略用于多阶段机器人抛光

Shuai Ke, Jiexin Zhang, Huan Zhao, Zhiao Wei, Yikun Guo, Tiange Wu, Guoqiang Guo, Haoyuan Zhou, Jie Pan, Han Ding

机构 * State Key Laboratory of Intelligent Manufacturing Equipment and Technology, Huazhong University of Science and Technology(华中科技大学智能制造装备与技术国家重点实验室) Shanghai Spaceflight Precision Machinery Institute(上海航天精密机械研究所)

AI总结 提出阶段感知与粗糙度约束扩散策略(SRDP),通过多模态观测推断阶段后验并约束去噪过程,实现无外部标签的阶段一致动作生成,提升多阶段抛光质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25277 2026-06-25 cs.RO cs.CV 新提交

An Integrated Hardware-Software Design for Low-Data Spatial Defect Detection in Robotic Visual Inspection with Hybrid Optoelectronic Neural Networks

面向机器人视觉检测中低数据空间缺陷检测的软硬件集成设计:混合光电神经网络

Chaoqing Tang, Jiaxuan Li, Huanze Zhuang, Guiyun Tian, Chao Wang, Yihao Ouyang, Wenzhong Liu

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) China Belt and Road Joint Lab on Measurement and Control(中国一带一路测量与控制联合实验室) School of Electric and Electrical Engineering, Chongqing University of Technology(重庆理工大学电气与电子工程学院) Department of Precision Instrument, Tsinghua University(清华大学精密仪器系)

AI总结 提出一种软硬件集成的光电架构,通过非成像低数据范式和传感器在环策略,结合压缩感知与CLIP引导注意力,实现缺陷检测数据量减少90%、计算量降低60%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23888 2026-06-24 eess.IV cs.AI cs.CV 新提交

E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis

E-MRL: 跨视图对齐的证据驱动多模态强化学习用于可靠的3D肿瘤分析

Sijing Li, Zhongwei Qiu, Zhuoya Wang, Boxiang Yun, Zhenyu Yi, Jianwei Xu, Wenqiao Zhang, Yingda Xia, Ling Zhang

机构 * Zhejiang University(浙江大学) DAMO Academy, Alibaba Group(阿里巴巴集团达摩院) Hupan Lab(华平实验室) Huazhong University of Science and Technology(华中科技大学) East China Normal University(华东师范大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 提出跨视图对齐的证据驱动多模态强化学习框架E-MRL,通过将生成过程建模为“诊断-定位-验证”的马尔可夫决策过程,并引入跨视图一致性奖励,减少视觉幻觉并提升3D CT肿瘤诊断准确性。

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23105 2026-06-23 cs.CV 新提交

Compression and Retrieval: Implicit Memory Retrieval for Video World Models

压缩与检索:视频世界模型的隐式记忆检索

Zhan Peng, Jie Ma, Huiqiang Sun, Chong Gao, Zhijie Xue, Zhiyu Pan, Zhiguo Cao, Jun Liang, Jing Li

机构 * Huazhong University of Science and Technology(华中科技大学) HUJING Digital Media & Entertainment Group(虎鲸数字媒体与娱乐集团) Sun Yat-sen University(中山大学)

AI总结 提出注意力驱动的隐式记忆检索机制CaR,通过位置编码注入视角信息实现灵活检索,并引入轻量级上下文压缩网络,在构建的SceneFly数据集上取得SOTA结果并展现强泛化性。

Comments Project page: 3DV-Team/CaR" target="_blank" rel="noopener">https://github.com/Orange-3DV-Team/CaR

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22883 2026-06-23 cs.AI 新提交

CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents

CLI-Universe:面向终端代理的可验证任务合成引擎

Zhanbo Hua, Yifan Yao, Weihao Xie, Yongchi Zhao, Minghao Liu, Ruizhi Qiu, Zhewei Huang, Zun Wang, Yiyan Ji, Yunhai Ye, Letian Zhu, Xinping Lei, Han Li, Zhiyuan Ma, Zili Wang, Zhaoxiang Zhang, Jiaheng Liu

机构 * Nanjing University(南京大学) StepFun(阶跃星辰) ZODA Shanghai AI Lab(上海人工智能实验室) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出CLI-Universe引擎,通过多维能力分类采样和证据引导研究生成终端代理任务,经多阶段可执行验证流水线筛选,构建高质量数据集CLI-Universe-6K,微调Qwen3-32B在Terminal-Bench 2.0上达到33.4%的SOTA。

Comments 20 pages, 5 figures, 3 tables. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22631 2026-06-23 cs.CV 新提交

4DVLT: Dynamic Scene Understanding with Worldline-Centered Vision-Language Tracking

4DVLT:以世界线为中心的视觉-语言跟踪动态场景理解

Chaoyue Li, Boxue Yang, Shengyao Zhou, Haoyang Wu, Rui Qian, Linfeng Zhang

机构 * Huazhong University of Science and Technology(华中科技大学) Shanghai Jiao Tong University(上海交通大学) Zhejiang University(浙江大学)

AI总结 提出世界线为中心的4D动态场景理解任务4DVLT及基准Instruct-4D,通过图条件世界线推理方法4DTrack实现指令跟踪,在指标上超越基线19.62点。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21956 2026-06-23 cs.CV 新提交

Denoising-Enhanced Coarse-to-Fine Infrared Small Target Detection with Attention Prior-Guided Knowledge Distillation

基于注意力先验引导知识蒸馏的增强去噪粗到细红外小目标检测

Houzhang Fang, Ruixuan Huang, Qiuhuan Chen, Xiaolin Wang, Yi Chang, Luxin Yan

机构 * Xidian University(西安电子科技大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出一种粗到细的红外小目标检测框架ECFNet,通过粗阶段的区域二分类网络和去噪辅助训练抑制背景,细阶段的轻量检测器与注意力先验知识蒸馏提升精度,在三个数据集上优于现有方法。

Comments Accepted by ECCV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.21244 2026-06-23 cs.CV cs.AI 新提交

ACE-GS: Acing the Trade-off with Accurate, Compact and Efficient 3D Gaussian Splatting

ACE-GS:精准、紧凑且高效的三维高斯泼溅权衡之道

Jijian Zhao

机构 * Huazhong University of Science and Technology(华中科技大学)

AI总结 提出ACE-GS渐进优化框架,通过动量一致性引导的致密化策略和统计灵敏度驱动的稀疏化机制实现精准基元管理,并引入跨维度残差频率补偿恢复细节,在保持紧凑表示的同时实现3.7倍训练加速和最高0.89 dB PSNR提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20590 2026-06-23 cs.NI cs.AI 新提交

Optimization-as-a-Service via Multi-Agent Large Language Model for Radio Access Networks

通过多智能体大语言模型实现无线接入网的优化即服务

Chaoqun You, Yueyue Dai, Xingqiu He, Yue Gao, Rahim Tafazolli, Yong Liang Guan

机构 * Fudan University(复旦大学) Huazhong University of Science and Technology(华中科技大学) University of Surrey(Surrey大学) Nanyang Technological University(南洋理工大学)

AI总结 提出将物理资源块分配问题作为大语言模型多智能体系统提供的优化即服务,通过场景理解、目标生成、求解器和反思智能体实现上下文感知的自校正公式,并引入单次反思蒸馏机制降低延迟。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20333 2026-06-19 cs.AI 新提交

SoftSkill: Behavioral Compression for Contextual Adaptation

SoftSkill: 用于上下文适应的行为压缩

Xijia Tao, Yihua Teng, Xinyu Fu, Ziru Liu, Kecheng Chen, Yuzhi Zhao, Suiyun Zhang, Rui Liu, Lingpeng Kong

机构 * The University of Hong Kong(香港大学) Huawei Research(华为研究院) City University of Hong Kong(香港城市大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出SoftSkill方法,通过可训练的软技能前缀压缩自然语言技能为紧凑连续向量,在冻结基模型上提升问答和数学任务性能,减少标记数量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20227 2026-06-19 cs.AI cs.SE 新提交

QMFOL: Benchmarking Large Language Model Reasoning via Quantifiable Monadic First-Order Logic Test Case Generation

QMFOL:通过可量化的一元一阶逻辑测试用例生成来基准测试大语言模型推理

Xinyi Zheng, Ling Shi, Tianlong Yu, Yongxin Zhao, Lorenz Goette, Kailong Wang

机构 * Huazhong University of Science and Technology(华中科技大学) Nanyang Technological University(南洋理工大学) Hubei University(湖北大学) East China Normal University(华东师范大学) National University of Singapore(新加坡国立大学)

AI总结 提出QMFOL框架,通过可控制复杂度的合取/析取模式生成一元一阶逻辑推理任务,并构建包含2880个实例的基准QMFOLBench,评估显示逻辑复杂度增加导致性能下降和计算开销上升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19736 2026-06-19 cs.CV 新提交

VFACamou: View-Fused Adversarial Camouflage for Environment-Adaptive Physical Evasion

VFACamou: 视图融合的对抗性伪装用于环境自适应物理规避

Shihui Yan, Hu Liu, Junyu Shi, Zihui Zhu, Ziqi Zhou, Yufei Song, Youming Geng, Minghui Li, Shengshan Hu

机构 * State Key Laboratory of Intelligent Vehicle Safety Technology(智能汽车安全技术国家重点实验室) School of Cyber Science and Engineering, Huazhong University of Science and Technology(华中科技大学网络空间安全学院) School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) Hebei Energy College of Vocation And Technology(河北能源职业技术学院)

AI总结 提出一种端到端框架,结合UV体积渲染与扩散纹理生成器,并引入照明颜色一致性估计器和多尺度动态训练策略,生成可穿戴对抗图案,在无人机侦察等动态视角和光照变化下实现稳定物理攻击。

Comments Accepted by ICME 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.20162 2026-06-19 cs.AI cs.IT cs.NI math.IT 新提交

Implicit Semantic-Aware Communication Based on Hypergraph Reasoning

基于超图推理的隐式语义感知通信

Yiwei Liao, Shurui Tu, Yong Xiao, Yingyu Li, Guangming Shi

机构 * China Electric Power Research Institute Co., Ltd(中国电力科学研究院有限公司) National Key Laboratory for Power Grid Environmental Protection(电网环境保护国家重点实验室) School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) Peng Cheng Laboratory(鹏城实验室) Pazhou Laboratory (Huangpu)(琶洲实验室(黄埔)) School of Mechanical Engineering and Electronic Information, China University of Geosciences(中国地质大学机械与电子信息学院)

AI总结 提出基于超图的隐式语义推理框架HISR,通过超图建模多实体高阶关系,在噪声信道下提升语义推理鲁棒性,准确率提升36.6%。

Comments This work is accepted at IEEE Transactions on Communications

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.16575 2026-06-19 cs.LG math-ph math.MP 新提交

RepNN: Tackling spectral bias in deep neural networks via parameter reparameterization

RepNet:通过参数重参数化解决深度神经网络中的谱偏差

Yong Wang, Tao Zhou, Xuhui Meng

机构 * Institute of Interdisciplinary Research for Mathematics and Applied Science, School of Mathematics and Statistics, Huazhong University of Science and Technology(华中科技大学数学与统计学院交叉科学与应用数学研究所) Institute of Computational Mathematics, Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国科学院数学与系统科学研究院计算数学研究所)

AI总结 针对深度神经网络在捕捉振荡和多尺度行为时的谱偏差问题,提出RepNet模型,通过重参数化第一隐藏层的权重和偏置,有效控制初始斜率尺度和分区点分布,实现自适应频率缩放,在函数逼近、PDE求解和算子学习中显著提升精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.19195 2026-06-18 cs.CV 新提交

Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance

Moebius: 0.2B轻量级图像修复框架,性能达10B级别

Kangsheng Duan, Ziyang Xu, Wenyu Liu, Xiaohu Ruan, Xiaoxin Chen, Xinggang Wang

机构 * Huazhong University of Science and Technology(华中科技大学) VIVO AI Lab(VIVO人工智能实验室)

AI总结 提出Moebius轻量级图像修复框架,通过局部-λ混合交互模块和自适应多粒度蒸馏策略,以0.22B参数实现与10B级模型FLUX.1-Fill-Dev相当甚至更优的生成质量,推理速度提升15倍以上。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17667 2026-06-17 cs.LG cs.AI 新提交

Handling Feature Heterogeneity with Learnable Graph Patches

处理特征异质性:可学习图块方法

Yifei Sun, Yang Yang, Xiao Feng, Zijun Wang, Haoyang Zhong, Chunping Wang, Lei Chen

机构 * Zhejiang University(浙江大学) Huazhong University of Science and Technology(华中科技大学) Finvolution Group(信也科技集团)

AI总结 提出可学习图块概念,将图分解为语义单元,通过补丁编码器和聚合器实现跨域图数据的可迁移预训练,提升下游任务性能。

Comments Accepted at KDD 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17449 2026-06-17 cs.CL cs.AI cs.CV cs.LG cs.MM 新提交

MODE-RAG: Manifold Outlier Diagnosis and Energy-based Retrieval-Augmented Generation Evaluation

MODE-RAG: 基于流形异常诊断和能量的检索增强生成评估

Zehang Wei, Jiaxin Dai, Jiamin Yan, Xiang Xiang

机构 * School of Computer Science & Tech, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) School of AI and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)

AI总结 提出MODE-RAG多智能体系统,利用变分自由能和内部注意力状态动态门控干预,结合蒙特卡洛树搜索和logit扰动减少多模态检索增强生成中的幻觉和逻辑捏造。

Comments To be presented at ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.15869 2026-06-16 cs.CV 新提交

Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation

Metis: 一种用于自动驾驶和城市导航的通用高效世界-动作模型

Jingyu Li, Zhe Liu, Dongnan Hu, Junjie Wu, Zipei Ma, Wenxiao Wu, Chao Han, Zhihui Hao, Zhikang Liu, Kun Zhan, Jiankang Deng, Xiatian Zhu, Li Zhang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) The University of Hong Kong(香港大学) Tongji University(同济大学) Li Auto Inc.(理想汽车) Huazhong University of Science and Technology(华中科技大学) Imperial College London(伦敦帝国理工学院) University of Surrey(萨里大学)

AI总结 提出Metis框架,通过解耦视频生成与动作预测,采用混合专家架构和不对称注意力掩码,实现高效推理与泛化,在多个导航基准上取得最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.14120 2026-06-15 eess.SP cs.AI cs.LG cs.SD eess.AS 新提交

FAConformer: Frequency-Aware Convolutional Transformer for Auditory Attention Decoding

FAConformer:用于听觉注意解码的频率感知卷积Transformer

Ziwei Wang, Xingyi He, Tianwang Jia, Hongbin Wang, Dongrui Wu

机构 * Hubei Key Laboratory of Brain-inspired Intelligent Systems, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(湖北脑启发智能系统重点实验室,人工智能与自动化学院,华中科技大学)

AI总结 提出FAConformer框架,通过频带特定编码和自适应跨频带交互,有效利用脑电图频域信息进行听觉注意解码,在公开数据集上超越现有最佳模型4.9%。

Comments 15 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.13392 2026-06-15 cs.AI 新提交

MiniMax Sparse Attention

MiniMax 稀疏注意力

Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Jinkai Hu, Jiayao Li, Rui Gao, Zekun Li, Songquan Zhu, Jingkai Zhou, Pengyu Zhao

机构 * MiniMax Peking University(北京大学) NVIDIA(英伟达) Zhejiang University(浙江大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出 MiniMax 稀疏注意力(MSA),一种基于分组查询注意力的块级稀疏注意力机制,通过轻量索引分支选择 Top-k 键值块,实现高效长上下文处理,在 109B 模型上以 1M 上下文减少 28.4 倍注意力计算,并带来 14.2 倍预填充和 7.6 倍解码加速。

Comments 30 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12295 2026-06-11 cs.CV cs.CL cs.IR 新提交

Findings of the MAGMaR 2026 Shared Task

MAGMaR 2026 共享任务结果

Alexander Martin, Dengjia Zhang, Joel Brogan, Francis Ferraro, Jeremy Gwinnup, Reno Kriz, Teng Long, Kenton Murray, Andrew Yates, Xiang Xiang

机构 * Johns Hopkins University(约翰霍普金斯大学) OpenAI University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校) Air Force Research Laboratory(空军研究实验室) Human Language Technology Center of Excellence, Johns Hopkins University(约翰霍普金斯大学人类语言技术卓越中心) University of Amsterdam(阿姆斯特丹大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 本文介绍MAGMaR 2026共享任务的结果,包括视频检索和基于检索视频的生成任务,所有提交系统均超越去年基线。

Comments Findings of the 2nd workshop on Multimodal Augmented Generation via Multimodal Retrieval (MAGMaR); Resources at this url: https://github.com/rekriz11/MAGMAR_2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11874 2026-06-11 cs.AI 新提交

AutoMine Solution for AV2 2026 Scenario Mining Challenge

AutoMine 解决方案:面向 AV2 2026 场景挖掘挑战

Songliang Cao, Jiele Zhao, Yuru Wang, Hao Li, Daqi Liu, Zehan Zhang, Fangzhen Li, Yu Wang, Yue Zhang, Bing Wang, Guang Chen, Hao Lu, Hangjun Ye

机构 * Xiaomi EV(小米汽车) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出基于 LLM 和 VLM 的自优化场景挖掘方法 AutoMine,通过语义保持提示增强、鲁棒轨迹原子函数与 VLM 函数结合以及执行反馈优化,在 CVPR 2026 挑战赛中取得领先性能。

Comments CVPR 2026 Scenario Mining Challenge (Temporal Track Winners)

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10899 2026-06-10 cs.RO 新提交

MV-Actor: Aligning Multi-View Semantics and Spatial Awareness for Bimanual Manipulation

MV-Actor:对齐多视角语义与空间感知以实现双臂操作

Yinchen Tian, Huan Li, Muyao Peng, Xi Wang, Yan Wang, You Yang

机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院) Institute for AI Industry Research (AIR), Tsinghua University(清华大学智能产业研究院) AIR Wuxi Innovation Center, Tsinghua University(清华大学智能产业研究院无锡创新中心)

AI总结 提出MV-Actor框架,通过多视角语义交互和语义-空间令牌交互统一语义与空间表示,并利用引导度量深度修复模块处理深度噪声,在PerAct2基准上达到87.8%平均成功率。

Comments 14 pages,9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10395 2026-06-10 cs.CV 新提交

Efficient RWKV-based Representation Learning for 3D Point Clouds

基于高效RWKV的三维点云表示学习

Yun Liu, Xuefeng Yan, Liangliang Nan, Xianzhi Li, Peng Li, Zhe Zhu, Honghua Chen, Mingqiang Wei

机构 * School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) Shenzhen Institute of Research, Nanjing University of Aeronautics and Astronautics(南京航空航天大学深圳研究院) Collaborative Innovation Center of Novel Software Technology and Industrialization(新型软件技术与产业化协同创新中心) Urban Data Science section, Delft University of Technology(代尔夫特理工大学城市数据科学部) Huazhong University of Science and Technology(华中科技大学)

AI总结 提出P-RWKV模块,通过局部感知扩展和空间上下文增强,将RWKV从序列建模适配到3D点云,实现线性复杂度的全局依赖建模,在多项任务中以更低计算成本取得竞争性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.10363 2026-06-10 cs.RO 新提交

HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation

HiMem-WAM: 用于机器人操作的分层记忆门控世界动作模型

Xiaoquan Sun, Ruijian Zhang, Chen Cao, Yihan Sun, Jiahui Chen, Zetian Xu, Bo Chen, Haijier Chen, Zhen Yang, Jiarun Zhu, Yijun Hong, JingZhe Xu, Jingrui Pang, Mingqi Yuan, Jiayu Chen

机构 * The University of Hong Kong(香港大学) INFIFORCE Huazhong University of Science and Technology(华中科技大学) Tsinghua University(清华大学) Wuhan University(武汉大学) Southern University of Science and Technology(南方科技大学)

AI总结 提出分层记忆门控世界动作模型HiMem-WAM,通过分层潜在动作框架和边界触发记忆更新,提升长时域机器人操作的任务相关记忆与泛化鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏