arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Science and Technology of China(中国科学技术大学)

共收录 2226
2603.16210 2026-03-18 cs.AI

MOSAIC: Composable Safety Alignment with Modular Control Tokens

MOSAIC:通过模块化控制令牌实现可组合的安全对齐

Jingyu Peng, Hongyu Chen, Jiancheng Dong, Maolin Wang, Wenxi Li, Yuchen Li, Kai Zhang, Xiangyu Zhao

机构 * University of Science and Technology of China(中国科学技术大学) City University of Hong Kong(香港城市大学) Baidu Inc(百度公司) Minzu University of China(民族大学)

AI总结 MOSAIC通过可学习的控制令牌实现可组合的安全对齐,解决传统方法在动态安全规则下的局限性,实验显示其在降低过拒绝率的同时保持模型效用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16129 2026-03-18 cs.CV

Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting

提升零样本物体计数的定量与空间意识

Da Zhang, Bingyu Li, Feiyu Wang, Zhiyuan Zhao, Junyu Gao

机构 * Northwestern Polytechnical University(西北工业大学) Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院) University of Science and Technology of China(中国科学技术大学) Fudan University(复旦大学)

AI总结 本文提出QICA框架,通过引入协同提示策略和成本聚合解码器,提升零样本物体计数的定量感知和空间聚合能力,实验表明其在多个数据集上具有优越的泛化性能。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14927 2026-03-18 cs.GR cs.LG

Masked BRep Autoencoder via Hierarchical Graph Transformer

基于分层图变换器的掩码BRep自编码器

Yifei Li, Kang Wu, Wenming Wu, Xiao-Ming Fu

机构 * University of Science and Technology of China(科学技术大学) Hefei University of Technology(合肥工业大学)

AI总结 本文提出一种自监督学习框架,通过自动学习CAD模型的表示以提升下游任务性能,采用掩码图自编码器和分层图变换器架构,实验表明模型在少量标注数据下表现优异。

Comments 27 pages, 11 figures. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12633 2026-03-18 cs.CV cs.AI

DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model

DiG:通过差异 grounding 提升多模态大语言模型的细粒度感知

Zhou Tao, Shida Wang, Yongxiang Hua, Haoyu Cao, Linli Xu

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)

AI总结 本文提出DiG框架,通过学习相似图像对的差异识别提升多模态大语言模型的细粒度感知能力,实验表明其在多个视觉感知基准上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07923 2026-03-18 cs.CV cs.AI

Exploring the Underwater World Segmentation without Extra Training

探索无需额外训练的水下世界分割

Bingyu Li, Tao Huo, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom, China(人工智能研究院(TeleAI),中国电信,中国) University of Science and Technology of China, China(中国科学技术大学,中国) Northwestern Polytechnical University, China(西北工业大学,中国)

AI总结 本文提出AquaOV255水下分割数据集及Earth2Ocean框架,通过几何引导视觉掩码生成和类别-视觉语义对齐模块,在无需额外训练的情况下实现高效的水下开放词汇分割。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15398 2026-03-18 cs.CV cs.AI

MARIS: Marine Open-Vocabulary Instance Segmentation with Geometric Enhancement and Semantic Alignment

MARIS: 基于几何增强与语义对齐的海洋开放词汇实例分割

Bingyu Li, Feiyu Wang, Da Zhang, Zhiyuan Zhao, Junyu Gao, Xuelong Li

机构 * University of Science and Technology of China(中国科学技术大学) Institute of Artificial Intelligence (TeleAI)(人工智能研究院(TeleAI)) China Telecom(中国电信) Fudan University(复旦大学) Northwestern Polytechnical University(西北工业大学)

AI总结 本文提出MARIS,首个大规模细粒度水下开放词汇分割基准,通过几何先验增强模块和语义对齐注入机制提升水下场景下的实例分割性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09551 2026-03-18 cs.CV

Label-supervised surgical instrument segmentation using temporal equivariance and semantic continuity

基于时间等变性和语义连续性的标注监督手术器械分割

Qiyuan Wang, Yanzhe Liu, Shang Zhao, Rong Liu, S. Kevin Zhou

机构 * School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China(中国科学技术大学生物医学工程学院,生命科学与医学系) Center for Medical Imaging, Robotics, Analytic Computing & Learning(MIRACLE), Suzhou Institute for Advanced Research, University of Science and Technology of China(中国科学技术大学苏州先进研究院医学影像中心、机器人与分析计算与学习中心) Key Laboratory of Precision and Intelligent Chemistry, University of Science and Technology of China(中国科学技术大学精密与智能化学重点实验室) Key Lab of Intelligent Information Processing of Chinese Academy of Sciences(CAS), Institute of Computing Technology, CAS(中国科学院智能信息处理重点实验室,计算技术研究所) Faculty of Hepato-Biliary-Pancreatic Surgery, The First Medical Center, Chinese PLA General Hospital(中国人民解放军总医院肝胆胰外科系)

AI总结 本文提出一种两阶段弱监督分割方法,通过时间等变性和语义连续性提升手术视频中器械分割的准确性,实验验证了其在两个手术数据集上的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15539 2026-03-17 cs.LG

Vib2ECG: A Paired Chest-Lead SCG-ECG Dataset and Benchmark for ECG Reconstruction

Vib2ECG:一种配对的胸导体SCG-ECG数据集和用于ECG重建的基准

Guorui Lu, Xiaohui Cai, Todor Stefanov, Qinyu Chen

机构 * Leiden Institute of Advanced Computer Science (LIACS), Leiden University(莱顿先进计算机科学研究所(LIACS)、莱顿大学) School of Computer Science and Technology, University of Science and Technology of China(中国科学技术大学计算机科学与技术学院)

AI总结 本文提出Vib2ECG数据集,用于从低成本振动信号重建ECG,展示了轻量U-Net在多导联ECG重建中的可行性,并分析了模型生成虚假ECG波形的问题。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15302 2026-03-17 cs.CV

Generative Video Compression with One-Dimensional Latent Representation

基于一维潜在表示的生成视频压缩

Zihan Zheng, Zhaoyang Jia, Naifu Xue, Jiahao Li, Bin Li, Zongyu Guo, Xiaoyi Zhang, Zhenghao Chen, Houqiang Li, Yan Lu

机构 * University of Science and Technology of China(中国科学技术大学) Communication University of China(通信大学) Microsoft Research Asia(微软亚洲研究院) University of Newcastle(新castle大学)

AI总结 本文提出GVC1D方法,通过一维潜在表示替代二维网格,减少空间和时间冗余,实现更高效的视频压缩。

Comments CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.15020 2026-03-17 cs.CV cs.CL

MER-Bench: A Comprehensive Benchmark for Multimodal Meme Reappraisal

MER-Bench:多模态表情包再解释的综合基准

Yiqi Nie, Fei Wang, Junjie Chen, Kun Li, Yudi Cai, Dan Guo, Chenglong Li, Meng Wang

机构 * School of Artificial Intelligence, Anhui University(安徽大学人工智能学院) Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) College of Information Technology, United Arab Emirates University(阿拉伯联合酋长国大学信息学院) Institute of Advanced Technology, University of Science and Technology of China(中国科学技术大学先进技术研究院)

AI总结 本文提出MER-Bench,一个用于多模态表情包再解释的综合基准,通过多模态大语言模型评估结构保持、语义一致性和情感转换,揭示现有系统在约束满足上的不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14976 2026-03-17 cs.MM cs.CV

Anchoring Emotions in Text: Robust Multimodal Fusion for Mimicry Intensity Estimation

在文本中锚定情绪:用于模仿强度估计的鲁棒多模态融合

Lingsi Zhu, Yuefeng Zou, Yunxiang Zhang, Naixiang Zheng, Guoyuan Wang, Jun Yu, Jiaen Liang, Wei Huang, Shengping Liu, Ximin Zheng

机构 * University of Science and Technology of China(中国科学技术大学) Unisound AI Technology Co., Ltd.(Unisound人工智能科技有限公司) Pingan Technology Co., Ltd.(平安科技有限公司)

AI总结 本文提出TAEMI框架,通过文本锚定机制融合多模态数据,提升在噪声和缺失数据下的情绪模仿强度估计性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14917 2026-03-17 eess.AS cs.AI cs.LG eess.SP

Spectrogram features for audio and speech analysis

基于频谱图的音频和语音分析特征

Ian McLoughlin, Lam Pham, Yan Song, Xiaoxiao Miao, Huy Phan, Pengfei Cai, Qing Gu, Jiang Nan, Haoyu Song, Donny Soh

机构 * Singapore Institute of Technology(新加坡理工学院) Austrian Institute of Technology(奥地利理工学院) The University of Science and Technology of China(中国科学技术大学) Meta Inc., Reality Labs(Meta公司,现实实验室)

AI总结 本文探讨频谱图在音频和语音分析中的应用,回顾其在不同任务中与后端分类器架构的匹配性,总结其在深度学习中的优势与挑战。

Comments 30 pages

Journal ref Analysis. Appl. Sci. 2026, 16, 572

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13361 2026-03-17 cs.CV cs.AI stat.ML

BrainCast: A Spatio-Temporal Forecasting Model for Whole-Brain fMRI Time Series Prediction

BrainCast:一种用于全脑fMRI时间序列预测的时空预测模型

Yunlong Gao, Jinbo Yang, Li Xiao, Haiye Huo, Yang Ji, Hao Wang, Aiying Zhang, Yu-Ping Wang

机构 * Institute of Advanced Technology, University of Science and Technology of China(科学技术大学先进技术研究院) MoE Key Laboratory of Brain-Inspired Intelligence Perception and Cognition, University of Science and Technology of China(科学技术大学脑启发智能感知与认知教育部重点实验室) School of Mathematics and Computer Sciences, Nanchang University(南昌大学数学与计算机科学学院) School of Data Science, University of Virginia(弗吉尼亚大学数据科学学院) Department of Biomedical Engineering, Tulane University(路易斯安那大学医学工程系)

AI总结 本文提出BrainCast模型,通过联合建模感兴趣区域内的时序动态和跨区域的空间交互,提升全脑fMRI时间序列预测性能,从而提升下游认知能力预测。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.13349 2026-03-17 cs.CV cs.AI

MURE: Hierarchical Multi-Resolution Encoding via Vision-Language Models for Visual Document Retrieval

MURE:基于视觉-语言模型的多分辨率编码框架用于视觉文档检索

Fengbin Zhu, Zijing Cai, Yuzhe Wang, Pengyang Shao, Wenjie Wang, Fuli Feng, Richang Hong, Tat-Seng Chua

机构 * National University of Singapore, Singapore(新加坡国立大学) University of Science and Technology of China, China(中国科学技术大学) Hefei University of Technology, China(合肥工业大学)

AI总结 本文提出MURE框架,通过多分辨率采样编码、跨粒度特征融合和语义感知聚类机制,提升视觉文档检索效率与效果,实验表明其优于现有基线方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01050 2026-03-17 cs.CV cs.AI cs.GR

EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos

EgoGrasp:从第一人称视频中估计世界空间手-物体交互

Hongming Fu, Wenjia Wang, Xiaozhen Qiao, Rolandos Alexandros Potamias, Taku Komura, Shuo Yang, Zheng Liu, Bo Zhao

机构 * Shanghai Jiao Tong University(上海交通大学) The University of Hong Kong(香港大学) University of Science and Technology of China(中国科学技术大学) Imperial College London(伦敦帝国学院) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Beijing Academy of Artificial Intelligence(北京人工智能研究院)

AI总结 EgoGrasp提出了一种从动态第一人称视频中重建世界空间手-物体交互的方法,支持开放词汇物体,通过多阶段框架实现鲁棒的3D重建与姿态估计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12648 2026-03-16 cs.CV

From Sparse to Dense: Multi-View GRPO for Flow Models via Augmented Condition Space

从稀疏到密集:通过增强条件空间的多视图GRPO用于流模型

Jiazi Bu, Pengyang Ling, Yujie Zhou, Yibin Wang, Yuhang Zang, Tianyi Wei, Xiaohang Zhan, Jiaqi Wang, Tong Wu, Xingang Pan, Dahua Lin

机构 * Shanghai Jiao Tong University(上海交通大学) S-Lab, Nanyang Technological University(南洋理工大学S实验室) University of Science and Technology of China(中国科学技术大学) Fudan University(复旦大学) The Chinese University of Hong Kong(香港中文大学) Shanghai AI Laboratory(上海人工智能实验室) Adobe Research(Adobe研究实验室) Stanford University(斯坦福大学) Shanghai Innovation Institute(上海创新研究院) CPII under InnoHK Project(InnoHK项目下的CPII)

AI总结 本文提出多视图GRPO,通过增强条件空间探索样本间关系,提升文本到图像流模型的对齐性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12557 2026-03-16 cs.LG cs.CV

Lyapunov Stable Graph Neural Flow

Lyapunov稳定图神经流

Haoyu Chu, Xiaotong Chen, Wei Zhou, Wenjun Cui, Kai Zhao, Shikui Wei, Qiyu Kang

机构 * School of Computer Science and Technology / School of Artificial Intelligence, China University of Mining and Technology(计算机科学与技术学院/人工智能学院,中国矿业大学) School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学)

AI总结 本文将控制理论与图神经网络结合,提出基于整数和分数阶Lyapunov稳定性的新防御框架,通过可学习的Lyapunov函数和新型投影机制提升GNN的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12108 2026-03-13 cs.CV

EvoTok: A Unified Image Tokenizer via Residual Latent Evolution for Visual Understanding and Generation

EvoTok:通过残差潜在进化实现统一的图像分词器用于视觉理解和生成

Yan Li, Ning Liao, Xiangyu Zhao, Shaofeng Zhang, Xiaoxing Wang, Yifan Yang, Junchi Yan, Xue Yang

机构 * University of Science and Technology of China(中国科学技术大学) Microsoft Corporation(微软公司)

AI总结 EvoTok通过残差潜在进化统一图像理解和生成,实现高效视觉表征建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11607 2026-03-13 cs.CV

DyWeight: Dynamic Gradient Weighting for Few-Step Diffusion Sampling

DyWeight: 动态梯度加权用于少步扩散采样

Tong Zhao, Mingkun Lei, Liangyu Yuan, Yanming Yang, Chenxi Song, Yang Wang, Beier Zhu, Chi Zhang

机构 * Zhejiang University(浙江大学) AGI Lab, Westlake University(西湖大学AGI实验室) TongJi University(同济大学) University of Science and Technology of China(中国科学技术大学)

AI总结 DyWeight通过动态梯度加权方法,提高了扩散采样效率,实现更少的函数评估和更稳定的生成效果。

Comments Code Link: see AGI-Lab/DyWeight" target="_blank" rel="noopener">https://github.com/Westlake-AGI-Lab/DyWeight

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11306 2026-03-13 cs.CV

Hierarchical Granularity Alignment and State Space Modeling for Robust Multimodal AU Detection in the Wild

层次粒度对齐与状态空间建模用于野外多模态面部动作单元检测

Jun Yu, Yunxiang Zhang, Naixiang Zheng, Lingsi Zhu, Guoyuan Wang

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出了一种基于层次粒度对齐和状态空间模型的多模态框架,通过强大的基础模型提取高保真视觉和音频表示,并引入视觉-马尔可夫模型和不对称交叉注意机制,实现对野外环境中面部动作单元的高效检测。

Comments 8 pages, 1 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18463 2026-03-13 cs.CV

Decoupling Perception from Reasoning for Hallucination-Resistant Video Understanding

解耦感知与推理以实现抗幻觉的视频理解

Bowei Pu, Chuanbin Liu, Yifan Ge, Peicheng Zhou, Yiwei Sun, Zhiying Lu, Zhangchi Hu, Hongtao Xie

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出DPL模型,通过解耦感知与推理以提升视频理解的抗幻觉能力,引入感知奖励和FAE评估器,有效提高训练后性能和数据效率。

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19459 2026-03-13 cs.LG cs.AI

Your Classifier Can Do More: Towards Balancing the Gaps in Classification, Robustness, and Generation

你的分类器可以做更多:朝着平衡分类、鲁棒性和生成之间的差距

Kaichao Jiang, He Wang, Xiaoshuai Hao, Xiulong Yang, Ajian Liu, Qi Chu, Yunfeng Diao, Richang Hong

机构 * Hefei University of Technology, China(合肥工业大学,中国) Jianghuai Advance Technology Center, China(江淮先进科技中心,中国) Anhui Provincial Key Laboratory of Humanoid Robots, China(安徽省人形机器人重点实验室,中国) AI Centre, University College London, UK(伦敦大学学院人工智能中心,英国) Xiaomi EV, China(小米电动车,中国) Central China Normal University, China(中部师范大学,中国) Institute of Automation, Chinese Academy of Sciences, China(中国科学院自动化研究所,中国) University of Science and Technology of China, China(中国科学技术大学,中国)

AI总结 EB-JDAT通过统一生成、判别和鲁棒性,实现分类、鲁棒性和生成能力的平衡,达到新的权衡前沿。

Comments accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10975 2026-03-12 cs.CV

VCR: Variance-Driven Channel Recalibration for Robust Low-Light Enhancement

VCR: 基于方差驱动的通道重校准用于鲁棒低光照增强

Zhixin Cheng, Fangwen Zhang, Xiaotian Yin, Baoqun Yin, Haodian Wang

机构 * School of Information Science and Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学) Institute of Advanced Technology, University of Science and Technology of China(先进科技研究院,中国科学技术大学) CHN Energy Digital Intelligence Technology Development (Beijing) Co., LTD.(中能数字智能技术发展(北京)有限公司)

AI总结 VCR通过方差驱动的通道重校准提升低光照图像增强的鲁棒性和感知质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10473 2026-03-12 cs.CL cs.AI

Aligning Large Language Models with Searcher Preferences

将大型语言模型对齐于搜索者偏好

Wei Wu, Peilun Zhou, Liyi Chen, Qimeng Wang, Chengqiang Lu, Yan Gao, Yi Wu, Yao Hu, Hui Xiong

机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(中国科学技术大学人工智能与数据科学学院) Xiaohongshu Inc.(小红书公司) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能研究所) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

AI总结 SearchLLM通过分层多维奖励系统提升开放式生成搜索的鲁棒性和用户需求对齐能力,实测有效消费率提升1.03%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13093 2026-03-12 cs.RO cs.LG

PvP: Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations

PvP: 通过体感优先对比表示实现数据高效的类人机器人学习

Mingqi Yuan, Tao Yu, Haolin Song, Bo Li, Xin Jin, Hua Chen, Wenjun Zeng

机构 * HK PolyU(香港理工大学) LimX Dynamics EIT, Ningbo(宁波工程学院) USTC(中国科学技术大学) ZJU-UIUC(浙江大学-UIUC) ZGCA(浙江工业大学)

AI总结 PvP通过体感优先对比学习提升类人机器人学习的样本效率和性能。

Comments 15 pages, 17 figures

Journal ref 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.12229 2026-03-12 cs.CV

Ultra-Low Bitrate Perceptual Image Compression with Shallow Encoder

超低比特率感知图像压缩与浅层编码

Tianyu Zhang, Dong Liu, Chang Wen Chen

机构 * University of Science and Technology of China(中国科学技术大学) The Hong Kong Polytechnic University(香港理工大学)

AI总结 本文提出了一种非对称极值图像压缩框架,利用浅层编码器和单步扩散解码器实现超低比特率下的高保真度重建,同时提升编码效率和解码速度。

Comments Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23719 2026-03-12 cs.CV

PD-Diag-Net: Clinical-Priors guided Network on Brain MRI for Auxiliary Diagnosis of Parkinson's Disease

PD-Diag-Net:基于临床先验知识的脑MRI辅助帕金森病诊断网络

Shuai Shao, Yan Wang, Shu Jiang, Shiyuan Zhao, Di Yang, Jiangtao Wang, Yutong Bai, Jianguo Zhang

机构 * Suzhou Institute for Advanced Research, University of Science and Technology of China(中国科学技术大学苏州市先进研究院) School of Artificial Intelligence and Data Science, University of Science and Technology of China(中国科学技术大学人工智能与数据科学学院) School of Electronic and Information Engineering, Beijing Jiaotong University(北京交通大学电子与信息工程学院) College of Control Science and Engineering, China University of Petroleum (East China)(中国石油大学(华东)控制科学与工程学院) School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院) Beijing Tiantan Hospital(北京天坛医院)

AI总结 PD-Diag-Net通过结合临床先验知识和MRI预处理,实现了基于脑MRI的帕金森病辅助诊断,准确率超过96%

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09877 2026-03-11 cs.CV

InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing

InternVL-U: 使统一多模态模型在理解、推理、生成和编辑方面更加普及

Changyao Tian, Danni Yang, Guanzhou Chen, Erfei Cui, Zhaokai Wang, Yuchen Duan, Penghao Yin, Sitao Chen, Ganlin Yang, Mingxin Liu, Zirun Zhu, Ziqian Fan, Leyao Gu, Haomin Wang, Qi Wei, Jinhui Yin, Xue Yang, Zhihang Zhong, Qi Qin, Yi Xin, Bin Fu, Yihao Liu, Jiaye Ge, Qipeng Guo, Gen Luo, Hongsheng Li, Yu Qiao, Kai Chen, Hongjie Zhang

机构 * Shanghai AI Laboratory(上海人工智能实验室) Fudan University(复旦大学) University of Science and Technology of China(中国科学技术大学) Shanghai Jiao Tong University(上海交通大学) South China University of Technology(华南理工大学) Nanjing University(南京大学) Xiamen University(厦门大学) CUHK MMLab(香港中文大学MMLab)

AI总结 InternVL-U通过轻量级统一多模态模型,在保持强大生成能力的同时,实现了在理解、推理、生成和编辑方面的高效平衡。

Comments technical report, 61 pages, https://github.com/OpenGVLab/InternVL-U

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09712 2026-03-11 cs.RO

Robotic Scene Cloning:Advancing Zero-Shot Robotic Scene Adaptation in Manipulation via Visual Prompt Editing

机器人场景克隆:通过视觉提示编辑推进零样本机器人场景适应在操作中

Binyuan Huang, Yuqing Wen, Yucheng Zhao, Yaosi Hu, Tiancai Wang, Chang Wen Chen, Haoqiang Fan, Zhenzhong Chen

机构 * Wuhan University(武汉大学) University of Science and Technology of China(中国科学技术大学) Dexmal The Hong Kong Polytechnic University(香港理工大学)

AI总结 机器人场景克隆通过视觉提示编辑提升零样本机器人场景适应能力,实现跨场景的高效操作性能

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09527 2026-03-11 cs.LG cs.AI

Efficiently Aligning Draft Models via Parameter- and Data-Efficient Adaptation

通过参数和数据高效适应实现草稿模型的高效对齐

Luxi Lin, Zhihang Lin, Zhanpeng Zeng, Yuhao Chen, Qingyu Zhang, Jixiang Luo, Xuelong Li, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(中国教育部多媒体可信感知与高效计算重点实验室,厦门大学) Shanghai Innovation Institute(上海创新研究院) Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究所(TeleAI)) University of Science and Technology of China (USTC)(中国科学技术大学)

AI总结 本文提出高效草稿适应框架EDA,通过解耦架构、数据再生策略和样本选择机制,实现草稿模型在特定领域微调时的高效适应,降低训练成本并提升推测解码性能。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏