arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Chinese Academy of Sciences(中国科学院大学)

共收录 1960
2506.23717 2025-12-02 cs.NE cs.AI cs.CV cs.LG

Towards Efficient and Accurate Spiking Neural Networks via Adaptive Bit Allocation

通过自适应位分配实现高效准确的脉冲神经网络

Xingting Yao, Qinghao Hu, Fei Zhou, Tielong Liu, Gang Li, Peisong Wang, Jian Cheng

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) School of Future Technology, University of Chinese Academy of Sciences(未来技术学院,中国科学院大学) China Electric Power Research Institute Co., Ltd(中国电力科学研究院有限公司)

AI总结 本文提出自适应位分配策略,通过改进神经元和机制优化,提升SNN的效率和精度,实现在ImageNet上2.69%的精度提升和4.16倍更低的比特预算。

Comments Neural Networks, In press

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20518 2025-12-02 cs.CV

Dynamic Attention Analysis for Backdoor Detection in Text-to-Image Diffusion Models

文本到图像扩散模型中后门检测的动态注意力分析

Zhongqi Wang, Jie Zhang, Shiguang Shan, Xilin Chen

机构 * Key Laboratory of AI Safety of CAS, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS), Beijing 100190, China, and also with the University of Chinese Academy of Sciences (UCAS), Beijing 100049, China(中国科学院人工智能安全重点实验室,计算技术研究所(ICT),中国科学院(CAS),北京100190,中国,以及中国科学院大学(UCAS),北京100049,中国)

AI总结 本研究提出动态注意力分析方法,通过分析扩散模型中注意力图的动态演变,有效检测文本到图像扩散模型中的后门攻击。

Comments Accepted by TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13047 2025-12-02 cs.CV

InsightDrive: Insight Scene Representation for End-to-End Autonomous Driving

InsightDrive: 用于端到端自动驾驶的洞察场景表示

Ruiqi Song, Xianda Guo, Yanlun Peng, Qinggong Wei, Hangbin Wu, Long Chen

机构 * College of Surveying and Geo-informatics, Tongji University(同济大学测绘与地理信息学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Computer Science, Wuhan University(武汉大学计算机学院) The School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Great Wall Motor(长城汽车) IAIR, Xi’an Jiaotong University(西安交通大学IAIR)

AI总结 InsightDrive通过结合显式和隐式场景表示,提升自动驾驶轨迹规划的鲁棒性和人类认知一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00342 2025-12-02 cs.LG cs.SY eess.SY

Adaptive prediction theory combining offline and online learning

结合离线与在线学习的自适应预测理论

Haizheng Li, Lei Guo

机构 * State Key Laboratory of Mathematical Sciences, Academy of Mathematics and Systems Science, Chinese Academy of Sciences, Beijing 100190, China(数学科学国家重点实验室,数学与系统科学研究院,中国科学院,北京) School of Mathematical Science, University of Chinese Academy of Sciences, Beijing 100049, China(数学科学学院,中国科学院大学,北京)

AI总结 本文提出一种结合离线和在线学习的两阶段框架,用于非线性随机动态系统的预测,通过理论分析和实验验证,提升了预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.17808 2025-12-02 cs.LG

Remote Sensing-Oriented World Model

面向遥感的世界模型

Yuxi Lu, Biao Wu, Zhidong Li, Kunqi Li, Chenya Huang, Huacan Wang, Qizhen Lan, Ronghao Chen, Ling Chen, Bin Liang

机构 * University of Technology Sydney (UTS)(悉尼技术大学) University of Chinese Academy of Sciences (UCAS)(中国科学院大学) University of Alabama at Birmingham(阿拉巴马大学伯明翰分校) Peking University(北京大学)

AI总结 本文提出首个面向遥感的世界模型框架,通过RemoteBAGEL模型在RSWISE基准上实现空间推理的高效外推。

Comments 10 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.23034 2025-12-01 cs.RO

LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models

LatBot: 从大规模物体操作视频中提炼通用潜在动作用于视觉-语言-动作模型

Zuolei Li, Xingyu Gao, Xiaofan Wang, Jianlong Fu

机构 * Institute of Microelectronics, Chinese Academy of Sciences(中国科学院微电子研究所) University of Chinese Academy of Sciences(中国科学院大学) Microsoft Research(微软研究院)

AI总结 LatBot通过整合动作预测和潜在动作分解,提升视觉-语言-动作模型在现实世界和模拟环境中的泛化与迁移能力。

Comments Project Page: https://mm-robot.github.io/distill_latent_action/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22892 2025-12-01 cs.CV cs.LG

ClearGCD: Mitigating Shortcut Learning For Robust Generalized Category Discovery

ClearGCD: 缓解捷径学习以实现稳健的通用类别发现

Kailin Lyu, Jianwei He, Long Xiao, Jianing Zeng, Liang Fan, Lin Shu, Jie Hao

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Loughborough University(洛桑大学)

AI总结 ClearGCD通过语义对齐和捷径抑制正则化缓解捷径学习,提升通用类别发现的稳健性和泛化能力。

Comments 5 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22434 2025-12-01 cs.CR cs.AI

FastFHE: Packing-Scalable and Depthwise-Separable CNN Inference Over FHE

FastFHE: 基于FHE的可扩展打包和深度可分离CNN推理

Wenbo Song, Xinxin Fan, Quanliang Jing, Shaoye Luo, Wenqi Wei, Chi Lin, Yunfeng Lu, Ling Liu

机构 * Institute of Computing Technology, CAS(中国科学院计算技术研究所) UCAS(中国科学院大学) Fordham University(福特汉姆大学) Dalian University of Technology(大连理工大学) Beihang University(北京航空航天大学) Georgia Institute of Technology(佐治亚理工学院)

AI总结 FastFHE通过可扩展加密打包、深度可分离卷积、BN点积融合和Legendre多项式近似,提升FHE下CNN推理效率与精度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22171 2025-12-01 cs.CV cs.GR

BrepGPT: Autoregressive B-rep Generation with Voronoi Half-Patch

BrepGPT: 基于Voronoi半块的自回归B-rep生成

Pu Li, Wenhao Zhang, Weize Quan, Biao Zhang, Peter Wonka, Dong-Ming Yan

机构 * MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) University of Chinese Academy of Sciences(中国科学院大学) King Abdullah University of Science and Technology(国王 Abdullah 科学与技术大学)

AI总结 BrepGPT通过Voronoi半块表示实现单阶段自回归B-rep生成,提升生成效率与模型紧凑性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22135 2025-12-01 cs.CV

EASL: Multi-Emotion Guided Semantic Disentanglement for Expressive Sign Language Generation

EASL: 多情绪引导的语义解耦用于表达性手语生成

Yanchao Zhao, Jihao Zhu, Yu Liu, Weizhuo Chen, Yuling Yang, Kun Peng

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) University of Chinese Academy of Sciences(中国科学院大学) University of Health and Rehabilitation Sciences(康复科学大学) The University of Aberdeen(阿伯丁大学)

AI总结 EASL通过多情绪引导的语义解耦架构,提升手语生成的表达性和情感表现力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21707 2025-12-01 cs.NI cs.AI

Sensing and Understanding the World over Air: A Large Multimodal Model for Mobile Networks

通过空气感知和理解世界:为移动网络设计的大型多模态模型

Zhuoran Duan, Yuhao Wei, Guoshun Nan, Zijun Wang, Yan Yan, Lihua Xiong, Yuhan Ran, Ji Zhang, Jian Li, Qimei Cui, Xiaofeng Tao, Tony Q. S. Quek

机构 * National Engineering Research Center for Mobile Network Technologies, Beijing University of Posts and Telecommunications (BUPT), Beijing(中国移动网络技术国家工程研究中心,北京邮电大学) Beiyou Shenzhen Institute(北邮深圳研究所) School of Cyber Security, University of Chinese Academy of Sciences (UCAS)(中国科学院大学网络安全学院) China Telecom Co., Ltd.(中国电信股份有限公司) Singapore University of Technology and Design (SUTD)(新加坡科技设计大学)

AI总结 本文提出了一种无线原生多模态模型,利用无线信号进行对比学习,验证了其在无线网络中的应用潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04182 2025-12-01 cs.CL cs.AI

COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs

COPO:面向MLLM幻觉的因果导向策略优化

Peizheng Guo, Jingyao Wang, Wenwen Qiang, Jiahuan Zhou, Changwen Zheng, Gang Hua

机构 * University of Chinese Academy of Sciences(中国科学院大学) Institute of Software Chinese Academy of Sciences(中国科学院软件研究所) Wangxuan Institute of Computer Technology, Peking University(北京大学王轩计算机技术研究所) Amazon.com, Inc.(亚马逊公司)

AI总结 COPO通过引入因果完整性奖励和GRPO框架,解决多模态大语言模型的幻觉问题,通过令牌级因果约束确保输出的正确性和证据基础。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01873 2025-12-01 cs.CV

DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization

DiffusionFF: 一种基于扩散的联合人脸伪造检测与细粒度特征定位框架

Siran Peng, Haoyuan Zhang, Li Gao, Tianshuo Zhang, Xiangyu Zhu, Bao Li, Weisong Zhao, Zhen Lei

机构 * MAIS, CASIA(CASIA人工智能研究所) SAI, UCAS(UCAS智能信息学院) CMFT(计算机视觉与模式识别技术研究所) IIE, CAS(中国科学院信息工程研究所) SCS, UCAS(UCAS安全学院) CAIR, HKISI, CAS(中国科学院自动化研究所) SCSE, FIE, M.U.S.T(慕苏尔科技大学安全与电子工程系)

AI总结 DiffusionFF通过结合预训练的伪造检测器和去噪扩散模型,实现人脸伪造检测与细粒度特征定位,提升检测能力和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19482 2025-12-01 cs.CL

KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models

KSHSeek: 通过数据驱动方法缓解和检测生成模型中的知识捷径幻觉

Zhongxin Liu, Zhiwei Wang, Jun Niu, Ying Li, Hongyu Sun, Meng Xu, He Wang, Gaofei Wu, Yuqing Zhang

机构 * Xidian University(西安电子科技大学) University of Chinese Academy of Sciences(中国科学院大学) Hainan University(海南大学) University of Waterloo(滑铁卢大学)

AI总结 KSHSeek通过数据驱动方法缓解和检测生成模型中的知识捷径幻觉,提升模型的鲁棒性和可靠性。

Comments 16 pages, 34 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.05538 2025-12-01 cs.CV

A Survey on Personalized Content Synthesis with Diffusion Models

扩散模型在个性化内容合成中的综述

Xulu Zhang, Xiaoyong Wei, Wentao Hu, Jinlin Wu, Jiaxin Wu, Wengyu Zhang, Zhaoxiang Zhang, Zhen Lei, Qing Li

机构 * Department of Computing(计算系) The Hong Kong Polytechnic University(香港理工大学) Center for Artificial Intelligence and Robotics(人工智能与机器人中心) Hong Kong Institute of Science & Innovation(香港科学创新研究院) Chinese Academy of Sciences(中国科学院) State Key Laboratory of Multimodal Artificial Intelligence Systems(多模态人工智能系统国家重点实验室) Chinese Academy of Sciences Institute of Automation(中国科学院自动化研究所) School of Artificial Intelligence(人工智能学院) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 本文综述了扩散模型在个性化内容合成中的应用,分析了测试时微调和预训练适应方法,探讨了个性化任务的挑战与创新,并提出了未来发展方向。

Journal ref Machine intelligence research, Oct. 2025, v. 22, no. 5, p. 817-848

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.08174 2025-12-01 cs.CV

OBSeg: Accurate and Fast Instance Segmentation Framework Using Segmentation Foundation Models with Oriented Bounding Box Prompts

OBSeg: 基于定向边界框提示的准确且快速的实例分割框架

Zhen Zhou, Junfeng Fan, Yunkai Ma, Sihan Zhao, Fengshui Jing, Min Tan

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)

AI总结 OBSeg通过引入定向边界框提示,提升了实例分割的准确性和速度,同时优化了轻量级基础模型的性能。

Journal ref Machine Intelligence Research 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21460 2025-11-27 cs.AI

MADRA: Multi-Agent Debate for Risk-Aware Embodied Planning

MADRA:多智能体辩论用于风险感知的具身规划

Junjian Wang, Lidan Zhao, Xi Sheryl Zhang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) University of Chinese Academy of Sciences, Nanjing(中国科学院大学南京校区)

AI总结 MADRA通过多智能体辩论和分层认知框架,在安全性和效率上超越现有方法,实现高安全任务拒绝率和低误拒率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21150 2025-11-27 cs.CV cs.AI

LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs

LLaVA-UHD v3:渐进式视觉压缩用于多模态大语言模型中的高效原分辨率编码

Shichu Sun, Yichen Zhang, Haolin Song, Zonghao Guo, Chi Chen, Yidan Zhang, Yuan Yao, Zhiyuan Liu, Maosong Sun

机构 * Tsinghua University(清华大学) University of Chinese Academy of Sciences(中国科学院大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航天信息研究所) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院)

AI总结 LLaVA-UHD v3通过渐进式视觉压缩方法,实现高效原分辨率编码,减少TTFT并提升多模态大语言模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21132 2025-11-27 cs.CV

DeepRFTv2: Kernel-level Learning for Image Deblurring

DeepRFTv2:基于核级的学习图像去模糊

Xintian Mao, Haofei Song, Yin-Nian Liu, Qingli Li, Yan Wang

机构 * Shanghai Key Laboratory of Multidimensional Information Processing, East China Normal University(上海多维信息处理重点实验室,华东师范大学) State Key Laboratory of Infrared Physics, Shanghai Institute of Technical Physics, Chinese Academy of Sciences(红外物理国家重点实验室,上海技术物理研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 DeepRFTv2通过核级学习提升图像去模糊性能,引入傅里叶核估计器和解耦多尺度架构,实现低复杂度的核级模糊学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20986 2025-11-27 cs.CV

Inversion-Free Style Transfer with Dual Rectified Flows

无反向过程的双修正流风格迁移

Yingying Deng, Xiangyu He, Fan Tang, Weiming Dong, Xucheng Yin

机构 * Department of Computer Science and Technology, University of Science and Technology Beijing(北京科技大学计算机科学与技术系) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)

AI总结 本文提出无反向过程的双修正流框架,通过并行预测和动态中点插值实现高效风格迁移,提升视觉保真度和计算效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09474 2025-11-27 cs.SE cs.AI cs.OS

MigGPT: Harnessing Large Language Models for Automated Migration of Out-of-Tree Linux Kernel Patches Across Versions

MigGPT:利用大语言模型实现跨版本的Linux内核非官方补丁自动化迁移

Pucheng Dang, Di Huang, Dong Li, Kang Chen, Yuanbo Wen, Qi Guo, Xing Hu

机构 * State Key Lab of Processors, Institute of Computing Technology, CAS(处理器国家重点实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) Zhongguancun Laboratory(中关村实验室) Tsinghua University(清华大学)

AI总结 MigGPT通过新颖的代码指纹结构和三个精心设计的模块,提高非官方内核补丁迁移的准确性和效率,平均完成率达74.07。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11381 2025-11-27 cs.CV cs.AI

Without Paired Labeled Data: End-to-End Self-Supervised Learning for Drone-view Geo-Localization

无需配对标注数据:无人机视角地理定位的端到端自监督学习

Zhongwei Chen, Zhao-Xu Yang, Hai-Jun Rong, Guoqi Li

机构 * State Key Laboratory for Strength and Vibration of Mechanical Structures(强度与振动机械结构国家重点实验室) Shaanxi Key Laboratory of Environment and Control for Flight Vehicle(陕西省飞行器环境与控制重点实验室) School of Aerospace Engineering(航空航天工程学院) Xi’an Jiaotong University(西安交通大学) Institute of Automation(自动化研究所) Chinese Academy of Sciences(中国科学院) School of Artificial Intelligence(人工智能学院) University of Chinese Academy of Sciences(中国科学院大学) Peng Cheng Laboratory(鹏城实验室)

AI总结 本文提出DMNIL方法,通过动态记忆和邻域信息学习提升无人机视角地理定位的自监督学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17464 2025-11-27 cs.LG cs.AI stat.ML

Data Valuation by Fusing Global and Local Statistical Information

通过融合全局和局部统计信息进行数据估值

Xiaoling Zhou, Ou Wu, Michael K. Ng, Hao Jiang

机构 * Peking University(北京大学) University of Chinese Academy of Sciences(中国科学院大学) Hong Kong Baptist University(香港 Baptist大学) Renmin University of China(中国人民大学)

AI总结 本文提出融合全局和局部统计信息的增强数据估值方法,通过正则化项优化Shapley值估计,并引入动态数据估值技术提升计算效率。

Comments 35 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20307 2025-11-26 cs.CV

TReFT: Taming Rectified Flow Models For One-Step Image Translation

TReFT: 修正流模型用于一步图像翻译

Shengqian Li, Ming Gao, Yi Liu, Zuzeng Lin, Feng Wang, Feng Dai

机构 * University of Chinese Academy of Sciences(中国科学院大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Beihang University(北航) Tianjin University(天津大学) CreateAI

AI总结 TReFT通过简化架构和优化损失函数,实现修正流模型在一步图像翻译中的高效实时应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20280 2025-11-26 cs.CV

Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement

通过VLM引导的迭代自优化提升物理导向的视频生成

Yang Liu, Xilin Zhao, Peisong Wen, Siran Dai, Qingming Huang

机构 * School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院) School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

AI总结 本文提出一种通过VLM引导的迭代自优化方法,提升视频生成的物理一致性,实验显示在PhyIQ基准上得分提升明显。

Comments ICCV 2025 Physics-IQ Challenge Third Place Solution

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19835 2025-11-26 cs.CV cs.AI

Rectified SpaAttn: Revisiting Attention Sparsity for Efficient Video Generation

校正SpaAttn:重新审视注意力稀疏性以实现高效的视频生成

Xuewen Liu, Zhikai Li, Jing Zhang, Mengjuan Chen, Qingyi Gu

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

AI总结 本文提出Rectified SpaAttn,通过校正注意力分配提升视频生成效率,实现显著速度提升且保持生成质量。

Comments Code at https://github.com/BienLuky/Rectified-SpaAttn

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.01007 2025-11-26 cs.CV

GMT: Effective Global Framework for Multi-Camera Multi-Target Tracking

GMT:多摄像头多目标跟踪的有效全局框架

Yihao Zhen, Mingyue Xu, Qiang Wang, Baojie Fan, Jiahua Dong, Tinghui Zhao, Huijie Fan

机构 * State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences(机器人与智能系统国家重点实验室,沈阳自动化研究所,中国科学院) Key Laboratory of Manufacturing Industrial Integrated Automation, Shenyang University(制造工业集成自动化重点实验室,沈阳大学) University of Chinese Academy of Sciences(中国科学院大学) Nanjing University of Posts and Telecommunications(南京邮电大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

AI总结 GMT提出一种全局框架,通过联合利用视内和视间线索,提升多摄像头多目标跟踪的性能,实验表明其在CVMA和CVIDF1指标上分别提升21.3%和17.2%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18711 2025-11-25 cs.CV cs.AI

Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation

模态协同低秩分解器用于少样本视频域适应

Yuyang Wanyan, Xiaoshan Yang, Weiming Dong, Changsheng Xu

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室,自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) PengCheng Laboratory(鹏城实验室)

AI总结 本文提出模态协同低秩分解器,通过分解不同领域偏移级别的模态特征,提升少样本视频域适应的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.17405 2025-11-25 cs.CL cs.AI

Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT

超越多项选择:可验证的开放问答用于鲁棒的视觉-语言 RFT

Yesheng Liu, Hao Li, Haiyu Xu, Baoqi Pei, Jiahao Wang, Mingxuan Zhao, Jingshu Zheng, Zheqi He, JG Yao, Bowen Qin, Xi Yang, Jiajun Zhang

机构 * Institute of Automation, CAS(中国科学院自动化研究所) School of Artificial Intelligence, UCAS(中国科学技术大学人工智能学院) BAAI FlagEval Team(百度AI旗评团队) BUAA(北京航空航天大学) PKU(北京大学) ZJU(浙江大学)

AI总结 ReVeL通过重写和验证多项选择问题为开放式问题,提升视觉-语言模型在鲁棒性与数据效率上的表现。

Comments Project url: https://flageval-baai.github.io/ReVeL/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18028 2025-11-25 cs.CV

MambaX: Image Super-Resolution with State Predictive Control

MambaX:基于状态预测控制的图像超分辨率

Chenyu Li, Danfeng Hong, Bing Zhang, Zhaojie Pan, Naoto Yokoya, Jocelyn Chanussot

机构 * Southeast University(东南大学) Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航空信息研究所) College of Resources and Environment, University of Chinese Academy of Sciences(中国科学院大学资源与环境学院) School of Mathematics, Southeast University(东南大学数学学院) Department of Complexity Science and Engineering, Graduate School of Frontier Sciences, the University of Tokyo(东京大学前沿科学研究院复杂科学与工程部门) Univ. Grenoble Alpes, Inria, CNRS, Grenoble INP, LJK(格勒诺布尔阿尔卑斯大学、Inria、CNRS、Grenoble INP、LJK)

AI总结 MambaX通过动态状态预测控制和多模态融合方法提升图像超分辨率性能。

详情

展开后加载摘要…

URL PDF HTML 收藏