arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

University of Chinese Academy of Sciences(中国科学院大学)

共收录 1952
2605.07414 2026-05-11 cs.MA cs.AI cs.CR

OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing

OrchJail:通过协调引导的模糊测试对工具调用文本到图像代理进行劫持

Jianming Chen, Yawen Wang, Junjie Wang, Zhe Liu, Qing Wang, Fanjiang Xu

机构 * Institute of Software, Chinese Academy of Sciences, Beijing, China Science \& Technology on Integrated Information System Laboratory, Beijing, China State Key Laboratory of Complex System Modeling University of Chinese Academy of Sciences, Beijing, China

AI总结 研究提出OrchJail框架,通过协调引导的模糊测试方法提升对工具调用文本到图像代理的劫持效果,实现更高的攻击成功率和图像质量,同时降低查询成本。

Journal ref ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07301 2026-05-11 cs.AI

SOM: Structured Opponent Modeling for LLM-based Agents via Structural Causal Model

基于结构因果模型的对手建模:通过结构因果模型实现基于大语言模型的智能体

Shiyue Cao, Pei Xu, Likun Yang, Lei Cui, Xiaotang Chen, Kaiqi Huang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences \& Institute of Automation, Chinese Academy of Sciences Beijing China National Key Laboratory of Cognition Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences Beijing China School of Artificial Intelligence, University of Chinese Academy of Sciences Beijing China Institute of Automation, Chinese Academy of Sciences Beijing China School of Artificial Intelligence, University of Chinese Academy of Sciences \& Institute of Automation, Chinese Academy of Sciences Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences School of Artificial Intelligence, University of Chinese Academy of Sciences Institute of Automation, Chinese Academy of Sciences

AI总结 本文提出SOM框架,通过结构因果模型明确分离对手建模与预测,提升多智能体环境中的预测准确性和适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07174 2026-05-11 cs.AI

Repeated Deceptive Path Planning against Learnable Observer

针对可学习观察者的重复欺骗路径规划

Shiyue Cao, Pei Xu, Likun Yang, Lei Cui, Shizhao Yu, Shiyu Zhang, Yongjian Ren, Xiaotang Chen, Kaiqi Huang

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) National Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institution of Automation, Chinese Academy of Sciences(中国科学院复杂系统认知与决策智能国家重点实验室)

AI总结 本文提出DeMP框架,通过双层优化应对可学习观察者的适应性,提升欺骗持续性,实验显示其在重复欺骗路径规划中表现优异。

Comments Full version of the extended abstract accepted at AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07153 2026-05-11 cs.CL

Beyond Reasoning: Reinforcement Learning Unlocks Parametric Knowledge in LLMs

超越推理:强化学习解锁大语言模型中的参数知识

Wanli Yang, Hongyu Zang, Junwei Zhang, Wenjie Shi, Du Su, Jingang Wang, Xueqi Cheng, Fei Sun

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, CAS University of Chinese Academy of Sciences(人工智能安全国家重点实验室,计算技术研究所,中国科学院,中国科学院大学)

AI总结 本文研究强化学习在零样本、单跳、封闭书 QA 任务中提升参数知识直接回忆的能力,发现其平均相对提升达27%,揭示 RL 通过重新分配概率质量解锁潜在知识而非获取新事实。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06676 2026-05-11 cs.LG cs.CL

LKV: End-to-End Learning of Head-wise Budgets and Token Selection for LLM KV Cache Eviction

LKV:端到端学习头部预算和令牌选择以LLM KV缓存淘汰

Enshuai Zhou, Yifan Hao, Chao Wang, Rui Zhang, Di Huang, Jiaming Guo, Xing Hu, Zidong Du, Qi Guo, Yunji Chen

机构 * University of Science and Technology of China(中国科学技术大学) State Key Lab of Processors, Institute of Computing Technology, CAS, Beijing, China(中国科学院计算技术研究所过程器重点实验室,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国)

AI总结 本文提出LKV,通过端到端可微优化问题实现KV缓存压缩,学习任务优化全局预算和内在KV重要性,提升长上下文推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05866 2026-05-11 cs.AI cond-mat.mtrl-sci cs.LG

XDecomposer: Learning Prior-Free Set Decomposition for Multiphase X-ray Diffraction

XDecomposer:学习无先验的多相X射线衍射集分解

Hanyu Gao, Bin Cao, Yunyue Su, Tong-Yi Zhang, Qiang Liu

机构 * New Laboratory of Pattern Recognition (NLPR), Institute of Automation, Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所新型模式识别实验室) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院) Guangzhou Municipal Key Laboratory of Materials Informatics, The Hong Kong University of Science and Technology (Guangzhou)(广州市材料信息学重点实验室,香港科技大学(广州)) Green Dynamics, Australia(澳大利亚绿动公司)

AI总结 XDecomposer通过无先验假设的方法,实现多相X射线衍射模式的联合分解与识别,无需候选相列表或结构模板,提升重建准确性和相识别能力。

Comments 28pages, 8figures, 6tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.08077 2026-05-11 cs.CV

AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding

AdaSpark: 为高效长视频理解设计的自适应稀疏性

Handong Li, Zikang Liu, Longteng Guo, Tongtian Yue, Yepeng Tang, Xinxin Zhu, Chuanyang Zheng, Ziming Wang, Zhibin Wang, Jun Song, Cheng Yu, Bo Zheng, Jing Liu

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Alibaba Group Holding Limited(阿里巴巴集团控股有限公司) Future Living Lab of Alibaba(阿里巴巴未来生活实验室)

AI总结 AdaSpark通过自适应稀疏框架减少计算负载达57% FLOPs,保持性能并保留细粒度长程依赖,适用于小时级视频基准测试。

Comments 8 pages, CVPR2026 Accept (Highlight)

Journal ref Proceedings of the IEEE/CVF Computer Vision and Pattern Recognition (CVPR), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09782 2026-05-11 cs.LG cs.AI cs.CL

Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective

在RLVR中采用梯度保持视角实现灵活熵控制

Kun Chen, Peng Shi, Fanfan Liu, Haibo Qiu, Zhixiong Zeng, Siqi Yang, Wenji Mao

机构 * School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS) Meituan(美团)

AI总结 本文从梯度保持裁剪角度出发,提出动态裁剪阈值机制和动态熵控制策略,有效缓解熵崩溃问题并提升多基准性能。

Comments https://github.com/Kwen-Chen/Flexible-Entropy-Control

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01752 2026-05-11 cs.CL cs.CR

WorldCup Sampling for Multi-bit LLM Watermarking

多比特LLM水印的WorldCup采样

Yidan Wang, Yubing Ren, Yanan Cao, Li Guo

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络与信息安全学院)

AI总结 本文提出WorldCup框架,通过结构化通信通道建模和分层竞争机制实现多比特水印嵌入,提升文本质量和解码鲁棒性,实验表明其在容量、检测性、鲁棒性等方面优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01166 2026-05-11 cs.RO

Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

隐式推理VLA:面向视觉-语言-动作模型的隐式思考与预测

Shuanghao Bai, Jing Lyu, Wanqi Zhou, Zhe Li, Dakai Wang, Lei Xing, Xiaoguang Zhao, Pengwei Wang, Zhongyuan Wang, Cheng Chi, Badong Chen, Shanghang Zhang

机构 * Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人工智能与机器人研究院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Institute of Automation, University of Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

AI总结 本文提出LaRA-VLA框架,通过连续潜在表示实现多模态推理,减少推理开销,提升机器人实时控制效率。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.02805 2026-05-11 cs.CL cs.AI

MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning

MemSearcher:通过端到端强化学习训练LLM进行推理、搜索和管理内存

Qianhao Yuan, Jie Lou, Zichao Li, Jiawei Chen, Yaojie Lu, Hongyu Lin, Le Sun, Debing Zhang, Xianpei Han

机构 * Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences, Beijing, China(中国科学院软件研究所信息处理实验室,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国) Xiaohongshu Inc(小红书公司)

AI总结 MemSearcher通过端到端强化学习训练LLM进行推理、搜索和内存管理,通过维护紧凑的记忆在多轮交互中保持上下文长度稳定,优于基于历史拼接的基线方法。

Comments Accepted to ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18715 2026-05-11 cs.CV

ChatSearch: a Dataset and a Generative Retrieval Model for General Conversational Image Retrieval

ChatSearch: 一个用于通用对话图像检索的数据集和生成检索模型

Zijia Zhao, Longteng Guo, Tongtian Yue, Erdong Hu, Shuai Shao, Zehuan Yuan, Hua Huang, Jing Liu

机构 * The Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Bytedance Inc.(字节跳动公司) School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院)

AI总结 本文提出ChatSearch数据集和生成检索模型,用于开放领域图像的对话式检索,通过多轮多模态对话上下文查询提升检索准确性。

Journal ref Pattern Recognition, 167 (2025) 111696

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06658 2026-05-08 cs.CV

Relit-LiVE: Relight Video by Jointly Learning Environment Video

Relit-LiVE: 通过联合学习环境视频实现视频照明

Weiqing Xiao, Hong Li, Xiuyu Yang, Houyuan Chen, Wenyi Li, Tianqi Liu, Shaocong Xu, Chongjie Ye, Hao Zhao, Beibei Wang

机构 * Nanjing University(南京大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology(香港科学与技术大学) University of Chinese Academy of Sciences(中国科学院大学) Huazhong University of Science and Technology(华中科技大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

AI总结 Relit-LiVE通过引入原始参考图像和环境视频预测方法,实现了无需相机姿态先验知识的物理一致视频照明,提升了真实场景下的照明效果和时间稳定性。

Comments Accepted at SIGGRAPH 2026. Project site: https://github.com/zhuxing0/Relit-LiVE

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06416 2026-05-08 cs.CL

MiA-Signature: Approximating Global Activation for Long-Context Understanding

MiA-Signature:近似全局激活以实现长上下文理解

Yuqing Li, Jiangnan Li, Mo Yu, Zheng Lin, Weiping Wang, Jie Zhou

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Pattern Recognition Center, WeChat AI, Tencent(微信AI模式识别中心) Hunyuan Team, Tencent(腾讯文心一言团队)

AI总结 本文提出MiA-Signature,通过压缩表示近似全局激活对下游处理的影响,提升长上下文理解任务的性能。

Comments This is a work in progress; we will continue to revise and improve the manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06279 2026-05-08 cs.SE cs.AI

Correct Code, Vulnerable Dependencies: A Large Scale Measurement Study of LLM-Specified Library Versions

正确的代码,易受攻击的依赖项:一项大规模的LLM指定库版本测量研究

Chengjie Wang, Jingzheng Wu, Xiang Ling, Tianyue Luo, Chen Zhao

机构 * Intelligent Software Research Center, Institute of Software, Chinese Academy of Sciences, Beijing, China(智能软件研究中心,软件研究所,中国科学院,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国) Key Laboratory of System Software (Chinese Academy of Sciences), Beijing, China(中国科学院系统软件重点实验室,北京,中国)

AI总结 研究LLM生成代码中库版本的安全性和兼容性风险,发现版本选择存在系统性偏差,且版本选择影响漏洞暴露和兼容性故障。

Comments 35 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06264 2026-05-08 cs.LG

Can Attribution Predict Risk? From Multi-View Attribution to Planning Risk Signals in End-to-End Autonomous Driving

属性能否预测风险?从多视图属性到端到端自动驾驶中的规划风险信号

Le Yang, Ruoyu Chen, Haijun Liu, Jiawei Liang, ShangQuan Sun, Xiaochun Cao

机构 * Sun Yat-sen University(中山大学) University of Chinese Academy of Sciences(中国科学院大学) Nanyang Technological University(南洋理工大学)

AI总结 本文提出了一种层次化属性框架,用于端到端规划中的风险预测,通过提取三种属性统计作为预测信号,提升了对轨迹误差和碰撞检测的预测能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06221 2026-05-08 cs.CL

UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification

UniPrefill: 通过块级动态稀疏化实现通用长上下文预填加速

Qihang Fan, Huaibo Huang, Zhiying Wu, Bingning Wang, Ran He

机构 * MAIS&NLPR, CASIA(MAIS与NLPR,中国科学院自动化研究所) UCAS(中国科学技术大学) WeChat, Tencent(微信,腾讯)

AI总结 本文提出UniPrefill框架,通过块级动态稀疏化提升长上下文预填效率,兼容各种模型架构,实现TTFT加速,支持连续批处理和vLLM调度策略。

Comments code: https://github.com/qhfan/UniPrefill.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06166 2026-05-08 cs.LG

One Algorithm, Two Goals: Dual Scoring for Parameter and Data Selection in LLM Fine-Tuning

一个算法,两个目标:在LLM微调中用于参数和数据选择的双评分

Xinrui Chen, Liu Yang, Ou Wu

机构 * Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences(中国科学院大学杭州高等研究院) College of Intelligence and Computing, Tianjin University(天津大学智能与计算学院)

AI总结 本文提出DualSFT算法,通过共享梯度统计生成参数掩码和数据子集,提升微调性能与稳定性-可塑性平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06096 2026-05-08 cs.CL cs.CV

Uncovering Entity Identity Confusion in Multimodal Knowledge Editing

揭示多模态知识编辑中的实体身份混淆

Shu Wu, Xiaotian Ye, Xinyu Mou, Dongsheng Liu, Xiaohan Wang, Mengqi Zhang

机构 * New Laboratory of Pattern Recognition (NLPR)(模式识别新实验室) State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS)(多模态人工智能系统国家重点实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Beijing University of Posts and Telecommunications(北京邮电大学) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院) Huazhong University of Science and Technology(华中科技大学) Shandong University(山东大学)

AI总结 本文研究多模态知识编辑中实体身份混淆问题,发现现有方法无法区分图像-实体绑定与实体-实体关系,导致模型错误关联。通过构建EC-Bench基准测试,提出约束编辑阶段以提升绑定准确性,从而减少混淆。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06080 2026-05-08 cs.CV

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation

MSD-Score:多尺度分布评分用于无参考图像描述评估

Shichao Kan, Xuyang Zhang, Haojie Zhang, Zhe Zhu, Yigang Cen, Yixiong Liang, Lianlei Shan, Linna Zhang, Zhe Qu, Jiazhi Xia

机构 * School of Computer Science and Engineering, Central South University(中南大学计算机科学与工程学院) School of Computer Science and Technology and the State Key Laboratory of Advanced Rail Autonomous Operation, Beijing Jiaotong University(北京交通大学计算机科学与技术学院和先进轨道交通自主运行国家重点实验室) School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院) School of Mechanical Engineering, Guizhou University(贵州大学机械工程学院)

AI总结 本文提出MSD-Score,一种无参考图像描述评估指标,通过多尺度分布评分方法量化语义差异,实现对局部 grounding 错误的透明诊断,优于现有无参考指标。

Comments Preprint. 17 pages, 10 figures. Code is available at: https://steinsgatesg.github.io/MSDScore/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.21877 2026-05-08 cs.LG cs.AI

P^2O: Joint Policy and Prompt Optimization

P²O:联合策略和提示优化

Xinyu Lu, Kaiqi Zhang, Jinglin Yang, Boxi Cao, Yaojie Lu, Hongyu Lin, Min He, Xianpei Han, Le Sun

机构 * Chinese Information Processing Laboratory, Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所信息处理实验室) University of Chinese Academy of Sciences(中国科学院大学) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) National Computer Network Emergency Response Technical Team/Coordination Center of China(中国国家计算机网络应急技术协调中心)

AI总结 本文提出P²O方法,通过交替更新连续策略与离散提示,解决RLVR在难样本上的优势崩溃问题,提升模型泛化能力并提升性能9.5%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12677 2026-05-08 cs.CL cs.AI

MetaKE: Meta-Learning for Knowledge Editing Toward a Better Accuracy-Editability Trade-off

MetaKE:面向更优准确率-可编辑性权衡的知识编辑的元学习

Shuxin Liu, Di Gao, Ou Wu

机构 * Hangzhou Institute for Advanced Study University of Chinese Academy of Sciences(杭州高等研究所中国科学院大学)

AI总结 MetaKE通过统一上游和下游阶段为双层优化问题,改进知识编辑的准确率与可编辑性权衡,引入结构梯度代理避免多层反向传播,实验显示优于现有基线。

Comments 37 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05833 2026-05-08 cs.AI

On the Role of Language Representations in Auto-Bidding: Findings and Implications

语言表示在自动出价中的作用:发现与启示

Guanyu Zhu, Jining Luan, Hanwen Du, Xinyu Fang, Sibo Xu, Ersheng Ni, Hongji Li, Jincheng Fang, Ronghao Chen, Huacan Wang, Xuanqi Lan, Yongxin Ni, Yiqi Sun, Youhua Li

机构 * City University of Hong Kong(香港城市大学) South China Agricultural University(华南农业大学) University of Electronic Science and Technology of China(电子科技大学) The Ohio State University, Columbus(俄亥俄州立大学) Hefei University of Technology(合肥工业大学) School of Economics and Management, Wuhan University(武汉大学经济管理学院) Faculty of Engineering, The University of Queensland(昆士兰大学工程学院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) The Hong Kong University of Science and Technology(香港科学大学) Peking University(北京大学) University of Chinese Academy of Sciences(中国科学院大学) Santa Clara University(圣克拉拉大学) Boston University(波士顿大学) The University of Hong Kong(香港大学)

AI总结 本文研究语言表示在自动出价中的作用,提出SemBid框架,通过整合语义与数值信息提升出价可控性和泛化能力,实验表明其在多种场景下优于现有方法。

Comments 19 page

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05797 2026-05-08 cs.RO cs.FL

Resource-Constrained Robotic Planning in the face of Mixed Uncertainty

面对混合不确定性的资源受限机器人规划

Yihao Yin, Pian Yu, Andrea Turrini, Zhiming Chi, Yong Li, Lijun Zhang

机构 * Hangzhou Institute for Advanced Study (HIAS), UCAS, China(杭州先进研究院(HIAS),UCAS,中国) Key Laboratory of System Software (Chinese Academy of Sciences), Institute of Software Chinese Academy of Sciences, Beijing, China(系统软件重点实验室(中国科学院),软件研究所,中国科学院,北京,中国) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学,北京,中国) University College London, London, UK(伦敦大学学院,伦敦,英国)

AI总结 本文提出基于CMDPST和LTLf的框架,解决机器人在资源受限条件下鲁棒策略合成问题,通过直接展开和优化方法提升效率,实验验证有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05722 2026-05-08 cs.CV

$\mathcal{B}^{3}$-Net: Controlled Posterior Bridge Learning for Multi-Task Dense Prediction

$\mathcal{B}^{3}$-Net:多任务密集预测中的受控后验桥学习

Meihua Zhou, Li Yang

机构 * Wannan Medical University(皖南医学院) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 本文提出$\mathcal{B}^{3}$-Net,通过可靠性估计、后验桥构建和受限再分配机制,改进多任务密集预测中任务证据的显式建模,提升共享表征的可靠性。

Comments 14 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13310 2026-05-08 cs.CV cs.AI

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension

视觉并行思考器:用于视觉理解的分而治之推理

Haoran Xu, Hongyu Wang, Jiaze Li, Shunpeng Chen, Zizhao Tong, Jianzhong Ju, Zhenbo Luo, Jian Luan

机构 * Zhejiang University(浙江大学) Hunan University(湖南大学) University of Chinese Academy of Sciences(中国科学院大学) Independent Researcher(独立研究者)

AI总结 本文提出视觉并行思考器,通过并行推理框架提升视觉理解能力,结合Pa-Attention和LPRoPE实现高效多模态处理,验证了并行推理在视觉领域的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.21471 2026-05-08 cs.AI

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition

SpatialBench: 多模态大语言模型空间认知的基准测试

Peiran Xu, Sudong Wang, Yao Zhu, Jianing Li, Gege Qi, Yunjian Zhang

机构 * Sun Yat-Sen University(中山大学) HKUST (GZ)(香港科技大学) Zhejiang University(浙江大学) Peking University(北京大学) CAICT(中国科学院电子技术研究所) UCAS(中国科学技术大学) CUC(中国科学技术大学)

AI总结 本文提出SpatialBench基准,通过五级空间认知框架评估多模态大语言模型的空间能力,揭示模型在感知与符号推理间的性能差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23629 2026-05-08 cs.AI cond-mat.dis-nn cond-mat.stat-mech cs.LG physics.soc-ph

Emergent Slow Thinking in LLMs as Inverse Tree Freezing

大语言模型中涌现的慢思考作为逆树冻结

Sihan Hu, Xiansheng Cai, Yuan Huang, Zhiyuan Yao, Linfeng Zhang, Pan Zhang, Youjin Deng, Kun Chen

机构 * Hefei National Laboratory, University of Science and Technology of China(中国科学技术大学合肥微尺度物质科学国家实验室) Hefei National Laboratory for Physical Sciences at the Microscale and Department of Modern Physics, University of Science and Technology of China(中国科学技术大学合肥微尺度物质科学国家实验室和现代物理系) Institute of Theoretical Physics, Chinese Academy of Sciences(中国科学院理论物理研究所) School of Fundamental Physics and Mathematical Sciences, Hangzhou Institute for Advanced Study, UCAS(杭州高等研究院基础物理与数学科学学院) DP Technology(DP技术) Lanzhou Center for Theoretical Physics, Key Laboratory of Theoretical Physics of Gansu Province, Key Laboratory of Quantum Theory and Applications of MoE, Gansu Provincial Research Center for Basic Disciplines of Quantum Physics, Lanzhou University(兰州理论物理中心、甘肃省理论物理重点实验室、教育部量子理论与应用重点实验室、甘肃省量子物理基础学科省重点实验室、兰州大学) AI for Science Institute, Beijing(北京人工智能科学研究院)

AI总结 本文通过统计物理视角揭示大语言模型中慢思考的涌现机制,提出逆树冻结结构,并提出Annealed-RLVR方法提升模型性能。

Comments 34 pages, 17 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01498 2026-05-05 cs.CV

Towards Visual Query Localization in the 3D World

面向3D世界的视觉查询定位

Liang Peng, Bohan Tan, Zhipeng Zhang, Haobo Li, Yifan Jiao, Xingping Dong, Libo Zhang

机构 * Wuhan University(武汉大学) AutoLab, SAI, Shanghai Jiao Tong University(AutoLab、SAI、上海交通大学) Anyverse Dynamics University of Chinese Academy of Sciences(中国科学院大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)

AI总结 本文首次提出3D多模态视觉查询定位基准3DVQL,包含2002个序列和17万帧数据,通过点云、RGB图像和深度图像多模态数据提升研究灵活性,并提出LaF算法提升性能。

Comments Accepted to CVPR 2026. 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01423 2026-05-05 hep-ex cs.AI cs.MA

HepScript: A Dual-Use DSL for Human-AI Collaborative Data Analysis Workflows in High-Energy Physics

HepScript:一种双用途DSL,用于高能物理中的人工智能协作数据分析工作流

Junkun Jiao, Tong Liu, Ke Li, Weimin Song, Yipu Liao, Bolun Zhang, Beijiang Liu, Chang-Zheng Yuan, Yue Sun

机构 * Jilin University(吉林大学) Institute of High Energy Physics, CAS(高能物理研究所) University of Chinese Academy of Sciences(中国科学院大学)

AI总结 HepScript通过双用途DSL提升高能物理数据分析效率,减少人工编写代码93%,AI可自主生成95%的成功执行规范。

详情

展开后加载摘要…

URL PDF HTML 收藏