arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Xi'an Jiaotong University(西安交通大学)

共收录 765
2605.11705 2026-05-13 cs.CV

CAST: Collapse-Aware multi-Scale Topology Fusion for Multimodal Coreset Selection

CAST:面向多模态聚类选择的坍缩感知多尺度拓扑融合

Boran Zhao, Hetian Liu, Zhenxian Hu, Yuqing Yuan, Yu Yan, Pengju Ren

机构 * School of Software Engineering, the National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, National Engineering Research Center for Visual Information and Applications, and Institute of Artificial Intelligence and Robotics(软件工程学院、人机混合增强智能国家重点实验室、视觉信息与应用国家工程研究中心、人工智能与机器人研究院) School of Software Engineering(软件工程学院) XJTU-POLIMI Joint School(西交大-波兰理工联合学院) Faculty of Electronic and Information Engineering(电子与信息工程学院) School of Human Settlements and Civil Engineering(人居与土木工程学院) the National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, National Engineering Research Center for Visual Information and Applications, and Institute of Artificial Intelligence and Robotics(人机混合增强智能国家重点实验室、视觉信息与应用国家工程研究中心、人工智能与机器人研究院)

AI总结 本文提出CAST框架,通过多尺度拓扑融合解决多模态数据集选择中的跨模态信息失衡和分布不匹配问题,提升跨架构泛化能力和能效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.11578 2026-05-13 cs.CV

The Midas Touch for Metric Depth

度量深度的Midas触感

Yu Ma, Zizhan Guo, Zuyi Xiong, Haoran Zhang, Yi Feng, Hongbo Zhao, Hanli Wang, Rui Fan

机构 * College of Electronic and Information Engineering, Tongji University(同济大学电子与信息工程学院) Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学上海智能自主系统研究所) National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi’an Jiaotong University(西安交通大学人机混合增强智能国家重点实验室)

AI总结 本文提出MTD方法,通过稀疏3D数据将相对深度转为度量深度,解决局部不一致和计算效率问题,提升深度估计精度并支持多种3D任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10993 2026-05-13 cs.RO

ECHO: Continuous Hierarchical Memory for Vision-Language-Action Models

ECHO:面向视觉-语言-动作模型的连续层次记忆

Yanbin Hu, Jin Cui, Jiayi Lu, Ruixuan Yang, Jun Ye, Boran Zhao, Xingyu Chen, Xuguang Lan, Pengju Ren

机构 * School of Software, Xi’an Jiaotong University(西安交通大学软件学院) School of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院) State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人机混合增强智能国家重点实验室,人工智能与机器人研究院)

AI总结 本文提出ECHO框架,通过连续层次空间提升视觉-语言-动作模型在长周期任务中的表现,实现高效经验检索与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10611 2026-05-12 cs.CR cs.AI

Re-Triggering Safeguards within LLMs for Jailbreak Detection

在大语言模型中重新触发安全机制以检测劫持

Zheng Lin, Zhenxing Niu, Haoxuan Ji, Yuzhe Huang, Haichang Gao

机构 * Xidian University(西安电子科技大学) Xi'an Jiaotong University(西安交通大学)

AI总结 本文提出一种大语言模型劫持提示检测方法,通过重新激活内部安全机制提升防御能力,实验表明其在白盒和黑盒环境下有效抵御最新劫持攻击。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10582 2026-05-12 cs.CR cs.AI

Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing

通过干扰-校正平滑实现保证的对抗防御

Zheng Lin, Zhenxing Niu, Haoxuan Ji, Haichang Gao

机构 * Xidian University(西安电子科技大学) Xi’an Jiaotong University(西安交通大学)

AI总结 本文提出了一种基于平滑的新型防御方法,通过干扰-校正机制提升大语言模型对劫持攻击的防御能力,理论分析提供了防御成功的紧界和干扰强度要求,实验表明其在安全性和有用性上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.03639 2026-05-12 cs.CV

Diffusion Masked Pretraining for Dynamic Point Cloud

动态点云的扩散掩码预训练

Zhuoyue Zhang, Jihua Zhu, Chaowei Fang, Jian Liu, Ajmal Saeed Mian

机构 * Xi’an Jiaotong University(西安交通大学) School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) University of Western Australia(西澳大学)

AI总结 本文提出DiMP框架,通过引入扩散建模解决动态点云预训练中位置泄露和运动监督问题,提升下游任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11674 2026-05-12 cs.RO cs.AI

AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation

AffordSim:一种可扩展的数据生成器和基准,用于面向 affordance 的机器人操作

Mingyang Li, Haofan Xu, Haowen Sun, Xinzhe Chen, Sihua Ren, Liqi Huang, Xinyang Sui, Chenyang Miao, Jiawei Ye, Qiongjie Cui, Zeyang Liu, Xingyu Chen, Xuguang Lan

机构 * School of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院)

AI总结 AffordSim 通过整合开放词汇 3D affordance 预测,解决了机器人操作中接触信息获取的挑战,实现了高成功率的轨迹生成和跨仿真到现实的迁移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09789 2026-05-12 cs.LG

When Less is More: The LLM Scaling Paradox in Context Compression

当少即多:在上下文压缩中的LLM扩展悖论

Ruishan Guo, Yibing Liu, Guoxin Ma, Yan Wang, Yueyang Zhang, Long Xia, Kecheng Chen, Zhiyuan Sun, Daiting Shi

机构 * Baidu Inc.(百度公司) Tsinghua University(清华大学) Xi’an Jiaotong University(西安交通大学) City University of Hong Kong(香港城市大学)

AI总结 研究揭示在上下文压缩中,增大压缩器规模反而降低重建上下文的忠实度,发现知识覆盖和语义漂移是主要因素,中等规模压缩器在忠实恢复上表现更优。

Comments 22 pages, 7 figures, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.09479 2026-05-12 eess.IV cs.CV cs.MM

ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality

ML-CLIPSim: 多层CLIP相似性用于机器导向的图像质量

Feng Ding, Haisheng Fu, Jie Liang, Qihan Xu, Siyu Zhu, Jingning Han

机构 * Simon Fraser University(西蒙弗雷泽大学) University of British Columbia(不列颠哥伦比亚大学) Eastern Institute of Technology(东部技术学院) Xi’an Jiaotong University(西安交通大学) Google Inc(谷歌公司)

AI总结 本文提出ML-CLIPSim,基于冻结的CLIP视觉编码器构建可微质量度量,通过聚合中间补丁-令牌相似性和全局图像嵌入,更符合机器导向偏好,同时在人类质量预测中保持竞争力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.03438 2026-05-12 cs.CV

Mantis: Mamba-native Tuning is Efficient for 3D Point Cloud Foundation Models

Mantis:Mamba原生微调在3D点云基础模型中的高效性

Zihao Guo, Jihua Zhu, Jian Liu, Ajmal Saeed Mian

机构 * Xi’an Jiaotong University(西安交通大学) School of Artificial Intelligence and Robotics, Hunan University(湖南大学人工智能与机器人学院) University of Western Australia(西澳大学)

AI总结 针对Mamba架构基础模型的微调问题,提出Mantis框架,通过引入状态感知适配器和双序列一致性蒸馏,实现高效且稳定的微调,仅需约5%的可训练参数。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00678 2026-05-12 cs.RO

Toward Reliable Sim-to-Real Predictability for MoE-based Robust Quadrupedal Locomotion

迈向基于MoE的稳健四足运动的可靠仿真到现实预测性

Tianyang Wu, Hanwei Guo, Yuhang Wang, Junshu Yang, Xinyang Sui, Jiayi Xie, Xingyu Chen, Zeyang Liu, Xuguang Lan

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人机混合增强智能国家重点实验室,人工智能与机器人研究院,西安交通大学)

AI总结 本文提出基于MoE的四足运动策略与RoboGauge评估框架,通过仿真测试实现可靠的真实世界迁移,实验显示在复杂地形上具备良好的鲁棒性和泛化能力。

Comments Accepted at Robotics Science and Systems (RSS), 2026. Project Page: https://robogauge.github.io/complete/

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08587 2026-05-12 cs.LG cs.AI

Kaczmarz Linear Attention

Kaczmarz线性注意力

Jiaxuan Zou, Ruifeng Ren, Yong Liu

机构 * School of Mathematics and Statistics(数学与统计学学院) Xi’an Jiaotong University(西安交通大学) Gaoling School of Artificial Intelligence(白洋学校人工智能学院) Renmin University of China(中国人民大学)

AI总结 本文提出Kaczmarz线性注意力(KLA),通过改进Gated DeltaNet的更新机制,提升长上下文语言模型的准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.08281 2026-05-12 cs.CV

Is Class Signal Clustered or Routed in Task-Induced Implicit Neural Representation Weight Spaces?

隐式神经表示中的类别信号是聚类还是路由?

Xinyi Guo, Mingyi He, Haobin Ding, Weiming Chen, Xinrui Chen, Jiawen Li, Di Zhang, Minxi Ouyang, Yizhi Wang, Xitong Ling

机构 * South China Normal University(南方科技大学) Beijing University of Chemical Technology(北京化工大学) Tsinghua University(清华大学) Xi’an Jiaotong University(西安交通大学)

AI总结 研究探讨了隐式神经表示中类别信号的几何结构,发现其并非简单聚类,而是通过读者路由实现可分类性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07393 2026-05-11 cs.AI

Offline Policy Optimization with Posterior Sampling

离线策略优化与后验采样

Hongqiang Lin, Dongxu Zhang, Yiding Sun, Mingzhe Li, Ning Yang, Haijun Zhang

机构 * Zhejiang University(浙江大学) Xi’an Jiaotong University(西安交通大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Science and Technology Beijing(北京科技大学)

AI总结 本文提出PSPO方法,通过将动态建模作为贝叶斯推断过程,结合后验采样与约束策略优化,提升离线强化学习的泛化能力与鲁棒性。

Comments 25 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07354 2026-05-11 eess.SP cs.CV

Task-Oriented Communication for Human Action Understanding via Edge-Cloud Co-Inference

面向任务的边缘-云协同通信用于人类动作理解

Jingyi Liu, Cheng Yuan, Lijun He, Jun Zhang, Jiawei Shao

机构 * Institute of Artificial Intelligence (TeleAI) of China Telecom(中国电信人工智能研究所) School of Information and Communications Engineering, Xi’an Jiaotong University(西安交通大学信息与通信工程学院) Department of Electronic and Computer Engineering, Hong Kong University of Science and Technology(香港科技大学电子与计算机工程系)

AI总结 本文提出TOAU框架,通过边缘-云协作减少传输负载和延迟,利用单目姿态估计器和VQ-VAE提取动作特征,提升动作理解精度。

Comments 12 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.07106 2026-05-11 cs.CL

Retrieve, Integrate, and Synthesize: Spatial-Semantic Grounded Latent Visual Reasoning

检索、整合与综合:空间-语义 grounded 的潜在视觉推理

Jin Cui, Xinyue Long, Xunyong Zhang, Yadong Zhang, Chuanchang Su, Jingye Gan, Boran Zhao, Pengju Ren

机构 * State Key Laboratory of Human-Machine Hybrid Augmented Intelligence, and Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人机混合增强智能国家重点实验室,人工智能与机器人研究院,西安交通大学)

AI总结 本文提出RIS框架,通过空间-语义 grounded 方法改进多模态大语言模型的视觉推理,通过构建逐步 grounded 数据集并引入短语言过渡token,提升潜在状态与词汇对齐的解码能力,实验显示在多个基准上优于现有基线。

Comments 19 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01862 2026-05-11 cs.LG

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL

QHyer: 用于离线目标条件强化学习的Q-条件混合注意力-门控Transformer

Xing Lei, Jincheng Wang, Xuetao Zhang, Donglin Wang

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) Institute of Artificial Intelligence(人工智能研究院) Robotics, Xi'an Jiaotong University(机器人技术,西安交通大学) Computer Science, University College London(计算机科学,伦敦大学学院) School of Engineering, Westlake University(工程学院,西湖大学)

AI总结 本文提出QHyer,通过引入状态条件化的Q估计器和混合注意力-门控架构,解决离线目标条件强化学习中长期依赖建模和稀疏奖励下的行为拼接问题,验证了其在非马尔可夫和马尔可夫数据集上的优越性能。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01166 2026-05-11 cs.RO

Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models

隐式推理VLA:面向视觉-语言-动作模型的隐式思考与预测

Shuanghao Bai, Jing Lyu, Wanqi Zhou, Zhe Li, Dakai Wang, Lei Xing, Xiaoguang Zhao, Pengwei Wang, Zhongyuan Wang, Cheng Chi, Badong Chen, Shanghang Zhang

机构 * Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(西安交通大学人工智能与机器人研究院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) Institute of Automation, University of Chinese Academy of Sciences(中国科学院自动化研究所) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)

AI总结 本文提出LaRA-VLA框架,通过连续潜在表示实现多模态推理,减少推理开销,提升机器人实时控制效率。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01003 2026-05-11 cs.LG cs.AI

ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning

ESSAM:一种用于内存高效的LLM微调的强化学习竞争进化策略方法

Zhishen Sun, Sizhe Dang, Guang Dai, Haishan Ye

机构 * Xi’an Jiaotong University(西安交通大学) SGIT AI Lab(SGIT人工智能实验室)

AI总结 本文提出ESSAM,结合进化策略的零阶搜索与SAM提升泛化能力,实现低内存高精度的LLM微调,实验显示其在GSM8K任务中表现优异,内存使用显著降低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06127 2026-05-08 cs.CV cs.AI

Continuous Expert Assembly: Instance-Conditioned Low-Rank Residuals for All-in-One Image Restoration

连续专家装配:用于全场景图像修复的实例条件低秩残差

Haisen He, Xiangyu Zou, SongLin Dong, Heng Li, Yihong Gong, Zhiheng Ma

机构 * Southern University of Science and Technology(南方科技大学) Shenzhen University(深圳大学) Shenzhen University of Advanced Technology(深圳大学先进技术学院) Xi’an Jiaotong University(西安交通大学)

AI总结 本文提出连续专家装配框架,通过轻量级交叉注意力超适配器生成实例条件低秩路由基和残差方向,实现无需外部提示的密集残差更新,提升图像修复的适应性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.06076 2026-05-08 cs.CL

Navigating by Old Maps: The Pitfalls of Static Mechanistic Localization in LLM Post-Training

用旧地图导航:在LLM微调中静态机械定位的陷阱

Hang Chen, Jiaying Zhu, Hongyang Chen, Hongxu Liu, Xinyu Yang, Wenya Wang

机构 * School of Computer Science and Technology(计算机科学与技术学院) Xi’an Jiaotong University(西安交通大学) School of Computer Science and Engineering(计算机科学与工程学院) The Chinese University of Hong Kong(香港中文大学) Shaanxi Co., Ltd(陕西有限公司) China Mobile Group(中国移动集团) College of Computing and Data Science(计算与数据科学学院) Nanyang Technological University(南洋理工大学)

AI总结 研究探讨了在LLM微调中静态机械定位的局限性,提出三种新指标分析电路演变,并强调需要前瞻性机制定位。

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05951 2026-05-08 cs.AI

HaM-World: Soft-Hamiltonian World Models with Selective Memory for Planning

HaM-World: 带选择性记忆的软哈密顿世界模型用于规划

Haoyun Tang, Haodong Cui, Keyao Xu, Kun Wang, Zhandong Mei

机构 * Xi’an Jiaotong University(西安交通大学) Huazhong University of Science and Technology(华中科技大学) Nankai University(南开大学) Nanyang Technological University(南洋理工大学)

AI总结 HaM-World通过分解潜在状态为规范子空间和上下文子空间,结合Mamba选择性状态空间记忆,提升规划稳定性与鲁棒性,实现高精度预测和任务完成。

Comments 22 pages, 5 figures. Code: https://github.com/HaoyunT/HaM_World

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01746 2026-05-08 cs.CV

Point-SRA: Self-Representation Alignment for 3D Representation Learning

点-SRA:用于3D表示学习的自表示对齐

Lintong Wei, Jian Lu, Haozhe Cheng, Jihua Zhu, Kaibing Zhang

机构 * School of Electronics and Information, Xi’an Polytechnic University(西安理工大学电子与信息学院) School of Software, Xi’an Jiaotong University(西安交通大学软件学院) School of Computer Science, Xi’an Polytechnic University(西安理工大学计算机科学学院)

AI总结 Point-SRA通过自蒸馏和概率建模对齐表示,改进3D表示学习,通过不同掩码比例和MeanFlow Transformer实现互补信息提取,优于Point-MAE并在多个任务中取得优异性能。

Comments This is an AAAI 2026 accepted paper titled "Point-SRA: Self-Representation Alignment for 3D Representation Learning", spanning 13 pages in total. The submission includes 7 figures (fig1 to fig7) that visually support the technical analysis

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 2026, Vol. 40, No. 13

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.05909 2026-05-08 cs.AI

Null Space Constrained Contrastive Visual Forgetting for MLLM Unlearning

空域约束对比遗忘用于MLLM反学习

Yuhang Wang, Zhenxing Niu, Haoxuan Ji, Guangyu He, Linlin Zhang, Haichang Gao

机构 * School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院) Xi’an Jiaotong University(西安交通大学)

AI总结 本文提出一种MLLM反学习方法,通过冻结LLM主干并微调视觉模块,在保留非目标视觉知识和全部文本知识的同时,有效遗忘目标视觉知识。

Comments 20 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01150 2026-05-08 cs.LG cs.AI cs.CR cs.CV math.OC

SMI: Statistical Membership Inference for Reliable Unlearned Model Auditing

SMI: 统计成员推断用于可靠未学习模型审计

Jialong Sun, Zeming Wei, Jiaxuan Zou, Jiacheng Gong, Jie Fu, Chengyang Dong, Heng Xu, Jialong Li, Bo Liu

机构 * Shenzhen University of Advanced Technology(深圳先进技术大学) Peking University(北京大学) Xi’an Jiaotong University(西安交通大学) Heilongjiang University(黑龙江大学) Stevens Institute of Technology(斯蒂文斯理工学院)

AI总结 本文提出SMI方法,通过统计成员混合比例评估未学习模型的遗忘率,无需训练影子模型,有效解决传统成员推断攻击的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01490 2026-05-05 cs.CV cs.AI cs.LG

CGFformer: Cluster-Guidance Frequency Transformer for Pansharpening

CGFformer:用于全色锐化中的簇引导频率变换器

Zijian Zhou, Jianing Zhang, Kai Sun, Xiangyu Zhao, Chunxia Zhang, Xiangyong Cao

机构 * College of Artificial Intelligence, Xi’an Jiaotong University(西安交通大学人工智能学院) School of Mathematics and Statistics, Xi’an Jiaotong University(西安交通大学数学与统计学学院) School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)

AI总结 CGFformer通过簇引导频率变换器解决全色锐化中频率分布和频率-空间交互问题,采用自适应分离模块和双流精炼模块提升去噪效果,实验表明其在多个基准数据集上表现优异。

Comments 35 pages, 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01368 2026-05-05 cs.RO

Assistance Without Interruption: A Benchmark and LLM-based Framework for Non-Intrusive Human-Robot Assistance

无中断协助:一种非侵入式人机协助的基准和基于大语言模型的框架

Yuedi Zhang, Shuanghao Bai, Wanqi Zhou, Haoran Zhang, Qi Zhang, Zhirong Luan, Badong Chen

机构 * Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) Xi’an Jiaotong University(西安交通大学) School of Electrical Engineering(电气工程学院)

AI总结 本文提出非侵入式协助的基准和框架,通过模拟基准和新指标评估,结合LLM与评分模型实现主动非侵入式协助,减少人类努力并保持任务有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01365 2026-05-05 cs.CV cs.RO

VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection

VoxAfford:多尺度体素-标记融合用于开放词汇3D affordance检测

Haowen Sun, Shaolong Zhang, Mingyang Li, Chengzhong Ma, Xinzhe Chen, Qiongjie Cui, Xingyu Chen, Zeyang Liu, Xuguang Lan

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Institute of Artificial Intelligence and Robotics, Xi’an Jiaotong University(人机混合增强智能国家级实验室,人工智能与机器人研究所,西安交通大学)

AI总结 本文提出VoxAfford,通过多尺度几何特征增强输出标记,提升3D affordance检测的定位精度,实验显示mIoU提升8%,并验证了零样本迁移能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.01338 2026-05-05 cs.AI

DiagramNet: An End-to-End Recognition Framework and Dataset for Non-Standard System-Level Diagrams

DiagramNet: 一种面向非标准系统级图表的端到端识别框架和数据集

Jincheng Lou, Ruohan Xu, Jiapeng Li, Junyin Pi, Runzhe Tao, Weijian Fan, Xiao Tan, Guojie Luo, Yibo Lin

机构 * School of IC, Peking University, Beijing, China(北京大学信息科学技术学院) College of Artificial Intelligence, Xi'an Jiaotong University, Xi'an, China(西安交通大学人工智能学院) Department of Precision Instruments, Tsinghua University, Beijing, China(清华大学精密仪器系) School of Software and Microelectronics, Peking University, Beijing, China(北京大学软件与微电子学院) School of Computer Science, Peking University, Beijing, China(北京大学计算机科学系) Institute of EDA, Peking University, Beijing, China(北京大学EDA研究院) Beijing Advanced Innovation Center for IC, Beijing, China(北京集成电路先进创新中心)

AI总结 本文提出DiagramNet数据集和训练框架,通过端到端方法在系统级图表识别中超越现有模型,提升多模态大语言模型的图表理解能力。

Comments 13 pages, 7 figures. Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.07885 2026-05-05 cs.CR cs.AI cs.SE

False Friends in the Shell: Unveiling the Emoticon Semantic Confusion in Large Language Models

壳中的假朋友:揭示大型语言模型中的表情符号语义混淆

Weipeng Jiang, Xiaoyu Zhang, Juan Zhai, Shiqing Ma, Chao Shen, Yang Liu

机构 * Xi’an Jiaotong University(西安交通大学) Nanyang Technological University(南洋理工大学) University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

AI总结 研究揭示大型语言模型对表情符号的语义混淆问题,通过构建数据集发现平均混淆率超38%,且多数混淆响应导致安全风险,呼吁开发有效缓解方法。

详情

展开后加载摘要…

URL PDF HTML 收藏