arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Shanghai Jiao Tong University(上海交通大学)

共收录 432
2606.31169 2026-07-01 cs.CV 新提交

Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

超越单字符:评估多模态大语言模型在句子级甲骨文理解中的表现

Ziqi Li, Zijian Chen, Tingzhu Chen, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学) Nanjing University of Science and Technology(南京理工大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

AI总结 提出S-OBI基准,通过字形替换与组合合成清晰标准的句子级甲骨文实例,设计语义匹配、语义槽提取和上下文推理任务,评估多模态大语言模型在句子级甲骨文理解中的表现,发现当前模型仍依赖字符级识别。

Comments 13 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.31086 2026-07-01 cs.CV 新提交

CasaMaestro: Multi-View Panoramas for House-Scale 3D Reconstruction

CasaMaestro:用于房屋级3D重建的多视角全景图

Yuzhou Ji, Xiaotian Yang, Zhipeng Zhang

机构 * AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院AutoLab)

AI总结 提出CasaMaestro模型,仅需20-50张稀疏多视角室内全景图,即可直接预测度量深度和相机位姿,实现全屋快速点云重建,是首个支持房屋级重建的多视角全景模型。

Comments Accepted to ECCV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30694 2026-07-01 cs.RO cs.AI 新提交

DSIP: A Dynamic Coordination Planner for Signal-Free Intersections using Diffusion-Model-Based Multi-Agent Motion Planning

DSIP: 基于扩散模型的多智能体运动规划的无信号交叉口动态协调规划器

Qian Hu, Haoyang Peng, Songan Zhang, Ming Yang, Hongtei Eric Tseng

机构 * Global Institute of Future Technology, Shanghai Jiao Tong University(上海交通大学全球未来技术学院) School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院) Department of Electrical Engineering, The University of Texas at Arlington(德克萨斯大学阿灵顿分校电气工程系)

AI总结 提出DSIP框架,利用扩散模型生成多车连续轨迹优化,替代传统信号相位控制,在SUMO仿真中显著降低平均延误并提高平均速度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27922 2026-07-01 cs.CV cs.AI 新提交

Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding

Reflect-R1:长视频理解中基于证据驱动的自纠正反思

Shuimu Chen, Yuteng Chen, Yuanshen Guan, Zebang Cheng, Zeyu Zhang, Shengqian Qin, Bin Xia, Jiaran Li, Wenming Yang, Fei Ma

机构 * Tsinghua University(清华大学) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东省人工智能与数字经济实验室(深圳)) Nanyang Technological University(南洋理工大学) University of Science and Technology of China(中国科学技术大学) Shenzhen University(深圳大学) University of California(加利福尼亚大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 提出Reflect-R1框架,通过直觉、验证、仲裁三阶段管道动态检索客观视觉证据,并设计阶段解耦强化学习算法SD-GRPO解决策略耦合,在长视频理解基准上达到最优性能。

Comments 2026 ECCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28215 2026-06-29 cs.CV cs.AI cs.GR 新提交

HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration

HAT-4D: 通过人机协作从单目视频提升4D多物体交互

Jiaxin Li, Yuxiang Wu, Zhenkai Zhang, Xinrui Shi, Haoyuan Wang, Yichen Zhao, Su Linxiang, Chenyang Yu, Mingyu Zhang, Yifan Ding, Boran Wen, Li Zhang, Ruiyang Liu, Yong-Lu Li

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai Innovation Institute(上海创新研究院) University of Science and Technology of China(中国科学技术大学) Math Magic

AI总结 提出HAT-4D框架,结合视觉大模型与多级人工反馈,从单目视频重建多物体的3D几何、时序动态和物理交互,解决遮挡和深度歧义,无需多相机系统。

Comments Accepted to ECCV 2026. 15 pages of main text and 39 pages of appendices. Project page: https://lijiaxin0111.github.io/HAT4D/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28094 2026-06-29 cs.CV cs.AI 新提交

OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal

OSOR: 一步扩散修复用于效果感知的对象移除

Qinming Zhou, Chenxi Sun, Deyang Kong, Junhao He, Xiangheng Tang, Peike Yu, Haotian Wu, Leilei Cao, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) Xidian University(西安电子科技大学) Tongji University(同济大学) University of Electronic Science and Technology of China(电子科技大学) Transsion(传音控股)

AI总结 提出一步扩散模型OSOR,通过占用引导判别器、alpha头和语义锚定验证管道,实现高效、效果感知且对不完美掩码鲁棒的对象移除,速度提升4-30倍。

Comments Code and resources are available at https://github.com/Zhouqm-Git/osor

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27880 2026-06-29 cs.CV 新提交

OrthoTryOn: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation

OrthoTryOn:面向无冲突统一时尚生成的几何正交化

Zhaotong Yang, Ying Tai, Jiahui Zhan, Yu Zheng, Jianjun Qian, Jian Yang

机构 * PCA Lab, School of Computer Science and Engineering, Nanjing University of Science and Technology(南京理工大学计算机科学与工程学院PCA实验室) PCA Lab, School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院PCA实验室) Shanghai Jiao Tong University(上海交通大学)

AI总结 提出OrthoTryOn框架,通过正交子空间投影和Fisher引导负向引导解决统一时尚生成中任务间梯度冲突,实现多任务协同优化。

Comments Accepted by ECCV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27872 2026-06-29 cs.RO cs.AI 新提交

S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation

S$^2$-VLA:面向长时程操作的基于状态空间的视觉-语言-动作模型

Zhipeng Xie, Zongyi Han, Xiangyi Wei, Shiliang Sun, Yang Li, Jing Zhao

机构 * School of Computer Science and Technology, East China Normal University(华东师范大学计算机科学与技术学院) State Key Laboratory of Submarine Geoscience, School of Automation and Intelligent Sensing, Shanghai Jiao Tong University(上海交通大学自动化与智能感知学院海底科学国家重点实验室)

AI总结 针对长时程操作任务中累积误差导致性能下降的问题,提出S$^2$-VLA框架,通过状态空间引导的自适应注意力机制动态融合视觉、语言和动作信息,在LIBERO等基准上超越7B模型。

Comments Accepted to IJCAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27679 2026-06-29 cs.CL cs.AI 新提交

From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models

从信号到迁移:基于探针的大语言模型不确定性估计的分解研究

Ponhvoan Srey, Xiaobao Wu, Cong-Duy Nguyen, Quang Minh Nguyen, Duc Anh Vu, Anh Tuan Luu

机构 * Nanyang Technological University(南洋理工大学) Shanghai Jiao Tong University(上海交通大学) VinUniversity KAIST(韩国科学技术院)

AI总结 通过分解研究,发现原始隐藏状态和注意力特征在域内表现优异,但结构化压缩特征在分布偏移下更鲁棒,提示和标签构建显著影响探针行为,并训练出可迁移的预训练探针。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27678 2026-06-29 cs.CV 新提交

Two-Stage Cross-Domain Cervical Abnormality Screening with Cytopathological Image Synthesis and Knowledge Distillation

两阶段跨域宫颈异常筛查:细胞病理图像合成与知识蒸馏

Jincheng Li, Yuzhi He, Yihui Zhan, Xinmei Zhang, Yifei Sun, Zelin Liu, Lichi Zhang, Minye Shao, Lili Zhao

机构 * School of Artificial Intelligence and Computer Science, Nantong University(南通大学人工智能与计算机科学学院) School of Telecommunications Engineering, Xidian University(西安电子科技大学通信工程学院) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) School of Biomedical Engineering, Shanghai Jiao Tong University(上海交通大学生物医学工程学院) Department of Computer Science, Durham University(杜伦大学计算机科学系)

AI总结 提出两阶段框架,先用SC-UNSB合成中间域缓解域偏移,再用知识蒸馏对齐特征,提升跨域宫颈细胞检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27295 2026-06-29 cs.RO 新提交

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

LA4VLA:通过语言-动作预训练实现无视觉行动学习

Tao Lin, Yuxin Du, Yiran Mao, Zewei Ye, Yilei Zhong, Bing Cheng, Yiming Wang, Jiting Liu, Yang Tian, Junchi Yan, Feiran Wu, Zenan Meng, Hu Wei, Yuqian Fu, Gen Li, Bo Zhao

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) Alibaba Group(阿里巴巴集团) Nanyang Technological University(南洋理工大学) KAUST(阿卜杜拉国王科技大学)

AI总结 提出LA4VLA框架,通过语言-动作预训练学习动作先验,减少对视觉线索的依赖,提升VLA策略的鲁棒性,在仿真和真实任务中成功率显著提升。

Comments Github: https://github.com/MINT-SJTU/LA4VLA

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25939 2026-06-29 cs.RO 新提交

DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning

DeformGen: 基于动力学的拓扑增强用于可变形操作策略学习

Zili Lin, Wenyao Zhang, Yuyang Zhang, Zekun Qi, Junyan Lin, Hanxin Zhu, Jiaolong Yang, Zhibo Chen, Yao Mu, Xiaokang Yang, Xin Jin, Wenjun Zeng

机构 * Shanghai Jiao Tong University(上海交通大学) Eastern Institute of Technology, Ningbo(宁波东方理工大学) Tsinghua University(清华大学) The Hong Kong Polytechnic University(香港理工大学) University of Science and Technology of China(中国科学技术大学) Zhongguancun Academy(中关村学院) Microsoft Research(微软研究院)

AI总结 提出DeformGen框架,通过局部物理扰动和动力学前向模拟生成拓扑一致的可变形状态,并利用变形场扭曲转移操作轨迹,联合增强状态分布与操作行为,提升可变形操作策略学习。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23686 2026-06-29 cs.RO 新提交

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models

LIBERO-Safety:视觉-语言-动作模型中物理与语义安全的综合基准

Rongxu Cui, Zongzheng Zhang, Jingrui Pang, Haohan Chi, Jinbang Guo, Saining Zhang, Shaoxuan Xie, Xin Jin, Yao Mu, Jiaolong Yang, Guocai Yao, Xianyuan Zhan, Ya-Qin Zhang, Hao Zhao

机构 * Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院) Beijing Academy of Artificial Intelligence (BAAI)(北京智源人工智能研究院) Beihang University(北京航空航天大学) Eastern Institute of Technology, Ningbo(宁波东方理工大学) Shanghai Jiao Tong University(上海交通大学) Microsoft Research Asia (MSRA)(微软亚洲研究院)

AI总结 针对视觉-语言-动作模型操作安全未验证的问题,提出参数化安全基准和关键帧驱动数据生成流水线,构建大规模无碰撞数据集,系统评估八种模型,揭示泛化-安全张力。

Comments Accepted by ECCV 2026, Project Page: https://libero-safety.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27036 2026-06-26 cs.RO 新提交

RelAfford6D: Relational 6D Affordance Graphs for Constraint-Driven Robotic Manipulation

RelAfford6D:用于约束驱动机器人操作的关系型6D可操作图

Guodong Zhang, Qichen He, Wenyuan Xie, Shaokai Wu, Yanbiao Ji, Qiuchang Li, Bayram Bayramli, Yue Ding, Hongtao Lu

机构 * School of Computer Science, Shanghai Jiaotong University(上海交通大学计算机科学学院)

AI总结 提出RelAfford6D框架,通过构建关系型6D可操作图将语义指令转化为SE(3)位姿约束,并利用运动学约束求解实现零样本操作,在仿真和真实环境中优于数据驱动方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27027 2026-06-26 cs.CR cs.AI 新提交

ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP

ShareLock: 针对MCP的隐蔽多工具阈值投毒攻击

Liwei Liu, Tianzhu Han, Zijian Liu, Zishu Dong, Na Ruan

机构 * Shanghai Jiao Tong University(上海交通大学)

AI总结 提出ShareLock框架,利用Shamir阈值方案将恶意指令分散到多个工具描述中,实现隐蔽性和容错性,平均攻击成功率超90%。

Comments 16 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26875 2026-06-26 cs.CL cs.AI 新提交

Information-Aware KV Cache Compression for Long Reasoning

面向长推理的信息感知KV缓存压缩

Jushi Kai, Zhuiri Xiao, Alexandra Birch, Zhouhan Lin

机构 * Shanghai Jiao Tong University(上海交通大学)

AI总结 提出InfoKV框架,通过结合预测不确定性与注意力分数进行KV缓存压缩,在长上下文推理中优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26874 2026-06-26 cs.AI 新提交

TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation

TAVR-VLM:面向抗幻觉报告生成的风险条件因果基础

Zhixiang Lu, Xiwei Liu, Sifan Song, Changkai Ji, Anh Nguyen, Jionglong Su, Imran Razzak, Jinfeng Wang

机构 * Xi’an Jiaotong-Liverpool University(西交利物浦大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Shanghai Jiao Tong University(上海交通大学) University of Liverpool(利物浦大学) Kunming University of Science and Technology(昆明理工大学)

AI总结 提出TAVR-VLM框架,通过风险条件因果注意力(R-CGA)建立“风险→区域→词”的结构化基础路径,在TAVR规划中减少诊断幻觉,实现新最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26694 2026-06-26 cs.CV 新提交

PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models

PhysEditWorld: 一个面向物理可编辑世界模型的大规模数据集

Bin Hu, Yanwen Ma, Jiehui Huang, Ziliang Zhang, Haoning Wu, Ruicheng Zhang, Yaokun Li, Zijun Wang, Yuechen Zhang, Chun-Mei Tseng, Hanhui Li, Shengju Qian, Jun Zhou, Kaipeng Zhang, Xiaodan Liang, Jiaya Jia, Xiu Li

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Beihang University(北京航空航天大学) The Hong Kong University of Science and Technology(香港科技大学) Independent Researcher(独立研究者) Shanghai Jiao Tong University(上海交通大学) Sun Yat-sen University(中山大学) The Chinese University of Hong Kong(香港中文大学) Tencent(腾讯)

AI总结 提出PhysEditWorld数据集,通过UE5重放渲染管线生成多重力配置下的动作条件视频,支持可控物理编辑,用于改进世界模型的动力学建模和一致性。

Comments Project page: https://yizhiqianbi.github.io/physeditworld/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26616 2026-06-26 cs.RO 新提交

A Closed-Form 4-DoF Inter-Robot Pose Estimator using Bearing-only Measurements

基于仅方位测量的闭式4自由度机器人间位姿估计器

Qixin De, Ao Zhuang, Yechen Zhang, Zhuozhou Qian, Danping Zou

机构 * Shanghai Jiao Tong University(上海交通大学)

AI总结 提出一种闭式4自由度机器人间位姿估计器,通过放松旋转估计的非线性约束和误差投影实现平移估计,并引入可观测性测试模块自主确定最优估计时刻,在降低计算成本的同时提高估计精度和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.27119 2026-06-26 quant-ph cs.AI cs.LG 新提交

Efficient foundation decoders for fault-tolerant quantum computing

用于容错量子计算的高效基础解码器

Ge Yan, Shanchuan Li, Shiyi Xiao, Pengyue Ma, Hanyan Cao, Feng Pan, Yuxuan Du

机构 * College of Computing and Data Science, Nanyang Technological University, Singapore(南洋理工大学计算与数据科学学院) Department of Electrical Engineering and Computer Science, Tokyo University of Agriculture and Technology, Koganei, Tokyo, Japan(东京农业与技术大学电气工程与计算机科学系) School of Artificial Intelligence, Shanghai Jiao Tong University, Shanghai, China(上海交通大学人工智能学院) Science, Mathematics and Technology Cluster, Singapore University of Technology and Design, Singapore(新加坡科技设计大学科学、数学与技术集群) School of Physical and Mathematical Science, Nanyang Technological University, Singapore(南洋理工大学物理与数学科学学院)

AI总结 提出神经迁移统一框架(NTU),通过代数结构对齐不同码距的解码任务,实现小码距知识加速大码距训练,NTU-Transformer在平面表面码和双变量自行车码上超越现有方法。

Comments 32 pages, 9 figures, comments are welcome

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.22437 2026-06-26 cs.CV cs.AI 新提交

MMGist: A Comprehensive Multimodal Benchmark for 2027

MMGist:面向2027年的综合多模态基准

Wenzhen Yuan, Jiacheng Ruan, Wutao Xiong, Chengping Zhao, Ting Liu, Yuzhuo Fu

机构 * SJTU(上海交通大学)

AI总结 针对现有视觉语言基准的视觉依赖不足、性能饱和及异常项问题,提出MMGist基准,通过三阶段过滤构建7,262项,覆盖7个能力维度,在保持模型排名高保真度(Spearman ρ=0.98)的同时减少69%评估项并提升78%跨模型区分度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26078 2026-06-25 cs.CV cs.AI 新提交

A cross-process welding penetration status prediction algorithm based on unsupervised domain adaptation in laser and TIG welding

基于无监督域适应的激光与TIG焊接跨工艺熔透状态预测算法

Sen Li, Haichao Cui, Chendong Shao, Yaqi Wang, Xinhua Tang

机构 * Shanghai Key Laboratory of Materials Laser Processing and Modification(上海材料激光加工与改性重点实验室) School of Materials Science and Engineering(材料科学与工程学院) Shanghai Jiao Tong University(上海交通大学)

AI总结 提出无监督域适应框架结合渐进源域扩展策略,解决激光与TIG焊接间域偏移问题,在跨工艺任务中准确率提升超43%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.26059 2026-06-25 cs.CV cs.AI 新提交

A welding penetration prediction model for laser welding process based on self-supervised learning using physics-informed neural networks

基于物理信息神经网络的自监督学习激光焊接过程熔透预测模型

Sen Li, Xiaoying Liu, Xiaojian Xu, Chendong Shao, Yaqi Wang, Ling Lan, Xinhua Tang, Haichao Cui

机构 * Shanghai Key Laboratory of Materials Laser Processing and Modification, School of Materials Science and Engineering, Shanghai Jiao Tong University(上海材料激光加工与改性重点实验室,上海交通大学材料科学与工程学院) Shanghai Shipbuilding Technology Research Institute(上海船舶工艺研究所)

AI总结 提出SimPhysNet模型,利用自监督学习嵌入物理先验,仅需少量标注图像即可高精度预测激光焊接熔透状态,准确率达96.06%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25907 2026-06-25 cs.CV 新提交

In-context Region-based Drag: Drag Any Region to Any Shape

基于上下文区域的拖拽:将任意区域拖拽至任意形状

Jiacheng Sui, Tianyu Hao, Bingjie Gao, Li Niu, Guangtao Zhai

机构 * Shanghai Jiao Tong University(上海交通大学)

AI总结 提出ICRDrag方法,通过上下文学习框架和两种注意力正则化,实现区域级拖拽编辑,在精度和视觉保真度上显著优于现有方法。

Comments Accepted by ECCV 2026. Dataset, code, and model are available at https://github.com/bcmi/ICRDrag-Region-Drag-Editing

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25819 2026-06-25 cs.CL cs.SE 新提交

Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliability

超越函数调用:工具环境不可靠性下的工具使用智能体基准测试

Yang Tian, Zhengpeng Shi, Yu Zhou, Bo Zhao

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院)

AI总结 提出ToolBench-X基准,在工具环境中注入五种可恢复可靠性风险,评估智能体任务完成能力,发现可靠工具下表现良好的智能体在可恢复风险下失败,失败主因是诊断和恢复能力不足。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25546 2026-06-25 cs.CV 新提交

Disease-Centric Vision-Language Pretraining with Hybrid Visual Encoding for 3D Computed Tomography

面向3D计算机断层扫描的疾病中心视觉语言预训练与混合视觉编码

Bowen Shi, Weiwei Cao, Ruifeng Yuan, Wanxing Chang, Wenrui Dai, Hongkai Xiong, Ling Zhang, Jianpeng Zhang

机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) Hupan Lab, 310023, Hangzhou, China(虎扑实验室,杭州,中国) Shanghai Jiao Tong University, China(上海交通大学,中国) Zhejiang University, China(浙江大学,中国) Fudan University, China(复旦大学,中国)

AI总结 提出一种结合CNN-ViT混合编码器、疾病级对比学习和诊断感知提示的视觉语言预训练框架,在CT-RATE和Rad-ChestCT上取得最优性能,并提升零样本诊断可靠性。

Comments ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23001 2026-06-25 cs.SE cs.LG cs.OS 新提交

EnerInfer: Energy-Aware On-Device LLM Inference

EnerInfer: 能量感知的设备端大语言模型推理

Bohua Zou, Nian Liu, Binqi Sun, Matteo Mascherin, Debayan Roy, Yutao Liu, Yu Peng, Ning Jia, Haibo Chen

机构 * Technical University of Munich(慕尼黑技术大学) Huawei Hilbert Research Center (Dresden)(华为希利伯特研究中心(德累斯顿)) Huawei Central Software Institute(华为中央软件研究所) Shanghai Jiao Tong University(上海交通大学)

AI总结 针对设备端LLM推理中能量与散热瓶颈,提出EnerInfer框架,通过解耦预测与在线反馈动态调整NPU/DDR频率,在保证服务质量前提下提升能效达65%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17945 2026-06-25 cs.AI 新提交

Small Initialization Matters for Large Language Models

小初始化对大语言模型至关重要

Liangkai Hang, Junjie Yao, Zhiyu Li, Feiyu Xiong, Hongkang Yang, Zhi-Qin John Xu

机构 * School of Mathematical Sciences, Shanghai Jiao Tong University(上海交通大学数学科学学院) Institute of Natural Sciences, Shanghai Jiao Tong University(上海交通大学自然科学研究院) MemTensor (Shanghai) Technology Co., Ltd.(上海记忆张量科技有限公司) Institute for Advanced Algorithms Research(先进算法研究所)

AI总结 本文发现减小初始化尺度能持续改善大语言模型预训练,尤其在推理任务上提升显著,并揭示了小初始化驱动参数从低复杂度结构向丰富表示演化的机制。

Comments 26 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23888 2026-06-24 eess.IV cs.AI cs.CV 新提交

E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis

E-MRL: 跨视图对齐的证据驱动多模态强化学习用于可靠的3D肿瘤分析

Sijing Li, Zhongwei Qiu, Zhuoya Wang, Boxiang Yun, Zhenyu Yi, Jianwei Xu, Wenqiao Zhang, Yingda Xia, Ling Zhang

机构 * Zhejiang University(浙江大学) DAMO Academy, Alibaba Group(阿里巴巴集团达摩院) Hupan Lab(华平实验室) Huazhong University of Science and Technology(华中科技大学) East China Normal University(华东师范大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 提出跨视图对齐的证据驱动多模态强化学习框架E-MRL,通过将生成过程建模为“诊断-定位-验证”的马尔可夫决策过程,并引入跨视图一致性奖励,减少视觉幻觉并提升3D CT肿瘤诊断准确性。

Comments 9 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.24775 2026-06-24 cs.CL cs.DB cs.IR 新提交

Are We Ready For An Agent-Native Memory System?

我们准备好构建原生智能体记忆系统了吗?

Wei Zhou, Xuanhe Zhou, Shaokun Han, Hongming Xu, Guoliang Li, Zhiyu Li, Feiyu Xiong, Fan Wu

机构 * Shanghai Jiao Tong University(上海交通大学) Tsinghua University(清华大学) MemTensor (Shanghai) Technology Co., Ltd(MemTensor(上海)科技有限公司)

AI总结 从数据管理视角系统研究智能体记忆,提出四模块分析框架,评估12种记忆系统,发现无单一架构占优,并揭示成本-性能权衡。

Comments Paper list available at: https://github.com/OpenDataBox/awesome-agent-memory. Source code available at: https://github.com/OpenDataBox/MemoryData

详情

展开后加载摘要…

URL PDF HTML 收藏