arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Tsinghua University(清华大学)

共收录 4547
2506.05171 2026-04-09 eess.SY cs.AI cs.SY

Towards provable probabilistic safety for scalable embodied AI systems

迈向可证明的可扩展具身AI系统的概率安全

Linxuan He, Lingxiang Fan, Qing-Shan Jia, Ang Li, Hongyan Sang, Ling Wang, Guanghui Wen, Jiwen Lu, Tao Zhang, Jie Zhou, Yi Zhang, Yisen Wang, Peng Wei, Zhongyuan Wang, Henry X. Liu, Shuo Feng

机构 * Department of Automation, Tsinghua University(清华大学自动化系) Beijing Academy of Artificial Intelligence(北京人工智能研究院) School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院) School of Computer, Liaocheng University(聊城大学计算机学院) Department of Automation, Southeast University(东南大学自动化学院) Department of Mechanical and Aerospace Engineering, George Washington University(乔治华盛顿大学机械与航空航天工程系) University of Michigan Transportation Research Institute(密歇根大学交通研究所) Department of Civil and Environmental Engineering, University of Michigan(密歇根大学土木与环境工程系)

AI总结 本文提出可证明的概率安全范式,旨在解决具身AI系统在复杂环境中安全验证的挑战,通过结合可证明保证与渐进式达成概率安全边界,提升系统可行性和可扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15325 2026-04-09 cs.CV

SoftHGNN: Soft Hypergraph Neural Networks for General Visual Recognition

SoftHGNN: 用于通用视觉识别的软超图神经网络

Mengqi Lei, Yihong Wu, Siqi Li, Xinhu Zheng, Juan Wang, Shaoyi Du, Yue Gao

机构 * Tsinghua University(清华大学) Yangtze Delta Region Institute of Tsinghua University, Zhejiang(清华大学长三角研究院(浙江)) Taiyuan University of Technology(太原理工大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 本文提出SoftHGNN,通过引入软超边和稀疏选择机制,提升视觉识别中高阶关联的建模能力,实现更高效的语义推理。

Comments This paper has been accepted by the International Journal of Computer Vision (IJCV)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.02129 2026-04-09 cs.LG cs.AI math.ST stat.TH

Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon

路径正则化:多层神经网络的近完整且最优的非渐近泛化理论及双下降现象

Hao Yu

机构 * Department of Automation, Tsinghua University(清华大学自动化系)

AI总结 本文提出一种非渐近泛化理论,针对路径正则化的多层神经网络,无需假设损失函数有界,考虑了近似误差,并揭示了双下降现象的潜在机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16240 2026-04-09 cs.LG

AFL: A Single-Round Analytic Approach for Federated Learning with Pre-trained Models

AFL:基于预训练模型的联邦学习单轮分析方法

Run He, Kai Tong, Di Fang, Han Sun, Ziqian Zeng, Haoran Li, Tianyi Chen, Huiping Zhuang

机构 * South China University of Technology(华南理工大学) Tsinghua University(清华大学) Beijing National Research Center for Information Science and Technology(北京信息科学与技术国家研究中心) The Hong Kong University of Science and Technology(香港科技大学) Microsoft(微软)

AI总结 本文提出AFL,一种利用分析解的联邦学习方法,通过单轮训练和绝对聚合法则,减少通信开销并提升收敛速度,具有数据分区不变性。

Comments Published in CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07250 2026-04-09 cs.CV

Geo-EVS: Geometry-Conditioned Extrapolative View Synthesis for Autonomous Driving

Geo-EVS:基于几何条件的外推视图合成用于自动驾驶

Yatong Lan, Rongkui Tang, Lei He

机构 * School of Vehicle and Mobility, Tsinghua University(清华大学车辆与运载学院) State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University(清华大学智能绿色车辆与移动国家重点实验室) School of Science, Minzu University(中央民族大学理学院) Chongqing Changan Automobile Co., Ltd.(重庆长安汽车股份有限公司)

AI总结 Geo-EVS通过几何条件化框架提升自动驾驶中稀疏视图合成的几何精度,尤其在高角度和低覆盖率场景中表现优异,并改进下游3D检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07171 2026-04-09 cs.LG

Smart Commander: A Hierarchical Reinforcement Learning Framework for Fleet-Level PHM Decision Optimization

智能指挥官:面向舰队级PHM决策优化的分层强化学习框架

Yong Si, Mingfei Lu, Jing Li, Yang Hu, Guijiang Li, Yueheng Song, Zhaokui Wang

机构 * School of Aerospace Engineering, Tsinghua University(清华大学航天航空学院) College of Information and Control Engineering, Xi'an University of Architecture and Technology(西安建筑科技大学信息与控制工程学院) Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院) First Aircraft Institute of Aviation Industry Corporation of China(中国航空工业集团公司第一飞机设计研究院) Science and Technology on Complex Aviation System Simulation Laboratory(复杂航空系统仿真重点实验室)

AI总结 本文提出Smart Commander框架,通过分层强化学习解决军事航空PHM中高维问题、稀疏反馈和随机任务配置带来的挑战,实现维护和后勤决策优化,展现出更高的效率和鲁棒性。

Comments 21 pages, 6 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05846 2026-04-08 cs.CL

AgentGL: Towards Agentic Graph Learning with LLMs via Reinforcement Learning

AgentGL: 通过强化学习实现基于LLM的代理图学习

Yuanfu Sun, Kang Li, Dongzhe Fan, Jiajin Liu, Qiaoyu Tan

机构 * New York University Shanghai(上海纽约大学) New York University(纽约大学) Tsinghua University(清华大学)

AI总结 AgentGL通过强化学习框架实现图学习,结合图导航与LLM推理,提升复杂关系环境下的自主导航与推理能力,优于现有基线方法。

Comments ACL 2026 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01328 2026-04-08 cs.LG

Efficient and Principled Scientific Discovery through Bayesian Optimization: A Tutorial

通过贝叶斯优化实现高效且原则性的科学发现:教程

Zhongwei Yu, Rasul Tutunov, Alexandre Max Maraval, Zikai Xie, Zhenzhi Tan, Jiankang Wang, Bin Cao, Zijing Li, Liangliang Xu, Qi Yang, Jun Jiang, Sanzhong Luo, Zhenxiao Guo, Tongyi Zhang, Haitham Bou-Ammar, Jun Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学) The University of Hong Kong(香港大学) Haihe Laboratory of Sustainable Chemical Transformations(海河可持续化学转化实验室) Shanghai University(上海大学) University College London(伦敦大学学院)

AI总结 本文通过贝叶斯优化框架,系统阐述了如何通过自动化科学发现循环提升实验效率,涵盖催化、材料科学等领域的案例研究及技术扩展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19894 2026-04-08 cs.SE cs.AI

TransAgent: Enhancing LLM-Based Code Translation via Fine-Grained Execution Alignment

TransAgent:通过细粒度执行对齐增强基于LLM的代码翻译

Zhiqiang Yuan, Weitong Chen, Hanlin Wang, Xin Peng, Zhenpeng Chen, Yiling Lou

机构 * Fudan University(复旦大学) Tsinghua University(清华大学)

AI总结 TransAgent通过细粒度执行对齐本地化错误代码块,提升LLM代码翻译的准确性与修复性能,优于现有方法33.3%和56.7%。

Comments Accepted by FSE'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05748 2026-04-08 cs.CV

SVC 2026: the Second Multimodal Deception Detection Challenge and the First Domain Generalized Remote Physiological Measurement Challenge

SVC 2026: 第二届多模态欺骗检测挑战赛及首个领域通用远程生理测量挑战

Dongliang Zhu, Zhiyi Niu, Bo Zhao, Jiajian Huang, Shuo Ye, Xun Lin, Hui Ma, Taorui Wang, Jiayu Zhang, Chunmei Zhu, Junzhe Cao, Yingjie Ma, Rencheng Song, Albert Clapés, Sergio Escalera, Dan Guo, Zitong Yu

机构 * Wuhan University(武汉大学) Great Bay University(大湾区大学) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学) Sun Yat-sen University(中山大学) Hefei University of Technology(合肥工业大学) University of Barcelona(巴塞罗那大学)

AI总结 本文提出SVC 2026挑战赛,旨在通过多模态欺骗检测和远程脉搏波测速估计任务,推动对细微视觉信号的鲁棒表示学习研究。

Comments Accepted by the SVC workshop @ CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05721 2026-04-08 cs.CV

GaussianGrow: Geometry-aware Gaussian Growing from 3D Point Clouds with Text Guidance

GaussianGrow:基于3D点云与文本引导的几何感知Gaussian生成

Weiqi Zhang, Junsheng Zhou, Haotian Geng, Kanle Shi, Shenkun Xu, Yi Fang, Yu-Shen Liu

机构 * School of Software, Tsinghua University(清华大学软件学院) Kuaishou Technology(快手科技) CAIR and CIDSAI, NYU Abu Dhabi(纽约大学阿布扎比分校CAIR和CIDSAI)

AI总结 本文提出GaussianGrow,通过学习从3D点云中生长Gaussian,结合多视角扩散模型和文本引导,提升生成几何精度与质量。

Comments Accepted by CVPR 2026. Project page: https://weiqi-zhang.github.io/GaussianGrow

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05719 2026-04-08 cs.CR cs.AI cs.SE

Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing

黑客还是幻觉?对基于LLM的自动化渗透测试的全面分析

Jiaren Peng, Zeqin Li, Chang You, Yan Wang, Hanlin Sun, Xuan Tian, Shuqiao Zhang, Junyi Liu, Jianguo Zhao, Renyang Liu, Haoran Ou, Yuqiang Sun, Jiancheng Zhang, Yutong Jiao, Kunshu Song, Chao Zhang, Fan Shi, Hongda Sun, Rui Yan, Cheng Huang

机构 * School of Cyber Science and Engineering, Sichuan University(四川大学网络空间安全学院) Institute for Network Sciences and Cyberspace, Tsinghua University(清华大学网络科学与网络空间研究院) College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算与数据科学学院) National University of Singapore(新加坡国立大学) College of Electronic Engineering, National University of Defense Technology(国防科技大学电子工程学院) Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院) School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)

AI总结 本文对基于LLM的自动化渗透测试框架进行系统分析和大规模评估,揭示其架构设计和性能差异,为未来研究提供结构化分类和基准。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05716 2026-04-08 cs.AI

Can Large Language Models Reinvent Foundational Algorithms?

大语言模型能否重新发明基础算法?

Jian Zhao, Haoren Luo, Yu Wang, Yuhan Cao, Pingyue Sheng, Tianxing He

机构 * Xiongan AI Institute(雄安人工智能研究院) Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院) Shanghai Qi Zhi Institute(上海期智研究院) Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) Beijing University of Posts and Telecommunications(北京邮电大学)

AI总结 本文研究大语言模型能否在受控环境中重新发明基础算法,通过Unlearn-and-Reinvent流程验证模型在无提示、低提示和高提示下的表现,发现生成验证器在推理过程中起到关键作用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05669 2026-04-08 stat.ML cs.LG

Efficient machine unlearning with minimax optimality

高效机器遗忘与最小最大最优性

Jingyi Xie, Linjun Zhang, Sai Li

机构 * Institute of Statistics and Big Data, Renmin University of China(中国人民大学统计与大数据研究院) Department of Statistics, Rutgers University(罗格斯大学统计系) Department of Statistics and Data Science, Tsinghua University(清华大学统计与数据科学系)

AI总结 本文提出一种通用损失函数的机器遗忘统计框架,针对平方损失开发了Unlearning Least Squares方法,证明其在仅使用预训练估计器、遗忘样本和少量剩余样本时的最小最大最优性,并建立了无需全量重训练的推断程序。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05449 2026-04-08 cs.CV

Not All Agents Matter: From Global Attention Dilution to Risk-Prioritized Game Planning

并非所有智能体都重要:从全局注意力稀释到风险优先的游戏规划

Kang Ding, Hongsong Wang, Jie Gui, Lei He

机构 * School of Cyberspace Security, Southeast University(东南大学网络空间安全学院) School of Computer Science and Engineering, Southeast University(东南大学计算机科学与工程学院) Vehicle and Mobility School, Tsinghua University(清华大学车辆与运载学院)

AI总结 本文提出风险优先游戏规划方法,通过GameAD框架提升自动驾驶轨迹安全性,实验表明其在复杂环境中表现更优。

Comments 14 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05405 2026-04-08 cs.CV

Weather-Conditioned Branch Routing for Robust LiDAR-Radar 3D Object Detection

针对恶劣天气的分支路由用于鲁棒的激光雷达-雷达3D目标检测

Hongsheng Li, Lingfeng Zhang, Zexian Yang, Liang Li, Rong Yin, Xiaoshuai Hao, Wenbo Ding

机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) College of Computer and Data Science, Fuzhou University(福州大学计算机与数据科学学院) Institute of Information Engineering, CAS(中国科学院信息工程研究所)

AI总结 本文提出一种基于天气条件的分支路由方法,通过动态调整模态偏好提升恶劣天气下3D目标检测的鲁棒性,实验表明其在K-Radar基准上达到最佳性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05398 2026-04-08 math.OC cs.LG

An Actor-Critic Framework for Continuous-Time Jump-Diffusion Controls with Normalizing Flows

一个用于连续时间跳跃扩散控制的Actor-Critic框架

Liya Guo, Ruimeng Hu, Xu Yang, Yi Zhu

机构 * Yau Mathematical Sciences Center, Tsinghua University(清华大学丘成桐数学科学中心) Department of Mathematics, Tsinghua University(清华大学数学系) Department of Mathematics, University of California, Santa Barbara(加州大学圣塔芭芭拉分校数学系) Department of Statistics and Applied Probability, University of California, Santa Barbara(加州大学圣塔芭芭拉分校统计与应用概率系) Yanqi Lake Beijing Institute of Mathematical Sciences and Applications(北京雁栖湖应用数学研究院)

AI总结 本文提出一个无网格求解器,用于解决熵正则化控制问题和具有跳跃的随机博弈,通过时间非齐性小q函数和适当的职业测度,结合条件归一化流参数化Actor,实现灵活的非高斯策略,并在多个金融和博弈场景中验证了方法的有效性。

Comments 29 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05323 2026-04-08 cs.CV cs.RO

VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success

VLA-InfoEntropy:一种无需训练的视觉-注意力信息熵方法,用于视觉-语言-动作模型的推理加速与成功

Chuhang Liu, Yayun He, Zuheng Kang, Xiaoyang Qu, Jianzong Wang

机构 * Ping An Technology (Shenzhen) Co., Ltd.(平安科技(深圳)有限公司) Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)

AI总结 本文提出VLA-InfoEntropy方法,通过图像熵和注意力熵结合时间步信息,动态调整模型关注区域,减少冗余并提升推理效率。

Comments Accepted to the 2026 IEEE International Conference on Multimedia and Expo (ICME 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25661 2026-04-08 cs.RO cs.CV

Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance

Fast-dVLA:加速离散扩散VLA以实现实时性能

Wenxuan Song, Jiayi Chen, Shuai Chen, Jingbo Wang, Pengxiang Ding, Han Zhao, Yikai Qin, Xinhu Zheng, Donglin Wang, Yan Wang, Haoang Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) ShanghaiTech University(上海科技大学) Shanghai Institute of Technical Physics, CAS(中国科学院上海技术物理研究所) AIR, Tsinghua University(清华大学智能产业研究院) Westlake University(西湖大学) Zhejiang University(浙江大学)

AI总结 本文提出一种方法,通过分离辅助任务训练目标,在参数空间中提升通用能力与任务特定动作分布,从而在减少计算开销的同时提高性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05765 2026-04-08 cs.AI

RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training

RL-VLA$^3$:一种灵活且异步的强化学习框架用于VLA训练

Haoran Sun, Yongjian Guo, Zhong Guan, Shuai Di, Xiaodong Bai, Jing Long, Tianyun Zhao, Mingxi Luo, Hongke Zhao, Likang Wu, Xiaotie Deng, Xu Chu, Xi Xiao, Sheng Wen, Yicheng Gong, Junwu Xiong

机构 * Peking University(北京大学) Tsinghua University(清华大学) Tianjin University(天津大学) JDT AI Infra(京东科技AI基础设施) Swinburne University of Technology(斯威本科技大学)

AI总结 本文提出RL-VLA$^3$,一种异步强化学习框架,通过动态批调度和灵活环境分片策略,提升VLA训练效率,实验显示其在8至256GPU上均取得85.2%的吞吐量提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.03741 2026-04-08 cs.CV

I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image Editing

I2E: 从图像像素到可操作的交互环境:用于文本引导的图像编辑

Jinghan Yu, Junhao Xiao, Chenyu Zhu, Jiaming Li, Jia Li, HanMing Deng, Xirui Wang, Guoli Jia, Jianjun Li, Xiang Bai, Bowen Zhou, Zhiyuan Ma

机构 * Huazhong University of Science and Technology(华中科技大学) Kuaishou Technology(快手科技) Central China Normal University(华中师范大学) Tsinghua University(清华大学) Shanghai AI Laboratory(上海人工智能实验室)

AI总结 I2E提出一种'分解-然后-行动'范式,通过分解器将无结构图像转换为可操作的对象层,并利用物理感知的视觉-语言-行动代理进行复杂指令解析,同时构建I2E-Bench基准测试集,实验表明其在处理复杂组合指令和保持物理合理性方面优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09425 2026-04-08 cs.LG stat.ML

Supporting Evidence for the Adaptive Feature Program across Diverse Models

支持适应性特征程序在多样模型中的证据

Yicheng Li, Qian Lin

机构 * Tsinghua University(清华大学)

AI总结 本文通过理论分析探讨神经网络优势,提出适应性特征程序并提供证据支持其在多种模型中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10238 2026-04-08 cs.CV cs.AI

ForgeryGPT: A Multimodal LLM for Interpretable Image Forgery Detection and Localization

ForgeryGPT: 一种多模态大语言模型用于可解释的图像伪造检测与定位

Fanrui Zhang, Jiawei Liu, Jiaying Zhu, Esther Sun, Dong Li, Qiang Zhang, Zheng-Jun Zha

机构 * University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学)

AI总结 ForgeryGPT通过多模态大语言模型捕捉伪造图像的高阶取证知识关联,实现可解释的图像伪造检测与定位,提升检测精度和交互能力。

Comments 13 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04642 2026-04-07 cs.RO

WaterSplat-SLAM: Photorealistic Monocular SLAM in Underwater Environment

WaterSplat-SLAM:水下环境中的光实单目SLAM

Kangxu Wang, Shaofeng Zou, Chenxing Jiang, Yixiang Dai, Siang Chen, Shaojie Shen, Guijin Wang

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) Key Laboratory of Marine Robotics, Shenyang(沈阳海洋机器人重点实验室) University of Chinese Academy of Sciences(中国科学院大学) Department of Electronic and Computer Engineering, The Hong Kong University of Science and Technology(香港科技大学电子及计算机工程学系)

AI总结 本文提出WaterSplat-SLAM,通过语义介质过滤和语义引导渲染实现水下高保真密集映射,提升水下单目SLAM的鲁棒性与可视化效果。

Comments 8 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04573 2026-04-07 cs.ET cs.LG

SAIL: Scene-aware Adaptive Iterative Learning for Long-Tail Trajectory Prediction in Autonomous Vehicles

SAIL:面向场景的自适应迭代学习用于自动驾驶车辆的长尾轨迹预测

Bin Rao, Haicheng Liao, Chengyue Wang, Keqiang Li, Zhenning Li, Hai Yang

机构 * State Key Laboratory of Internet of Things for Smart City and Department of Civil and Environmental Engineering, University of Macau(澳门大学智慧城市物联网国家重点实验室及土木与环境工程系) State Key Laboratory of Internet of Things for Smart City and Department of Computer and Information Science, University of Macau(澳门大学智慧城市物联网国家重点实验室及计算机与信息科学系) Department of Automotive Engineering, Tsinghua University(清华大学汽车工程系) State Key Laboratory of Internet of Things for Smart City and Departments of Civil and Environmental Engineering and Computer and Information Science, University of Macau(澳门大学智慧城市物联网国家重点实验室及土木与环境工程系和计算机与信息科学系) Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology(香港科技大学土木与环境工程系)

AI总结 本文提出SAIL框架,通过定义和建模轨迹的三个关键属性维度,结合属性引导增强与自适应对比学习策略,提升自动驾驶车辆在长尾场景下的轨迹预测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04502 2026-04-07 cs.RO

Veo-Act: How Far Can Frontier Video Models Advance Generalizable Robot Manipulation?

Veo-Act:前沿视频模型在通用机器人操作中能走多远?

Zhongru Zhang, Chenghan Yang, Qingzhou Lu, Yanjiang Guo, Jianke Zhang, Yucheng Hu, Jianyu Chen

机构 * Tsinghua University(清华大学)

AI总结 本文探讨前沿视频模型如Veo-3在通用机器人操作中的应用,提出Veo-Act框架结合高阶运动规划与低阶执行器,提升指令跟随性能。

Comments 16 pages, 12 figures. Equal contribution by Zhongru Zhang, Chenghan Yang, Qingzhou Lu and Yanjiang Guo. Project lead: Yanjiang Guo

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04487 2026-04-07 cs.CV

Training-Free Image Editing with Visual Context Integration and Concept Alignment

无需训练的图像编辑与视觉上下文整合与概念对齐

Rui Song, Guo-Hua Wang, Qing-Guo Chen, Weihua Luo, Tongda Xu, Zhening Liu, Yan Wang, Zehong Lin, Jun Zhang

机构 * Alibaba International Digital Commerce Group(阿里巴巴国际数字商业集团) Institute for AI Industry Research (AIR), Tsinghua University(清华大学智能产业研究院(AIR)) School of Data Science, Lingnan University(岭南大学数据科学学院)

AI总结 本文提出VicoEdit,一种无需训练和反向扩散的图像编辑方法,通过视觉上下文直接转换源图像至目标图像,提升编辑一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04419 2026-04-07 cs.CV

BoxComm: Benchmarking Category-Aware Commentary Generation and Narration Rhythm in Boxing

BoxComm:基于 boxing 的类别感知评论生成与叙述节奏基准测试

Kaiwen Wang, Kaili Zheng, Rongrong Deng, Yiming Shi, Chenyi Guo, Ji Wu

机构 * Tsinghua University(清华大学) Beijing Sport University(北京体育大学)

AI总结 本文提出 BoxComm 数据集,包含 445 场世界拳击锦标赛视频及 52000 多条专业评论,通过分类注解和两项新评估方法,解决拳击评论生成中的节奏与类别准确性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04145 2026-04-07 cs.AI

Solar-VLM: Multimodal Vision-Language Models for Augmented Solar Power Forecasting

Solar-VLM:用于增强太阳能发电预测的多模态视觉语言模型

Hang Fan, Haoran Pei, Runze Liang, Weican Liu, Long Cheng, Wei Wei

机构 * North China Electric Power University(华北电力大学) Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学)

AI总结 本文提出Solar-VLM框架,通过融合时序观测、卫星图像和文本天气信息,提升太阳能发电预测的准确性,采用多模态编码器和图注意力网络捕捉空间依赖性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.28532 2026-04-07 cs.LG cs.AI stat.AP

Detecting low left ventricular ejection fraction from ECG using an interpretable and scalable predictor-driven framework

通过可解释且可扩展的预测驱动框架从ECG检测低左心室射血分数

Ya Zhou, Tianxiang Hao, Ziyi Cai, Haojie Zhu, Kejun He, Jia Liu, Xiaohan Fan, Jing Yuan

机构 * Fuwai Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(中国医学科学院阜外医院、北京协和医学院) The Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Institute of Statistics and Big Data, Renmin University of China(中国人民大学统计与大数据研究院) National Center for Cardiovascular Diseases(国家心血管病中心) Function Test Center, Fuwai Hospital(阜外医院功能检测中心)

AI总结 本文提出ECGPD-LEF框架,结合基础模型诊断概率与可解释模型,有效检测ECG中的低左心室射血分数,优于现有方法,且具备高可解释性。

Comments This version includes minor typographical corrections. The results and conclusions remain unchanged

详情

展开后加载摘要…

URL PDF HTML 收藏