arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Tsinghua University(清华大学)

2026-08-18 至 2026-08-18 共收录 10
2608.09311 2026-08-18 cs.CV 版本更新

Degraded Infrared Small Object Detection via Degradation-Adapted Physics-Guided Restoration

面向退化红外小目标检测的退化自适应物理引导恢复方法

Xinkai Lu, Wenjun Chen, Yi Li, Yi Chang, Luxin Yan

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院) School of Ocean Engineering, Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院海洋工程学院)

AI总结 针对红外小目标检测中退化泛化性差的问题,提出DAISOD框架,结合退化识别、专用分支处理与物理引导恢复,构建对应数据集,实验表明其性能优于现有方法。

Comments Accept by ICIG2026 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27924 2026-08-18 cs.LG cs.CV cs.RO 版本更新

ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

ODEWorld:一种基于物理时间流的连续预测架构

Dongxiu Liu, Haoyi Niu, Peng Cheng, Yuan Gao, Xirui Kang, Sangli Teng, Koushil Sreenath, Xianyuan Zhan

机构 * Institute for AI Industry Research (AIR), Tsinghua University(清华大学人工智能产业研究院(AIR)) Berkeley Artificial Intelligence Research (BAIR), University of California, Berkeley(加州大学伯克利分校伯克利人工智能研究院(BAIR))

AI总结 研究针对现有世界建模机器学习范式局限于离散时间预测的问题,提出基于PT-Flow的连续时间潜在世界模型ODEWorld,解决表示崩溃问题,在视频生成和机器人控制任务中表现出色。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.10745 2026-08-18 cs.CL 版本更新

The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese

首个中文BabyLM挑战:训练数据高效且认知合理的中文语言模型

Siyuan Song, Zhiheng Qian, Yunhao Zhang, Linyang He, Xiaozhe Ji, Yingxin Lin, Hongao Zhu, Chongtian Shao, Chuhan Lang, Luan Li, Rui Wang, Renfen Hu, Shaonan Wang, Hai Hu

机构 * Princeton University(普林斯顿大学) Shanghai Jiao Tong University(上海交通大学) Chinese Academy of Sciences(中国科学院) Columbia University(哥伦比亚大学) Beijing Normal University(北京师范大学) Tsinghua University(清华大学) University of California San Diego(加利福尼亚大学圣地亚哥分校) The Hong Kong Polytechnic University(香港理工大学)

AI总结 首个中文BabyLM挑战将在2026年自然语言处理与中文计算会议举办,要求用1亿中文词元从头训练语言模型,在自然语言理解、认知对齐和汉字知识三轨道评估,不限分词器、模型架构和训练轮数。

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12485 2026-08-18 cs.LG cs.AI 版本更新

Speculative Rollback Correction for Quality-Diverse Web Agent Imitation

面向质量多样性的Web智能体模仿的推测性回滚修正

Longkun Hao, Hongyu Lin, Hao Li, Zhuowen Liu, Zhichao Yang, Haojie Hao, Dongshuo Huang, Haitao Yang, Hongyu Ge, Ming jie Xie, Yanjun Wu, Zi Hao Yin, Yan Bai, Yihang Lou

机构 * Beihang University(北京航空航天大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) The Hong Kong University of Science and Technology(香港科技大学) Northwestern Polytechnical University(西北工业大学) Tsinghua University(清华大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Peking University(北京大学)

AI总结 提出推测性回滚修正(SRC)框架,通过固定视野分支审查和回滚机制,在减少教师查询的同时保持轨迹多样性,在WebArena-Infinity上收集了977条通过验证的轨迹和9183个下一步动作示例。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.21858 2026-08-18 cs.CL 版本更新

Hypergraph as Language

超图作为语言

Mengqi Lei, Guohuan Xie, Shihui Ying, Shaoyi Du, Jun-Hai Yong, Chuan Shi, Ling Tian, Siqi Li, Yue Gao

机构 * Tsinghua University(清华大学) Yangtze Delta Region Institute(长江三角洲研究院) Shanghai Institute of Applied Mathematics and Mechanics(上海应用数学和力学研究所) State Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) National Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究所)

AI总结 本文提出了一种基于超图的语言模型对齐框架Hyper-Align,通过将超图结构转换为可被大语言模型理解的超图令牌,以更有效地处理高阶关联关系,从而在结构建模任务中取得显著优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02637 2026-08-18 cs.CL 版本更新

Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training

通过角色扮演训练LLM:探索AI素养对说服力的影响

Qihui Fan, Min Ge, Chenyan Jia, Weiyan Shi

机构 * Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(法国国家信息与自动化研究所巴黎-罗康库尔中心) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕默研究实验室)

AI总结 本文通过角色扮演训练LLMimic,提升用户AI素养,减少AI说服力效果,并增强诚实与社会责任感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.05513 2026-08-18 cs.RO cs.AI 版本更新

DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter

DECO:解耦多模态扩散变压器用于配备插件触觉适配器的双臂灵巧操作

Xukun Li, Yu Sun, Lei Zhang, Bosheng Huang, Yibo Peng, Yuan Meng, Haojun Jiang, Shaoxuan Xie, Guocai Yao, Alois Knoll, Zhenshan Bing, Xinlong Wang, Zhenguo Sun

机构 * Beijing Academy of Artificial Intelligence, Beijing, China(北京人工智能研究院) School of Computation, Information and Technology, Technical University of Munich, Garching, Germany(慕尼黑技术大学计算与信息学院) Department of Shenyang Institute of Computing Technology, University of Chinese Academy of Sciences, Beijing, China(中国科学院沈阳计算技术研究所部门) State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China(新型软件技术国家重点实验室) Department of Computer Science and Technology, Tsinghua University, Beijing, China(清华大学计算机科学与技术系)

AI总结 DECO通过解耦多模态输入和触觉适配器,实现了双臂灵巧操作的高效整合与高成功率

Comments 25 pages, 8 figures. Project Page: this https URL (https://baai-humanoid.github.io/DECO-webpage/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.13185 2026-08-18 cs.CV cs.GR 版本更新

FlexAM: Flexible Appearance-Motion Decomposition for Versatile Video Generation Control

FlexAM: 一种灵活的外观-运动分解方法用于多功能视频生成控制

Mingzhi Sheng, Zekai Gu, Peng Li, Cheng Lin, Hao-Xiang Guo, Ying-Cong Chen, Yuan Liu

机构 * HKUST(GZ)(香港科技大学(广州)) HKUST(香港科技大学) MUST(澳门大学) Tsinghua University(清华大学)

AI总结 FlexAM通过引入新型3D控制信号,实现外观与运动的解耦,提升视频生成任务的灵活性和生成质量。

Comments Codes: this https URL (https://github.com/IGL-HKUST/FlexAM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.18005 2026-08-18 cs.CV 版本更新

UrbanWorld2.0: A Multimodal Agentic Framework for Reality-Aligned 3D World Generation at City-Scale

RAISECity: 一种用于城市级现实对齐3D世界生成的多模态代理框架

Shengyuan Wang, Zhiheng Zheng, Yu Shang, Lixuan He, Yangcheng Yu, Fan Hangyu, Jie Feng, Qingmin Liao, Yong Li

机构 * College of AI, Tsinghua University(人工智能学院,清华大学) Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学) Department of Electronic Engineering, BNRist, Tsinghua University(电子工程系,北京研究院,清华大学)

AI总结 RAISECity通过多模态代理框架实现城市级3D世界生成,提升现实对齐、精度和性能,适用于沉浸媒体和具身智能应用。

Comments Accepted by ACM MM 2026, the code is available at: this https URL (https://github.com/tsinghua-fib-lab/UrbanWorld2.0)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23863 2026-08-18 cs.LG cs.AI 版本更新

PhyxMamba: Chaotic System Reconstruction from Short Context Observations with Generative State-Space Models

PhyxMamba:基于生成式状态空间模型的短上下文观测混沌系统重构

Chang Liu, Bohao Zhao, Jingtao Ding, Huandong Wang, Yong Li

机构 * Department of Electronic Engineering, BNRist Tsinghua University(电子工程系,清华大学)

AI总结 该研究提出PhyxMamba框架,结合Mamba状态空间模型与物理知情原理,在短观测数据下实现混沌系统重构,在Lorenz96系统上性能优于基线,鲁棒性强。

详情

展开后加载摘要…

URL PDF HTML 收藏