In-Video Instructions: Visual Signals as Generative Control
视频中的指令:将视觉信号作为生成控制
机构 * National University of Singapore(新加坡国立大学)
AI总结 本研究提出通过视频中嵌入的视觉信号作为指令,实现可控的图像到视频生成,通过空间感知的指令分配提升多对象场景下的生成可靠性。
高校专区
视频中的指令:将视觉信号作为生成控制
机构 * National University of Singapore(新加坡国立大学)
AI总结 本研究提出通过视频中嵌入的视觉信号作为指令,实现可控的图像到视频生成,通过空间感知的指令分配提升多对象场景下的生成可靠性。
被掩盖的扩散模型实际上是秘密学习的顺序自回归模型
机构 * Aalto University(奥卢大学) ; NUS, Singapore(新加坡国立大学) ; IIT Bombay(印度理工学院班加罗尔)
AI总结 本文揭示被掩盖的扩散模型实际上是一种具有可学习顺序的自回归模型,通过优化解码顺序提升生成性能。
Comments Accepted at EurIPS 2025 Workshop on Principles of Generative Modeling (PriGM)
最大损失非质心聚类中的核心可能为空
机构 * TU Clausthal(图鲁斯大学) ; National University of Singapore(新加坡国立大学)
AI总结 该研究证明在最大损失目标下非质心聚类中核心可能为空,并给出了相关理论界和构造。
用负责任的AI考虑来防御大型语言模型对抗劫持攻击
机构 * National University of Singapore(新加坡国立大学)
AI总结 本文提出三种防御策略,通过提示级、logit引导和领域特定代理方法,有效降低大型语言模型的劫持攻击成功率。
Comments 20 pages including appendix; technical report; NeurIPS 2024 style
概念而非文档:基于AMR的概念熵进行上下文压缩
机构 * University of Southern Queensland(南方昆士兰大学) ; University of Technology Sydney(技术悉尼大学) ; The Hong Kong Polytechnic University(香港理工大学) ; Wuhan University of Technology(武汉理工大学) ; National University of Singapore(新加坡国立大学) ; The Education University of Hong Kong(香港教育大学)
AI总结 本文提出基于AMR的概念熵方法,通过压缩上下文保留核心语义,提升RAG任务的准确性和效率。
Edit2Perceive: 图像编辑扩散模型是强大的密集感知器
机构 * Peking University(北京大学) ; Show Lab, National University of Singapore(新加坡国立大学Show实验室)
AI总结 Edit2Perceive通过统一的扩散框架,利用图像编辑模型实现更高效的密集感知任务,展示了在深度、法线和磨边任务上的最新成果。
通过视频推理:首次评估视频模型在迷宫解决任务中的推理能力
机构 * DeepWisdom ; Tsinghua University(清华大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Renmin University of China(中国人民大学) ; University of Oxford(牛津大学) ; National University of Singapore(新加坡国立大学) ; Xiamen University(厦门大学) ; Hong Kong University of Science and Technology (GuangZhou)(香港科技大学(广州))
AI总结 本文首次评估视频模型在迷宫解决任务中的推理能力,提出VR-Bench基准,展示视频生成在空间推理中的潜力。
FOCUS: 长视频理解中的高效关键帧选择
机构 * National University of Singapore(新加坡国立大学) ; TikTok
AI总结 FOCUS通过两阶段探索-利用策略,在严格标记预算下高效选择关键帧,提升长视频理解的准确性。
卷积计算的几何学:VCNet中的流形解缠与预测动态
机构 * Department of Computer Science University of Wisconsin-Madison(计算机科学系 威斯康星大学麦迪逊分校) ; Department of Computer Science National University of Singapore(计算机科学系 新加坡国立大学)
AI总结 VCNet通过融合神经科学原理和几何框架,实现了更高效且鲁棒的视觉计算,展示了在图像分类任务中优于现有模型的性能。
Comments Published in the proceedings of the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Symmetry and Geometry in Neural Representations (NeurReps). Additionally accepted for presentation in NeurIPS 2025 Workshop: Interpreting Cognition in Deep Learning Models (CogInterp)
传达计划,而非感知:基于具身世界模型的可扩展多智能体协调
机构 * Department of Computer Science University of Wisconsin-Madison(计算机科学系 明尼苏达大学) ; Department of Computer Science National University of Singapore(计算机科学系 新加坡国立大学)
AI总结 本文提出基于具身世界模型的意图通信方法,通过端到端学习与工程化设计对比,展示在复杂环境下更优的协调能力。
Comments Published in the Proceedings of the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Scaling Environments for Agents (SEA). Additionally accepted for presentation in the NeurIPS 2025 Workshop: Embodied World Models for Decision Making (EWM) and the NeurIPS 2025 Workshop: Optimization for Machine Learning (OPT)
4D-VGGT:一种具有时空意识的通用基础模型,用于动态场景几何估计
机构 * National Key Lab of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(多谱信息智能处理国家实验室,人工智能与自动化学院,华中科技大学) ; School of Computing, National University of Singapore(计算学院,新加坡国立大学)
AI总结 4D-VGGT通过分而治之的时空表示方法,提升动态场景几何估计的准确性和通用性。
SciEducator: 基于Deming循环多智能体系统的科学视频理解与教育
机构 * Jinan University(济南大学) ; National University of Singapore(新加坡国立大学) ; Nanyang Technological University(南洋理工大学) ; Peking University(北京大学) ; University of Electronic Science and Technology of China(电子科技大学) ; South China University of Technology(华南理工大学) ; Guangming Laboratory(光明实验室) ; Zhejiang University(浙江大学)
AI总结 SciEducator通过Deming循环多智能体系统实现科学视频的自演化理解与教育,生成多模态教学内容并超越现有大语言模型和视频智能体。
MamTiff-CAD: 多尺度潜在扩散与Mamba+用于复杂参数序列
机构 * Northwestern Polytechnical University(西北工业大学) ; National University of Singapore(新加坡国立大学) ; Shanghai Jiao Tong University(上海交通大学) ; Nanchang University(南昌大学)
AI总结 MamTiff-CAD通过结合Mamba+和Transformer的多尺度潜在扩散模型,有效生成复杂CAD参数序列,实现长序列生成任务的高性能表现。
Comments ICCV 2025 Conference
RELEAP: 通过强化学习增强的标签高效主动表型分析用于电子健康记录
机构 * Department of Biostatistics and Bioinformatics, Duke University(生物统计学与生物信息学系,杜克大学) ; Cancer Prevention and Control Research Program, Duke Cancer Institute(癌症预防与控制研究计划,杜克癌症研究所) ; Department of Population Health Sciences, Duke University School of Medicine(流行病学与公共卫生科学系,杜克大学医学院) ; Centre for Quantitative Medicine, Duke-NUS Medical School(定量医学中心,杜克-新加坡医学学校) ; Programme in Health Services and Systems Research, Duke-NUS Medical School(健康服务与系统研究计划,杜克-新加坡医学学校) ; Department of Statistics and Data Science, National University of Singapore(统计与数据科学系,新加坡国立大学) ; Department of Biostatistics, Peking University Health Science Center(生物统计学系,北京大学医学部) ; Beijing International Center for Mathematical Research, Peking University(北京国际数学研究中心,北京大学)
AI总结 RELEAP通过强化学习增强标签效率,利用下游预测性能反馈优化电子健康记录表型修正,提升风险预测可靠性。
Comments 20 pages, 5 figures, 1 table. Includes supplementary material. Submitted to JAMIA Open. † These authors contributed equally. *Corresponding author: Chuan Hong
注意差距:通过用户需求对齐知识库以增强心理健康检索
机构 * Princeton University(普林斯顿大学) ; National University of Singapore(国立新加坡大学) ; MOH Office for Healthcare Transformation(卫生部医疗转型办公室)
AI总结 通过用户需求对齐知识库,提升心理健康检索性能,减少内容创建需求,实现高质量信息检索。
Comments 25 pages, 3 figures, submitted to NeurIPS 2025 GenAI4Health
AutoHFormer:高效的层次自回归变换器用于时间序列预测
机构 * School of Computing and Data Science, The University of Hong Kong (HKU)(计算与数据科学学院,香港大学) ; Zhejiang Key Laboratory of Intelligent Education Technology and Application, Zhejiang Normal University (ZJNU)(智能教育技术与应用浙江省重点实验室,浙江师范大学) ; Department of Computer Science, National University of Singapore (NUS)(计算机科学系,新加坡国立大学) ; Department of Computer Science, Aalborg University (AU)(计算机科学系,奥尔堡大学) ; Department of Computer Science and Technology, Cambridge University (Cambridge)(计算机科学与技术系,剑桥大学)
AI总结 AutoHFormer通过层次时间建模、动态窗口注意力和自适应时间编码,实现了高效且精确的时间序列预测,训练速度提升10.76倍,内存减少6.06倍。
Comments 14 pages
Journal ref ICDE'2026
文化消逝之处:揭示文本到图像生成中的文化差距
机构 * The University of Sydney(悉尼大学) ; Nanjing University of Science and Technology(南京理工大学) ; Central South University(中南大学) ; Nanjing University(南京大学) ; National University of Singapore(新加坡国立大学)
AI总结 本文提出了一种解决多语言文本到图像生成中文化一致性问题的方法,通过局部化文化敏感信号并改进模型对文化相关层的激活,提升生成图像的文化一致性。
VLA-4D: 将4D意识嵌入视觉-语言-动作模型中以实现时空一致的机器人操作
机构 * School of Computing, National University of Singapore(新加坡国立大学计算机学院) ; School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
AI总结 VLA-4D通过4D意识增强视觉-语言-动作模型,实现时空一致的机器人操作,提升动作执行的空间平滑性和时间一致性。
大语言模型中用于知识存储的参数专业化趋势的兴起
机构 * Alibaba Group(阿里巴巴集团) ; New York University(纽约大学) ; National University of Singapore(新加坡国立大学) ; Singapore Management University(新加坡管理大学) ; Singapore University of Technology and Design(新加坡科技设计大学)
AI总结 本研究分析了大语言模型中MLP参数的知识存储方式,发现参数专业化提高了知识利用效率,并通过实验验证了其重要性。
Comments Accepted in NeurIPS 2025
V-ReasonBench:面向视频生成模型的统一推理基准套件
机构 * NUS(新加坡国立大学) ; HKUST(GZ)(香港科技大学(广州)) ; HKU(香港大学) ; USYD(澳大利亚悉尼大学) ; CUHK(香港中文大学) ; LIGHTSPEED Project(LIGHTSPEED项目)
AI总结 V-ReasonBench通过四个维度评估视频生成模型的推理能力,提供可验证任务以衡量结构化问题解决、空间认知、模式推理和物理动态,促进更可靠的推理模型发展。
Comments Project Page: https://oahzxl.github.io/VReasonBench
VisPlay: 从图像中自我进化视觉-语言模型
机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; Washington University in St. Louis(华盛顿大学圣路易斯分校) ; University of Maryland(马里兰大学) ; National University of Singapore(新加坡国立大学)
AI总结 VisPlay通过自我进化强化学习框架,利用未标注图像数据提升视觉-语言模型的推理能力,实现多模态智能的可扩展发展。
BanditSpec: 通过多臂老虎机算法实现自适应推测解码
机构 * National University of Singapore(国立新加坡大学) ; Sea AI Lab(Sea人工智能实验室) ; Singapore Management University(新加坡管理学院) ; Yale University(耶鲁大学)
AI总结 BanditSpec通过多臂老虎机算法实现自适应推测解码,优化超参数选择以提升LLM推理效率。
Comments 35 pages, 4 figures, accepted to ICML, typos and affiliations are corrected
超越雾霾:生成式夜间图像去雾
机构 * National University of Singapore(新加坡国立大学) ; Microsoft Research Asia(微软亚洲研究院)
AI总结 BeyondHaze通过生成式模型提升夜间图像去雾效果,结合强背景先验和引导训练,实现雾霾和光晕下的清晰重建。
似曾相识:LLM在识别已见过的文件时表现出确定性
机构 * Huazhong University of Science and Technology(华中科技大学) ; National University of Singapore(国立新加坡大学) ; Macquarie University(麦考瑞大学)
AI总结 COPYCHECK利用LLM的不确定性信号,通过双策略检测训练数据中的受版权内容,实现高准确率的版权检测。
CaKE:电路感知编辑实现通用知识学习
机构 * Zhejiang University(浙江大学) ; National University of Singapore(新加坡国立大学) ; University of California, Los Angeles(美国加州大学洛杉矶分校)
AI总结 CaKE通过电路感知编辑提升LLMs对更新知识的多跳推理能力,实现20%的准确率提升并降低内存消耗
Comments EMNLP 2025
计算机使用代理作为生成用户界面的法官
机构 * University of Oxford(牛津大学) ; Show Lab, National University of Singapore(新加坡国立大学Show实验室) ; Microsoft(微软公司)
AI总结 本研究提出Coder-CUA协作框架,通过代理作为法官与编码模型协作,提升自动GUI设计的效率和可靠性。
Comments Project: https://showlab.github.io/AUI Github: https://github.com/showlab/AUI
机构 * Nanjing University of Science and Technology, China(南京理工大学,中国) ; NExT++ Research Centre, National University of Singapore, Singapore(NExT++研究中心,新加坡国立大学)
机构 * Hangzhou International Innovation Institute, Beihang University, China(北京航空航天大学杭州国际创新研究院) ; CoreControl Inc, Hangzhou, China(杭州核心控制公司) ; Department of Mechanical Engineering, National University of Singapore, Singapore(新加坡国立大学机械工程系)
Comments Accepted for presentation at the 2025 IEEE International Conference on Robotics and Automation (ICRA)
Journal ref 2025 IEEE International Conference on Robotics and Automation ICRA pp. 1-7
机构 * South China University of Technology(华南理工大学) ; Sun Yat-sen University(中山大学) ; Hangzhou Dianzi University(杭州电子科技大学) ; Zhejiang University of Finance & Economics(浙江财经大学) ; National University of Singapore(新加坡国立大学) ; Shenzhen Research Institute of Big Data(深圳大数据研究院) ; City University of Hong Kong(香港城市大学)
Comments CVPR 2026 Under Review
机构 * Department of Computer Science, School of Computing, National University of Singapore(新加坡国立大学计算机科学系) ; Transport, Health and Urban Systems Research Lab, Melbourne School of Design, University of Melbourne(墨尔本大学设计学院交通、健康与城市系统研究实验室) ; School of Civil Engineering, Faculty of Engineering, University of Sydney(悉尼大学土木工程学院) ; Delft Institute of Applied Mathematics, Delft University of Technology(代尔夫特理工大学应用数学研究所)
Comments Accepted for publication at the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2026