arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

National University of Singapore(新加坡国立大学)

2026-06-30 至 2026-06-30 共收录 18
2606.30639 2026-06-30 cs.AI cs.CL

Self-Evolving World Models for LLM Agent Planning

面向LLM智能体规划的自我进化世界模型

Xuan Zhang, Wenxuan Zhang, See-Kiong Ng, Yang Deng

机构 * National University of Singapore(国立新加坡大学) Singapore University of Technology and Design(新加坡科技设计大学) Singapore Management University(新加坡管理学院)

AI总结 提出WorldEvolver框架,通过情景记忆、语义记忆和选择性预见三个模块在测试时修正世界模型,提升预测准确性和下游规划成功率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30632 2026-06-30 cs.RO cs.AI cs.CV

GROW$^2$: Grounding Which and Where for Robot Tool Use

GROW$^2$:为机器人工具使用确定哪个物体和哪个部位

Yuhong Deng, Yuyao Liu, David Hsu

机构 * National University of Singapore(新加坡国立大学) Massachusetts Institute of Technology(麻省理工学院)

AI总结 提出GROW$^2$方法,通过分层语义和几何接地实现开放世界工具选择与部位定位,在零样本泛化中优于基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30626 2026-06-30 cs.AI

DOPD: Dual On-policy Distillation

DOPD:双重在线策略蒸馏

Xinlei Yu, Gen Li, Qingyi Si, Guibin Zhang, Yuqi Xu, Congcong Wang, Shuai Dong, Kaiwen Tuo, Xiangyu Zeng, Kaituo Feng, Qunzhong Wang, Yang Shi, Xiaobin Hu, Xiangyu Yue, Jiaqi Wang, Shuicheng Yan

机构 * NUS(新加坡国立大学) MMLab, CUHK(香港中文大学多媒体实验室) PKU(北京大学) Explore Academy, JD(京东探索研究院)

AI总结 提出DOPD方法,通过优势感知的双重蒸馏范式动态分配令牌级监督,缓解特权幻觉,在LLM和VLM上优于传统在线策略蒸馏。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30491 2026-06-30 cs.CL cs.AI

SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation

SIMAX: 一种可扩展且可解释的多保真度标注医患对话模拟框架

Zhuhan Bao, Rui Yang, Bohao Yang, Zhiyi Liu, Sicheng Shu, Ruio Heerschap, Le Li, Doris Yang, Elisabeth Bond, Haoyuan Wang, Nicoleta Economou-Zavlanos, Joshua M. Biro, Matthew McDermott, Nan Liu, Anand Chowdhury, Kai Sun, Kathryn Pollak, Ed Hammond, Chuan Hong

机构 * Department of Biostatistics and Bioinformatics, Duke University School of Medicine(杜克大学医学学院生物统计学与生物信息学系) Duke-NUS AI + Medical Sciences Initiative, Duke-NUS Medical School(杜克-新加坡国立大学医学科学院AI+医学科学计划) Centre for Biomedical Data Science, Duke-NUS Medical School(杜克-新加坡国立大学医学学院生物医学数据科学中心) Department of Statistical Science, Duke University(杜克大学统计科学系) Leiden University Medical Centre(莱顿大学医学中心) Department of Mathematics, University of Texas at Austin(德克萨斯大学奥斯汀分校数学系) Department of Internal Medicine, Yale School of Medicine(耶鲁医学院内科学系) Department of Biostatistics, Epidemiology and Informatics, Perelman School of Medicine, University of Pennsylvania(宾夕法尼亚大学佩尔曼医学院生物统计学、流行病学与信息学系) The Graduate Group in Applied Mathematics and Computational Science, School of Arts and Sciences, University of Pennsylvania(宾夕法尼亚大学艺术与科学学院应用数学与计算科学联合组) Medstar Health National Center for Human Factors in Healthcare, Washington, DC, USA(Medstar健康国家人因工程中心,华盛顿特区,美国) Department of Biomedical Informatics, Columbia University(哥伦比亚大学生物医学信息学系) Cancer Prevention and Control, Duke Cancer Institute, Durham, NC, USA(杜克癌症研究所癌症预防与控制部,达勒姆,北卡罗来纳州,美国) Department of Population Health Sciences, Duke University School of Medicine(杜克大学医学学院流行病学与公共卫生系) Division of Rheumatology and Immunology, Duke University School of Medicine(杜克大学医学学院风湿病学与免疫学系) Pre-hospital and Emergency Research Centre, Health Services Research and Population Health, Duke-NUS Medical School(杜克-新加坡国立大学医学学院院前急救与应急研究中心,健康服务研究与人口健康) NUS Artificial Intelligence Institute, National University of Singapore(新加坡国立大学人工智能研究所) Division of Pulmonary, Allergy and Critical Care Medicine, Duke University School of Medicine(杜克大学医学学院呼吸科、过敏科与危重医学系) Duke Center for Health Informatics, Duke University(杜克大学健康信息学中心) Duke Clinical Research Institute, Durham, NC, USA(杜克临床研究中心,达勒姆,北卡罗来纳州,美国)

AI总结 提出SIMAX框架,通过预定义场景、角色和沟通行为生成可控医患对话,自动评估显示语音自然度和转录保真度良好,可用于开发和验证沟通编码系统。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30474 2026-06-30 cs.RO

Grasp-Oriented Non-Prehensile Manipulation via Learning a Graspability Field

基于抓取能力场学习的面向抓取的非抓取操作

Licheng Zhong, Gim Hee Lee

机构 * Department of Computer Science, National University of Singapore(新加坡国立大学计算机科学系)

AI总结 提出通过学习物体配置的抓取能力场,将非抓取操作优化为最大化抓取成功概率,而非达到预设位姿,实现闭环操作到抓取的单一策略。

Comments European Conference on Computer Vision (ECCV), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30367 2026-06-30 cs.RO

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation

FutureNav: 面向视觉与语言导航的统一世界-动作建模

Lingfeng Zhang, Zeying Gong, Xiaoshuai Hao, Haoxiang Fu, Qiang Zhang, Mingliang Zhou, Hangjun Ye, Xiaojun Liang, Junwei Liang, Wenbo Ding

机构 * Tsinghua University(清华大学) Pengcheng Laboratory(鹏城实验室) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Xiaomi EV(小米电动车) National University of Singapore(新加坡国立大学)

AI总结 提出FutureNav,一种基于VLM的统一世界-动作建模框架,通过联合优化动作策略、逆/前向动力学和未来状态预测四个目标,在4B规模骨干上实现多VLN基准的最优性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.30290 2026-06-30 cs.RO

X-Morph: Human Motion Priors for Scalable Robot Learning Across Morphologies

X-Morph:跨形态可扩展机器人学习的人类运动先验

Ritwik Sharma, Shivam Sood, Arhaan Jain, Shyam Charan Kesavamoorthi, Chengyang He, Guillaume Sartoretti

机构 * National University of Singapore(新加坡国立大学)

AI总结 提出X-Morph流水线,将人类运动转化为适用于多种非人形腿式机器人(四足、六足、四足操作器)的可部署运动与操作策略,通过跨形态重定向和强化学习实现,支持视频遥操作等下游任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29976 2026-06-30 cs.CV

Learning Efficient 4D Gaussian Representations from Monocular Videos with Flow Splatting

利用流溅射从单目视频学习高效4D高斯表示

Shengjun Zhang, Jinzhao Li, Xin Fei, Yueqi Duan

机构 * Tsinghua University(清华大学) National University of Singapore(新加坡国立大学)

AI总结 提出Flow Splatting方法,通过构建速度场并扩展溅射技术渲染光流,从单目视频中高效学习4D高斯表示,实现高质量动态场景重建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29593 2026-06-30 cs.LG cs.AI cs.NA math.NA math.OC stat.ML

How AI settled the complexity of the oldest SGD algorithm

AI 如何解决了最古老 SGD 算法的复杂度问题

Michał Dereziński, Xiaoyu Dong

机构 * University of Michigan(密歇根大学) National University of Singapore(新加坡国立大学)

AI总结 本文利用现代 AI 模型(如 ChatGPT 和 Gemini)协作,发现了 Kaczmarz 算法(最早 SGD 算法)的最坏情况复杂度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29020 2026-06-30 cs.CV cs.AI cs.ET cs.MM

Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Video Synthesis

语义感知、物理信息、几何基础的天气视频合成

Chenghao Qian, Nedko Savov, Lingdong Kong, Yeying Jin, Rui Song, Wenjing Li, Zhun Zhong, Jiaqi Ma, Gustav Markkula, Luc Van Gool

机构 * University of Leeds, UK(利兹大学,英国) National University of Singapore, Singapore(新加坡国立大学) University of California, Los Angeles, USA(美国加州大学洛杉矶分校) Hefei University of Technology, China(合肥工业大学)

AI总结 提出一种语义感知、物理信息、几何基础的框架,通过语义、动力学和几何三个条件信号引导视频编辑器合成多样且真实的天气效果,并提升自动驾驶语义分割的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28826 2026-06-30 cs.CV

RefGlass-GS: A UAV-Enabled Fusion Framework for Photorealistic, Semantic and Interactive Digitization of Reflective Glass Facades via Gaussian Splatting

RefGlass-GS:一种通过高斯泼溅实现反射玻璃幕墙逼真、语义和交互式数字化的无人机融合框架

Zhenyu Liang, Xiao Zhang, Boyu Wang, Zhaolun Liang, Ang Li, Jeff Chak Fu Chan, Mingzhu Wang, Jack C. P. Cheng

机构 * Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology(香港科技大学土木与环境工程系) Department of Civil and Environmental Engineering, University of Maryland(马里兰大学土木与环境工程系) Department of the Built Environment, National University of Singapore(新加坡国立大学环境与建筑系) Department of Architecture and Civil Engineering, City University of Hong Kong(香港城市大学建筑与土木工程系)

AI总结 提出RefGlass-GS框架,结合最大后验估计分割玻璃面板、无人机视角规划优化、反射MLP与延迟着色改进高斯泼溅,实现反射玻璃幕墙的高质量数字化,在面板分割、视角合成和建模质量上显著优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.28390 2026-06-30 cs.CV cs.AI

Automated Quality Assessment of Geospatial Vector Data: A GeoAI Approach using Spatial Representation Learning

地理空间矢量数据自动化质量评估:基于空间表示学习的GeoAI方法

Hao Li, Chen Chu, Filip Biljecki, Cyrus Shahabi, Wenwen Li

机构 * Department of Geography, National University of Singapore, Singapore(新加坡国立大学地理系) University of Southern California(南加州大学) Department of Architecture, National University of Singapore, Singapore(新加坡国立大学建筑系) Department of Real Estate, National University of Singapore, Singapore(新加坡国立大学房地产系)

AI总结 提出Topo4Vec框架,通过拓扑错误模拟和空间表示学习,自动化评估地理空间矢量数据质量,在三个城市数据集上达到0.99的建筑物重叠检测准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.11025 2026-06-30 cs.LG 新提交

Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models

Flow-DPPO:用于流匹配模型的散度近端策略优化

Bowen Ping, Xiangxin Zhou, Penghui Qi, Minnan Luo, Liefeng Bo, Tianyu Pang

机构 * Xi’an Jiaotong University(西安交通大学) Tencent Hunyuan(腾讯混元) National University of Singapore(新加坡国立大学)

AI总结 针对流匹配模型中PPO比率裁剪的结构性缺陷,提出Flow-DPPO方法,利用高斯策略的精确KL散度计算实现散度近端约束,提升奖励和训练稳定性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.31603 2026-06-30 cs.CV cs.AI

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models

Lumos-Nexus: 面向视频统一模型的高效频率桥接与同质潜在空间

Jiazheng Xing, Hangjie Yuan, Lingling Cai, Xinyu Liu, Yujie Wei, Fei Du, Tao Feng, Hai Ci, Jiasheng Tang, Weihua Chen, Fan Wang, Yong Liu

机构 * Zhejiang University(浙江大学) DAMO Academy, Alibaba Group(阿里云达摩院) Hupan Lab(虎扑实验室) National University of Singapore(新加坡国立大学) Hong Kong University of Science and Technology(香港科技大学) Fudan University(复旦大学) Tsinghua University(清华大学)

AI总结 提出Lumos-Nexus框架,通过两阶段训练和渐进频率桥接,在保持推理能力的同时显著提升视频生成保真度。

Comments ECCV 2026 Camera-Ready Version. Project page (https://jiazheng-xing.github.io/nexus-lumos-home/) and Code (https://github.com/alibaba-damo-academy/Lumos-Custom/) are available

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.14252 2026-06-30 cs.LG cs.AI

Not All Timesteps Matter Equally: Selective Alignment Knowledge Distillation for Spiking Neural Networks

并非所有时间步都同等重要:用于脉冲神经网络的选择性对齐知识蒸馏

Kai Sun, Peibo Duan, Yongsheng Huang, Guowei Zhang, Benjamin Smith, Nanxu Gong, Levin Kuhlmann

机构 * Faculty of Information Technolody, Monash University, Australia(墨尔本大学信息科技学院,澳大利亚) School of Software, Northeastern University, China(东北大学软件学院,中国) Department of Medicine, National University of Singapore, Singapore(新加坡国立大学医学部,新加坡)

AI总结 本文提出Selective Alignment Knowledge Distillation方法,通过选择性对齐类别和时间知识,改进SNN性能,实验证明在静态图像和神经形态事件数据集上优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.16506 2026-06-30 cs.CV

VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations

VIEW2SPACE:从稀疏观察研究多视角视觉推理

Fucai Ke, Zhixi Cai, Boying Li, Long Chen, Beibei Lin, Weiqing Wang, Pari Delir Haghighi, Gholamreza Haffari, Hamid Rezatofighi

机构 * Monash University(莫纳什大学) National University of Singapore(新加坡国立大学)

AI总结 本文提出VIEW2SPACE基准,通过物理仿真生成高保真3D场景,评估多视角视觉推理的现状,发现多数模型表现欠佳,提出Grounded Chain-of-Thought with Visual Evidence提升性能并推广至现实数据。

Journal ref European Conference on Computer Vision (ECCV), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06732 2026-06-30 cs.CL cs.AI cs.IR

Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization

LLMs的排名可靠性如何?通过两阶段令牌优化实现排名操纵

Tiancheng Xing, Jerry Li, Yixuan Du, Xiyang Hu

机构 * National University of Singapore(新加坡国立大学) University of Southern California(南加州大学) Georgetown University(乔治城大学) Arizona State University(亚利桑那州立大学)

AI总结 本文提出RAF方法,通过两阶段令牌优化生成自然语言提示,提升目标项在LLM排名中的位置,同时保持语言自然性,实验显示其在提升排名和保持自然性方面比现有方法更鲁棒。

Comments ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02804 2026-06-30 cs.CL

Multimodal Mathematical Reasoning with Diverse Solving Perspective

多模态数学推理与多样化的解题视角

Wenhao Shi, Zhiqiang Hu, Yi Bin, Guoqing Wang, Xing Xu, Yang Yang, See-Kiong Ng

机构 * Tongji University(同济大学) National University of Singapore(新加坡国立大学) Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究所) Tencent(腾讯) Center for Future Media, University of Electronic Science and Technology of China(电子科技大学未来媒体中心)

AI总结 本文提出MathV-DP数据集和Qwen-VL-DP模型,通过多样化解题视角提升多模态数学推理的准确性和生成多样性。

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏