arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

National University of Singapore(新加坡国立大学)

共收录 2374
2505.22143 2025-12-08 cs.CV

3D Question Answering via only 2D Vision-Language Models

仅通过2D视觉-语言模型实现3D问答

Fengyun Wang, Sicheng Yu, Jiawei Wu, Jinhui Tang, Hanwang Zhang, Qianru Sun

机构 * Nanyang Technological University, Singapore Singapore Management University, Singapore National University of Singapore, Singapore Nanjing University of Science \& Technology, Nanjing, China

AI总结 本文提出cdViews方法,通过仅使用2D视觉-语言模型,实现3D问答任务的高性能表现。

Comments ICML2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.12868 2025-12-08 cs.HC cs.AI

As Confidence Aligns: Exploring the Effect of AI Confidence on Human Self-confidence in Human-AI Decision Making

信心相辅:探索AI信心对人类自信心在人类-AI决策中的影响

Jingshu Li, Yitian Yang, Q. Vera Liao, Junti Zhang, Yi-Chieh Lee

机构 * National University of Singapore(新加坡国立大学) Microsoft Research(微软研究院)

AI总结 本研究发现人类自信心与AI置信度相辅相成,且这种一致性在AI退出后仍持续,同时实时反馈会降低一致性,表明两者并非独立。

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.06060 2025-12-08 cs.HC cs.AI

Wild Narratives: Exploring the Effects of Animal Chatbots on Empathy and Positive Attitudes toward Animals

野性叙事:探讨动物聊天机器人对同理心和对动物积极态度的影响

Jingshu Li, Aaditya Patwari, Yi-Chieh Lee

机构 * Computer Science, National University of Singapore(新加坡国立大学计算机科学系)

AI总结 本研究探讨了动物聊天机器人如何通过情感表达和真实细节提升用户对动物的同理心和积极态度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04918 2025-12-05 cs.LG

Multi-Agent Reinforcement Learning for Intraday Operating Rooms Scheduling under Uncertainty

多代理强化学习用于在日手术室调度中的不确定性处理

Kailiang Liu, Ying Chen, Ralf Borndörfer, Thorsten Koch

机构 * Department of Mathematics, National University of Singapore(新加坡国立大学数学系) Department of Mathematics & Risk Management Institute & Center for Quantitative Finance, National University of Singapore(新加坡国立大学数学系及风险管理研究所及量化金融中心) Department of Mathematics and Computer Science, Freie Universität Berlin(柏林自由大学数学与计算机科学系) Department of Network Optimization, Zuse Institute Berlin(柏林Zuse研究所网络优化部) Faculty of Mathematics and Natural Sciences, Technische Universität Berlin(柏林技术大学数学与自然科学学院) Department of Applied Algorithmic Methods, Zuse Institute Berlin(柏林Zuse研究所应用算法方法部)

AI总结 本文提出基于多代理强化学习的手术室调度方法,通过集中训练和分布式执行策略,有效应对不确定性,提升调度效率和及时性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04537 2025-12-05 cs.CV

X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale

X-Humanoid:将人类视频机器人化以生成大规模人形视频

Pei Yang, Hai Ci, Yiren Song, Mike Zheng Shou

机构 * Show Lab, National University of Singapore(新加坡国立大学Show实验室)

AI总结 X-Humanoid通过生成式视频编辑方法,将人类视频机器人化,生成大规模人形视频数据集,提升人形机器人研究的训练效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04381 2025-12-05 cs.RO

FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination

FALCON:基于基础模型的解耦视觉-运动政策用于具有基础模型协调的移动- manipulation

Chengyang He, Ge Sun, Yue Bai, Junkai Lu, Jiadong Zhao, Guillaume Sartoretti

机构 * Department of Mechanical Engineering, College of Design and Engineering, National University of Singapore(机械工程系,设计与工程学院,新加坡国立大学)

AI总结 FALCON通过基于基础模型的解耦视觉-运动策略,实现移动- manipulation任务中更高效的协调与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15555 2025-12-05 cs.LG cs.AI stat.ML

Bayesian Concept Bottleneck Models with LLM Priors

具有LLM先验的贝叶斯概念瓶颈模型

Jean Feng, Avni Kothari, Luke Zier, Chandan Singh, Yan Shuo Tan

机构 * University of California, San Francisco(加州大学旧金山分校) Microsoft Research(微软研究院) National University of Singapore(新加坡国立大学)

AI总结 本文提出BC-LLM模型,利用贝叶斯框架和LLM作为先验,实现高效的概念提取和可解释性提升。

Comments 2025 Conference on Neural Information Processing Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03564 2025-12-04 cs.LG cs.CR

Towards Irreversible Machine Unlearning for Diffusion Models

向扩散模型的不可逆机器遗忘学习迈进

Xun Yuan, Zilong Zhao, Jiayu Li, Aryan Pasikhani, Prosanta Gope, Biplab Sikdar

机构 * Department of Electrical and Computer Engineering, College of Design and Engineering, National University of Singapore(新加坡国立大学电子与计算机工程系,设计与工程学院) Betterdata, Singapore(新加坡Betterdata公司) Department of Computer Science, University of Sheffield(谢菲尔德大学计算机科学系)

AI总结 本文提出DiMRA攻击可逆转基于微调的扩散模型机器遗忘学习方法,并提出DiMUM方法通过记忆替代数据提升模型鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03058 2025-12-04 cs.LG

Dynamical Properties of Tokens in Self-Attention and Effects of Positional Encoding

自注意力机制中令牌的动力学特性及位置编码的影响

Duy-Tung Pham, An The Nguyen, Viet-Hoang Tran, Nhan-Phu Chung, Xin T. Tong, Tan M. Nguyen, Thieu N. Vo

机构 * FPT Software AI Center(FPT软件AI中心) National University of Singapore(国立新加坡大学) Ho Chi Minh University of Economics(胡志明经济大学)

AI总结 本文研究了Transformer中令牌的动力学特性,揭示了位置编码对模型性能的影响,并提出了改进方法以缓解收敛问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18980 2025-12-04 cs.CL cs.AI cs.CV

IW-Bench: Evaluating Large Multimodal Models for Converting Image-to-Web

IW-Bench: 评估大规模多模态模型的图像到网页转换能力

Hongcheng Guo, Wei Zhang, Junhao Chen, Yaonan Gu, Jian Yang, Junjia Du, Shaosheng Cao, Binyuan Hui, Tianyu Liu, Jianxin Ma, Chang Zhou, Zhoujun Li

机构 * Beihang University(北京航空航天大学) Alibaba Group(阿里巴巴集团) Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学) National University of Singapore(新加坡国立大学)

AI总结 IW-Bench提出了一种新的基准,用于评估大规模多模态模型在图像到网页转换中的性能,通过元素准确率和布局准确率以及五跳提示方法来提升评估效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02874 2025-12-03 cs.CL

Think in Parallel, Answer as One: Logit Averaging for Open-Ended Reasoning

并行思考,一答为终:用于开放性推理的Logit平均

Haonan Wang, Chao Du, Kenji Kawaguchi, Tianyu Pang

机构 * National University of Singapore(新加坡国立大学) Sea AI Lab(Sea AI实验室)

AI总结 ThinkMerge通过并行推理轨迹的logit平均提升开放式推理性能,适用于代码生成和网络深度研究等任务。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02870 2025-12-03 cs.CV

Taming Camera-Controlled Video Generation with Verifiable Geometry Reward

驯服受控摄像机视频生成的可验证几何奖励

Zhaoqing Wang, Xiaobo Xia, Zhuolin Bie, Jinlin Liu, Dongdong Yu, Jia-Wang Bian, Changhu Wang

机构 * AIsphere National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

AI总结 本文提出了一种在线强化学习框架,通过可验证的几何奖励提升摄像机控制的视频生成精度与一致性。

Comments 11 pages, 4 figures, 7 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02777 2025-12-03 cs.RO cs.MA

CogDrive: Cognition-Driven Multimodal Prediction-Planning Fusion for Safe Autonomy

CogDrive: 基于认知的多模态预测-规划融合用于安全自主性

Heye Huang, Yibin Yang, Mingfeng Fan, Haoran Wang, Xiaocong Zhao, Jianqiang Wang

机构 * Singapore-MIT Alliance for Research and Technology (SMART), Singapore(新加坡-麻省理工联合研究技术联盟) Department of Urban Studies and Planning, Massachusetts Institute of Technology, USA(麻省理工学院城市研究与规划系) School of Vehicle and Mobility, Tsinghua University, China(清华大学车辆与移动系统学院) Department of Mechanical Engineering, National University of Singapore, Singapore(新加坡国立大学机械工程系) Key Laboratory of Road and Traffic Engineering, Ministry of Education, Tongji University, China(同济大学交通工程重点实验室)

AI总结 CogDrive通过结合认知多模态预测与安全导向规划,实现了在复杂交通中的安全自主性,提升了轨迹预测和适应性行为。

Comments 25 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19785 2025-12-03 cs.LG cs.AI

medDreamer: Model-Based Reinforcement Learning with Latent Imagination on Complex EHRs for Clinical Decision Support

medDreamer:基于复杂电子病历的潜在想象的模型驱动强化学习用于临床决策支持

Qianyi Xu, Gousia Habib, Feng Wu, Dilruk Perera, Mengling Feng

机构 * National University of Singapore(新加坡国立大学)

AI总结 medDreamer通过结合潜在想象和自适应特征集成模块,实现基于复杂电子病历的模型驱动强化学习,以提升个性化治疗推荐的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15450 2025-12-03 cs.CL

SkyLadder: Better and Faster Pretraining via Context Window Scheduling

SkyLadder: 通过上下文窗口调度实现更优更高效的预训练

Tongyao Zhu, Qian Liu, Haonan Wang, Shiqi Chen, Xiangming Gu, Tianyu Pang, Min-Yen Kan

机构 * National University of Singapore(新加坡国立大学) Sea AI Lab(Sea AI实验室) City University of Hong Kong(香港城市大学)

AI总结 SkyLadder通过上下文窗口调度策略,在保持基准性能的同时,提升了长上下文任务的表现,并加快了训练速度。

Comments Accepted to NeurIPS 2025. 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02299 2025-12-03 cs.CL cs.AI

HealthContradict: Evaluating Biomedical Knowledge Conflicts in Language Models

HealthContradict: 评估语言模型在生物医学知识中的矛盾

Boya Zhang, Alban Bornet, Rui Yang, Nan Liu, Douglas Teodoro

机构 * Faculty of Medicine, University of Geneva(日内瓦大学医学院) Department of Biomedical Informatics, Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学 Yong Loo Lin 医学院生物医学信息学系) Department of Biostatistics & Bioinformatics, Duke University Artificial Intelligence Institute, National University of Singapore(新加坡国立大学人工智能研究所生物统计学与生物信息学系)

AI总结 HealthContradict数据集评估语言模型在处理生物医学知识矛盾时的推理能力,揭示模型在正确上下文利用与错误上下文抵抗方面的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01913 2025-12-02 eess.IV cs.CV

Disentangling Progress in Medical Image Registration: Beyond Trend-Driven Architectures towards Domain-Specific Strategies

解构医学图像配准的进步:超越趋势驱动的架构,迈向领域特定的策略

Bailiang Jian, Jiazhen Pan, Rohit Jena, Morteza Ghahremani, Hongwei Bran Li, Daniel Rueckert, Christian Wachinger, Benedikt Wiestler

机构 * Technical University of Munich(慕尼黑技术大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Imperial College London(伦敦帝国理工学院) University of Pennsylvania(宾夕法尼亚大学) National University of Singapore(新加坡国立大学)

AI总结 本文通过模块化框架解构医学图像配准中趋势驱动架构与领域特定设计的影响,发现后者在性能提升上优于前者,推动研究重点转向领域特定原则。

Comments Submitted to Medical Image Analysis. Journal Extension of arXiv:2407.19274

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01831 2025-12-02 cs.LG cs.AI

Deconstructing Generative Diversity: An Information Bottleneck Analysis of Discrete Latent Generative Models

解构生成多样性:基于信息瓶颈理论的离散潜在生成模型分析

Yudi Wu, Wenhao Zhao, Dianbo Liu

机构 * Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学)

AI总结 本文提出基于信息瓶颈理论的框架,分析离散潜在生成模型中生成多样性的来源,并揭示了三种不同的策略:多样性优先、压缩优先和解耦。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01728 2025-12-02 cs.CL

Reasoning About the Unsaid: Misinformation Detection with Omission-Aware Graph Inference

对未言之的推理:基于遗漏感知图推断的虚假信息检测

Zhengjia Wang, Danding Wang, Qiang Sheng, Jiaying Wu, Juan Cao

机构 * Media Synthesis and Forensics Lab, Institute of Computing Technology, Chinese Academy of Sciences(媒体合成与取证实验室,计算技术研究所,中国科学院) University of Chinese Academy of Sciences(中国科学院大学) National University of Singapore(新加坡国立大学)

AI总结 本文提出OmiGraph框架,通过省略感知图推断实现虚假信息检测,提升检测性能。

Comments AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01657 2025-12-02 cs.CV

DB-KAUNet: An Adaptive Dual Branch Kolmogorov-Arnold UNet for Retinal Vessel Segmentation

DB-KAUNet: 一种自适应双分支Kolmogorov-Arnold UNet用于视网膜血管分割

Hongyu Xu, Panpan Meng, Meng Wang, Dayu Hu, Liming Liang, Xiaoqi Sheng

机构 * School of Computer Science and Software Engineering, Southwest University(西南大学计算机科学与软件工程学院) Innovation Centre of Ministry of Education for Development and Diseases, the Sixth Affiliated Hospital, School of Medicine, South China University of Technology(华南理工大学医学院附属第六医院) Centre for Innovation and Precision Eye Health, Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学 Yong Loo Lin 医学院创新与精准眼科健康中心) Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore(新加坡国立大学 Yong Loo Lin 医学院眼科部) College of Medicine and Biological Information Engineering, Northeastern University(东北大学医学院与生物信息工程学院) School of Electrical Engineering Automation, Jiangxi University of Science and Technology(江西理工大学电气工程自动化学院) School of Future Technology, South China University of Technology(华南理工大学未来技术学院)

AI总结 DB-KAUNet通过自适应双分支结构结合CNN和Transformer,提升视网膜血管分割的准确性和鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01514 2025-12-02 cs.LG

Label Forensics: Interpreting Hard Labels in Black-Box Text Classifier

标签取证:解读黑盒文本分类器中的硬标签

Mengyao Du, Gang Yang, Han Fang, Quanjun Yin, Ee-chien Chang

机构 * National University of Defense Technology(国防科技大学) National University of Singapore(新加坡国立大学)

AI总结 本文提出标签取证框架,通过语义邻域采样和迭代优化,推断黑盒文本分类器中标签的语义意义,实现高一致性标签解释和负责任的AI审计。

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10194 2025-12-02 cs.CV

B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding

B2N3D: 从二元关系到N元关系的3D物体接地的渐进式学习

Feng Xiao, Hongbin Xu, Hai Ci, Wenxiong Kang

机构 * School of Automation Science and Engineering, South China University of Technology(自动化科学与工程学院,华南理工大学) ByteDance Seed(字节跳动种子) Show Lab, National University of Singapore(Show Lab,新加坡国立大学)

AI总结 B2N3D通过引入N元关系学习提升3D物体接地的准确性,利用分组监督损失和混合注意力机制实现更精确的多模态关系建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11842 2025-12-02 cs.CV cs.AI cs.LG

MoH: Multi-Head Attention as Mixture-of-Head Attention

MoH:多头注意力作为专家混合注意力

Peng Jin, Bo Zhu, Li Yuan, Shuicheng Yan

机构 * School of Electronic and Computer Engineering, Shenzhen Graduate School, Peking University, Shenzhen, China(电子与计算机工程系,深圳研究生院,北京大学,深圳,中国) Pengcheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国) School of AI for Science, Shenzhen Graduate School, Peking University, Shenzhen, China(科学人工智能学院,深圳研究生院,北京大学,深圳,中国) National University of Singapore, Singapore(新加坡国立大学,新加坡)

AI总结 MoH通过将注意力头视为专家,提升推理效率并优化性能,仅使用部分注意力头即可超越传统多头注意力。

Comments Accepted by ICML 2025, code: https://github.com/SkyworkAI/MoH

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01335 2025-12-02 cs.CR cs.AI cs.CL

EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations

EmoRAG:评估RAG对符号扰动的鲁棒性

Xinyun Zhou, Xinfeng Li, Yinan Peng, Ming Xu, Xuanwang Zhang, Miao Yu, Yidong Wang, Xiaojun Jia, Kun Wang, Qingsong Wen, XiaoFeng Wang, Wei Dong

机构 * ZJU Hangzhou China(浙江大学杭州校区) NTU Singapore(南洋理工大学) Hengxin Tech. Singapore(新加坡恒心科技) NUS Singapore(国立新加坡大学) NJU Nanjing China(南京大学) PKU Beijing China(北京大学)

AI总结 EmoRAG研究揭示RAG系统对细微表情符号扰动的鲁棒性问题,发现单个表情符号可导致检索严重误导,并提出针对性防御措施。

Comments Accepted to ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01314 2025-12-02 cs.CV

TokenPure: Watermark Removal through Tokenized Appearance and Structural Guidance

TokenPure: 通过令牌化外观和结构引导进行水印移除

Pei Yang, Yepeng Liu, Kelly Peng, Yuan Gao, Yiren Song

机构 * First Intelligence University of California, Santa Barbara(加州大学圣巴巴拉分校) National University of Singapore(新加坡国立大学)

AI总结 TokenPure通过令牌化外观和结构引导,实现高效且一致的水印移除,显著提升水印移除的保真度和一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.14712 2025-12-02 cs.CV

FreeSwim: Revisiting Sliding-Window Attention Mechanisms for Training-Free Ultra-High-Resolution Video Generation

FreeSwim: 重新审视滑动窗口注意力机制以实现无训练超高清视频生成

Yunfeng Wu, Jiayi Song, Zhenxiong Tan, Zihao He, Songhua Liu

机构 * School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院) Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学) Xi’an Jiaotong University(西安交通大学) National University of Singapore(新加坡国立大学)

AI总结 FreeSwim通过无训练滑动窗口注意力机制实现超高清视频生成,提升效率并保持视觉细节。

Comments 23 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03657 2025-12-02 cs.CV

Dynamic Multimodal Prototype Learning in Vision-Language Models

视觉-语言模型中的动态多模态原型学习

Xingyu Zhu, Shuo Wang, Beier Zhu, Miaoge Li, Yunfan Li, Junfeng Fang, Zhicai Wang, Dongsheng Wang, Hanwang Zhang

机构 * University of Science and Technology of China(中国科学技术大学) Nanyang Technological University(南洋理工大学) The Hong Kong Polytechnic University(香港理工大学) Sichuan University(四川大学) National University of Singapore(新加坡国立大学) Shenzhen University(深圳大学)

AI总结 本文提出ProtoMM框架,通过动态更新视觉粒子和多模态原型学习,提升视觉-语言模型在测试时的适应性能。

Comments ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22643 2025-12-02 cs.CV

SPIRAL: Semantic-Aware Progressive LiDAR Scene Generation and Understanding

SPIRAL: 基于语义的渐进式LiDAR场景生成与理解

Dekai Zhu, Yixuan Hu, Youquan Liu, Dongyue Lu, Lingdong Kong, Slobodan Ilic

机构 * Technical University of Munich(慕尼黑技术大学) Fudan University(复旦大学) National University of Singapore(新加坡国立大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

AI总结 Spiral提出了一种基于范围视图的LiDAR扩散模型,能够同时生成深度、反射图像和语义地图,实现高效的3D场景生成与理解。

Comments NeurIPS 2025; 24 pages, 10 figures, 9 tables; Code at https://github.com/worldbench/SPIRAL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08594 2025-12-02 cs.CV cs.LG

PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction

PointNSP: 通过下一尺度细节预测实现自回归3D点云生成

Ziqiao Meng, Qichao Wang, Zhiyang Dou, Zixing Song, Zhipeng Zhou, Irwin King, Peilin Zhao

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学) University of Hong Kong(香港大学) University of Cambridge(剑桥大学) The Chinese University of Hong Kong(香港中文大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 PointNSP通过下一尺度细节预测实现自回归3D点云生成,首次在自回归范式中达到最先进的生成质量,并在参数、训练和推理效率上超越扩散基线。

Comments 24 pages; Previously this version appeared as arXiv:2510.05613 which was submitted as a new work by accident

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.17725 2025-12-02 cs.RO

Differentiable Contact Dynamics for Stable Object Placement Under Geometric Uncertainties

可微接触动力学用于在几何不确定性下稳定物体放置

Linfeng Li, Gang Yang, Lin Shao, David Hsu

机构 * School of Computing, National University of Singapore(计算学院,新加坡国立大学) Smart Systems Institute, National University of Singapore(智能系统研究所,新加坡国立大学)

AI总结 本文提出了一种基于可微接触动力学的方法,通过梯度下降估计几何不确定性,实现稳定物体放置。

详情

展开后加载摘要…

URL PDF HTML 收藏