arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

Nanjing University(南京大学)

2026-04-09 至 2026-04-09 共收录 8
2604.06777 2026-04-09 cs.CV

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization

行走与言说:通过多模态代理策略优化弥合图像推理与行动之间的差距

Wenhao Yang, Yu Xia, Jinlong Huang, Shiyin Lu, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, Yuchen Zhou, Xiaobo Xia, Yuanyu Wan, Lijun Zhang, Tat-Seng Chua

机构 * National Key Laboratory for Novel Software Technology, Nanjing University, China(南京大学计算机软件新技术国家重点实验室) School of Artificial Intelligence, Nanjing University, China(南京大学人工智能学院) AI Business, Alibaba Group(阿里巴巴集团智能计算研究院) School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, China(上海交通大学自动化与智能感知学院) School of Intelligent Systems Engineering, Sun Yat-sen University, Shenzhen, China(中山大学智能工程学院(深圳)) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学技术学院) School of Software Technology, Zhejiang University, China(浙江大学软件学院) School of Computing, National University of Singapore, Singapore, Singapore(新加坡国立大学计算机学院)

AI总结 本文提出MAPO方法,通过强制生成视觉内容的显式文本描述,结合语义对齐与任务奖励,提升多模态推理能力,实验表明其在多个视觉推理基准上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06725 2026-04-09 cs.CV

Enhancing MLLM Spatial Understanding via Active 3D Scene Exploration for Multi-Perspective Reasoning

通过主动3D场景探索增强MLLM空间理解以实现多视角推理

Jiahua Chen, Qihong Tang, Weinong Wang, Qi Fan

机构 * Tsinghua University(清华大学) Nanjing University(南京大学) Tencent(腾讯)

AI总结 本文提出一种无需训练的框架,通过显式3D重建机制提升MLLM的空间理解能力,通过主动探索生成新视角,优于专门空间模型和通用MLLM。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06603 2026-04-09 cs.CL cs.AI

Scientific Knowledge-driven Decoding Constraints Improving the Reliability of LLMs

基于科学知识的解码约束:提升大语言模型的可靠性

Maotian Ma, Zheni Zeng, Zhenghao Liu, Yukun Yan

机构 * Nanjing University(南京大学) Tsinghua University(清华大学) Northeastern University(东北大学)

AI总结 本文提出SciDC方法,通过整合领域知识与强约束提升LLM在科学任务中的可靠性,实验显示在工业配方设计、肿瘤诊断和逆合成规划中平均准确率提升12%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19433 2026-04-09 cs.CV

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models

dMLLM-TTS:用于扩散多模态大语言模型的自验证和高效测试时间扩展

Yi Xin, Siqi Luo, Tianxiang Xu, Qi Qin, Haoxing Chen, Kaiwen Zhu, Zhiwei Zhang, Yangfan He, Rongchao Zhang, Jinbin Bai, Shuo Cao, Bin Fu, Junjun He, Yihao Liu, Yuewen Cao, Xiaohong Liu

机构 * Nanjing University(南京大学) Shanghai Innovation Institute(上海创新研究院) Shanghai AI Lab(上海人工智能实验室) Shanghai Jiao Tong University(上海交通大学) Peking University(北京大学) National University of Singapore(新加坡国立大学)

AI总结 本文提出dMLLM-TTS框架,通过轨迹探索扩展和迭代细化扩展两个互补的扩展轴,提升生成多样性和稳定性,同时通过自验证机制提高效率,实验表明在GenEval基准上生成质量显著提升且效率提高6倍。

Comments Project page: https://github.com/Alpha-VLLM/Lumina-DiMOO

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19365 2026-04-09 cs.CV cs.AI

DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation

DeCo:频率解耦的像素扩散用于端到端图像生成

Zehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang, Qi Tian

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室) Nanjing University(南京大学) Huawei Inc.(华为公司)

AI总结 DeCo通过解耦高频与低频成分生成,提升像素扩散效率,实现更高效的端到端图像生成,实验表明其在ImageNet上取得优于其他模型的性能。

Comments Accepted to CVPR2026. Project Page: https://zehong-ma.github.io/DeCo. Code Repository: https://github.com/Zehong-Ma/DeCo

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20340 2026-04-09 cs.SE cs.AI cs.PL

Once4All: Skeleton-Guided SMT Solver Fuzzing with LLM-Synthesized Generators

Once4All: 基于骨架引导的SMT求解器模糊测试与LLM合成生成器

Maolin Sun, Yibiao Yang, Yuming Zhou

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室)

AI总结 Once4All通过LLM合成生成器生成可重用的逻辑表达式,解决传统模糊测试中生成公式语法无效和计算开销大的问题,有效发现SMT求解器中的43个已确认bug。

Comments Accepted at ASPLOS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.07210 2026-04-09 cs.CV

VersaVogue: Visual Expert Orchestration and Preference Alignment for Unified Fashion Synthesis

VersaVogue: 视觉专家协作与偏好对齐的统一时尚合成

Jian Yu, Fei Shen, Cong Wang, Yi Xin, Si Shen, Xiaoyu Du, Jinhui Tang

机构 * Nanjing University of Science and Technology(南京理工大学) National University of Singapore(新加坡国立大学) Nanjing University(南京大学) Nanjing Forestry University(南京林业大学)

AI总结 本文提出VersaVogue框架,通过 trait-routing attention 模块和自动化多视角偏好优化管道,实现多条件可控的时尚合成,提升视觉真实性和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21105 2026-04-09 cs.CV

BrepGaussian: CAD reconstruction from Multi-View Images with Gaussian Splatting

BrepGaussian:从多视角图像中通过高斯点划进行CAD重建

Jiaxing Yu, Dongyang Ren, Hangyu Xu, Zhouyuxiao Yang, Yuanqi Li, Jie Guo, Zhengkang Zhou, Yanwen Guo

机构 * State Key Laboratory of Novel Software Technology, Nanjing University(南京大学计算机软件新技术国家重点实验室) Nanjing Urban Construction Tunnel& Bridge Intelligent Management Co., Ltd.(南京城建隧道桥梁智能管理有限公司)

AI总结 本文提出BrepGaussian框架,通过学习2D图像中的3D参数化表示,实现从多视角图像中重建CAD模型,其核心方法是结合可学习特征的高斯点划渲染器和特定拟合策略,有效分离了几何重建与特征学习。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏