arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高校专区

University of Science and Technology of China(中国科学技术大学)

2026-03-19 至 2026-03-19 共收录 13
2603.18001 2026-03-19 cs.CV

EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding

EchoGen:基于循环一致性的统一布局-图像生成与理解学习

Kai Zou, Hongbo Liu, Dian Zheng, Jianxiong Gao, Zhiwei Zhao, Bin Liu

机构 * School of Cyber Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院) Tongji University(同济大学) Sun Yat-sen University(中山大学) Fudan University(复旦大学) Hefei University of Technology(合肥工业大学) Anhui Province Key Laboratory of Digital Security(安徽省数字安全重点实验室)

AI总结 本文提出EchoGen框架,通过统一模型联合训练布局到图像生成和图像定位任务,提升生成图像的布局准确性与文本描述一致性,同时增强图像定位的鲁棒性。

Comments 9 pages, Accepted at the 40th AAAI Conference on Artificial Intelligence (AAAI 2026)

Journal ref Proceedings of the AAAI Conference on Artificial Intelligence, 40(16), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17746 2026-03-19 cs.CV

Concept-to-Pixel: Prompt-Free Universal Medical Image Segmentation

概念到像素:无提示通用医学图像分割

Haoyun Chen, Fenghe Tang, Wenxin Ma, Shaohua Kevin Zhou

机构 * School of Biomedical Engineering, Division of Life Sciences Medicine, University of Science Technology of China (USTC), Hefei, Anhui 230026, China Center for Medical Imaging, Robotics, Analytic Computing \& Learning (MIRACLE), Suzhou Institute for Advanced Research, USTC, Suzhou, Jiangsu 215123, China Jiangsu Provincial Key Laboratory of Multimodal Digital Twin Technology, Suzhou Jiangsu, 215123, China State Key Laboratory of Precision

AI总结 本文提出C2P框架,通过分离解剖学知识为几何和语义表示,利用多模态大语言模型生成语义令牌,并引入几何令牌约束,实现无提示的通用医学图像分割,实验表明其在多种模态和数据集上表现优异。

Comments 32 pages, code is available at: https://github.com/Yundi218/Concept-to-Pixel

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17718 2026-03-19 cs.CV

DiffVP: Differential Visual Semantic Prompting for LLM-Based CT Report Generation

DiffVP:基于LLM的CT报告生成的微分视觉语义提示

Yuhe Tian, Kun Zhang, Haoran Ma, Rui Yan, Yingtai Li, Rongsheng Wang, Shaohua Kevin Zhou

机构 * Department of Electronic Engineering Information Science, School of Information Science Technology, University of Science Technology of China (USTC), Hefei, Anhui 230026, China School of Biomedical Engineering, Division of Life Sciences Medicine, University of Science Technology of China (USTC), Hefei, Anhui 230026, China Center for Medical Imaging, Robotics, Analytic Computing \& Learning (MIRACLE), Suzhou Institute for Advanced Research, University of Science Technology of China (USTC), Suzhou, Jiangsu 215123, China Jiangsu Provincial Key Laboratory of Multimodal Digital Twin Technology, University of Science Technology of China (USTC), Suzhou, Jiangsu 215123, China State Key Laboratory of Precision Intelligent Chemistry, University of Science

AI总结 DiffVP通过微分视觉提示方法,利用扫描与参考之间的高阶语义差异指导LLM生成CT报告,提升报告准确性,实验表明其在BLEU和临床效果上均优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11158 2026-03-19 cs.IR cs.AI

Role-Augmented Intent-Driven Generative Search Engine Optimization

角色增强的意图驱动生成搜索引擎优化

Xiaolu Chen, Haojie Wu, Jie Bao, Zhen Chen, Yong Liao, Hu Huang

机构 * School of Cyber Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院)

AI总结 本文提出角色增强的意图驱动生成搜索引擎优化方法,通过反思性细化不同信息角色来建模搜索意图,提升生成搜索引擎中的内容可见性。

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.05305 2026-03-19 cs.CV cs.AI

Frequency Autoregressive Image Generation with Continuous Tokens

基于连续标记的频率自回归图像生成

Hu Yu, Hao Luo, Hangjie Yuan, Yu Rong, Jie Huang, Feng Zhao

机构 * University of Science and Technology of China(中国科学技术大学) Alibaba Group, DAMO Academy(阿里巴巴集团,达摩院)

AI总结 本文提出频率自回归(FAR)范式,通过连续标记器构建图像,利用频谱依赖性提升生成效果,并在ImageNet上验证了其在图像生成中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17554 2026-03-19 cs.CV

Prompt-Free Universal Region Proposal Network

无提示通用区域提议网络

Qihong Tang, Changhan Liu, Shaofeng Zhang, Wenbin Li, Qi Fan, Yang Gao

机构 * Nanjing University(南京大学) University of Science and Technology of China(中国科学技术大学)

AI总结 本文提出无提示通用区域提议网络PF-RPN,通过自适应查询嵌入和级联自提示模块实现无需外部提示的对象定位,适用于多种目标检测场景。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17546 2026-03-19 cs.CV

ProGVC: Progressive-based Generative Video Compression via Auto-Regressive Context Modeling

ProGVC:基于渐进式的生成视频压缩 via 自动回归上下文建模

Daowen Li, Ruixiao Dong, Ying Chen, Kai Li, Ding Ding, Li Li

机构 * Alibaba Group(阿里巴巴集团) University of Science and Technology of China(中国科学技术大学)

AI总结 ProGVC提出一种统一渐进传输、高效熵编码和细节合成的视频压缩框架,通过多尺度自回归上下文模型实现低比特率下的高质量视频压缩与可扩展性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17408 2026-03-19 cs.CV cs.AI

Joint Degradation-Aware Arbitrary-Scale Super-Resolution for Variable-Rate Extreme Image Compression

联合退化感知任意尺度超分辨率用于可变速率极值图像压缩

Xinning Chai, Zhengxue Cheng, Xin Li, Rong Xie, Li Song

机构 * School of Information Science and Electronic Engineering, Shanghai Jiao Tong University(上海交通大学信息科学与电子工程学院) Department of Electronic Engineer and Information Science, University of Science and Technology of China(中国科学技术大学电子工程与信息科学系) MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(上海交通大学人工智能教育部重点实验室)

AI总结 本文提出ASSR-EIC框架,利用任意尺度超分辨率支持可变速率极值图像压缩,通过联合退化感知解码器实现率自适应重建,提升低比特率下的图像恢复质量。

Comments Accepted by IEEE Transactions on BroadCasting

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17385 2026-03-19 cs.LG

The Causal Uncertainty Principle: Manifold Tearing and the Topological Limits of Counterfactual Interventions

因果不确定性原理:流形撕裂与反事实干预的拓扑极限

Rui Wu, Hong Xie, Yongjun Li

机构 * University of Science and Technology of China(中国科学技术大学)

AI总结 本文探讨了反事实干预的几何挑战,提出了反事实事件地平线和流形撕裂定理,建立了因果不确定性原理,并引入了Geometry-Aware Causal Flow算法。

Comments 33 pages, 6 figures. Submitted to the Journal of Machine Learning Research (JMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.17384 2026-03-19 cs.LG

Cohomological Obstructions to Global Counterfactuals: A Sheaf-Theoretic Foundation for Generative Causal Models

上同调障碍与全局反事实:一种基于层论的生成因果模型基础

Rui Wu, Hong Xie, Yongjun Li

机构 * School of Management, University of Science and Technology of China(管理学院,中国科学技术大学) School of Computer Science and Engineering, University of Science and Technology of China(计算机科学与工程学院,中国科学技术大学)

AI总结 本文基于层论构建生成因果模型,揭示因果图非平凡同调导致的上同调障碍,并提出熵正则化和熵沃尔什因果层拉普拉斯方程,实现高维数据反事实导航。

Comments 34 pages, 5 figures. Submitted to JMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15662 2026-03-19 cs.AI

Stepwise Think-Critique: A Unified Framework for Robust and Interpretable LLM Reasoning

逐步思考-批判:一种用于鲁棒且可解释的大语言模型推理的统一框架

Jiaqi Xu, Cuiling Lan, Xuejin Chen, Yan Lu

机构 * University of Science and Technology of China(中国科学技术大学) Microsoft Research Asia(微软亚洲研究院)

AI总结 本文提出STC框架,通过在推理与自我批判间交替进行,提升大语言模型的鲁棒性和可解释性,实验显示其在数学推理任务中表现出色。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.20095 2026-03-19 cs.CV

WPT: World-to-Policy Transfer via Online World Model Distillation

WPT: 通过在线世界模型蒸馏实现世界到策略的迁移

Guangfeng Jiang, Yueru Luo, Jun Liu, Yi Huang, Yiyao Zhu, Zhan Qu, Dave Zhenyu Chen, Bingbing Liu, Xu Yan

机构 * University of Science and Technology of China(科学技术大学) CUHK-SZ(香港中文大学(深圳)) HKUST(香港理工大学) Huawei Foundation Model Department(华为基金会模型部)

AI总结 本文提出WPT方法,通过在线世界模型蒸馏实现策略迁移,提升规划性能并保持实时部署能力,实验表明其在开放环和闭合环基准上表现优异。

Comments CVPR2026 Accepted

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25758 2026-03-19 cs.AI

TheraMind: A Strategic and Adaptive Agent for Longitudinal Psychological Counseling

TheraMind:一种用于长期心理辅导的战略和自适应代理

He Hu, Chiyuan Ma, Qianning Wang, Lin Liu, Yucheng Zhou, Laizhong Cui, Fei Ma, Qi Tian

机构 * Shenzhen University(深圳大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Auckland University of Technology(奥克兰理工大学) University of Science and Technology of China(中国科学技术大学) SKL-IOTSC, CIS, University of Macau(澳门科技研究所、CIS、澳门大学) Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ)(广东人工智能与数字经济实验室(深圳))

AI总结 本文提出TheraMind,一种用于可信在线长期心理辅导的战略和自适应代理,通过双循环架构提升对话管理和治疗规划能力,实验证明其在多会话指标上优于其他方法。

详情

展开后加载摘要…

URL PDF HTML 收藏