arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

The Hong Kong University of Science and Technology(香港科技大学)

共收录 2796
2603.12743 2026-03-16 cs.CV cs.AI cs.CL

MoKus: Leveraging Cross-Modal Knowledge Transfer for Knowledge-Aware Concept Customization

MoKus: 利用跨模态知识迁移实现知识感知的概念定制

Chenyang Zhu, Hongxiang Li, Xiu Li, Long Chen

机构 * Tsinghua University(清华大学) HKUST(香港科技大学)

AI总结 本文提出MoKus框架,通过跨模态知识迁移实现知识感知的概念定制,解决稀有token在预训练数据中缺失导致的性能不稳定问题,并引入KnowCusBench基准测试,验证了方法在概念定制任务中的优越性。

Comments Project Page: https://chenyangzhu1.github.io/MoKus/

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12636 2026-03-16 math.OC cs.LG

Weakly Time-Coupled Approximation of Markov Decision Processes

弱时间耦合的马尔可夫决策过程近似

Negar Soheili, Selvaprabu Nadarajah, Bo Yang

机构 * Information and Decision Sciences Department, University of Illinois at Chicago(伊利诺伊大学芝加哥分校信息与决策科学系) Industrial Engineering and Decision Analytics, Hong Kong University of Science and Technology(香港科学与技术大学工业工程与决策分析系)

AI总结 本文提出弱时间耦合近似方法,用于解决高维外生不确定性马尔可夫决策过程的可扩展性问题,通过独立于时间跨度的交叉阶段依赖关系,改进了ALP和PO方法的上界估计。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12482 2026-03-16 cs.CV

CalliMaster: Mastering Page-level Chinese Calligraphy via Layout-guided Spatial Planning

CalliMaster:通过布局引导的空间规划掌握页面级中文书法

Tianshuo Xu, Tiantian Hong, Zhifei Chen, Fei Chao, Ying-cong Chen

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Faculty of Engineering and IT, University of Technology Sydney(悉尼大学工程与信息学院) Xiamen University(厦门大学) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 本文提出CalliMaster框架,通过解耦空间规划与内容合成,解决页面级书法生成中精度与布局的平衡问题,支持可控生成与编辑,扩展至文物修复与鉴定。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12397 2026-03-16 cs.CL

Not Just the Destination, But the Journey: Reasoning Traces Causally Shape Generalization Behaviors

不只是终点,而是旅程:推理痕迹因果地塑造泛化行为

Pengcheng Wen, Yanxu Zhu, Jiapeng Sun, Han Zhu, Yujin Zhou, Chi-Min Chan, Sirui Han, Yike Guo

机构 * Hong Kong University of Science and Technology(香港科技大学) Beijing Jiaotong University(北京交通大学)

AI总结 研究探讨推理痕迹对模型泛化行为的因果影响,通过设计控制实验发现不同推理类型导致不同行为模式,证明推理本身具有独立信号。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12271 2026-03-16 cs.CL cs.AI cs.LG

Diagnosing Retrieval Bias Under Multiple In-Context Knowledge Updates in Large Language Models

在大型语言模型中多上下文知识更新下的检索偏见诊断

Boyu Qiao, Sean Guo, Xian Yang, Kun Li, Wei Zhou, Songlin Hu, Yunya Song

机构 * Institute of Information Engineering, Chinese Academy of Sciences(信息工程研究所,中国科学院) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络与安全学院) Hong Kong University of Science and Technology(香港科学与技术大学) The University of Manchester(曼彻斯特大学)

AI总结 研究探讨了多次上下文知识更新对大型语言模型检索偏的影响,发现更新次数增加时偏见加剧,且最新状态准确性显著下降,通过分析发现模型在处理最新更新时存在识别困难。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23790 2026-03-16 cs.CV

Fourier Angle Alignment for Oriented Object Detection in Remote Sensing

基于傅里叶角对齐的遥感定向目标检测

Changyu Gu, Linwei Chen, Lin Gu, Ying Fu

机构 * Beijing Institute of Technology(北京理工大学) University of Hong Kong(香港大学) Tohoku University(东北大学) The Hong Kong University of Science and Technology(香港科学与技术大学) Hong Kong Generative AI Research and Development Center(香港生成式人工智能研究与发展中心)

AI总结 本文提出傅里叶角对齐方法,通过频谱分析角度信息并对其对齐,提升遥感中旋转目标检测的性能,实验显示在DOTA数据集上取得新的SOTA结果。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12257 2026-03-13 cs.CV

DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning

DreamVideo-Omni: 基于潜在身份强化学习的多主体视频定制与 Omni-运动控制

Yujie Wei, Xinyu Liu, Shiwei Zhang, Hangjie Yuan, Jinbo Xing, Zhekai Chen, Xiang Wang, Haonan Qiu, Rui Zhao, Yutong Feng, Ruihang Chu, Yingya Zhang, Yike Guo, Xihui Liu, Hongming Shan

机构 * Fudan University(复旦大学) The Hong Kong University of Science and Technology(香港科技大学) Tongyi Lab, Alibaba Group(阿里云实验室) Zhejiang University(浙江大学) MMLab, The University of Hong Kong(香港大学MMLab) Nanyang Technological University(南洋理工大学) Show Lab, National University of Singapore(新加坡国立大学Show Lab)

AI总结 DreamVideo-Omni通过渐进式两阶段训练范式,实现多主体视频定制与Omni-运动控制,采用条件感知嵌入和分层运动注入策略,提升身份保持与运动控制精度。

Comments Project Page: https://dreamvideo-omni.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12166 2026-03-13 cs.CV

LatentGeo: Learnable Auxiliary Constructions in Latent Space for Multimodal Geometric Reasoning

LatentGeo: 在潜在空间中学习可学习的辅助构造以进行多模态几何推理

Haiying Xu, Zihan Wang, Song Dai, Zhengxuan Zhang, Kairan Dou, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Nankai University(南开大学) Communication University of China(中国传媒大学)

AI总结 LatentGeo通过学习潜在空间中的连续视觉表示,解决多模态几何推理中辅助构造的表示问题,采用三阶段课程和强化学习方法提升几何推理任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12155 2026-03-13 cs.CV cs.AI

GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows

GlyphBanana: 通过代理工作流推进精确文本渲染

Zexuan Yan, Jiarui Jin, Yue Ma, Shijian Wang, Jiahui Hu, Wenxiang Jiao, Yuan Lu, Linfeng Zhang

机构 * Shanghai Jiao Tong University(上海交通大学) Xiaohongshu Inc.(小红书公司) Hong Kong University of Science and Technology(香港科技大学) Southeast University(东南大学) South China University of Technology(华南理工大学)

AI总结 GlyphBanana通过代理工作流整合辅助工具,提升复杂文本和公式渲染的精度,适用于多种文本到图像模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11915 2026-03-13 cs.CL

CoMMET: To What Extent Can LLMs Perform Theory of Mind Tasks?

CoMMET:大型语言模型在理论思维任务中能发挥多大作用?

Ruirui Chen, Weifeng Jiang, Chengwei Qin, Cheston Tan

机构 * Agency for Science, Technology and Research (A*STAR)(科技研究局) Nanyang Technological University(南洋理工大学) Hong Kong University of Science and Technology (Guangzhou), China(香港科技大学(广州),中国)

AI总结 本文提出CoMMET多模态基准数据集,用于评估大型语言模型在多轮对话中的理论思维能力,通过扩展心理状态评估和引入多轮测试,分析模型优劣势并指明未来改进方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11798 2026-03-13 cs.AI

DocSage: An Information Structuring Agent for Multi-Doc Multi-Entity Question Answering

DocSage:一个多文档多实体问答的信息结构代理

Teng Lin, Yizhang Zhu, Zhengxuan Zhang, Yuyu Luo, Nan Tang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

AI总结 DocSage通过动态模式发现、结构化信息提取和模式感知推理,解决多文档多实体问答中的跨文档证据链构建和实体关系推断问题,实现超过27%的准确率提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11632 2026-03-13 cs.HC cs.RO

From Pets to Robots: MojiKit as a Data-Informed Toolkit for Affective HRI Design

从宠物到机器人:MojiKit作为影响性人机交互设计的数据驱动工具包

Liwen He, Pingting Chen, Ziheng Tang, Yixiao Liu, Jihong Jeung, Teng Han, Xin Tong

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Tsinghua University(清华大学) Beijing Institute of Technology(北京理工大学) Sichuan University(四川大学) Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 MojiKit通过结合结构化参考卡片、拟人化机器人原型和行为控制工作室,为影响性人机交互设计提供数据驱动的系统化工具,帮助用户设计更多样化的机器人影响性行为。

Comments 25 pages, 11 figures, Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11193 2026-03-13 cs.CL

DeReason: A Difficulty-Aware Curriculum Improves Decoupled SFT-then-RL Training for General Reasoning

DeReason: 一种考虑难度的课程改进了解耦的SFT-然后-RL训练以促进一般推理

Hanxu Hu, Yuxuan Wang, Maggie Huan, Jannis Vamvas, Yinya Huang, Zhijiang Guo, Rico Sennrich

机构 * University of Zurich(苏黎世大学) University of Pennsylvania(宾夕法尼亚大学) ETH Zurich(苏黎世联邦理工学院) HKUST (GZ)(香港科技大学(广州))

AI总结 DeReason通过基于难度的数据解耦策略,优化SFT与RL的训练分配,提升一般推理能力。

Comments 13 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10682 2026-03-12 cs.RO

OnFly: Onboard Zero-Shot Aerial Vision-Language Navigation toward Safety and Efficiency

OnFly:面向安全与效率的机载零样本空中视觉语言导航

Guiyong Zheng, Yueting Ban, Mingjie Zhang, Juepeng Zheng, Boyu Zhou

机构 * School of Artificial Intelligence, Sun Yat-Sen University(中山大学人工智能学院) Southern University of Science and Technology(南方科技大学) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 OnFly通过共享感知双智能体架构和混合记忆机制,实现高效稳定的零样本空中视觉语言导航,提升安全性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10473 2026-03-12 cs.CL cs.AI

Aligning Large Language Models with Searcher Preferences

将大型语言模型对齐于搜索者偏好

Wei Wu, Peilun Zhou, Liyi Chen, Qimeng Wang, Chengqiang Lu, Yan Gao, Yi Wu, Yao Hu, Hui Xiong

机构 * School of Artificial Intelligence and Data Science, University of Science and Technology of China(中国科学技术大学人工智能与数据科学学院) Xiaohongshu Inc.(小红书公司) Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)人工智能研究所) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(香港科技大学计算机科学与工程系)

AI总结 SearchLLM通过分层多维奖励系统提升开放式生成搜索的鲁棒性和用户需求对齐能力,实测有效消费率提升1.03%。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10379 2026-03-12 cs.LG cs.AI

Optimal Expert-Attention Allocation in Mixture-of-Experts: A Scalable Law for Dynamic Model Design

混合专家模型中专家-注意力计算分配的最优性:动态模型设计的可扩展定律

Junzhuo Li, Peijie Jiang, Changxin Tian, Jia Liu, Zhiqiang Zhang, Xuming Hu

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) The Hong Kong University of Science and Technology(香港科学与技术大学) Ant Group(蚂蚁集团)

AI总结 本文提出了一种新的MoE模型缩放定律,通过研究专家-注意力计算比例的最优分配,为动态模型设计提供了可扩展的理论框架。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08017 2026-03-12 cs.HC cs.AI

Alignment-Process-Outcome: Rethinking How AIs and Humans Collaborate

对齐-过程-结果:重新思考AI和人类如何协作

Haichang Li, Anjun Zhu, Arpit Narechania

机构 * George Mason University(乔治·马歇尔大学) Simon Fraser University(西蒙·弗雷泽大学) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 本文提出通过任务和意图两个视角重新理解AI与人类协作的结构关系,揭示对齐、过程和结果之间的动态联系。

Comments Accepted by Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems (CHI EA 26), Barcelona, Spain, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.19442 2026-03-12 cs.CV

UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment

UrbanAlign: 域内任务中VLM与人类偏好对齐的后处理语义校准

Yecheng Zhang, Rong Zhao, Zhizhou Sha, Yong Li, Lei Wang, Ce Hou, Wen Ji, Hao Huang, Yunshan Wan, Jian Yu, Junhao Xia, Yuru Zhang, Chunlei Shi

机构 * Tsinghua University(清华大学) University College London(伦敦大学学院) Hong Kong University of Science and Technology(香港科学与技术大学) Peking University(北京大学) Southwest Jiaotong University(西南交通大学) Zhejiang University(浙江大学) University of Texas at Austin(德克萨斯大学奥斯汀分校) Renmin University of China(中国人民大学) Southeast University(东南大学)

AI总结 UrbanAlign通过后处理方法实现VLM与人类偏好的对齐,无需修改模型权重,提升领域任务的准确性和可解释性。

Comments 26 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10955 2026-03-12 cs.CR cs.AI

Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents

超越最大令牌:通过工具调用链实现隐秘的资源放大

Kaiyu Zhou, Yongsen Zheng, Yicheng He, Meng Xue, Xueluan Gong, Yuji Wang, Xuanye Zhang, Kwok-Yan Lam

机构 * Nanyang Technological University, Singapore(南洋理工大学,新加坡) University of Illinois Urbana-Champaign, United States(伊利诺伊大学厄巴纳-香槟分校,美国) The Hong Kong University of Science and Technology, Hong Kong(香港科学与技术大学,香港) Shanghai Jiao Tong University, China(上海交通大学,中国)

AI总结 本文提出了一种基于工具调用链的隐秘DoS攻击,通过多轮交互显著提升LLM的资源消耗,挑战现有安全防护机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12918 2026-03-12 cs.CL

ThinkPatterns-21k: A Systematic Study on the Impact of Thinking Patterns in LLMs

ThinkPatterns-21k: 对LLMs中思维模式影响的系统研究

Pengcheng Wen, Jiaming Ji, Chi-Min Chan, Juntao Dai, Donghai Hong, Yaodong Yang, Sirui Han, Yike Guo

机构 * Hong Kong University of Science and Technology(香港科技大学) Peking University(北京大学) Zhejiang University(浙江大学)

AI总结 ThinkPatterns-21k研究了LLMs中思维模式的影响,通过系统分析发现结构化思维对小模型有益,而大模型使用结构化思维会降低性能,独白思维在各类模型中均有效。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09968 2026-03-11 cs.CV

ReCoSplat: Autoregressive Feed-Forward Gaussian Splatting Using Render-and-Compare

ReCoSplat: 基于渲染与比较的自回归前馈高斯点云生成

Freeman Cheng, Botao Ye, Xueting Li, Junqi You, Fangneng Zhan, Ming-Hsuan Yang

机构 * University of California Merced(加州大学默塞德分校) ETH Zurich(苏黎世联邦理工学院) NVIDIA Shanghai Jiao Tong University(上海交通大学) Hong Kong University of Science and Technology(香港科技大学)

AI总结 ReCoSplat通过引入渲染与比较模块,解决高斯点云生成中姿态误差问题,实现高效稳定的视角合成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09809 2026-03-11 cs.CV

RA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source Understanding

RA-SSU:迈向细粒度音频视觉学习的区域感知声音源理解

Muyi Sun, Yixuan Wang, Hong Wang, Chen Su, Man Zhang, Xingqun Qi, Qi Li, Zhenan Sun

机构 * School of Artificial Intelligence, Beijing University of Posts and Telecommunications(北京邮电大学人工智能学院) Academy of Interdisciplinary Studies, The Hong Kong University of Science and Technology(香港科技大学跨学科研究院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

AI总结 RA-SSU通过构建细粒度音频视觉数据集和SSUFormer模型,实现区域感知的声音源理解,提升音频视觉学习的细粒度表现。

Comments Accepted by IEEE TMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09496 2026-03-11 cs.CV

SurgFed: Language-guided Multi-Task Federated Learning for Surgical Video Understanding

SurgFed: 基于语言引导的多任务联邦学习用于手术视频理解

Zheng Fang, Ziwei Niu, Ziyue Wang, Zhu Zhuo, Haofeng Liu, Shuyang Qian, Jun Xia, Yueming Jin

机构 * National University of Singapore(国立新加坡大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Zhejiang University(浙江大学) Nanyang Technological University(南洋理工大学)

AI总结 SurgFed通过语言引导的通道选择和超聚合方法,提升手术视频多任务联邦学习的跨站点和跨任务适应性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09400 2026-03-11 cs.CL

Reward Prediction with Factorized World States

基于分解世界状态的奖励预测

Yijun Shen, Delong Chen, Xianming Hu, Jiaming Mi, Hongbo Zhao, Kai Zhang, Pascale Fung

机构 * East China Normal University(华东师范大学) HKUST(香港科技大学)

AI总结 本文提出StateFactory方法,通过分解世界状态实现跨领域的准确奖励预测,提升智能体规划性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01068 2026-03-11 cs.RO cs.LG

Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level Composition

制定你的策略!通过测试时的分布级组合改进扩散型或流型机器人策略

Jiahang Cao, Yize Huang, Hanzhong Guo, Rui Zhang, Mu Nan, Weijian Mai, Jiaxu Wang, Hao Cheng, Jingkai Sun, Gang Han, Wen Zhao, Qiang Zhang, Yijie Guo, Qihao Zheng, Chunfeng Song, Xiao Li, Ping Luo, Andrew F. Luo

机构 * The University of Hong Kong(香港大学) Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心) Shanghai AI Lab(上海人工智能实验室) Shanghai Jiaotong University(上海交通大学) The Hong Kong University of Science and Technology(香港科技大学)

AI总结 本研究提出无需训练的通用策略组合方法,通过测试时分布级组合提升机器人策略性能。

Comments Accepted to ICLR 2026. Project Page: https://sagecao1125.github.io/GPC-Site/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18722 2026-03-11 cs.AI

VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraft

VistaWise: 构建低成本代理的跨模态知识图谱用于Minecraft

Honghao Fu, Junlong Ren, Qi Chai, Deheng Ye, Yujun Cai, Hao Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) University of Queensland(昆士兰大学) Tencent(腾讯)

AI总结 VistaWise通过整合跨模态知识图谱和专用模型,实现低成本、高效率的Minecraft代理构建。

Comments Accepted by EMNLP 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.01611 2026-03-11 cs.CV cs.AI cs.LG

DRUPI: Dataset Reduction Using Privileged Information

DRUPI: 利用特权信息进行数据集缩减

Shaobo Wang, Youxin Jiang, Tianle Niu, Yantai Yang, Ruiji Zhang, Shuhao Hu, Shuaiyu Zhang, Chenghao Sun, Weiya Li, Conghui He, Xuming Hu, Linfeng Zhang

机构 * EPIC Lab, SJTU(EPIC实验室,上海交通大学) ICBC Shanghai AI Lab(上海人工智能实验室) HKUST(GZ)(香港科技大学(广州))

AI总结 本文提出DCPI方法,通过合成特权信息提升数据集凝练效果,在多个数据集上实现性能提升。

Comments 21 pages, 5 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09083 2026-03-11 cs.RO

Provably Safe Trajectory Generation for Manipulators Under Motion and Environmental Uncertainties

在运动和环境不确定性下为机械臂提供可证明安全的轨迹生成

Fei Meng, Zijiang Yang, Xinyu Mao, Haobo Liang, Max Q. -H. Meng

机构 * Hong Kong Center for Construction Robotics, The Hong Kong University of Science and Technology(香港建设机器人中心,香港理工大学) Department of Mechanical and Automation Engineering, The Chinese University of Hong Kong(机械与自动化工程系,中国香港大学) Department of Electronic Engineering, The Chinese University of Hong Kong(电子工程系,中国香港大学) Shenzhen Key Laboratory of Robotics Perception and Intelligence and the Department of Electronic and Electrical Engineering at Southern University of Science and Technology(深圳机器人感知与智能重点实验室和南方科技大学电子与电气工程系) Department of Electronic Engineering at The Chinese University of Hong Kong(电子工程系,中国香港大学)

AI总结 本文提出了一种可证明安全的轨迹生成框架,通过整合深度随机Koopman操作符模型和分层验证方法,有效应对运动和环境不确定性下的碰撞风险问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08812 2026-03-11 cs.CV

VisionCreator-R1: A Reflection-Enhanced Native Visual-Generation Agentic Model

VisionCreator-R1: 一种增强反思的原生视觉生成代理模型

Jinxiang Lai, Wenzhe Zhao, Zexin Lu, Hualei Zhang, Qinyu Yang, Rongwei Quan, Zhimin Li, Shuai Shao, Song Guo, Qinglin Lu

机构 * Tencent Hunyuan(腾讯文言) Hong Kong University of Science and Technology(香港科学与技术大学)

AI总结 VisionCreator-R1通过引入反射-计划联合优化方法,提升视觉生成代理在单图和多图任务中的表现,优于现有模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08589 2026-03-10 cs.CV

CARE-Edit: Condition-Aware Routing of Experts for Contextual Image Editing

CARE-Edit:基于条件的专家路由用于上下文图像编辑

Yucheng Wang, Zedong Wang, Yuetong Wu, Yue Ma, Dan Xu

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

AI总结 CARE-Edit通过条件感知的专家路由提升上下文图像编辑的性能和稳定性

Comments Accepted by CVPR 2026. Project page: https://care-edit.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏