Towards Robust Text-to-Image Person Retrieval: Multi-View Reformulation for Semantic Compensation
面向鲁棒的文本到图像人物检索:多视图重新公式化用于语义补偿
机构 * Beihang University(北航)
AI总结 本文提出基于大语言模型的语义补偿框架,通过多视图语义重新公式化和特征补偿提升跨模态表示一致性,解决文本到图像检索中的表达漂移问题,实验表明其在三个数据集上表现优异。
高校专区
面向鲁棒的文本到图像人物检索:多视图重新公式化用于语义补偿
机构 * Beihang University(北航)
AI总结 本文提出基于大语言模型的语义补偿框架,通过多视图语义重新公式化和特征补偿提升跨模态表示一致性,解决文本到图像检索中的表达漂移问题,实验表明其在三个数据集上表现优异。
ComPASS:通过工具增强的陪伴实现个性化代理社交支持
机构 * Renmin University of China(中国人民大学) ; Beihang University(北京航空航天大学)
AI总结 本文提出ComPASS框架,通过外部工具增强代理能力,构建首个个性化社交支持基准,训练出ComPASS-Qwen模型,提升响应质量。
基于几何的3D视觉标记修剪用于视频-语言模型
机构 * Beihang University(北京航空航天大学)
AI总结 本文提出Geo3DPruner框架,通过几何感知全局注意力建模跨帧相关性,并采用两阶段修剪过程,保留90%的视觉标记同时保持90%原始性能,优于现有文本引导和视觉引导修剪方法。
Comments Accepted by CVPR 2026
使用平行反平行四边形腱驱动手腕实现手帕纺任务的周期稳态控制
机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) ; School of Biomedical Engineering, Tsinghua University(清华大学生物医学工程系) ; School of Artificial Intelligence, Beihang University(北航人工智能学院) ; Institute of Nuclear and New Energy Technology, Tsinghua University(清华大学核能与新能量技术研究院) ; Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; School of Automation, Nanjing University of Science and Technology(南京理工大学自动化学院) ; Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学上海智能自主系统研究院) ; Department of Mechanical Engineering, Tsinghua University(清华大学机械工程系) ; School of Computation, Information and Technology, Technical University of Munich(慕尼黑工业大学计算、信息与技术学院) ; State Key Laboratory for Novel Software Technology and the School of Science and Technology, Nanjing University (Suzhou Campus)(南京大学软件新技术国家重点实验室(苏州校区)) ; Hepato-pancreato-biliary Center, Beijing Tsinghua Changgung Hospital(北京清华长庚医院肝胆外科中心) ; Key Laboratory of Digital Intelligence Hepatology (Ministry of Education), Beijing, China(教育部数字智能肝病重点实验室(北京)) ; School of Clinical Medicine, Tsinghua Medicine, Tsinghua University(清华大学医学部临床医学学院)
AI总结 本文提出了一种平行反平行四边形腱驱动手腕,结合分层控制方案和粒子弹簧模型,实现了高动态手帕纺任务的高展开比和高精度跟踪。
Comments ICRA2026
ImpRIF:更强的隐式推理导致更好的复杂指令遵循
机构 * ByteDance China(字节跳动中国) ; School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院) ; State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室)
AI总结 本文提出ImpRIF方法,通过构建可验证推理图提升大语言模型对隐式推理指令的理解,从而增强复杂指令遵循能力,在五个基准测试中表现优异。
Comments Accepted at ACL 2026 Main Conference
PiERN:基于令牌级别的路由以整合高精度计算与推理
机构 * Peking University(北京大学) ; Peking University Changsha Institute for Computing and Digital Economy(北京大学长沙计算与数字经济研究院) ; Beihang University(北京航空航天大学) ; Tsinghua University(清华大学)
AI总结 PiERN通过内生整合计算能力,提升语言模型在复杂系统中的精度与效率,实现高准确性、低延迟和可扩展性。
SVGDreamer:基于扩散模型的文本引导SVG生成
机构 * Beihang University(北航) ; The University of Hong Kong(香港大学)
AI总结 SVGDreamer提出一种基于扩散模型的文本引导SVG生成方法,通过引入语义驱动的图像向量化过程和向量粒子基于的分数蒸馏技术,提升生成的可编辑性、视觉质量和多样性。
Comments Accepted by CVPR 2024. Project Page: https://ximinng.github.io/SVGDreamer-project/
通过时间推理聚合实现高效的测试时扩展
机构 * School of Computer Science and Engineering, Beihang University(北航计算机科学与工程学院) ; Qingdao Research Institute, Beihang University(北航青岛研究院) ; Hangzhou Innovation Institute, Beihang University(北航杭州创新研究院)
AI总结 本文提出TRACE框架,通过时间聚合多步骤证据实现高效测试时扩展,减少推理token使用25-30%,保持准确性,优于现有动态推理方法。
Comments Accepted to Findings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026)
LookasideVLN: 基于方向的空中视觉与语言导航
机构 * Sun Yat-sen University(中山大学) ; Peng Cheng Laboratory(鹏城实验室) ; The Chinese University of Hong Kong(香港中文大学) ; Centre for Perceptual and Interactive Intelligence(感知与交互智能中心) ; Cardiff University(卡迪夫大学) ; Beihang University(北航) ; Guangdong Key Laboratory of Big Data Analysis and Processing(广东大数据分析与处理重点实验室)
AI总结 本文提出LookasideVLN,通过利用自然语言中的方向线索提升空中视觉与语言导航的准确性和效率,采用方向感知图、空间地标知识库和方向MLLM导航代理三种核心组件。
Comments Accepted by CVPR 2026
通过道德攻击 jailbreak 大型语言模型
机构 * South China University of Technology(南方科技大学) ; HKUST(香港科技大学) ; Beihang University(北航大学)
AI总结 本文通过道德攻击研究大型语言模型的内部道德价值观,构建了包含10300个实例的道德数据集,提出了四种对抗攻击方法,揭示了LLM和防护模型在道德感知攻击中的关键漏洞。
Comments 27 pages, 6 figures, 18 tables. Accepted by ACL 2026 Findings
FedOBP: 通过云-边元素解耦实现联邦最优脑个性化
机构 * Beijing Key Laboratory of MEMS Technology and Device Reliability for Industrial Internet, the School of Intelligent Engineering and Automation, Beijing University of Posts and Telecommunications(北京工业互联网MEMS技术与器件可靠性重点实验室,智能工程与自动化学院,北京邮电大学) ; School of Business Administration, Southwestern University of Finance and Economics(西南财经大学商学院) ; Institute of Artificial Intelligence and the School of Computer Science, Beihang University(北京航空航天大学人工智能研究院和计算机学院) ; School of Electronic Engineering, Dublin City University(都柏林城市大学电子工程学院) ; DreamSoul
AI总结 本文提出FedOBP算法,通过量化阈值机制和元素重要性评分,结合最优脑损伤理论,解决联邦学习中个性化参数选择问题,实验表明其在多种数据集和异构场景中表现优异。
HalluSAE:通过稀疏自编码器检测大型语言模型中的幻觉
机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, School of Artificial Intelligence, Beihang University(北京未来区块链与隐私计算先进研究院,北京航空航天大学人工智能学院) ; State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室) ; Beijing Academy of Blockchain and Edge Computing(北京区块链与边缘计算研究院) ; Renmin University of China(中国人民大学)
AI总结 HalluSAE通过稀疏自编码器和几何势能度量定位潜在相变区域,利用对比logit归因识别幻觉相关稀疏特征,通过线性探针检测因果幻觉,实验表明其在Gemma-2-9B上达到最先进的幻觉检测性能。