arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Harbin Institute of Technology(哈尔滨工业大学)

共收录 1439
2603.12912 2026-03-16 cs.CV cs.AI

FedBPrompt: Federated Domain Generalization Person Re-Identification via Body Distribution Aware Visual Prompts

FedBPrompt: 通过身体分布感知的视觉提示实现联邦领域泛化的人重识别

Xin Xu, Weilong Li, Wei Liu, Wenke Huang, Zhixi Yu, Bin Yang, Xiaoying Liao, Kui Jiang

机构 * School of Computer Science and Technology, Wuhan University of Science and Technology, China(武汉理工大学计算机科学与技术学院,中国) NERC for Multimedia Software, School of Computer Science, Wuhan University, China(武汉大学计算机学院,中国) Nanyang Technological University, Singapore(新加坡南洋理工大学) Changsha Bus Group, China(中国长沙公交集团) Central South University of Forestry and Technology, China(中国林业科技大学) Harbin Institute of Technology Zhengzhou Research Institute, China(哈尔滨工业大学郑州研究院,中国)

AI总结 本文提出FedBPrompt,通过引入可学习的视觉提示引导Transformer关注行人区域,解决联邦领域泛化重识别中域偏移问题,同时设计轻量级提示微调策略降低通信开销。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.12208 2026-03-13 cs.CV

ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models

ForensicZip: 更多令牌更好但并非必要在取证视觉-语言模型中

Yingxin Lai, Zitong Yu, Jun Wang, Linlin Shen, Yong Xu, Xiaochun Cao

机构 * Great Bay University(大湾大学) Shenzhen University(深圳大学) Harbin Institute of Technology(哈尔滨工业大学) School of Cyber Science and Technology, Sun Yat-sen University(中山大学网络安全科学与技术学院)

AI总结 ForensicZip通过引入无需训练的框架,从伪造驱动的角度重新定义令牌压缩,利用最优传输问题量化物理不连续性,实现高压缩率下的取证证据分离与高效检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.11620 2026-03-13 cs.LG

Personalized Federated Learning via Gaussian Generative Modeling

通过高斯生成建模实现个性化联邦学习

Peng Hu, Jianwei Ma

机构 * School of Mathematics, Harbin Institute of Technology(哈尔滨工业大学数学学院) Institute for Artificial Intelligence and the School of Earth and Space Sciences, Peking University(北京大学人工智能研究所和地球和空间科学学院) Institute for Artificial Intelligence and the School of Mathematics, Harbin Institute of Technology(哈尔滨工业大学人工智能研究所和数学学院)

AI总结 pFedGM通过高斯生成建模实现个性化联邦学习,结合全局优化导航器和本地分布统计提取器,利用双尺度融合框架提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02816 2026-03-12 cs.CV cs.AI

BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation

BrandFusion: 一种多智能体框架,用于文本到视频生成中的无缝品牌整合

Zihao Zhu, Ruotong Wang, Siwei Lyu, Min Zhang, Baoyuan Wu

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) Shenzhen Loop Area Institute(深圳河套学院) State University of New York at Buffalo(纽约州立大学布法罗分校) Harbin Institute of Technology(哈尔滨工业大学)

AI总结 BrandFusion通过多智能体框架实现文本到视频生成中的无缝品牌整合,提升品牌辨识度和内容自然度,推动T2V的可持续商业化应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13879 2026-03-12 cs.MM cs.CL cs.CV

Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring

链式推理压缩不应盲目:通过双路径锚定实现高效的多模态推理的V-Skip

Dongxu Zhang, Yiding Sun, Cheng Tan, Wenbiao Yan, Ning Yang, Jihua Zhu, Haijun Zhang

机构 * School of Software Engineering, Xi’an Jiaotong University(西安交通大学软件工程学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Harbin Institute of Technology, Shenzhen(深圳哈尔滨工业大学) University of Science and Technology Beijing(北京科技大学)

AI总结 V-Skip通过双路径锚定机制解决多模态推理中令牌剪枝的盲目性问题,实现高效的推理速度提升与精度保持。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.09691 2026-03-11 cs.CL cs.AI

ESAinsTOD: A Unified End-to-End Schema-Aware Instruction-Tuning Framework for Task-Oriented Dialog Modeling

ESAinsTOD: 一种统一的端到端模式感知指令微调框架用于任务导向对话建模

Dechuan Teng, Chunlin Lu, Libo Qin, Wanxiang Che

机构 * Research Center for Social Computing and Information Retrieval, Harbin Institute of Technology(社会科学与信息检索研究中心,哈尔滨工业大学) School of Computer Science and Engineering, Central South University(计算机科学与工程学院,中南大学)

AI总结 ESAinsTOD提出一种统一的端到端模式感知指令微调框架,提升任务导向对话建模的性能与泛化能力。

Comments Published at International Journal of Machine Learning and Cybernetics (IJMLC)

Journal ref Int. J. Mach. Learn. & Cyber. 17, 127 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12130 2026-03-11 cs.CL

PRISM of Opinions: A Persona-Reasoned Multimodal Framework for User-centric Conversational Stance Detection

观点PRISM:一种基于人设的多模态框架用于以用户为中心的对话立场检测

Bingbing Wang, Zhixin Bai, Zhengda Jin, Zihan Wang, Xintong Song, Jingjie Lin, Sixuan Li, Jing Li, Ruifeng Xu

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) Peng Cheng Laboratory, Shenzhen, China(鹏城实验室) Guangdong Provincial Key Laboratory of Novel Security Intelligence Technologies(广东省新型安全智能技术重点实验室) The Hong Kong Polytechnic University, Hong Kong, China(香港理工大学) Macau University of Science and Technology, Hong Kong, China(澳门科学技术大学)

AI总结 PRISM通过多模态和用户人设分析,提升对话立场检测的准确性与实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08387 2026-03-11 cs.CV

Recognition-Synergistic Scene Text Editing

识别协同场景文本编辑

Zhengyao Fang, Pengyuan Lyu, Jingjing Wu, Chengquan Zhang, Jun Yu, Guangming Lu, Wenjie Pei

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Peng Cheng Laboratory(鹏城实验室) Tencent(腾讯) Department of Computer Vision Technology, Baidu Inc.(百度公司计算机视觉技术部门)

AI总结 RS-STE通过整合文本识别与编辑的协同效应,提出了一种统一框架,实现高效且一致的场景文本编辑,提升了复杂场景下的性能。

Comments accepted by CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08453 2026-03-10 cs.LG cs.AI cs.CL

LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing

LycheeCluster: 基于结构感知分块和分层KV索引的高效长上下文推理

Dongfang Li, Zixuan Liu, Gang Lin, Baotian Hu, Min Zhang

机构 * Harbin Institute of Technology(哈尔滨工业大学) University of Electronic Science and Technology of China(电子科技大学)

AI总结 LycheeCluster通过结构感知分块和分层KV索引,实现高效长上下文推理,提升推理速度3.6倍且性能损失极小。

Comments 17 pages, 12 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08361 2026-03-10 cs.CV

$Δ$VLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation

$Δ$VLA:通过世界知识变化引导的视觉-语言-动作模型

Yijie Zhu, Jie He, Rui Shao, Kaishen Yuan, Tao Tan, Xiaochen Yuan, Zitong Yu

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Great Bay University(大湾大学) The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Macao Polytechnic University(澳门理工学院)

AI总结 $Δ$VLA通过建模世界知识变化,结合先验引导和潜在空间学习,提升机器人动作生成的效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15920 2026-03-10 cs.LG

ScaleGNN: Towards Scalable Graph Neural Networks via Adaptive High-order Neighboring Feature Fusion

ScaleGNN: 通过自适应高阶邻域特征融合实现可扩展的图神经网络

Xiang Li, Jianpeng Qi, Haobing Liu, Yuan Cao, Guoqing Chao, Zhongying Zhao, Junyu Dong, Xinwang Liu, Yanwei Yu

机构 * Ocean University of China(海洋大学) Harbin Institute of Technology(哈尔滨工业大学) Shandong University of Science and Technology(山东科技大学) School of Computer, National University of Defense Technology(国防科技大学计算机学院)

AI总结 ScaleGNN通过自适应融合多跳节点特征,解决大规模图神经网络的可扩展性和过平滑问题,提升预测准确性和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08024 2026-03-10 cs.CL

ConflictBench: Evaluating Human-AI Conflict via Interactive and Visually Grounded Environments

ConflictBench: 通过交互式和视觉 grounded 环境评估人类-人工智能冲突

Weixiang Zhao, Haozhen Li, Yanyan Zhao, xuda zhi, Yongbo Huang, Hao He, Bing Qin, Ting Liu

机构 * Harbin Institute of Technology(哈尔滨工业大学) SERES

AI总结 ConflictBench通过交互式和视觉 grounded 环境评估人类-人工智能冲突,揭示智能体在不同风险情境下的行为差异及对齐失败问题。

Comments 29 pages, 20 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07966 2026-03-10 cs.CV

Listening with the Eyes: Benchmarking Egocentric Co-Speech Grounding across Space and Time

用眼睛倾听:跨时空的自体视觉共指基准测试

Weijie Zhou, Xuantang Xiong, Zhenlin Hu, Xiaomeng Zhu, Chaoyang Zhao, Honghui Dong, Zhengyou Zhang, Ming Tang, Jinqiao Wang

机构 * Beijing Jiaotong University(北京交通大学) Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences (CASIA)(基础模型研究中心、自动化研究所、中国科学院(CASIA)) Tencent Robotics X(腾讯机器人X) Department of Computer Science and Engineering, The Hong Kong University of Science and Technology (HKUST)(计算机科学与工程系、香港科学与技术大学(HKUST)) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳学院)

AI总结 本文提出EcoG-Bench,通过严格测试语音-手势绑定,揭示多模态接口可能限制时间对齐线索的可观察性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07911 2026-03-10 cs.CV

Beyond Heuristic Prompting: A Concept-Guided Bayesian Framework for Zero-Shot Image Recognition

超越启发式提示:一种基于概念的贝叶斯框架用于零样本图像识别

Hui Liu, Kecheng Chen, Jialiang Wang, Xianming Liu, Wenya Wang, Haoliang Li

机构 * City University of Hong Kong(香港城市大学) Harbin Institute of Technology(哈尔滨工业大学) Nanyang Technological University(南洋理工大学)

AI总结 本文提出一种基于概念的贝叶斯框架,通过整合类特定概念提升零样本图像识别性能,采用多阶段概念合成和自适应软修剪似然,实现更鲁棒和高效的分类效果。

Comments 19 pages, Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07699 2026-03-10 cs.RO

C$^2$-Explorer: Contiguity-Driven Task Allocation with Connectivity-Aware Task Representation for Decentralized Multi-UAV Exploration

C$^2$-Explorer: 基于连通性的任务分配与连接感知的任务表示用于分布式多无人机探索

Xinlu Yan, Mingjie Zhang, Yuhao Fang, Yanke Sun, Jun Ma, Youmin Gong, Boyu Zhou, Jie Mei

机构 * School of Intelligence Science and Engineering, Harbin Institute of Technology, Shenzhen, Guangdong, China(哈尔滨工业大学深圳研究院) Robotics and Autonomous Systems Thrust, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)机器人与自主系统方向) Department of Mechanical and Energy Engineering, Southern University of Science and Technology, Shenzhen, Guangdong, China(南方科技大学机械与能源工程系)

AI总结 C$^2$-Explorer通过连接感知的任务表示和基于连续性的分配策略,提升多无人机探索效率,减少探索时间和路径长度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01763 2026-03-10 cs.CV

HiconAgent: History Context-aware Policy Optimization for GUI Agents

HiconAgent: 历史上下文感知的GUI代理策略优化

Xurui Zhou, Gongwei Chen, Yuquan Xie, Zaijing Li, Kaiwen Zhou, Shuai Wang, Shuo Yang, Zhuotao Tian, Rui Shao

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Shenzhen Loop Area Institute(深圳河套学院) Huawei Noah’s Ark Lab(华为诺亚实验室)

AI总结 HiconAgent通过历史上下文感知策略优化,高效利用历史信息,在GUI导航任务中实现更高的准确率和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10934 2026-03-10 cs.DB cs.LG

Towards Practical Benchmarking of Data Cleaning Techniques: On Generating Authentic Errors via Large Language Models

面向数据清洗技术实用基准测试的探索:通过大语言模型生成真实错误

Xinyuan Liu, Jiahui Chen, Bocheng Hu, Yu Sun, Xinyang Chen, Shaoxu Song, Yongxin Tong

机构 * Nankai University(南开大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Tsinghua University(清华大学) Beihang University(北京航空航天大学)

AI总结 TableEG利用大语言模型生成真实错误,通过三元组表示和表格微调策略,提升数据清洗技术的基准测试效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09156 2026-03-10 cs.CV

LEL: Lipschitz Continuity Constrained Ensemble Learning for Efficient EEG-Based Intra-subject Emotion Recognition

LEL:基于Lipschitz连续性约束的集成学习用于高效EEG基的个体情感识别

Shengyu Gong, Yueyang Li, Zijian Kang, Bo Chai, Weiming Zeng, Hongjie Yan, Zhiguo Zhang, Wai Ting Siok, Nizhuan Wang

机构 * Lab of Digital Image and Intelligent Computation(数字图像与智能计算实验室) Shanghai Maritime University(上海 Maritime 大学) Department of Chinese and Bilingual Studies(中文与双语研究部) The Hong Kong Polytechnic University(香港理工大学) Department of Neurology(神经病学部) Affiliated Lianyungang Hospital of Xuzhou Medical University(徐州医学院附属连云港医院) The Institute of Computing and Intelligence(计算与智能研究所) Harbin Institute of Technology Shenzhen(哈尔滨工业大学深圳研究院)

AI总结 LEL通过引入Lipschitz连续性约束的集成学习方法,提升EEG基个体情感识别的稳定性和鲁棒性,实现高效准确的情绪识别。

Journal ref IEEE Sensors Journal, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07098 2026-03-10 cs.CV

NuNext: Reframing Nucleus Detection as Next-Point Detection

NuNext:将核检测重新表述为下一个点检测

Zhongyi Shui, Honglin Li, Xiaozhong Ji, Ye Zhang, Zijiang Yang, Chenglu Zhu, Yuxuan Sun, Kai Yao, Conghui He, Cheng Tan

机构 * Zhejiang University(浙江大学) Westlake University(西湖大学) Nanjing University(南京大学) Harbin Institute of Technology(哈尔滨工业大学) University of Science and Technology Beijing(北京科技大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)

AI总结 NuNext通过多模态大语言模型直接输出核质心,采用空间感知软监督和强化微调策略提升检测性能,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.07048 2026-03-10 cs.CV cs.AI

Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation

回望与前行:用于多图像幻觉缓解的跨图像注意力校准与关注偏好学习

Xiaochen Yang, Hao Fang, Jiawei Kong, Yaoxin Mao, Bin Chen, Shu-Tao Xia

机构 * Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) Harbin Institute of Technology(哈尔滨工业大学) Beijing Institute of Technology(北京理工大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校)

AI总结 本文提出CAPL框架,通过跨图像注意力校准和偏好学习缓解多图像幻觉问题,提升模型对跨图像关联的建模能力,并在多个任务中取得稳定性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06666 2026-03-10 cs.CV

SJD-PV: Speculative Jacobi Decoding with Phrase Verification for Autoregressive Image Generation

SJD-PV: 基于短语验证的推测雅可比解码用于自回归图像生成

Zhehao Yu, Baoquan Zhang, Bingqi Shan, Xinhao Liu, Dongliang Zhou, Guotao Liang, Guangming Ye, Yunming Ye

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳学院)

AI总结 SJD-PV通过短语级推测验证提升自回归图像生成的解码效率,减少函数评估次数并提升生成质量。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.02951 2026-03-10 cs.LG cs.CV

CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuning

CGL: 通过强化微调推进连续GUI学习

Zhenquan Yao, Zitong Huang, Yihan Zeng, Jianhua Han, Hang Xu, Chun-Mei Feng, Jianwei Ma, Wangmeng Zuo

机构 * Harbin Institute of Technology(哈尔滨工业大学) Huawei Noah’s Ark Lab(华为诺亚实验室) University College Dublin(都柏林大学) Peking University(北京大学)

AI总结 CGL通过强化微调与监督微调的协同优化,解决GUI连续学习中遗忘旧任务的问题,提出AndroidControl-CL基准验证方法有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.01111 2026-03-10 cs.CV

DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles

DeAR: 通过分解注意力头角色实现细粒度VLM适应

Yiming Ma, Hongkun Yang, Lionel Z. Wang, Bin Chen, Weizhi Xian, Jianzhi Teng

机构 * Chongqing Research Institute of Harbin Institute of Technology(哈尔滨工业大学重庆研究所) Ocean University of China(中国海洋大学) The Hong Kong Polytechnic University(香港理工大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

AI总结 DeAR通过分解注意力头角色实现细粒度VLM适应,有效平衡任务适应与泛化能力。

Comments Accepted by CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.16160 2026-03-10 cs.CV

Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning

Video2Layout: 重建用于空间推理的度量基础认知图

Yibin Huang, Wang Xu, Wanyue Zhang, Helu Zhi, Jingjing Huang, Yangbin Xu, Yangang Sun, Conghui Zhu, Tiejun Zhao

机构 * Faculty of Computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) Tsinghua University(清华大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Institute of Microelectronics of the Chinese Academy of Sciences(中国科学院微电子研究所)

AI总结 Video2Layout通过重建度量基础的空间布局,提升多模态大语言模型在空间推理中的性能,其方法结合监督微调与强化微调,实现更精确的空间认知图构建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23051 2026-03-10 cs.LG

SwiftTS: A Swift Selection Framework for Time Series Pre-trained Models via Multi-task Meta-Learning

SwiftTS: 一种基于多任务元学习的时序预训练模型快速选择框架

Tengxue Zhang, Biao Ouyang, Yang Shu, Xinyang Chen, Chenjuan Guo, Bin Yang

机构 * East China Normal University(华东师范大学) Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))

AI总结 SwiftTS通过多任务元学习方法,提出了一种高效的时序预训练模型快速选择框架,提升了模型在不同数据集和时间范围下的泛化能力和鲁棒性。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24850 2026-03-10 cs.CV

PHASE-Net: Physics-Grounded Harmonic Attention System for Efficient Remote Photoplethysmography Measurement

PHASE-Net:基于物理的谐波注意力系统用于高效的远程光体积脉搏波测记测量

Bo Zhao, Dan Guo, Junzhe Cao, Yong Xu, Bochao Zou, Tao Tan, Yue Sun, Zitong Yu

机构 * Great Bay University(大湾大学) Hefei University of Technology(合肥工业大学) Macao Polytechnic University(澳门 polytechnic 大学) University of Science and Technology Beijing(北京科技大学) Harbin Institute of Technology, Shenzhen(哈尔滨工业大学深圳分校) Dongguan Key Laboratory for Intelligence and Information Technology(东莞智能与信息科技重点实验室)

AI总结 PHASE-Net基于物理原理设计,通过谐波注意力系统实现高效的远程光体积脉搏波测记测量,提升鲁棒性和可解释性。

Comments Accepted by CVPR2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13816 2026-03-10 cs.RO

Agile in the Face of Delay: Asynchronous End-to-End Learning for Real-World Aerial Navigation

在延迟面前敏捷:面向现实世界空中导航的异步端到端学习

Yude Li, Zhexuan Zhou, Huizhe Li, Youmin Gong, Jie Mei

机构 * School of Intelligence and Engineering(智能与工程学院) Guangdong Provincial Key Laboratory of Intelligent Morphing Mechanisms and Adaptive Robotics(广东省智能变形机制与自适应机器人重点实验室) Harbin Institute of Technology(哈尔滨工业大学)

AI总结 本文提出异步强化学习框架,解决端到端导航中感知与控制频率冲突问题,实现高频率控制与异步感知融合,提升现实环境中航空器的鲁棒性和敏捷性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14824 2026-03-10 cs.CV

Prototype Perturbation for Relaxing Alignment Constraints in Backward-Compatible Learning

原型扰动用于缓解后向兼容学习中的对齐约束

Zikun Zhou, Yushuai Sun, Wenjie Pei, Xin Li, Yaowei Wang

机构 * Harbin Institute of Technology(哈尔滨工业大学) Pengcheng Laboratory(鹏城实验室)

AI总结 本文提出通过扰动旧特征原型来缓解后向兼容学习中的对齐约束,提升新模型的判别能力。

Comments Accept to IEEE TMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06213 2026-03-09 cs.CV cs.AI

Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Events

直击核心:一种无需训练的多模态摘要方法 via 事件链

Xiaoxing You, Qiang Huang, Lingyu Li, Xiaojun Chang, Jun Yu

机构 * School of Computer Science, Hangzhou Dianzi University(杭州电子科技大学计算机科学学院) School of Intelligence Science and Engineering, Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)智能科学与工程学院) School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学与技术学院)

AI总结 无需训练的多模态摘要方法CoE通过事件链和层次事件图实现跨模态整合与时间推理,提升摘要质量与领域适应性。

Comments Accepted to CVPR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.06090 2026-03-09 cs.CV cs.CL

DeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal Model

DeepSight: 通过深度驱动的多模态模型连接深度图与语言

Hao Yang, Hongbo Zhang, Yanyan Zhao, Bing Qin

机构 * Harbin Institute of Technology(哈尔滨工业大学)

AI总结 DeepSight通过深度驱动的多模态模型提升三维场景理解,利用深度图像特性增强空间推理能力,并通过新数据集和模型改进实现了深度感知和任务性能的显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏