arXivDaily arXiv每日学术速递 周一至周五更新

高校专区

Huazhong University of Science and Technology(华中科技大学)

共收录 702
2509.20278 2026-01-08 cs.CL

Quantifying LLM Biases Across Instruction Boundary in Mixed Question Forms

量化混合问题形式中跨指令边界的LLM偏见

Zipeng Ling, Shuliang Liu, Yuehao Tang, Chen Huang, Gaoyang Jiang, Shenghong Fu, Junqi Yang, Yao Wan, Jiawan Zhang, Kejia Huang, Xuming Hu

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) University of Pennsylvania(宾夕法尼亚大学) Huazhong University of Science and Technology(华中科技大学) Hong Kong Polytechnic University(香港理工大学) Nanjing University of Posts and Telecommunications(南京邮电大学)

AI总结 本文提出BiasDetector基准,用于评估LLM在混合问题形式数据集下对稀疏标签混合的识别能力,揭示用户指令对LLM偏见的影响。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.02020 2026-01-06 cs.CV

Adapting Depth Anything to Adverse Imaging Conditions with Events

在恶劣成像条件下适应Depth Anything以应对事件

Shihan Peng, Yuyang Xiong, Hanyu Zhou, Zhiwei Shi, Haoyue Liu, Gang Chen, Luxin Yan, Yi Chang

机构 * National Key Lab of Multispectral Information Intelligent Processing Technology, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(multispectral information intelligent processing technology 国家重点实验室,人工智能与自动化学院,华中科技大学) School of Computing, National University of Singapore(computing 学院,新加坡国立大学) School of Computer Science and Engineering, Sun Yat-Sen University(computer science and engineering 学院,中山大学)

AI总结 本文提出ADAE框架,通过熵感知空间融合和运动引导时间校正,提升Depth Anything在恶劣成像条件下的深度估计性能。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.01192 2026-01-06 cs.CV

Crowded Video Individual Counting Informed by Social Grouping and Spatial-Temporal Displacement Priors

受社交分组和时空位移先验信息启发的拥挤视频个体计数

Hao Lu, Xuhui Zhu, Wenjing Zhang, Yanan Li, Xiang Bai

机构 * State Key Laboratory of Multispectral Information Intelligent Processing Technology(多谱信息智能处理技术国家重点实验室;人工智能与自动化学院,华中科技大学) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(湖北省智能机器人重点实验室;计算机科学与工程人工智能学院,武汉理工大学) Hubei Key Laboratory of Intelligent Robot(软件工程学院,华中科技大学) School of Computer Science & Engineering Artificial Intelligence, Wuhan Institute of Technology School of Software Engineering, Huazhong University of Science and Technology

AI总结 本文提出OMAN++方法,通过引入社交分组和时空位移先验信息,提升拥挤场景下视频个体计数的准确性。

Comments Journal Extension of arXiv:2506.13067

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00588 2026-01-06 cs.CL

CSSBench: Evaluating the Safety of Lightweight LLMs against Chinese-Specific Adversarial Patterns

CSSBench: 评估轻量级大语言模型对中文特定对抗模式的安全性

Zhenhong Zhou, Shilinlu Yan, Chuanpu Liu, Qiankun Li, Kun Wang, Zhigang Zeng

机构 * Nanyang Technological University(南洋理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Huazhong University of Science and Technology(华中科技大学)

AI总结 CSSBench通过评估中文特定对抗模式,揭示轻量级大语言模型在中文环境下的安全挑战,为实际应用提供安全评估框架。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00322 2026-01-05 cs.CV

Depth-Synergized Mamba Meets Memory Experts for All-Day Image Reflection Separation

深度协同Mamba遇见记忆专家用于全天图像反射分离

Siyan Fang, Long Peng, Yuntao Wang, Ruonan Wei, Yuehuan Wang

机构 * Huazhong University of Science and Technology(华中科技大学) University of Science and Technology of China(中国科学技术大学)

AI总结 DMDNet通过深度感知扫描和记忆专家补偿模块,提升全天图像反射分离性能,尤其在夜间表现更优。

Comments This paper has been accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.04000 2026-01-01 cs.IT cs.LG math.IT

Distributed Information Bottleneck Theory for Multi-Modal Task-Aware Semantic Communication

多模态任务感知语义通信的分布式信息瓶颈理论

Yujie Zhou, Cheng Peng, Rulong Wang, Yong Xiao, Yingyu Li, Guangming Shi, Ping Zhang

机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(电子信息与通信学院,华中科技大学) Peng Cheng Laboratory(鹏城实验室) Pazhou Laboratory (Huangpu)(琶洲实验室(黄埔)) School of Mechanical Engineering and Electronic Information, China University of Geosciences(机械工程与电子信息学院,中国地质大学) State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学)

AI总结 本文提出了一种多模态任务感知语义通信的分布式信息瓶颈框架,通过量化模态贡献来优化资源利用,提升通信效率与任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24321 2026-01-01 cs.CV cs.RO

UniAct: Unified Motion Generation and Action Streaming for Humanoid Robots

UniAct:面向人形机器人的统一动作生成与动作流

Nan Jiang, Zimo He, Wanhe Yu, Lexi Pang, Yunhao Li, Hongjie Li, Jieming Cui, Yuhan Li, Yizhou Wang, Yixin Zhu, Siyuan Huang

机构 * Institute for AI, Peking University(人工智能研究院,北京大学) Beijing Institute for General Artificial Intelligence (BIGAI)(北京通用人工智能研究院) School of Psychological and Cognitive Sciences, Peking University(心理学与认知科学学院,北京大学) School of Computer Science, Peking University(计算机科学学院,北京大学) Yuanpei College, Peking University(元培学院,北京大学) School of Foreign Languages, Peking University(外语学院,北京大学) School of EECS, Peking University(电子工程与科学学院,北京大学) Huazhong University of Science and Technology(华中科技大学) State Key Lab of General AI(通用人工智能国家重点实验室) Nat’l Eng. Research Center of Visual Technology(视觉技术国家工程研究中心) Beijing Key Laboratory of Behavior and Mental Health, Peking University(北京行为与心理健康重点实验室,北京大学) Embodied Intelligence Lab, PKU-Wuhan Institute for Artificial Intelligence(具身智能实验室,北京大学-武汉人工智能研究院)

AI总结 UniAct通过统一感知与控制框架,实现人形机器人在多模态指令下的高效动作生成与实时执行,提升零样本跟踪成功率19%。

Comments Project page: https://jnnan.github.io/uniact/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24112 2026-01-01 cs.RO

RflyUT-Sim: A Simulation Platform for Development and Testing of Complex Low-Altitude Traffic Control

RflyUT-Sim:一种用于复杂低空交通控制开发和测试的仿真平台

Zonghan Li, Tianwen Tao, Rao Fu, Liang Wang, Dongyuan Zhang, Quan Quan

机构 * School of Automation Science and Electrical Engineering, Beihang University(自动化科学与电气工程学院,北京航空航天大学) Beijing Intelligent Token Technology Co., Ltd.(北京智能令牌科技有限公司) School of Traffic and Transportation, Beijing Jiaotong University(交通与运输学院,北京交通大学) School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学)

AI总结 RflyUT-Sim是一款集成高保真度低空UAV交通仿真平台,整合RflySim/AirSim和Unreal Engine 5,提供全状态模型和3D地图,支持灵活的场景定制和开源代码,用于低空交通控制的开发与测试。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24016 2026-01-01 cs.CV

FitControler: Toward Fit-Aware Virtual Try-On

FitControler: 向拟合感知的虚拟试衣

Lu Yang, Yicheng Liu, Yanan Li, Xiang Bai, Hao Lu

机构 * Huazhong University of Science and Technology(华中科技大学) Wuhan Institute of Technology(武汉理工大学)

AI总结 FitControler通过拟合感知的布局生成器和多尺度注入器,实现对虚拟试衣中服装合身度的定制化控制,并构建了包含13,000个配对的数据集。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21614 2026-01-01 cs.CL

A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond

大推理模型高效推理的综述:语言、多模态与更远的探索

Xiaoye Qu, Yafu Li, Zhao-Chen Su, Weigao Sun, Jianhao Yan, Dongrui Liu, Ganqu Cui, Daizong Liu, Shuxian Liang, Junxian He, Peng Li, Wei Wei, Jing Shao, Chaochao Lu, Yue Zhang, Xian-Sheng Hua, Bowen Zhou, Yu Cheng

机构 * Shanghai AI Laboratory(上海人工智能实验室) Soochow University(苏州大学) Westlake University(西湖大学) Peking University(北京大学) Tongji University(同济大学) The Hong Kong University of Science and Technology(香港科技大学) Tsinghua University(清华大学) Huazhong University of Science and Technology(华中科技大学) The Chinese University of Hong Kong(香港中文大学)

AI总结 本文综述了大推理模型在提升推理效率方面的最新研究,聚焦于语言、多模态及未来方向,旨在推动该领域的发展。

Comments Update recent RL papers. Project page: https://github.com/XiaoYee/Awesome_Efficient_LRM_Reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.05499 2025-12-30 cs.CR cs.AI cs.CL cs.SE

Prompt Injection attack against LLM-integrated Applications

针对集成大语言模型应用的提示注入攻击

Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Zihao Wang, Xiaofeng Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, Leo Yu Zhang, Yang Liu

机构 * Griffith University(格里菲斯大学) Nanyang Technological University(南洋理工大学) University of New South Wales(新南威尔士大学) Huazhong University of Science and Technology(华中科技大学) Indiana University at Bloomington(印第安纳大学布卢明顿分校) Southern University of Science and Technology(南方科技大学) Tianjin University(天津大学)

AI总结 本研究提出HouYi技术,揭示LLM集成应用中提示注入攻击的潜在风险及缓解方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22629 2025-12-30 cs.AI cs.IR

DICE: Discrete Interpretable Comparative Evaluation with Probabilistic Scoring for Retrieval-Augmented Generation

DICE:基于概率评分的离散可解释比较评估用于检索增强生成

Shiyan Liu, Jian Ma, Rui Qu

机构 * School of Computer Science and Technology(计算机科学与技术学院) Huazhong University of Science and Technology(华中科技大学)

AI总结 DICE提出一种基于概率评分的两阶段框架,提升RAG评估的可解释性和稳健性,通过透明判断和系统性错误诊断,实现高效且可信的评估。

Comments Accepted at ResponsibleFM @ NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03332 2025-12-30 cs.LG cs.AI

Exploring Layer-wise Information Effectiveness for Post-Training Quantization in Small Language Models

探索分层信息有效性以实现小语言模型的后训练量化

He Xiao, Qingyao Yang, Dirui Xie, Wendong Xu, Zunhai Su, Runming yang, Wenyong Zhou, Haobo Liu, Zhengwu Liu, Ngai Wong

机构 * The University of Hong Kong(香港大学) Huazhong University of Science and Technology(华中科技大学) Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生学院)

AI总结 LieQ通过分层信息有效性量化方法,在亚8B模型中实现高效低比特压缩,减少精度损失并提升边缘设备部署可行性。

Comments low-bit quantization

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21691 2025-12-29 cs.CV

Analyzing the Mechanism of Attention Collapse in VGGT from a Dynamics Perspective

从动力学角度分析VGGT中注意力崩溃的机制

Huan Li, Longjun Luo, Yuling Shi, Xiaodong Gu

机构 * Huazhong University of Science and Technology(华中科技大学) Guangdong University of Technology(广东工业大学) Shanghai Jiao Tong University(上海交通大学)

AI总结 VGGT中注意力崩溃现象从动力学角度进行数学解释,揭示token特征流收敛规律及token合并方法对延迟崩溃的作用。

Comments 8 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23574 2025-12-29 math.OC cs.LG

Online Convex Optimization with Memory and Limited Predictions

具有记忆和有限预测的在线凸优化

Zhengmiao Wang, Zhi-Wei Liu, Ming Chi, Xiaoling Wang, Housheng Su, Lintao Ye

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) College of Automation and College of Artificial Intelligence, Nanjing University of Posts and Telecommunications(自动化学院和人工智能学院,南京邮电大学)

AI总结 本文提出了一种具有记忆和有限预测的在线凸优化算法,通过截断高斯平滑技术实现指数衰减的动态遗憾,并在一般凸优化中达到线性收敛速度。

Comments 35 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17019 2025-12-25 cs.CV cs.AI cs.CY

Let Androids Dream of Electric Sheep: A Human-Inspired Image Implication Understanding and Reasoning Framework

让安卓梦见电羊:一种受人类启发的图像隐喻理解和推理框架

Chenhao Zhang, Yazhe Niu

机构 * Shanghai AI Laboratory(上海人工智能实验室) Huazhong University of Science and Technology(华中科技大学) The Chinese University of Hong Kong(香港中文大学)

AI总结 本研究提出LAD框架,通过三阶段方法解决图像隐喻理解问题,在多个基准测试中取得优异成绩,推动视觉语言推理和人机交互发展。

Comments 19 pages, 9 figures, 7 tables. Code & Dataset: https://github.com/MING-ZCH/Let-Androids-Dream-of-Electric-Sheep

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05671 2025-12-24 eess.AS cs.CL

Low-Resource Domain Adaptation for Speech LLMs via Text-Only Fine-Tuning

通过文本仅微调实现语音大模型的低资源领域适应

Yangui Fang, Jing Peng, Xu Li, Yu Xi, Chengwei Zhang, Guohui Zhong, Kai Yu

机构 * Huazhong University of Science and Technology, School of Electronic Information and Communications(华中科技大学电子信息与通信学院) MoE Key Lab of Artificial Intelligence, AI Institute, X-LANCE Lab, Shanghai Jiao Tong University, Shanghai, China(人工智能MoE重点实验室、AI研究院、X-LANCE实验室、上海交通大学、上海,中国) Jiangsu Key Lab of Language Computing, Suzhou, China(江苏省语言计算重点实验室、苏州,中国) AISpeech Co., Ltd., Suzhou, China(AISpeech公司、苏州,中国)

AI总结 本文提出通过文本仅微调实现语音大模型在低资源环境下的领域适应,有效提升识别性能并减少领域迁移中的性能下降。

Comments This paper has been ACCEPTED for publication in ASRU

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24347 2025-12-24 cs.CL eess.AS

Fewer Hallucinations, More Verification: A Three-Stage LLM-Based Framework for ASR Error Correction

更少的幻觉,更多的验证:一种基于LLM的三阶段框架用于语音识别错误纠正

Yangui Fang, Baixu Chen, Jing Peng, Xu Li, Yu Xi, Chengwei Zhang, Guohui Zhong

机构 * Huazhong University of Science and Technology, School of Electronic Information and Communications(华中科技大学电子信息与通信学院) MoE Key Lab of Artificial Intelligence, AI Institute, X-LANCE Lab, Shanghai Jiao Tong University, Shanghai, China(人工智能联合实验室、AI研究院、X-LANCE实验室、上海交通大学) AISpeech Ltd, Suzhou, China(上海苏州AISpeech有限公司)

AI总结 本文提出一种基于LLM的三阶段框架,用于减少语音识别错误纠正中的幻觉问题,通过预检测、链式思维修正和推理验证提高纠正准确性。

Comments This paper has been ACCEPTED for publication in ASRU

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19331 2025-12-23 cs.CV

DeltaMIL: Gated Memory Integration for Efficient and Discriminative Whole Slide Image Analysis

DeltaMIL: 基于门控记忆的高效且判别性全滑片图像分析

Yueting Zhu, Yuehao Song, Shuai Zhang, Wenyu Liu, Xinggang Wang

机构 * School of Electronic Information and Communications, Huazhong University of Science and Technology(电子信息学院,华中科技大学)

AI总结 DeltaMIL通过门控记忆机制提升全滑片图像分析的效率与判别性,实验显示其在生存预测和分类任务中均取得显著性能提升。

Comments 11 pages,7 figures,8 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17296 2025-12-23 cs.CV

Towards Pixel-Wise Anomaly Location for High-Resolution PCBA via Self-Supervised Image Reconstruction

面向高分辨率PCBA的像素级异常定位方法:基于自监督图像重建

Wuyi Liu, Le Jin, Junxian Yang, Yuanchao Yu, Zishuo Peng, Jinfeng Xu, Xianzhi Li, Jun Zhou

机构 * Huazhong University of Science and Technology(华中科技大学) Siemens AG(西门子股份公司) University of Electronic Science and Technology of China(电子科技大学) Peking University(北京大学)

AI总结 本文提出HiSIR-Net,通过自监督图像重建实现高分辨率PCBA的像素级异常定位,结合SIR-Gate和ROPS方案,提升定位精度与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.10609 2025-12-23 cs.CV

MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling

MSTAR: 无框多查询场景文本检索与注意力回收

Liang Yin, Xudong Xie, Zhang Li, Xiang Bai, Yuliang Liu

机构 * Huazhong University of Science and Technology(华中科技大学)

AI总结 MSTAR提出了一种无框多查询场景文本检索方法,通过渐进视觉嵌入和多实例匹配模块提升检索性能,首次构建MQTR数据集并优于现有模型。

Comments Neurips 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04162 2025-12-23 cs.IR cs.AI

Semantic Retrieval Augmented Contrastive Learning for Sequential Recommendation

语义检索增强对比学习用于序列推荐

Ziqiang Cui, Yunpeng Weng, Xing Tang, Xiaokun Zhang, Shiwei Li, Peiyang Liu, Bowei He, Dugang Liu, Weihong Luo, Xiuqiang He, Chen Ma

机构 * City University of Hong Kong(香港城市大学) Huazhong University of Science and Technology(华中科技大学) Tencent(腾讯) Shenzhen Technology University(深圳技术大学) Peking University(北京大学) Shenzhen University(深圳大学)

AI总结 SRA-CL通过利用LLMs的语义能力生成高质量对比对,提升序列推荐模型的性能。

Comments Accepted by NeurIPS 2025. Code is available at: https://github.com/ziqiangcui/SRA-CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12220 2025-12-23 cs.RO

Hybrid Dynamics Modeling and Trajectory Planning for a Cable-Trailer System with a Quadruped Robot

混合动力学建模与电缆-拖车系统四足机器人的轨迹规划

Wentao Zhang, Shaohang Xu, Gewei Zuo, Bolin Li, Jingbo Wang, Lijun Zhu

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(人工智能与自动化学院,华中科技大学) Department of Data Science, City University of Hong Kong(数据科学系,城市大学) Embodied AI Center of Shanghai AI Lab(上海人工智能实验室具身智能中心)

AI总结 本文提出了一种混合动力学模型和搜索算法,用于四足机器人与电缆-拖车系统的轨迹规划,通过几何多边形碰撞避免约束实现安全高效的运动控制。

Comments 8 pages, 8 figures, Accept by RA-L 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.10101 2025-12-23 cs.RO

Agile and Safe Trajectory Planning for Quadruped Navigation with Motion Anisotropy Awareness

具备运动各向异性意识的四足机器人导航轨迹规划

Wentao Zhang, Shaohang Xu, Peiyuan Cai, Lijun Zhu

机构 * School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, China(华中科技大学人工智能与自动化学院)

AI总结 本文提出了一种考虑四足机器人运动各向异性的导航框架,通过轨迹生成与优化提升导航效率和安全性。

Comments 8 pages, 6 figures, submitted to 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18245 2025-12-23 cs.CV cs.AI

Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image

光谱偏差与跨模态语义一致性学习用于超光谱图像的目标检测

Xiao He, Chang Tang, Xinwang Liu, Wei Zhang, Zhimin Gao, Chuankun Li, Shaohua Qiu, Jiangfeng Xu

机构 * School of Computer, Wuhan University(武汉大学计算机学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件学院) school of computer, National University of Defense Technology(国防科技大学计算机学院) Shandong Provincial Key Laboratory of Computer Networks, Shandong Computer Science Center (National Supercomputing Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences)(山东省计算机网络重点实验室、山东省计算机科学中心(国家超级计算中心济南中心)、齐鲁工业大学(山东省科学院)) School of Computer and Artificial Intelligence, Zhengzhou University(郑州大学计算机与人工智能学院) School of Information and Communication Engineering, North University of China(北方大学信息与通信工程学院) National Key Laboratory of Electromagnetic Energy, Naval University of Engineering(电磁能国家重点实验室、海军工程大学) Hexagon AB

AI总结 本文提出SDCM网络,通过光谱偏差与跨模态语义一致性学习,提升超光谱图像目标检测的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17545 2025-12-22 cs.CV cs.AI

ClothHMR: 3D Mesh Recovery of Humans in Diverse Clothing from Single Image

ClothHMR: 从单张图像中恢复穿着多样服装人类的3D网格

Yunqi Gao, Leyuan Liu, Yuhan Li, Changxin Gao, Yuanyuan Liu, Jingying Chen

机构 * Central China Normal University(中央财经大学) Huazhong University of Science and Technology(华中科技大学) China University of Geosciences (WuHan)(中国地质大学(武汉))

AI总结 ClothHMR通过定制服装和基础模型视觉信息提升,实现从单张图像中准确恢复多样化服装人类的3D网格。

Comments 15 pages,16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.15110 2025-12-22 cs.CV

Is Nano Banana Pro a Low-Level Vision All-Rounder? A Comprehensive Evaluation on 14 Tasks and 40 Datasets

Nano Banana Pro是否是低级视觉的全能选手?对14项任务和40个数据集的全面评估

Jialong Zuo, Haoyou Deng, Hanyu Zhou, Jiaxin Zhu, Yicheng Zhang, Yiwei Zhang, Yongxin Yan, Kaixing Huang, Weisen Chen, Yongtai Deng, Rui Jin, Nong Sang, Changxin Gao

机构 * National Key Laboratory of Multispectral Information Intelligent Processing Technology(多谱信息智能处理技术国家重点实验室) School of Artificial Intelligence and Automation(人工智能与自动化学院) Huazhong University of Science and Technology(华中科技大学)

AI总结 本文评估了Nano Banana Pro在14项低级视觉任务上的表现,发现其在主观视觉质量上优于专业模型,但在定量指标上表现欠佳,揭示了生成模型在保持像素一致性方面的挑战。

Comments Technical Report; 65 Pages, 36 Figures, 17 Tables; Poject Page: https://lowlevelbanana.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00720 2025-12-22 cs.LG cs.CY stat.ML

Fairness via Independence: A (Conditional) Distance Covariance Framework

通过独立性实现公平性:一种(条件)距离协方差框架

Ruifan Huang, Haixia Liu

机构 * School of Mathematics and Statistics, Huazhong University of Science and Technology, Wuhan, Hubei, China(华中科技大学数学与统计学院) Institute of Interdisciplinary Research for Mathematics and Applied Science & Hubei Key Laboratory of Engineering Modeling and Scientific Computing, Huazhong University of Science and Technology, Wuhan, Hubei, China(华中科技大学数学与应用科学交叉研究 institute 及湖北省工程建模与科学计算重点实验室)

AI总结 本文提出一种基于距离协方差的公平性提升方法,通过在模型训练中引入惩罚项并优化计算效率,有效缩小了机器学习中的公平性差距。

Comments 25 pages, 4 figures. The old title is "Bridging Fairness Gaps: A (Conditional) Distance Covariance Perspective in Fairness Learning"

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16584 2025-12-19 cs.CV

Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs

Sketch-in-Latents: 在潜在空间中实现多模态统一推理

Jintao Tong, Jiaqi Gu, Yujing Lou, Lubin Fan, Yixiong Zou, Yue Wu, Jieping Ye, Ruixuan Li

机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) Alibaba Cloud Computing(阿里巴巴云计算)

AI总结 SkiLa通过在潜在空间中实现多模态统一推理,扩展MLLMs的自回归能力,生成连续视觉嵌入,提升视觉任务性能和多模态泛化能力。

Comments 14 pages, 11 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16485 2025-12-19 cs.CV cs.AI

Smile on the Face, Sadness in the Eyes: Bridging the Emotion Gap with a Multimodal Dataset of Eye and Facial Behaviors

脸上微笑,眼中悲伤:通过眼和面部行为的多模态数据集弥合情感差距

Kejun Liu, Yuanyuan Liu, Lin Wei, Chang Tang, Yibing Zhan, Zijing Chen, Zhe Chen

机构 * School of Computer Science, China University of Geosciences (Wuhan)(中国地质大学(武汉)计算机科学学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院) School of Computer Science, Wuhan University(武汉大学计算机科学学院) School of Computing, Engineering and Mathematical Sciences, La Trobe University(拉筹伯大学计算科学、工程与数学科学学院)

AI总结 本文通过构建包含眼行为和面部行为的多模态数据集EMER,提出EMERT模型以提升情感识别的鲁棒性。

Comments Accepted by TMM

详情

展开后加载摘要…

URL PDF HTML 收藏