arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

共收录 1944
2409.14072 2025-08-06 cs.CV

Dynamic 2D Gaussians: Geometrically Accurate Radiance Fields for Dynamic Objects

Shuai Zhang, Guanjun Wu, Zhoufeng Xie, Xinggang Wang, Bin Feng, Wenyu Liu

机构 * School of EIC Huazhong University of Science(电子信息学院华中科技大学) School of CS Huazhong University of Science(计算机学院华中科技大学) Huazhong University of Science(华中科技大学)

Comments Accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03437 2025-08-06 cs.CV cs.AI

Spatial Imputation Drives Cross-Domain Alignment for EEG Classification

Hongjun Liu, Chao Yao, Yalan Zhang, Xiaokun wang, Xiaojuan Ban

机构 * School of Intelligence Science and Technology(智能科学与技术学院) University of Science and Technology Beijing(科学技术大学) School of Computer and Communication Engineering(计算机与通信工程学院)

Comments ACMMM 2025 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03277 2025-08-06 cs.CV

Efficient Multi-Slide Visual-Language Feature Fusion for Placental Disease Classification

Hang Guo, Qing Zhang, Zixuan Gao, Siyuan Yang, Shulin Peng, Xiang Tao, Ting Yu, Yan Wang, Qingli Li

Comments Accepted by ACMMM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03091 2025-08-06 cs.AI cs.CR cs.CV

T2UE: Generating Unlearnable Examples from Text Descriptions

Xingjun Ma, Hanxun Huang, Tianwei Song, Ye Sun, Yifeng Gao, Yu-Gang Jiang

机构 * Fudan University(复旦大学) The University of Melbourne(墨尔本大学)

Comments To appear in ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03038 2025-08-06 cs.AI

Tree-of-Reasoning: Towards Complex Medical Diagnosis via Multi-Agent Reasoning with Evidence Tree

Qi Peng, Jialin Cui, Jiayuan Xie, Yi Cai, Qing Li

机构 * South China University of Technology(华南理工大学) Key Laboratory of Big Data and Intelligent Robot (SCUT), Ministry of Education(大数据与智能机器人重点实验室(SCUT)) Hong Kong Polytechnic University(香港理工大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22676 2025-08-06 cs.CL cs.MM

Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment

Jia Li, Yang Wang, Wenhao Qian, Jialong Hu, Zhenzhen Hu, Richang Hong, Meng Wang

机构 * Hefei University of Technology(合肥工业大学)

Comments 8 pages, 4 figures, ACM MM 2025. github:https://github.com/MSA-LMC/365Aspects

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20756 2025-08-06 cs.CL cs.AI cs.CV cs.LG cs.MM

ADS-Edit: A Multimodal Knowledge Editing Dataset for Autonomous Driving Systems

Chenxi Wang, Jizhan Fang, Xiang Chen, Bozhong Tian, Ziwen Xu, Huajun Chen, Ningyu Zhang

机构 * Zhejiang University,\ University - Ant \ Joint Laboratory of Knowledge Graph Hangzhou China Zhejiang University,\ University - Ant \ Joint Laboratory of Knowledge Graph

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18375 2025-08-06 cs.CV eess.IV

Individual Content and Motion Dynamics Preserved Pruning for Video Diffusion Models

Yiming Wu, Zhenghao Chen, Huan Wang, Dong Xu

机构 * The University of Hong Kong(香港大学) University of Newcastle(新castle大学) Westlake University(西湖大学)

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17471 2025-08-06 cs.LG cs.CR cs.CV

Learning New Concepts, Remembering the Old: Continual Learning for Multimodal Concept Bottleneck Models

Songning Lai, Mingqian Liao, Zhangyi Hu, Jiayu Yang, Wenshuo Chen, Hongru Xiao, Jianheng Tang, Haicheng Liao, Yutao Yue

机构 * HKUST(GZ) Deep Interdisciplinary Intelligence Lab(香港科技大学(广州)深度跨学科智能实验室) Wuhan University(武汉大学) Shandong University(山东大学) Tongji University(同济大学) Peking University(北京大学) University of Macau(澳门大学) HKUST(GZ) Institute of Deep Perception Technology, JITRI Deep Interdisciplinary Intelligence Lab(香港科技大学(广州)感知技术研究所,JITRI深度跨学科智能实验室)

Journal ref ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02538 2025-08-05 cs.IR

Hubness Reduction with Dual Bank Sinkhorn Normalization for Cross-Modal Retrieval

Zhengxin Pan, Haishuai Wang, Fangyu Wu, Peng Zhang, Jiajun Bu

Comments ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02374 2025-08-05 cs.CV cs.IR cs.LG

Uni-Layout: Integrating Human Feedback in Unified Layout Generation and Evaluation

Shuo Lu, Yanyin Chen, Wei Feng, Jiahao Fan, Fengheng Li, Zheng Zhang, Jingjing Lv, Junjie Shen, Ching Law, Jian Liang

机构 * School of AI, UCAS(人工智能学院,中国科学院自动化研究所)

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02243 2025-08-05 cs.CV cs.IR

I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking

Ziyan Liu, Junwen Li, Kaiwen Li, Tong Ruan, Chao Wang, Xinyan He, Zongyu Wang, Xuezhi Cao, Jingping Liu

机构 * East China University of Science and Technology(东华大学) South China University of Technology(华南理工大学) Shanghai University(上海大学)

Comments 10 pages, 6 figures, accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01644 2025-08-05 cs.MM cs.AI cs.CV cs.SD eess.AS

DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition

Peiyuan Jiang, Yao Liu, Qiao Liu, Zongshun Zhang, Jiaye Yang, Lu Liu, Daibing Yao

机构 * School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院)

Comments Published in ACM Multimedia 2025. 10 pages, 4 figures

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01064 2025-08-05 eess.IV cs.CV

Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentation

Fenghe Tang, Bingkun Nian, Jianrui Ding, Wenxin Ma, Quan Quan, Chengqi Dong, Jie Yang, Wei Liu, S. Kevin Zhou

机构 * University of Science and Technology of China(中国科学技术大学) Suzhou Institute for Advanced Research, USTC(苏州先进研究院) Shanghai Jiao Tong University(上海交通大学) Harbin Institute of Technology(哈尔滨工业大学) State Grid Hunan ElectricPower Corporation Limited Research Institute(国网湖南电力有限公司研究院)

Comments Accepted by ACM Multimedia 2025. Code: https://github.com/FengheTan9/Mobile-U-ViT

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17349 2025-08-05 cs.CV cs.IR

DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition

Yiyan Xu, Wuqiang Zheng, Wenjie Wang, Fengbin Zhu, Xinting Hu, Yang Zhang, Fuli Feng, Tat-Seng Chua

机构 * University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

Comments Accepted for publication in ACM MM'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02113 2025-08-05 cs.CV eess.IV

DeflareMamba: Hierarchical Vision Mamba for Contextually Consistent Lens Flare Removal

Yihang Huang, Yuanfei Huang, Junhui Lin, Hua Huang

机构 * School of Artificial Intelligence, Beijing Normal University(北京师范大学人工智能学院) Engineering Research Center of Intelligent Technology and Educational Application, Ministry of Education(教育部智能技术与教育应用工程研究中心)

Comments Accepted by ACMMM 2025

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27--31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02050 2025-08-05 cs.IR

Why Generate When You Can Transform? Unleashing Generative Attention for Dynamic Recommendation

Yuli Liu, Wenjun Kong, Cheng Luo, Weizhi Ma

Comments Accepted at ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01723 2025-08-05 cs.RO

OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping

Danyang Li, Zenghui Yang, Guangpeng Qi, Songtao Pang, Guangyong Shang, Qiang Ma, Zheng Yang

机构 * School of Software, Tsinghua University(清华大学软件学院) School of computer science and engineering, Central South University(中南大学计算机科学与工程学院) Inspur Yunzhou Industrial Internet Co., Ltd(Inspur Yunzhou工业互联网有限公司) Tsinghua University(清华大学)

Comments ACM MM '25

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01650 2025-08-05 cs.CV

StrandDesigner: Towards Practical Strand Generation with Sketch Guidance

Na Zhang, Moran Li, Chengming Xu, Han Feng, Xiaobin Hu, Jiangning Zhang, Weijian Cao, Chengjie Wang, Yanwei Fu

机构 * Fudan University(复旦大学) Tencent YouTu Lab(腾讯YouTu实验室)

Comments Accepted to ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01558 2025-08-05 cs.CV

EvoVLMA: Evolutionary Vision-Language Model Adaptation

Kun Ding, Ying Wang, Shiming Xiang

机构 * Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

Comments This paper has been accepted by ACM Multimedia 2025 (ACM MM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01525 2025-08-05 cs.CV cs.AI

MiraGe: Multimodal Discriminative Representation Learning for Generalizable AI-Generated Image Detection

Kuo Shi, Jie Lu, Shanshan Ye, Guangquan Zhang, Zhen Fang

机构 * University of Technology Sydney(悉尼技术大学)

Comments Accepted to ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01282 2025-08-05 cs.HC

ExplorAR: Assisting Older Adults to Learn Smartphone Apps through AR-powered Trial-and-Error with Interactive Guidance

Jiawei Li, Linjie Qiu, Zhiqing Wu, Qiongyan Chen, Ziyan Wang, Mingming Fan

Comments 10 pages, 5 figures, Proceedings of the 33rd ACM International Conference on Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01250 2025-08-05 cs.CV

DisFaceRep: Representation Disentanglement for Co-occurring Facial Components in Weakly Supervised Face Parsing

Xiaoqin Wang, Xianxu Hou, Meidan Ding, Junliang Chen, Kaijun Deng, Jinheng Xie, Linlin Shen

机构 * School of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of AI and Advanced Computing, Xi’an Jiaotong-Liverpool University(西安交通大学利物浦大学人工智能与先进计算学院) Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University(香港理工大学电子与电气工程系) Department of Electrical and Computer Engineering, National University of Singapore(新加坡国立大学电子与计算机工程系) Computer Vision Institute, School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院计算机视觉研究所) Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen University(广东省智能信息处理重点实验室,深圳大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01236 2025-08-05 cs.CV

Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models

Mingyu Fu, Wei Suo, Ji Ma, Lin Yuanbo Wu, Peng Wang, Yanning Zhang

机构 * Northwestern Polytechnical University(西北工业大学) National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(集成空天地海大数据应用技术国家工程实验室) Swansea University(斯旺西大学)

Comments accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21177 2025-08-05 cs.CR cs.LG

FedBAP: Backdoor Defense via Benign Adversarial Perturbation in Federated Learning

Xinhai Yan, Libing Wu, Zhuangzhuang Zhang, Bingyi Liu, Lijuan Huo, Jing Wang

机构 * School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院) Department of Computer and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机与人工智能学院) School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)

Comments Accepted to ACM Multimedia 2025

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27--31, 2025, Dublin, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19821 2025-08-05 cs.CV cs.MM

LAVA: Language Driven Scalable and Versatile Traffic Video Analytics

Yanrui Yu, Tianfei Zhou, Jiaxin Sun, Lianpeng Qiao, Lizhong Ding, Ye Yuan, Guoren Wang

机构 * Beijing Institute of Technology(北京理工大学)

Comments Accepted by ACM MM 2025, code: https://github.com/yuyanrui/LAVA

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19958 2025-08-05 cs.CV

UltraVSR: Achieving Ultra-Realistic Video Super-Resolution with Efficient One-Step Diffusion Space

Yong Liu, Jinshan Pan, Yinchuan Li, Qingji Dong, Chao Zhu, Yu Guo, Fei Wang

机构 * Xi'an Jiaotong University(西安交通大学) Nanjing University of Science and Technology(南京理工大学) Huawei Noah's Ark Lab(华为诺亚实验室)

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12799 2025-08-05 cs.CV

TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors

Mingwei Li, Pu Pang, Hehe Fan, Hua Huang, Yi Yang

机构 * Zhejiang University(浙江大学) Xi'an Jiaotong University(西安交通大学) Beijing Normal University(北京师范大学)

Comments Project page: https://longxiang-ai.github.io/TSGS/ . Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06098 2025-08-05 eess.IV

ELFATT: Efficient Linear Fast Attention for Vision Transformers

Chong Wu, Maolin Che, Renjie Xu, Zhuoheng Ran, Hong Yan

Comments Accepted by ACM International Conference on Multimedia (MM '25) [The version includes Appendix]

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18277 2025-08-05 cs.CV cs.LG

Towards Modality Generalization: A Benchmark and Prospective Analysis

Xiaohao Liu, Xiaobo Xia, Zhuo Huang, See-Kiong Ng, Tat-Seng Chua

机构 * National University of Singapore(新加坡国立大学) The University of Sydney(悉尼大学)

Comments Accepted by ACM MM 2025 (CR)

详情

展开后加载摘要…

URL PDF HTML 收藏