arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

共收录 1944
2504.09451 2025-11-05 cs.CV

FractalForensics: Proactive Deepfake Detection and Localization via Fractal Watermarks

Tianyi Wang, Harry Cheng, Ming-Hui Liu, Mohan Kankanhalli

机构 * National University of Singapore(新加坡国立大学) Shandong University(山东大学)

Comments ACM Multimedia 2025 Oral

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16800 2025-11-04 cs.CV

Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception

Jiajing Lin, Zhenzhong Wang, Dejun Xu, Shu Jiang, YunPeng Gong, Min Jiang

机构 * School of Informatics, Xiamen University(厦门大学信息学院)

Comments Accepted by ACM MM 2025. Project Page: https://jiajinglin.github.io/Phys4DGen

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05227 2025-11-04 cs.RO cs.CV cs.LG cs.MM cs.SY eess.SY

NavigScene: Bridging Local Perception and Global Navigation for Beyond-Visual-Range Autonomous Driving

Qucheng Peng, Chen Bai, Guoxiang Zhang, Bo Xu, Xiaotong Liu, Xiaoyin Zheng, Chen Chen, Cheng Lu

机构 * Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学)

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.26759 2025-10-31 eess.IV cs.CV cs.MM

MORE: Multi-Organ Medical Image REconstruction Dataset

Shaokai Wu, Yapan Guo, Yanbiao Ji, Jing Tong, Yuxiang Lu, Mei Li, Suizhi Huang, Yue Ding, Hongtao Lu

机构 * Shanghai Jiao Tong University(上海交通大学) Suzhou Xiangcheng People’s Hospital(苏州湘城人民医院)

Comments Accepted to ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05696 2025-10-31 cs.CV

MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory

Ana Carolina Condez, Diogo Tavares, João Magalhães

机构 * NOVA LINCS, NOVA School of Science and Technology(NOVA LINCS,NOVA科学与技术学院)

Comments Updated version: corresponds to the ACM MM '25 published paper and includes full appendix material

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.14334 2025-10-31 cs.CV cs.HC cs.LG

Evaluating the Evaluators: Towards Human-aligned Metrics for Missing Markers Reconstruction

Taras Kucherenko, Derek Peristy, Judith Bütepage

机构 * SEED, Electronic Arts(SEED,电子艺名)

Comments Accepted at the ACM International Conference on Multimedia 2025 (ACM MM'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.01701 2025-10-31 cs.CV

Signal-SGN: A Spiking Graph Convolutional Network for Skeletal Action Recognition via Learning Temporal-Frequency Dynamics

Naichuan Zheng, Yuchen Du, Hailun Xia, Zeyu Liang

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.25163 2025-10-30 cs.CV

Target-Guided Bayesian Flow Networks for Quantitatively Constrained CAD Generation

Wenhao Zheng, Chenwei Sun, Wenbo Zhang, Jiancheng Lv, Xianggen Liu

机构 * College of Computer Science, Sichuan University(四川大学计算机科学学院) School of Computer Science and Technology, Xidian University(西安电子科技大学计算机科学与技术学院) Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education, Chengdu, China(教育部机器学习与工业智能工程研究中心)

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (2025) 3330-3339

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22481 2025-10-30 eess.IV cs.AI cs.CV cs.MM

Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework

Tianyi Liu, Kejun Wu, Chen Cai, Yi Wang, Kim-Hui Yap, Lap-Pui Chau

机构 * School of EEE, Nanyang Technological University(南洋理工大学电子工程系) School of EIC, Huazhong University of Science and Technology(华中科技大学电子信息学院) Dept. of EEE, The Hong Kong Polytechnic University(香港理工大学电子工程系)

Comments 10 pages, 5 figures, accepted by ACMMM 2025

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19085 2025-10-30 cs.LG

Clustering-Oriented Generative Attribute Graph Imputation

Mulin Chen, Bocheng Wang, Jiaxin Zhong, Zongcheng Miao, Xuelong Li

机构 * School of Artificial Intelligence, OPtics and ElectroNics (iOPEN)(人工智能学院(iOPEN)) Northwestern Polytechnical University(西北工业大学) Institute of Artificial Intelligence (TeleAI)(人工智能研究所(TeleAI))

Comments Accepted by ACM MM'25

Journal ref ACM MM (2025), pages 1092-1101

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20077 2025-10-29 cs.RO cs.CV cs.HC

Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning

Xun Li, Rodrigo Santa Cruz, Mingze Xi, Hu Zhang, Madhawa Perera, Ziwei Wang, Ahalya Ravendran, Brandon J. Matthews, Feng Xu, Matt Adcock, Dadong Wang, Jiajun Liu

机构 * CSIRO(澳大利亚联邦科学与工业研究组织)

Journal ref MM '25: Proceedings of the 33rd ACM International Conference on Multimedia (2025) Pages 12492 - 12500

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17939 2025-10-29 cs.CV cs.AI

GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning

Bo Liu, Xiangyu Zhao, Along He, Yidi Chen, Huazhu Fu, Xiao-Ming Wu

机构 * The Hong Kong Polytechnic University(香港理工大学) Shenzhen University(深圳大学) West China Hospital of Sichuan University(四川大学华西医院) IHPC, Agency for Science, Technology and Research(科技研究局IHPC)

Comments Accepted at ACM MM 2025 (also known as GEMeX-ThinkVG)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09193 2025-10-28 eess.IV cs.CV

BiECVC: Gated Diversification of Bidirectional Contexts for Learned Video Compression

Wei Jiang, Junru Li, Kai Zhang, Li Zhang

机构 * Bytedance(字节跳动)

Comments Accepted to ACMMM 2025

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia, pp.7248-7257, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22683 2025-10-28 cs.CV

Estimation of Fireproof Structure Class and Construction Year for Disaster Risk Assessment

Hibiki Ayabe, Kazushi Okamoto, Koki Karube, Atsushi Shibata, Kei Harada

机构 * The University of Electro-Communications(电通大学)

Journal ref Workshop on Visual and Signal Communication Technologies in Design of Housing, Urban Spaces, Local Communities, and Human Behavior in conjunction with ACM Multimedia Asia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22513 2025-10-28 cs.LG cs.AI

Toward Robust Signed Graph Learning through Joint Input-Target Denoising

Junran Wu, Beng Chin Ooi, Ke Xu

机构 * National University of Singapore(新加坡国立大学) Zhejiang University(浙江大学) Beihang University(北京航空航天大学)

Comments ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17924 2025-10-28 cs.CV

Gaze into the Heart: A Multi-View Video Dataset for rPPG and Health Biomarkers Estimation

Konstantin Egorov, Stepan Botman, Pavel Blinov, Galina Zubkova, Anton Ivaschenko, Alexander Kolsanov, Andrey Savchenko

机构 * Sber AI Lab(Sber AI实验室) Samara State Medical University(萨马拉州医学大学) ISP RAS Research Center for Trusted Artificial Intelligence(俄罗斯科学院信息与通信技术研究所可信人工智能研究中心)

Comments Accepted to ACMMM 2025, Datasets track

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22828 2025-10-28 cs.CV cs.AI

CapRecover: A Cross-Modality Feature Inversion Attack Framework on Vision Language Models

Kedong Xiu, Sai Qian Zhang

机构 * New York University(纽约大学)

Comments 9 pages, accepted by the 2025 ACM Multimedia Conference. Code is available at https://jus1mple.github.io/Image2CaptionAttack

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17394 2025-10-28 cs.CV cs.AI

HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs

Zhaolin Cai, Fan Li, Ziwei Zheng, Yanjun Qin

机构 * Xinjiang University(新疆大学) Xi'an Jiaotong University(西安交通大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04705 2025-10-28 cs.CV

Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations

Yuji Wang, Moran Li, Xiaobin Hu, Ran Yi, Jiangning Zhang, Han Feng, Weijian Cao, Yabiao Wang, Chengjie Wang, Lizhuang Ma

机构 * Shanghai Jiao Tong University, Tencent Youtu Lab(上海交通大学,腾讯云图实验室) Tencent Youtu Lab(腾讯云图实验室) Shanghai Jiao Tong University(上海交通大学) Tencent(腾讯)

Comments ACM Multimedia 2025; code URL: https://github.com/rain152/IPVG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19549 2025-10-28 cs.CV

DEEMO: De-identity Multimodal Emotion Recognition and Reasoning

Deng Li, Bohao Xing, Xin Liu, Baiqiang Xia, Bihan Wen, Heikki Kälviäinen

机构 * Lappeenranta-Lahti University of Technology LUT(拉普兰塔-拉赫蒂技术大学) Nanyang Technological University(南洋理工大学) Brno University of Technology(布拉格技术大学)

Comments Accepted by ACMMM 2025

Journal ref Proceedings of the 33rd ACM International Conference on Multimedia (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04273 2025-10-28 cs.IR cs.CV cs.MM cs.SD eess.AS

Audio Does Matter: Importance-Aware Multi-Granularity Fusion for Video Moment Retrieval

Junan Lin, Daizong Liu, Xianke Chen, Xiaoye Qu, Xun Yang, Jixiang Zhu, Sanyuan Zhang, Jianfeng Dong

机构 * Zhejiang University(浙江大学) Peking University(北京大学) Zhejiang Gongshang University(浙江工商大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Science and Technology of China(中国科学技术大学)

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06278 2025-10-28 cs.RO cs.HC

Robust Understanding of Human-Robot Social Interactions through Multimodal Distillation

Tongfei Bian, Mathieu Chollet, Tanaya Guha

机构 * University of Glasgow(格拉斯哥大学) University of Glasgow School of Computer Science(格拉斯哥大学计算机科学学院)

Comments Accepted by ACM Multimedia 2025, camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20393 2025-10-24 cs.CV cs.MM

Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval

Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim

机构 * Singapore Management University(新加坡管理大学)

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20327 2025-10-24 cs.LG cs.AI

LEGO: A Lightweight and Efficient Multiple-Attribute Unlearning Framework for Recommender Systems

Fengyuan Yu, Yuyuan Li, Xiaohua Feng, Junjie Fang, Tao Wang, Chaochao Chen

机构 * Zhejiang University(浙江大学) Hangzhou Dianzi University(杭州电子科技大学)

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19451 2025-10-23 cs.CV cs.MM

Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis

Xueqi Ma, Yanbei Jiang, Sarah Erfani, James Bailey, Weifeng Liu, Krista A. Ehinger, Jey Han Lau

机构 * The University of Melbourne(墨尔本大学) China University of Petroleum (East China)(中国石油大学(华东))

Comments Accepted by ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02266 2025-10-23 cs.CV cs.HC

NeuroSwift: A Lightweight Cross-Subject Framework for fMRI Visual Reconstruction of Complex Scenes

Shiyi Zhang, Dong Liang, Yihang Zhou

机构 * Department of Electronic and Electrical Engineering, Southern University of Science and Technology(电子与电气工程系,南方科技大学) Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(深圳先进技术研究院,中国科学院)

Journal ref ACM Multimedia Asia (MMAsia), 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09617 2025-10-22 cs.CV

WMamba: Wavelet-based Mamba for Face Forgery Detection

Siran Peng, Tianshuo Zhang, Li Gao, Xiangyu Zhu, Haoyuan Zhang, Kai Pang, Zhen Lei

机构 * CMFT Co., Ltd.(CMFT公司) Guangzhou Pixel Solutions Co., Ltd.(广州像素解决方案公司)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16463 2025-10-21 cs.CV

HGC-Avatar: Hierarchical Gaussian Compression for Streamable Dynamic 3D Avatars

Haocheng Tang, Ruoke Yan, Xinhui Yin, Qi Zhang, Xinfeng Zhang, Siwei Ma, Wen Gao, Chuanmin Jia

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University.(信息处理国家重点实验室,北京大学计算机科学学院) School of Computer Science and Technology, University of Chinese Academy of Sciences.(中国科学院大学计算机科学与技术学院) WICT, State Key Laboratory of Multimedia Information Processing, Peking University.(信息处理国家重点实验室,北京大学WICT) Institute for Clarity in Documentation(文档清晰性研究所) Inria Paris-Rocquencourt(巴黎- Rocquencourt 分部,Inria) Rajiv Gandhi University(拉贾·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(Palmer研究实验室)

Comments ACM International Conference on Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15752 2025-10-20 cs.CV cs.AI

NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image Generation

Yitong Sun, Yao Huang, Ruochen Zhang, Huanran Chen, Shouwei Ruan, Ranjie Duan, Xingxing Wei

机构 * Institute of Artificial Intelligence, Beihang University(北航人工智能研究院) College of AI, Tsinghua University(清华人工智能学院) Security Group, Alibaba Group(阿里集团安全组)

Comments 10 pages, 8 figures, accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12493 2025-10-20 cs.CV

BSGS: Bi-stage 3D Gaussian Splatting for Camera Motion Deblurring

An Zhao, Piaopiao Yu, Zhe Zhu, Mingqiang Wei

机构 * Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

Comments Accept by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏