arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

共收录 1944
2508.20546 2025-08-29 cs.MM cs.AI

MM-HSD: Multi-Modal Hate Speech Detection in Videos

Berta Céspedes-Sarrias, Carlos Collado-Capell, Pablo Rodenas-Ruiz, Olena Hrynenko, Andrea Cavallaro

机构 * EPFL(苏黎世联邦理工学院) Idiap Research Institute(日内瓦研究所)

Comments Accepted at ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20530 2025-08-29 cs.CV

Enhancing Pseudo-Boxes via Data-Level LiDAR-Camera Fusion for Unsupervised 3D Object Detection

Mingqian Ji, Jian Yang, Shanshan Zhang

机构 * PCA Lab, School of Computer Science and Engineering, Nanjing University of Science and Technology(PCA实验室,计算机科学与工程学院,南京理工大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17232 2025-08-29 cs.MM cs.AI cs.CL

A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task

Mashiro Toyooka, Kiyoharu Aizawa, Yoko Yamakata

机构 * The University of Tokyo(东京大学)

Comments Accepted to ACM Multimedia 2025. The dataset are publicly available at: https://huggingface.co/datasets/mashi6n/nhkrecipe-100-anno-1

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10977 2025-08-29 cs.HC cs.CV

SMARTe-VR: Student Monitoring and Adaptive Response Technology for e-Learning in Virtual Reality

Roberto Daza, Lin Shengkai, Aythami Morales, Julian Fierrez, Katashi Nagao

机构 * BiDA Lab(BiDA实验室) Universidad Autonoma de Madrid(马德里自治大学) Nagao Laboratory(Nagao实验室) Nagoya University(名古屋大学)

Comments 10 pages, 3 figures. Published in ACM Intl. Conf. on Multimedia Workshops (ACM MM Workshops 2025, I2M-MM 25). Also presented at IEEE/CVF Conf. on Computer Vision and Pattern Recognition Workshops (CVPRW) and AAAI Workshop on Artificial Intelligence for Education (AI4EDU)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.15842 2025-08-28 cs.CV cs.GR

DiffArtist: Towards Structure and Appearance Controllable Image Stylization

Ruixiang Jiang, Changwen Chen

机构 * The Hong Kong Polytechnic University(香港理工大学)

Comments Accepted to ACM MM 2025, Homepage: https://DiffusionArtist.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.18904 2025-08-27 cs.CV

Event-Enriched Image Analysis Grand Challenge at ACM Multimedia 2025

Thien-Phuc Tran, Minh-Quang Nguyen, Minh-Triet Tran, Tam V. Nguyen, Trong-Le Do, Duy-Nam Ly, Viet-Tham Huynh, Khanh-Duy Le, Mai-Khiem Tran, Trung-Nghia Le

机构 * University of Science(科学大学) University of Dayton(代顿大学)

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14996 2025-08-27 cs.MM cs.CV cs.HC eess.IV

adder-viz: Real-Time Visualization Software for Transcoding Event Video

Andrew C. Freeman, Luke Reinkensmeyer

机构 * Baylor University(贝勒大学)

Comments Accepted to the Open-Source Track at ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06905 2025-08-27 cs.CV

MultiRef: Controllable Image Generation with Multiple Visual References

Ruoxi Chen, Dongping Chen, Siyuan Wu, Sinan Wang, Shiyun Lang, Petr Sushko, Gaoyang Jiang, Yao Wan, Ranjay Krishna

机构 * Zhejiang Wanli University(浙江万里大学) University of Washington(华盛顿大学) Huazhong University of Science and Technology(华中科技大学) Allen Institute for AI(人工智能研究院)

Comments Accepted to ACM MM 2025 Datasets

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03055 2025-08-27 cs.CV cs.AI

Uncertainty-Guided Face Matting for Occlusion-Aware Face Transformation

Hyebin Cho, Jaehyup Lee

机构 * Korea Advanced Institute of Science \& Technology School of Electrical Engineering Daejeon Republic of Korea Kyungpook National University School of Computer Science Korea Advanced Institute of Science \& Technology Kyungpook National University

Comments Accepted to ACM MM 2025. 9 pages, 8 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18372 2025-08-27 cs.CV

OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding

Hieu Nguyen, Phuc-Tan Nguyen, Thien-Phuc Tran, Minh-Quang Nguyen, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le

机构 * University of Science(科学大学) University of Dayton(戴维森大学)

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16801 2025-08-27 cs.CV

Decoupled Global-Local Alignment for Improving Compositional Understanding

Xiaoxing Hu, Kaicheng Yang, Jun Wang, Haoran Xu, Ziyong Feng, Yupei Wang

机构 * Beijing Institute of Technology(北京理工大学) Zhejiang University(浙江大学)

Comments ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17857 2025-08-26 cs.CV cs.AI

VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference

Pengfei Jiang, Hanjun Li, Linglan Zhao, Fei Chao, Ke Yan, Shouhong Ding, Rongrong Ji

机构 * Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(多媒体可信感知与高效计算重点实验室,教育部,厦门大学) Tencent Youtu Lab(腾讯优图实验室)

Comments Accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17166 2025-08-26 cs.MM eess.IV

Generative Flow Networks for Personalized Multimedia Systems: A Case Study on Short Video Feeds

Yili Jin, Ling Pan, Rui-Xiao Zhang, Jiangchuan Liu, Xue Liu

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17163 2025-08-26 cs.MM eess.IV

Generative AI for Multimedia Communication: Recent Advances, An Information-Theoretic Framework, and Future Opportunities

Yili Jin, Xue Liu, Jiangchuan Liu

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.09912 2025-08-26 cs.CV

E-4DGS: High-Fidelity Dynamic Reconstruction from the Multi-view Event Cameras

Chaoran Feng, Zhenyu Tang, Wangbo Yu, Yatian Pang, Yian Zhao, Jianbin Zhao, Li Yuan, Yonghong Tian

机构 * School of Electronic and Computer Engineering, Peking University(电子与计算机工程学院,北京大学) National University of Singapore(新加坡国立大学) School of Future Technology, Dalian University of Technology(未来技术学院,大连理工大学)

Comments 16 pages, 10 figures, 5 Tables, accepted by ACMMM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10072 2025-08-26 cs.CV

Frequency Regulation for Exposure Bias Mitigation in Diffusion Models

Meng Yu, Kun Zhan

机构 * School of Information Science and Engineering(信息科学与工程学院)

Comments ACM Multimedia 2025 accepted!

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04958 2025-08-26 cs.CV cs.MM

Boosting Temporal Sentence Grounding via Causal Inference

Kefan Tang, Lihuo He, Jisheng Dang, Xinbo Gao

机构 * School of Electronic Engineering, Xidian University Xi'an China School of Information Science \& Engineering, Lanzhou University Lanzhou China Xidian University Lanzhou University

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16859 2025-08-26 cs.CV

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark

Jinpeng Hu, Hongchang Shi, Chongyuan Dai, Zhuo Li, Peipei Song, Meng Wang

机构 * Hefei University of Technology(合肥工业大学) The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Science and Technology of China(中国科学技术大学) Institute of Artificial Intelligence (IAI), Hefei Comprehensive National Science Center(人工智能研究院(IAI),合肥综合性国家科学中心)

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16217 2025-08-25 cs.CV

PromptFlare: Prompt-Generalized Defense via Cross-Attention Decoy in Diffusion-Based Inpainting

Hohyun Na, Seunghoo Hong, Simon S. Woo

机构 * Sungkyunkwan University(成均馆大学)

Comments Accepted to ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16148 2025-08-25 cs.IR cs.CL cs.MM

Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering

Ao Zhou, Zebo Gu, Tenghao Sun, Jiawen Chen, Mingsheng Tu, Zifeng Cheng, Yafeng Yin, Zhiwei Jiang, Qing Gu

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学) Chongqing University of Posts and Telecommunications(重庆邮电大学)

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16147 2025-08-25 cs.IR

Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity Prediction

Ao Zhou, Mingsheng Tu, Luping Wang, Tenghao Sun, Zifeng Cheng, Yafeng Yin, Zhiwei Jiang, Qing Gu

Comments This paper has been accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15429 2025-08-25 cs.SD

AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation

Yulin Sun, Qisheng Xu, Yi Su, Qian Zhu, Yong Dou, Xinwang Liu, Kele Xu

机构 * College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学)

Comments 8 pages, 5 figures, accepted in ACM MM 2025 dataset track

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12415 2025-08-22 cs.CV

TiP4GEN: Text to Immersive Panorama 4D Scene Generation

Ke Xing, Hanwen Liang, Dejia Xu, Yuyang Yin, Konstantinos N. Plataniotis, Yao Zhao, Yunchao Wei

机构 * Institute of Information Science, Beijing Jiaotong University(信息科学研究院,北京交通大学) Visual Intelligence + X International Joint Laboratory(视觉智能+X国际联合实验室) University of Toronto(多伦多大学) The University of Texas at Austin(德克萨斯大学奥斯汀分校)

Comments Accepted In Proceedings of the 33rd ACM International Conference on Multimedia (MM' 25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15535 2025-08-22 cs.CV

Multi-Object Sketch Animation with Grouping and Motion Trajectory Priors

Guotao Liang, Juncheng Hu, Ximing Xing, Jing Zhang, Qian Yu

机构 * School of Software(软件学院) Beihang University(北航) Qingdao Research Institute(青岛研究院)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15272 2025-08-22 cs.CV

RATopo: Improving Lane Topology Reasoning via Redundancy Assignment

Han Li, Shaofei Huang, Longfei Xu, Yulu Gao, Beipeng Mu, Si Liu

机构 * School of Artificial Intelligence Beihang University Beijing China(人工智能学院 北航 北京 中国) Zhongguancun Academy Beijing China(中关村学院 北京 中国) Hangzhou International Innovation Institute Beihang University Hangzhou China(杭州国际创新研究院 北航 杭州 中国) Beihang University(北航) Zhongguancun Academy(中关村学院) University of Macau(澳门大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15232 2025-08-22 cs.CV

AeroDuo: Aerial Duo for UAV-based Vision and Language Navigation

Ruipu Wu, Yige Zhang, Jinyu Chen, Linjiang Huang, Shifeng Zhang, Xu Zhou, Liang Wang, Si Liu

机构 * Beihang University(北京航空航天大学) Sangfor Technologies Inc.(深信服科技有限公司) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15183 2025-08-22 cs.IR

Generating Negative Samples for Multi-Modal Recommendation

Yanbiao Ji, Dan Luo, Chang Liu, Shaokai Wu, Jing Tong, Qicheng He, Deyi Ji, Hongtao Lu, Yue Ding

Comments Accepted by ACM Multimedia

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14609 2025-08-21 cs.CV

AnchorSync: Global Consistency Optimization for Long Video Editing

Zichi Liu, Yinggui Wang, Tao Wei, Chao Ma

机构 * MoE Key Lab of Artificial Intelligence, AI Institute Shanghai Jiao Tong University Shanghai China(人工智能联合实验室,人工智能研究院,上海交通大学)

Comments ACM MM 2025; Code is released at https://github.com/VISION-SJTU/AnchorSync

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02906 2025-08-21 cs.CL cs.AI

Boosting Chart-to-Code Generation in MLLM via Dual Preference-Guided Refinement

Zhihan Zhang, Yixin Cao, Lizi Liao

机构 * Singapore Management University(新加坡国立管理学院) Fudan University(复旦大学)

Comments Accepted by ACM MM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14058 2025-08-21 cs.IR cs.AI

Dual-Phase Playtime-guided Recommendation: Interest Intensity Exploration and Multimodal Random Walks

Jingmao Zhang, Zhiting Zhao, Yunqi Lin, Jianghong Ma, Tianjun Wei, Haijun Zhang, Xiaofeng Zhang

机构 * Harbin Institute of Technology(哈尔滨工业大学) Nanyang Technological University(南洋理工大学)

Comments Accepted for publication at ACM Multimedia (ACM MM) 2025. 10 pages, 5 figures. Code and dataset: https://github.com/zqxwcevrtyui/DP2Rec

详情

展开后加载摘要…

URL PDF HTML 收藏