arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

共收录 1944
2412.05293 2024-12-10 cs.CV cs.LG

FodFoM: Fake Outlier Data by Foundation Models Creates Stronger Visual Out-of-Distribution Detector

Jiankang Chen, Ling Deng, Zhiyong Gan, Wei-Shi Zheng, Ruixuan Wang

Comments 13 pages, 7 figures

Journal ref Proceedings of the 32nd ACM International Conference on Multimedia, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08935 2024-12-10 cs.CV

SDDNet: Style-guided Dual-layer Disentanglement Network for Shadow Detection

Runmin Cong, Yuchen Guan, Jinpeng Chen, Wei Zhang, Yao Zhao, Sam Kwong

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08930 2024-12-10 cs.CV

Point-aware Interaction and CNN-induced Refinement Network for RGB-D Salient Object Detection

Runmin Cong, Hongyu Liu, Chen Zhang, Wei Zhang, Feng Zheng, Ran Song, Sam Kwong

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.08924 2024-12-10 cs.CV

Frequency Perception Network for Camouflaged Object Detection

Runmin Cong, Mengyao Sun, Sanyi Zhang, Xiaofei Zhou, Wei Zhang, Yao Zhao

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01343 2024-12-03 cs.CV

MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models

Xiaomin Li, Xu Jia, Qinghe Wang, Haiwen Diao, Mengmeng Ge, Pengxiang Li, You He, Huchuan Lu

Comments Accepted by ACM MM 2024, code will be released in https://github.com/XiaominLi1997/MoTrans

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00683 2024-12-03 cs.CV

DMFourLLIE: Dual-Stage and Multi-Branch Fourier Network for Low-Light Image Enhancement

Tongshun Zhang, Pingping Liu, Ming Zhao, Haotian Lv

Comments Accepted to ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.15829 2024-12-03 cs.CV

SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization

Sicheng Liu, Lintao Wang, Xiaogang Zhu, Xuequan Lu, Zhiyong Wang, Kun Hu

Comments 8 pages, 5 figures, submitted to ACM Multimedia Asia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2212.00459 2024-12-03 cs.MM cs.CV

Disparity-based Stereo Image Compression with Aligned Cross-View Priors

Yongqi Zhai, Luyang Tang, Yi Ma, Rui Peng, Ronggang Wang

Comments 10 pages, 8 figures, published to ACM Multimedia 2022

Journal ref Proceedings of the 30th ACM International Conference on Multimedia (MM '22), October 10--14, 2022, Lisboa, Portugal

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17209 2024-11-27 cs.CV

LampMark: Proactive Deepfake Detection via Training-Free Landmark Perceptual Watermarks

Tianyi Wang, Mengxiao Huang, Harry Cheng, Xiao Zhang, Zhiqi Shen

Comments Accepted to ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15758 2024-11-26 cs.AI cs.CY cs.SI

Decoding Urban Industrial Complexity: Enhancing Knowledge-Driven Insights via IndustryScopeGPT

Siqi Wang, Chao Liang, Yunfan Gao, Yang Liu, Jing Li, Haofen Wang

Comments 9 pages, 6 figures, the 32nd ACM International Conference on Multimedia

Journal ref In Proceedings of the 32nd ACM International Conference on Multimedia, pp. 4757-4765 (2024, October)

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.08352 2024-11-26 cs.CV

Show Me What I Like: Detecting User-Specific Video Highlights Using Content-Based Multi-Head Attention

Uttaran Bhattacharya, Gang Wu, Stefano Petrangeli, Viswanathan Swaminathan, Dinesh Manocha

Comments 14 pages, 5 figures, 7 tables

Journal ref In Proceedings of the 30th ACM International Conference on Multimedia, 2022, Lisboa, Portugal

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.00262 2024-11-26 cs.MM cs.LG

Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning

Uttaran Bhattacharya, Elizabeth Childs, Nicholas Rewkowski, Dinesh Manocha

Comments 11 pages, 4 figures, 2 tables. Proceedings of the 29th ACM International Conference on Multimedia, October 20-24, 2021, Virtual Event, China

Journal ref In Proceedings of the 29th ACM International Conference on Multimedia, 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14242 2024-11-22 cs.CV cs.MM

Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images

Bo Yuan, Danpei Zhao, Zhuoran Liu, Wentao Li, Tian Li

Comments Accepted in ACMMM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.14923 2024-11-20 cs.CV

RayFormer: Improving Query-Based Multi-Camera 3D Object Detection via Ray-Centric Strategies

Xiaomeng Chu, Jiajun Deng, Guoliang You, Yifan Duan, Yao Li, Yanyong Zhang

Comments Accepted by ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10773 2024-11-19 eess.IV cs.CV

An End-to-End Real-World Camera Imaging Pipeline

Kepeng Xu, Zijia Ma, Li Xu, Gang He, Yunsong Li, Wenxin Yu, Taichu Han, Cheng Yang

Comments accept by ACMMM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10742 2024-11-19 cs.CV

It Takes Two: Accurate Gait Recognition in the Wild via Cross-granularity Alignment

Jinkai Zheng, Xinchen Liu, Boyue Zhang, Chenggang Yan, Jiyong Zhang, Wu Liu, Yongdong Zhang

Comments 12 pages, 9 figures; Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03472 2024-11-19 cs.CL

PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction

Zening Lin, Jiapeng Wang, Teng Li, Wenhui Liao, Dayi Huang, Longfei Xiong, Lianwen Jin

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.16824 2024-11-15 cs.CV

V2A-Mark: Versatile Deep Visual-Audio Watermarking for Manipulation Localization and Copyright Protection

Xuanyu Zhang, Youmin Xu, Runyi Li, Jiwen Yu, Weiqi Li, Zhipei Xu, Jian Zhang

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01781 2024-11-12 cs.CV

MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation

Duc Dang Trung Tran, Byeongkeun Kang, Yeejin Lee

Comments 14 pages, 9 figures, 7 tables, conference

Journal ref ACM Multimedia 2024, pages 1467-1475

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05322 2024-11-11 cs.MM cs.CV

Rate-aware Compression for NeRF-based Volumetric Video

Zhiyu Zhang, Guo Lu, Huanxiong Liang, Zhengxue Cheng, Anni Tang, Li Song

Comments Accepted by ACM MM 2024 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.15721 2024-11-11 cs.AI cs.CL

Hal-Eval: A Universal and Fine-grained Hallucination Evaluation Framework for Large Vision Language Models

Chaoya Jiang, Hongrui Jia, Wei Ye, Mengfan Dong, Haiyang Xu, Ming Yan, Ji Zhang, Shikun Zhang

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14789 2024-11-08 cs.CV

Revisiting Surgical Instrument Segmentation Without Human Intervention: A Graph Partitioning View

Mingyu Sheng, Jianan Fan, Dongnan Liu, Ron Kikinis, Weidong Cai

Comments Accepted by The 32nd ACM International Conference on Multimedia (ACM MM 2024) Workshop on Multimedia Computing for Health and Medicine (MCHM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.16591 2024-11-08 cs.CV

CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification

Qijie Wang, Guandu Liu, Bin Wang

Comments ACM Multimedia 2024 Poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.03595 2024-11-07 cs.MM

Investigating Conceptual Blending of a Diffusion Model for Improving Nonword-to-Image Generation

Chihaya Matsuhira, Marc A. Kastner, Takahiro Komamizu, Takatsugu Hirayama, Ichiro Ide

Comments Paper accepted at ACM MM 2024 (doi: 10.1145/3664647.3681202) with supplementary materials concatenated

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10229 2024-11-06 cs.CL

Generative Text Steganography with Large Language Model

Jiaxuan Wu, Zhengxian Wu, Yiming Xue, Juan Wen, Wanli Peng

Comments 9 pages, 4 figures, accepted at ACM Multimedia 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.01545 2024-11-05 cs.CV

Towards Small Object Editing: A Benchmark Dataset and A Training-Free Approach

Qihe Pan, Zhen Zhao, Zicheng Wang, Sifan Long, Yiming Wu, Wei Ji, Haoran Liang, Ronghua Liang

Comments 9 pages, 8 figures, Accepted by ACMMM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12711 2024-11-05 cs.CV cs.AI

Unsupervised Visible-Infrared Person ReID by Collaborative Learning with Neighbor-Guided Label Refinement

De Cheng, Xiaojian Huang, Nannan Wang, Lingfeng He, Zhihui Li, Xinbo Gao

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.12673 2024-11-05 cs.CV cs.AI

Efficient Bilateral Cross-Modality Cluster Matching for Unsupervised Visible-Infrared Person ReID

De Cheng, Lingfeng He, Nannan Wang, Shizhou Zhang, Zhen Wang, Xinbo Gao

Comments Accepted by ACM MM 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01437 2024-11-01 cs.CV cs.AI

Kvasir-VQA: A Text-Image Pair GI Tract Dataset

Sushant Gautam, Andrea Storås, Cise Midoglu, Steven A. Hicks, Vajira Thambawita, Pål Halvorsen, Michael A. Riegler

Comments to be published in VLM4Bio 2024, part of the ACM Multimedia (ACM MM) conference 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.22983 2024-10-31 cs.LG

Dual-Optimized Adaptive Graph Reconstruction for Multi-View Graph Clustering

Zichen Wen, Tianyi Wu, Yazhou Ren, Yawen Ling, Chenhang Cui, Xiaorong Pu, Lifang He

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏