arXivDaily arXiv每日学术速递 周一至周五更新

期刊&会议

ACM International Conference on Multimedia · 会议 · Multimedia

2026-08-14 至 2026-08-14 共收录 4
2603.21783 2026-08-14 cs.CV 版本更新

SHARP: Spectrum-aware Highly-dynamic Adaptation for Resolution Promotion in Remote Sensing Synthesis

SHARP: 为遥感合成促进分辨率提升的频谱感知高动态适应

Bingxuan Zhao, Qing Zhou, Chuang Yang, Junyu Gao, Qi Wang

机构 * School of Computer Science, Northwestern Polytechnical University, Xi'an, China(计算机科学学院,西北工业大学,西安,中国) School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University, Xi'an, China(人工智能学院,光学与电子学(iOPEN),西北工业大学,西安,中国)

AI总结 本文提出SHARP方法,通过频谱感知的高动态适应策略,在无需训练的情况下提升遥感图像分辨率,优于现有无训练基线,在CLIP分数、审美分数和HPSv2上表现更优。

Comments Accepted by the 34th ACM International Conference on Multimedia (ACM MM 2026)

Journal ref Proceedings of the 34th ACM International Conference on Multimedia (ACM MM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21628 2026-08-14 cs.CR cs.LG 版本更新

Noise as a Probe: Membership Inference Attacks on Diffusion Models Leveraging Initial Noise

噪声作为探针:利用初始噪声的扩散模型成员推断攻击

Puwei Lian, Yujun Cai, Songze Li, Bingkun Bao

机构 * Southeast University(东南大学) The University of Queensland(昆士兰大学) Nanjing University of Posts and Telecommunications(南京邮电大学)

AI总结 本文提出利用扩散模型初始噪声中的残余语义信息进行成员推断攻击,揭示了微调模型在隐私保护方面的脆弱性。

Comments Accepted to 34th ACM International Conference on Multimedia (MM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22282 2026-08-14 cs.CV cs.AI cs.CL 版本更新

CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning

CityRiSE:基于强化学习的大视觉语言模型城市社会经济地位推理框架

Tianhui Liu, Hetian Pang, Xin Zhang, Jie Feng, Pan Hui, Yong Li

机构 * Information Hub, The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)信息中心) Department of Electronic Engineering, BNRist, Tsinghua University(清华大学电子工程系)

AI总结 本研究提出CityRiSE框架,结合强化学习与大视觉语言模型,提升城市社会经济感知的预测准确性与泛化能力,尤其在未见城市和指标上表现优异。

Comments Accepted by ACM MM 2026, https://github.com/tsinghua-fib-lab/CityRiSE

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08199 2026-08-14 cs.CV 版本更新

Generalizable Operating Room Expert with Multimodal Enhancement

具备多模态增强能力的可泛化手术室专家

Peiqi He, Zhenhao Zhang, Yixiang Zhang, Jiaxin Liu, Xiongjun Zhao, Shaoliang Peng

AI总结 针对手术室空间建模的局限,提出仅用RGB图像推理的多模态大语言模型OR-Expert,通过内部推导空间线索实现三维推理,在手术室基准上达SOTA且泛化性良好。

Comments Accepted by ACM Multimedia 2026

详情

展开后加载摘要…

URL PDF HTML 收藏