机构
*
The Hong Kong University of Science and Technology (Guangzhou), China(香港科技大学(广州)中国)
;
The Hong Kong University of Science and Technology, Hong Kong SAR, China(香港科技大学,香港特别行政区,中国)
专题命中
视觉问答
:multimodal large language model(abstract);分类 cs.CV
机构
*
Key Laboratory of Intelligent Information Processing, Institute of Computing Technology (ICT), Chinese Academy of Sciences (CAS)(智能信息处理重点实验室,计算技术研究所(ICT),中国科学院(CAS))
;
Jilin University (JLU)(吉林大学(JLU))
;
Beijing University of Posts and Telecommunications (BUPT)(北京邮电大学(BUPT))
;
University of the Chinese Academy of Sciences(中国科学院大学)
CommentsAccepted by ACL 2026 Main. 17 pages, 7 figures, 8 tables. TL;DR: We propose MM-Mem, a cognition-inspired, dual-trace hierarchical memory framework for long-horizon video understanding grounded in Fuzzy-Trace Theory. It features adaptive memory compression via the Information Bottleneck and employs an entropy-driven top-down retrieval to access fine-grained details only when necessary
机构
*
University of Amsterdam(阿姆斯特丹大学)
;
University of Bristol(布里斯托大学)
;
Amazon AGI(亚马逊人工智慧)
;
College of Business and Economics(商学院和经济学学院)
;
University of Johannesburg(约翰内斯堡大学)
专题命中
视觉定位与Grounding
:grounding(summary_cn,abstract);MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.AI
AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison
AD-Copilot:通过视觉上下文比较实现工业异常检测的视觉-语言助手
Xi Jiang, Yue Guo, Jian Li, Yong Liu, Bin-Bin Gao, Hanqiu Deng, Jun Liu, Heng Zhao, Chengjie Wang, Feng Zheng
机构
*
Department of Computer Science and Engineering, Southern University of Science and Technology(计算机科学与工程系,南方科技大学)
;
Macau University of Science and Technology(澳门科技大学)
;
Tencent YouTu Lab(腾讯YouTu实验室)
;
Nanjing University(南京大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Centre for Frontier AI Research, A*STAR(前沿人工智能研究中心,A*STAR)
专题命中
视觉定位与Grounding
:MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
Align then Refine: Text-Guided 3D Prostate Lesion Segmentation
对齐后再细化:基于文本的3D前列腺病变分割
Cuiling Sun, Linkai Peng, Adam Murphy, Elif Keles, Hiten D. Patel, Ashley Ross, Frank Miller, Baris Turkbey, Andrea Mia Bejar, Halil Ertugrul Aktas, Gorkem Durak, Ulas Bagci
机构
*
Department of Radiology, Northwestern University, Chicago, USA(放射科,西北大学,芝加哥,美国)
;
Department of Urology, Northwestern University, Chicago, USA(泌尿科,西北大学,芝加哥,美国)
;
Center for Cancer Research, National Cancer Institute, Bethesda, USA(癌症研究中心,国家癌症研究所,贝塞斯达,美国)
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
OmniParser V2:结构化思维点用于统一的视觉文本解析及其在多模态大语言模型中的通用性
Wenwen Yu, Zhibo Yang, Jianqiang Wan, Sibo Song, Jun Tang, Wenqing Cheng, Yuliang Liu, Xiang Bai
机构
*
School of Information Science and Engineering, East China University of Science and Technology(东华大学信息科学与工程学院)
;
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology(华中科技大学人工智能与自动化学院)
;
School of Software Engineering, Huazhong University of Science and Technology(华中科技大学软件工程学院)
;
School of Electronic Information and Communications, Huazhong University of Science and Technology(华中科技大学电子信息与通信学院)
;
Alibaba Group(阿里巴巴集团)
专题命中
文档图表理解
:multimodal large language model(title,abstract);分类 cs.CV
Fine-tuning DeepSeek-OCR-2 for Molecular Structure Recognition
为分子结构识别微调DeepSeek-OCR-2
Haocheng Tang, Xingyu Dang, Junmei Wang
机构
*
Department of Computer Science, Princeton University, Princeton, NJ, USA(普林斯顿大学计算机科学系)
;
School of Pharmacy, University of Pittsburgh, Pittsburgh, PA, USA(匹兹堡大学药学院)
;
Khoury College of Computer Science, Northeastern University, Boston, MA, USA(东北大学Khoury计算机科学学院)
;
Computational Chemical Genomics Screening Center, University of Pittsburgh, Pittsburgh, PA, USA(匹兹堡大学计算化学基因组筛选中心)