Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
在自进化大型多模态模型中更加关注视觉标记
Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar, Rao Muhammad Anwer, Hisham Cholakkal, Salman Khan, Fahad Khan
机构
*
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Aalto University(阿尔托大学)
;
Australian National University(澳大利亚国立大学)
;
Linköping University(林雪平大学)
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models
从结构到协同:多模态大语言模型中视觉-语言感知范式演进综述
Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao, Jiancheng Lv
机构
*
School of Computer Science, Sichuan University(四川大学计算机学院)
;
School of Electronic and Computer Engineering, Peking University Shenzhen Graduate School(北京大学深圳研究生院电子与计算机工程学院)
;
Institute of Artificial Intelligence (TeleAI), China Telecom and Northwestern Polytechnical University(中国电信与西北工业大学人工智能研究院(TeleAI))
专题命中
视觉推理
:multimodal large language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.CV、cs.AI、cs.LG
机构
*
Central South University(中南大学)
;
Tsinghua University(清华大学)
;
South China Normal University(华南师范大学)
;
ByteDance Inc(字节跳动公司)
;
University of Zaragoza(阿拉维达大学)
;
CosmosMind
;
Wuhan University(武汉大学)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
Southeast University(东南大学)
;
Tencent(腾讯公司)
;
Nankai University(南开大学)
;
Supermicro Computer Inc(Supermicro计算机公司)
;
Huazhong University of Science and Technology(华中科技大学)
Position Rebinding Cache Reuse: Replay-Free Visual Revisiting for Interleaved Multimodal Reasoning
位置重绑定缓存复用:交错多模态推理中无重放的视觉重访
Mengzhao Wang, Yanli Ji, Wangmeng Zuo, Peng Ye, Chongjun Tu
机构
*
Sun Yat-sen University(中山大学)
;
Shenzhen Loop Area Institute(深圳河套学院)
;
Harbin Institute of Technology (HIT)(哈尔滨工业大学)
;
Fudan University(复旦大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong(香港中文大学)
Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning
Visual-OPSD:用于高效统一多模态推理的跨模态在策略自蒸馏
Pengyu Li, Zhitao Gao, Lingling Zhang, Muye Huang, Yuanming Li, Fangzhi Xu, Jun Liu
机构
*
Xi’an Jiaotong University(西安交通大学)
;
MOE KLINNS Lab(MOE KLINNS实验室)
;
Shaanxi Province Key Laboratory of Big Data Knowledge Engineering(陕西省大数据知识工程重点实验室)
;
Sun Yat-sen University(中山大学)
TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation
TAVR-VLM:面向抗幻觉报告生成的风险条件因果基础
Zhixiang Lu, Xiwei Liu, Sifan Song, Changkai Ji, Anh Nguyen, Jionglong Su, Imran Razzak, Jinfeng Wang
机构
*
Xi’an Jiaotong-Liverpool University(西交利物浦大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
University of Liverpool(利物浦大学)
;
Kunming University of Science and Technology(昆明理工大学)
专题命中
视觉定位与Grounding
:VLM(title,title_cn);grounding(title,abstract);multimodal large language model(abstract);分类 cs.AI
KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction
KG-TRACE:一种用于抗菌素耐药性预测的机制性基础神经符号框架
Naman Garg, Sarika Jain, Sourav Yadav, Bharat K. Bhargava, Ghanapriya Singh, Abhishek Srivastava, Parimal Kar
机构
*
National Institute of Technology Kurukshetra(国立库鲁克舍特拉理工学院)
;
Indian Institute of Information Technology Manipur(印度曼尼普尔信息技术学院)
;
Purdue University(普渡大学)
;
Indian Institute of Technology Indore(印度印多尔理工学院)
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Shenzhen University(深圳大学)
;
Beijing University of Technology(北京工业大学)
;
JD Explore Academy(京东探索研究院)
See & Sniff: Learning Visuo-Olfactory Representations
See & Sniff: 学习视觉-嗅觉表征
Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Son Chung, Arda Senocak
机构
*
Korea Advanced Institute of Science and Technology(韩国科学技术院)
;
Hankuk University of Foreign Studies(韩国外国语大学)
;
Ulsan National Institute of Science and Technology(蔚山科学技术院)
PressMimic: Pressure-Guided Motion Capture and Control for Humanoid Robot Imitation
PressMimic: 压力引导的动作捕捉与控制用于人形机器人模仿
Yi Lu, Shenghao Ren, Tianyu Xiong, Zhaoxiang Li, Jiaqi Li, He Zhang, Tao Yu, Qiu Shen, Xun Cao
机构
*
School of Electronic Science and Engineering, Nanjing University(南京大学电子科学与工程学院)
;
Key Laboratory of Optoelectronic Devices and Systems with Extreme Performances of MOE, Nanjing University(南京大学极端性能光电技术与系统教育部重点实验室)
;
BNRist, Tsinghua University(清华大学北京信息科学与技术国家研究中心)
MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation
MKG-RAG-Bench:多模态知识图谱增强生成中的检索基准
Xiaochen Wang, Bao Hoang, Han Liu, Ting Wang, Fenglong Ma
机构
*
The Pennsylvania State University(宾夕法尼亚州立大学)
;
Michigan State University(密歇根州立大学)
;
Dalian University of Technology(大连理工大学)
;
Stony Brook University(石溪大学)