CommentsThis is the extended version of the paper accepted in ICASSP'26, which will be publicly available in May. Authors' contributions may vary among the versions
Seeing the Forest and the Trees: Query-Aware Tokenizer for Long-Video Multimodal Language Models
看清森林与树木:面向长视频多模态语言模型的查询感知分词器
Siyou Li, Huanan Wu, Juexi Shao, Yinghao Ma, Yujian Gan, Yihao Luo, Yuwei Wang, Dong Nie, Lu Wang, Wenqing Wu, Le Zhang, Massimo Poesio, Juntao Yu
机构
*
Queen Mary University of London(伦敦女王学院)
;
University of Sheffield(谢菲尔德大学)
;
Imperial College London(伦敦帝国学院)
;
Pengcheng Laboratory(鹏城实验室)
;
Meta Inc(Meta公司)
;
Meituan Inc(美团公司)
;
Nanjing University of Science(南京理工大学)
;
University of Birmingham(伯明翰大学)
;
Utrecht University(乌得勒支大学)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
WeaveTime: 从早期帧中流式传输以形成涌现记忆在视频LLMs中
Yulin Zhang, Cheng Shi, Sibei Yang
机构
*
ShanghaiTech University(上海科技大学)
;
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
;
School of Computing and Data Science, The University of Hong Kong(香港大学计算科学与数据科学学院)
Primary-Fine Decoupling for Action Generation in Robotic Imitation
动作生成中的主-细分离
Xiaohan Lei, Min Wang, Wengang Zhou, Xingyu Lu, Houqiang Li
机构
*
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知国家重点实验室,中国科学技术大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(人工智能研究院,合肥综合性国家科学中心)
V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval
V-Retrver: 以证据驱动的代理推理用于通用多模态检索
Dongyang Chen, Chaoyang Wang, Dezhao Su, Xi Xiao, Zeyu Zhang, Jing Xiong, Qing Li, Yuzhang Shang, Shichao Kan
机构
*
Tsinghua University(清华大学)
;
University of Central Florida(中央佛罗里达大学)
;
Fudan University(复旦大学)
;
The Australian National University(澳大利亚国立大学)
;
The University of Hong Kong(香港大学)
;
Pengcheng Laboratory(鹏城实验室)
;
Central South University(中南大学)
The Design Space of Tri-Modal Masked Diffusion Models
三模态掩码扩散模型的设计空间
Louis Bethune, Victor Turrisi, Bruno Kacper Mlodozeniec, Pau Rodriguez Lopez, Lokesh Boominathan, Nikhil Bhendawade, Amitis Shidani, Joris Pelemans, Theo X. Olausson, Devon Hjelm, Paul Dixon, Joao Monteiro, Pierre Ablin, Vishnu Banna, Arno Blaas, Nick Henderson, Kari Noriy, Dan Busbridge, Josh Susskind, Marco Cuturi, Irina Belousova, Luca Zappella, Russ Webb, Jason Ramapuram
机构
*
Apple(苹果公司)
;
Google Deepmind(谷歌DeepMind)
;
University of Cambridge(剑桥大学)
;
MIT(麻省理工学院)
WildSVG: Towards Reliable SVG Generation Under Real-Word Conditions
WildSVG:迈向真实世界条件下可靠的SVG生成
Marco Terral, Haotian Zhang, Tianyang Zhang, Meng Lin, Xiaoqing Xie, Haoran Dai, Darsh Kaushik, Pai Peng, Nicklas Scharpff, David Vazquez, Joan Rodriguez
机构
*
QuiverAI
;
Columbia University(哥伦比亚大学)
;
Illinois Institute of Technology(伊利诺伊理工学院)
;
Mila - Quebec Artificial Intelligence Institute(魁北克人工智能研究所)
;
University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
;
ServiceNow Research(ServiceNow研究)
机构
*
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Tsinghua University(清华大学)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
University of Science and Technology of China(中国科学技术大学)
Architecture-Agnostic Curriculum Learning for Document Understanding: Empirical Evidence from Text-Only and Multimodal
文档理解的架构无关课程学习:来自纯文本和多模态的实证证据
Mohammed Hamdan, Vincenzo Dentamaro, Giuseppe Pirlo, Mohamed Cheriet
机构
*
1 Synchromedia Laboratory, \' E cole de Technologie Sup\' e rieure (\' E TS), 1100 Notre-Dame St W, Montreal, QC H3C 1K3, Canada
;
2 Department of Computer Science, University of Bari Aldo Moro, Via Orabona 4, 70125 Bari, Italy