A Benchmark for Omni-Modal Reasoning in Long Videos
长视频全模态推理基准
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Jinxing Zhou, Sahal Shaji Mullappilly, Mohammad Almansoori, Noor Ahsan, Beknur Kalmakhanbet, Sambal Shikhar, Rishabh Lalla, Jean Lahoud, Mariette Awad, Fahad Shahbaz Khan, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal
机构
*
Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
American University of Beirut(贝鲁特美国大学)
;
Linköping University(利尔贝里大学)
机构
*
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Nanyang Technological University(南洋理工大学)
;
Qwen Team, Alibaba Group(阿里巴巴集团Qwen团队)
机构
*
institutetext: MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models Yue Wu Changyuan Wang Zixuan Wang Shilin Ma Yansong Tang(机构文本:MorphoQuant:多模态大语言模型的模态感知量化 Yue Wu 王昌元 王梓轩 马世林 唐彦松)
Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets
Med-CRAFT:通过知识图谱遍历自动构建可解释的多跳视频工作负载
Shenxi Liu, Kan Li, Mingyang Zhao, Yuhang Tian, Bin Li
机构
*
Beijing Institute of Technology(北京理工大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所)
机构
*
School of Computer Science and Artificial Intelligence, Wuhan University of Technology(武汉理工大学计算机科学与人工智能学院)
;
School of Engineering, Swinburne University of Technology(斯威本科技大学工程学院)
;
School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
;
School of Computing, Engineering and Mathematical Sciences, La Trobe University(拉筹伯大学计算、工程与数学科学学院)
Multimodal Ambivalence/Hesitancy Recognition in Videos for Personalized Digital Health Interventions
视频中矛盾/犹豫识别用于个性化数字健康干预
Manuela González-González, Soufiane Belharbi, Muhammad Osama Zeeshan, Masoumeh Sharafi, Muhammad Haseeb Aslam, Lorenzo Sia, Nicolas Richet, Marco Pedersoli, Alessandro Lameiras Koerich, Simon L Bacon, Eric Granger
机构
*
LIVIA, Dept. of Systems Engineering, ETS Montreal, Canada(ETS蒙特利尔大学系统工程系LIVIA实验室)
;
LIVIA, Dept. of Software and IT Engineering, ETS Montreal, Canada(ETS蒙特利尔大学软件与信息工程系LIVIA实验室)
;
Dept. of Health, Kinesiology, & Applied Physiology, Concordia University, Montreal, Canada(康科迪亚大学健康、运动科学与应用生理学系)
;
Montreal Behavioural Medicine Centre, CIUSSS Nord-de-l’Ile-de-Montréal, Canada(蒙特利尔行为医学中心,蒙特利尔北岛卫生与社会服务局)
MLCR: Multi-Level Cue Refinement for Long-Term Multimodal Action Quality Assessment
PIDNet: 逐步隐式解耦网络用于多模态动作质量评估
Qiqi Li, Pengfei Wang, Hongyu Chen, Nenggan Zheng
机构
*
Qiushi Academy for Advanced Studies (QAAS), Zhejiang University(浙江大学启斯特先进研究院)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
School of Software Technology, Zhejiang University(浙江大学软件学院)
;
State Key Lab of Brain-Machine Intelligence(脑机智能国家重点实验室)
;
Collaborative Innovation Center for Artificial Intelligence by MOE and Zhejiang Provincial Government (ZJU)(教育部-浙江省人工智能协同创新中心)
;
Zhejiang Lab(浙江实验室)
机构
*
Applied Artificial Intelligence and Intelligent Systems (AAIINS) Laboratory(应用人工智能与智能系统实验室)
;
Department of Computer Science and Engineering(计算机科学与工程系)
;
Department of Data Science and Artificial Intelligence(数据科学与人工智能系)
;
Department of Software Systems & Cybersecurity(软件系统与网络安全系)
;
Energy and Resources Institute, Faculty of Science and Technology(能源与资源研究所,科学与技术学院)
;
Faculty of Science and Technology(科学与技术学院)
;
School of Engineering and Energy(工程与能源学院)
OpenRC: An Open-Source Robotic Colonoscopy Framework for Multimodal Data Acquisition and Autonomy Research
OpenRC:一种用于多模态数据采集和自主性研究的开源机器人结肠镜框架
Siddhartha Kapuria, Mohammad Rafiee Javazm, Naruhiko Ikoma, Joga Ivatury, Mohammad Ali Nasseri, Nassir Navab, Farshid Alambeigi
机构
*
Walker Department of Mechanical Engineering, The University of Texas at Austin(德克萨斯大学奥斯汀分校沃克机械工程系)
;
Department of Surgical Oncology, Division of Surgery, The University of Texas MD Anderson Cancer Center(德克萨斯大学MD安德森癌症中心外科肿瘤学系)
;
School of Medicine and Health, Technical University of Munich(慕尼黑工业大学医学与健康学院)
MoCA: Multi-modal Cross-masked Autoencoder for Time Series in Digital Health
MoCA:用于数字健康测量的多模态交叉掩码自编码器
Howon Ryu, Yuliang Chen, Yacun Wang, Andrea Z. LaCroix, Chongzhi Di, Loki Natarajan, Yu Wang, Jingjing Zou
机构
*
Herbert Wertheim School of Public Health and Human Longevity Science, University of California, San Diego, San Diego, CA, USA(赫伯特·韦瑟姆公共卫生与人类长寿科学学院,加州大学圣地亚哥分校)
;
Halıcıoğlu Data Science Institute, University of California San Diego, San Diego, CA, USA(哈利奇奥卢数据科学研究所,加州大学圣地亚哥分校)
;
Department of Computer Science and Engineering, University of California San Diego, San Diego, CA, USA(计算机科学与工程系,加州大学圣地亚哥分校)
;
Division of Public Health Sciences, Fred Hutchinson Cancer Center, Seattle, WA, USA(公共卫生科学部,Fred Hutchinson癌症中心)
Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach
第十一届ABAW竞赛中的团队RAS:多模态矛盾情绪识别方法
Elena Ryumina, Maxim Markitantov, Alexandr Axyonov, Fedor Shchetinin, Timur Abdulkadirov, Dmitry Ryumin, Alexey Karpov
机构
*
St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS)(俄罗斯科学院圣彼得堡联邦研究中心)
;
HSE University(圣彼得堡高等经济学院)
;
ITMO University(圣彼得堡国立信息技术机械与光学大学)
CommentsProceedings of the 39th Annual Conference on Neural Information Processing Systems, ARLET Workshop (Aligning Reinforcement Learning Experimentalists and Theorists)
Journal refTransactions on Machine Learning Research, Vol. 2026, June 2026