M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition
M2S-AVSR:面向鲁棒视听语音识别的模态感知多视角自监督表示
Fei Su, Cancan Li, Ming Li, Juan Liu
机构
*
School of Artificial Intelligence and the School of Computer Science, Wuhan University, China(人工智能学院和计算机科学学院,武汉大学,中国)
;
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(人工智能学院,香港中文大学(深圳),中国)
;
School of Artificial Intelligence, Wuhan University, China(人工智能学院,武汉大学,中国)
机构
*
institutetext: MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models Yue Wu Changyuan Wang Zixuan Wang Shilin Ma Yansong Tang(机构文本:MorphoQuant:多模态大语言模型的模态感知量化 Yue Wu 王昌元 王梓轩 马世林 唐彦松)
Physics Guided Conditional Diffusion Framework for Generative Inverse Design of Manufacturable Metasurface based Absorbers
基于物理引导的条件扩散模型的超材料吸收体逆向设计
Vineetha Joy, Jamshed Palai, Satwik Sahu, Anshuman Kumar, Amit Sethi, Hema Singh
机构
*
Centre for Electromagnetics, CSIR-National Aerospace Laboratories(电磁研究中心,国家航空航天实验室)
;
Birla Institute of Technology and Science, Pilani(比拉理工学院,皮兰)
;
Indian Institute of Technology, Bombay(孟买印度理工学院)
MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval
MCERF:通过增强检索推进工程文档的多模态大语言模型评估
Kiarash Naghavi Khanghah, Hoang Anh Nguyen, Anna C. Doris, Amir Mohammad Vahedi, Daniele Grandi, Faez Ahmed, Hongyi Xu
机构
*
School of Mechanical, Aerospace, and Manufacturing Engineering, University of Connecticut, Storrs, CT 06269(机械、航空航天与制造工程学院,康涅狄格大学,斯托尔斯,CT 06269)
;
Department of Mechanical Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139, USA(机械工程系,麻省理工学院,剑桥,MA 02139,美国)
OGA-AID: Clinician-in-the-loop AI Report Drafting Assistant for Multimodal Observational Gait Analysis in Post-Stroke Rehabilitation
OGA-AID:用于中风后康复多模态观察性步态分析的临床医生在环AI报告起草助手
Khoi T. N. Nguyen, Nghia D. Nguyen, Hui Yu Koh, Patrick W. H. Kwong, Karen Sui Geok Chua, Ananda Sidarta, Baosheng Yu
机构
*
Rehabilitation Research Institute of Singapore, Nanyang Technological University, Singapore(新加坡康复研究中心,南洋理工大学,新加坡)
;
Lee Kong Chian School of Medicine, Nanyang Technological University, Singapore(李光前医学院,南洋理工大学,新加坡)
;
The Grainger College of Engineering, University of Illinois Urbana-Champaign, United States(伊利诺伊大学厄巴纳-香槟分校格雷格学院,美国)
;
Department of Rehabilitation Sciences, The Hong Kong Polytechnic University, Hong Kong(香港理工大学康复科学系,香港)
;
VinUni-Illinois Smart Health Center, VinUniversity, Vietnam(Vin大学Vin-伊利诺伊智能健康中心,越南)
;
Institute of Rehabilitation Excellence, Tan Tock Seng Hospital, NHG Health, Singapore(卓越康复研究所,坦托克桑格医院,NHG健康,新加坡)
Comments11 pages, 9 figures. Accepted as a demo paper at ICAIL 2026. This arXiv version includes an appendix, new results, bug fixes, and presentation improvements beyond the earlier preprint; consequently, some reported numbers differ
Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents
面向视觉原生多模态深度搜索智能体的在策略数据演化
Shijue Huang, Hangyu Guo, Guanting Dong, Chenxin Li, Junting Lu, Xinyu Geng, Zhaochen Su, Zhenyu Li, Shuang Chen, Hongru Wang, Yi R. Fung
机构
*
Hong Kong University of Science and Technology(香港理工大学)
;
Renmin University of China(中国人民大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Peking University(北京大学)
;
Tsinghua University(清华大学)
;
University of Edinburgh(爱丁堡大学)
Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents
Agent规划基准:LLM Agent规划能力的诊断框架
Haoyu Sun, Wenxuan Wang, Mingyang Song, Jujie He, Weinan Zhang, Yang Liu, Yang Yang, Yu Cheng
机构
*
Tongji University(同济大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Fudan University(复旦大学)
;
Skywork AI
;
University of California, Santa Cruz(加州大学圣克鲁兹分校)
;
Shanghai Jiao Tong University(上海交通大学)
;
The Chinese University of Hong Kong(香港中文大学)
机构
*
School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院)
;
School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)
;
Independent Researcher(独立研究者)
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference
MACS: 模态感知容量缩放用于高效多模态MoE推理
Bo Li, Chuan Wu, Shaolin Zhu
机构
*
School of Software, Tsinghua University, Beijing, China(清华大学软件学院,北京,中国)
;
TJUNLP Lab, School of Computer Science and Technology, Tianjin University, China(天津大学计算机科学与技术学院,中国)
;
School of New Media and Communication, Tianjin University, China(天津大学新媒体与传播学院,中国)
TokaMind: A Multi-Modal Transformer Foundation Model for Tokamak Plasma Dynamics
TokaMind: 用于托卡马克等离子体动力学的多模态Transformer基础模型
Tobia Boschi, Andrea Loreti, Nicola C. Amorisco, Rodrigo H. Ordonez-Hurtado, Cécile Rousseau, George K. Holt, Eszter Székely, Alexander Whittle, Samuel Jackson, Adriano Agnello, Stanislas Pamela, Alessandra Pascale, Robert Akers, Juan Bernabe Moreno, Vassil Alexandrov, Mykhaylo Zayats
机构
*
IBM Research(IBM研究院)
;
UK Atomic Energy Authority(英国原子能局)
;
STFC Hartree Centre(科学与技术设施研究中心哈特ree中心)