Physics Guided Conditional Diffusion Framework for Generative Inverse Design of Manufacturable Metasurface based Absorbers
基于物理引导的条件扩散模型的超材料吸收体逆向设计
Vineetha Joy, Jamshed Palai, Satwik Sahu, Anshuman Kumar, Amit Sethi, Hema Singh
机构
*
Centre for Electromagnetics, CSIR-National Aerospace Laboratories(电磁研究中心,国家航空航天实验室)
;
Birla Institute of Technology and Science, Pilani(比拉理工学院,皮兰)
;
Indian Institute of Technology, Bombay(孟买印度理工学院)
MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models
MMBU: 大规模多模态生物医学理解基准,用于探测视觉语言模型的感知能力
Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun, Daniel Vela Jarquin, Min Woo Sun, Josiah Aklilu, James Burgess, Yuhui Zhang, Ryan Nayebi, Paola Avila, Robayo, Jin Ye, Ming Hu, Zhongying Deng, Junjun He, Xin Chen, Yue Yao, Robert Tibshirani, Jeffrey J. Nirschl, Serena Yeung-Levy
机构
*
Stanford University(斯坦福大学)
;
University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
;
Instituto Tecnológico de Monterrey(蒙特雷技术学院)
;
Monash University(墨尔本大学)
;
University of Cambridge(剑桥大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
Shandong University(山东大学)
MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval
MCERF:通过增强检索推进工程文档的多模态大语言模型评估
Kiarash Naghavi Khanghah, Hoang Anh Nguyen, Anna C. Doris, Amir Mohammad Vahedi, Daniele Grandi, Faez Ahmed, Hongyi Xu
机构
*
School of Mechanical, Aerospace, and Manufacturing Engineering, University of Connecticut, Storrs, CT 06269(机械、航空航天与制造工程学院,康涅狄格大学,斯托尔斯,CT 06269)
;
Department of Mechanical Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139, USA(机械工程系,麻省理工学院,剑桥,MA 02139,美国)
From Vision to Text: A Compact Multimodal Approach for Robust, Cross-Domain Presentation Attack Detection on ID Cards
从视觉到文本:一种用于身份证件跨域鲁棒演示攻击检测的紧凑多模态方法
Qingwen Zeng, Juan E. Tapia, Sneha Das, Christoph Busch
机构
*
da/sec-Biometrics and Security Research Group, Hochschule Darmstadt(da/sec生物安全研究组,达姆施塔特应用技术大学)
;
Technical University of Denmark (DTU)(丹麦技术大学(DTU))
Evidence-Based Intelligent Diagnostic and Therapeutic Visualization System with Large Language Models: Multi-Turn Interaction and Multimodal Treatment Plan Generation
基于证据的智能诊断与治疗可视化系统与大语言模型:多轮交互与多模态治疗方案生成
Yunhan Wang, Yuda Wang, Zhiying Tu, Mingqiang Song, Li Song, Kun Li, Dianhui Chu, Bolin Zhang
机构
*
Harbin Institute of Technology, Weihai(哈尔滨工业大学(威海))
;
Harbin Institute of Technology (Weihai) Qingdao Research Institute(哈尔滨工业大学(威海)青岛研究院)
;
Shandong Key Laboratory of Digital Service Computing Technology and Systems(山东省数字服务计算技术与系统重点实验室)
;
Weihai Municipal Hospital(威海市人民医院)
;
Shanghai Taizhu Technology Co., Ltd(上海泰山技术有限公司)
;
Tianjin Zhifu Qihuang Medical Technology Co., Ltd(天津中孚启黄医疗技术有限公司)
OGA-AID: Clinician-in-the-loop AI Report Drafting Assistant for Multimodal Observational Gait Analysis in Post-Stroke Rehabilitation
OGA-AID:用于中风后康复多模态观察性步态分析的临床医生在环AI报告起草助手
Khoi T. N. Nguyen, Nghia D. Nguyen, Hui Yu Koh, Patrick W. H. Kwong, Karen Sui Geok Chua, Ananda Sidarta, Baosheng Yu
机构
*
Rehabilitation Research Institute of Singapore, Nanyang Technological University, Singapore(新加坡康复研究中心,南洋理工大学,新加坡)
;
Lee Kong Chian School of Medicine, Nanyang Technological University, Singapore(李光前医学院,南洋理工大学,新加坡)
;
The Grainger College of Engineering, University of Illinois Urbana-Champaign, United States(伊利诺伊大学厄巴纳-香槟分校格雷格学院,美国)
;
Department of Rehabilitation Sciences, The Hong Kong Polytechnic University, Hong Kong(香港理工大学康复科学系,香港)
;
VinUni-Illinois Smart Health Center, VinUniversity, Vietnam(Vin大学Vin-伊利诺伊智能健康中心,越南)
;
Institute of Rehabilitation Excellence, Tan Tock Seng Hospital, NHG Health, Singapore(卓越康复研究所,坦托克桑格医院,NHG健康,新加坡)
Closed-Form Spectral Regularization for Multi-Task Model Merging
多任务模型融合的闭式谱正则化
Yongxian Wei, Runxi Cheng, Xingxuan Zhang, Li Shen, Chun Yuan, Peng Cui, Dacheng Tao
机构
*
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Sun Yat-sen University(中山大学)
;
Nanyang Technological University(南洋理工大学)
Comments11 pages, 9 figures. Accepted as a demo paper at ICAIL 2026. This arXiv version includes an appendix, new results, bug fixes, and presentation improvements beyond the earlier preprint; consequently, some reported numbers differ
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Beijing University of Chemical Technology(北京化工大学)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
;
Beijing Institute of Technology (Zhuhai)(北京理工大学(珠海))
;
Tencent Hy(腾讯(深圳))
;
Peng Cheng Laboratory(鹏城实验室)
Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents
面向视觉原生多模态深度搜索智能体的在策略数据演化
Shijue Huang, Hangyu Guo, Guanting Dong, Chenxin Li, Junting Lu, Xinyu Geng, Zhaochen Su, Zhenyu Li, Shuang Chen, Hongru Wang, Yi R. Fung
机构
*
Hong Kong University of Science and Technology(香港理工大学)
;
Renmin University of China(中国人民大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Peking University(北京大学)
;
Tsinghua University(清华大学)
;
University of Edinburgh(爱丁堡大学)
Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents
Agent规划基准:LLM Agent规划能力的诊断框架
Haoyu Sun, Wenxuan Wang, Mingyang Song, Jujie He, Weinan Zhang, Yang Liu, Yang Yang, Yu Cheng
机构
*
Tongji University(同济大学)
;
Shanghai AI Laboratory(上海人工智能实验室)
;
Harbin Institute of Technology(哈尔滨工业大学)
;
Fudan University(复旦大学)
;
Skywork AI
;
University of California, Santa Cruz(加州大学圣克鲁兹分校)
;
Shanghai Jiao Tong University(上海交通大学)
;
The Chinese University of Hong Kong(香港中文大学)
机构
*
School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院)
;
School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)
;
Independent Researcher(独立研究者)
Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy Optimization
教方法而非答案:用于多模态策略优化的特权辅导蒸馏
Shizhe Xiang, Ke An, Wenlong Yu, Yue Liu, Jian Luan, Pei Fu, Qilong Wang
机构
*
Tianjin University(天津大学)
;
Beijing Institute of Technology(北京理工大学)
;
Singapore Management University(新加坡国立大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Xiaomi Inc(小米公司)
机构
*
Division of Computational Pathology, Department of Pathology and Laboratory Medicine, Indiana University School of Medicine(计算病理学部,病理学与实验室医学部,印第安纳大学医学院)
;
IU Melvin and Bren Simon Comprehensive Cancer Center(印第安纳大学Melvin和Bren Simon综合癌症中心)
;
Departments of Biostatistics and Health Data Science(生物统计学与健康数据科学部)
;
Radiology and Imaging Sciences(放射学与影像科学部)
;
Neurological Surgery(神经外科)
;
Indiana University School of Medicine(印第安纳大学医学院)
;
Department of Computer Science, Luddy School of Informatics, Computing, and Engineering(计算机科学部,Luddy信息、计算与工程学院)
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference
MACS: 模态感知容量缩放用于高效多模态MoE推理
Bo Li, Chuan Wu, Shaolin Zhu
机构
*
School of Software, Tsinghua University, Beijing, China(清华大学软件学院,北京,中国)
;
TJUNLP Lab, School of Computer Science and Technology, Tianjin University, China(天津大学计算机科学与技术学院,中国)
;
School of New Media and Communication, Tianjin University, China(天津大学新媒体与传播学院,中国)