EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
EvoLMM:具有连续奖励的自进化大型多模态模型
Omkar Thawakar, Shravan Venkatraman, Ritesh Thawkar, Abdelrahman Shaker, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, Fahad Khan
机构
*
Mohamed bin Zayed University of AI(Mohamed bin Zayed人工智能大学)
;
Aalto University(阿alto大学)
;
Australian National University(澳大利亚国立大学)
;
Linköping University(林肯大学)
Physics-informed generative AI for semiconductor manufacturing: Enforcing hard physical constraints in generative models by construction
物理信息驱动的生成式AI在半导体制造中的应用:通过构造强制生成模型中的硬物理约束
Yaser Mike Banad, Sarah Sharif
机构
*
School of Electrical and Computer Engineering, University of Oklahoma(俄克拉荷马大学电气与计算机工程学院)
;
Center for Quantum Research and Technology, University of Oklahoma(俄克拉荷马大学量子研究与技术中心)
;
Intelligent Neuromorphic and Quantum Understanding for Innovative Research and Engineering (INQUIRE) Laboratory(创新研究与工程智能神经形态与量子理解实验室)
;
Material Science and Engineering Program, University of Oklahoma, Norman, OK 73019 USA(俄克拉荷马大学材料科学与工程项目,Norman, OK 73019 USA)
专题命中
多模态生成
:multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI
机构
*
KTH Royal Institute of Technology(皇家理工学院)
;
Swiss Federal Institute of Technology Lausanne(洛桑联邦理工学院)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)
;
CISPA Helmholtz Center for Information Security(信息安全赫尔姆霍兹中心)
;
RISE Research Institutes of Sweden(瑞典RISE研究机构)
;
Halmstad University(哈马碧大学)
Cross-Modal Benchmarking for Robotic Perception in Natural Environments
自然环境中机器人感知的跨模态基准测试
David Hall, Joshua Knights, Mark Cox, Peyman Moghadam
机构
*
CSIRO Robotics, CSIRO, Australia(CSIRO机器人研究所,CSIRO,澳大利亚)
;
University of Sydney (USyd), Australia(悉尼大学(USyd),澳大利亚)
;
Queensland University of Technology (QUT), Australia(昆士兰理工大学(QUT),澳大利亚)
M4FC: a Multimodal, Multilingual, Multicultural, Multitask Real-World Fact-Checking Dataset
M4FC:一个多模态、多语言、多文化、多任务的真实世界事实验证数据集
Jiahui Geng, Jonathan Tonglet, Iryna Gurevych
机构
*
Mohamed bin Zayed University of Artificial Intelligence(Mohamed bin Zayed人工智能大学)
;
Ubiquitous Knowledge Processing Lab(ubiquitous知识处理实验室)
;
Department of Computer Science, TU Darmstadt(TU Darmstadt计算机科学系)
;
National Research Center for Applied Cybersecurity ATHENE(应用网络安全国家研究中心ATHENE)
;
Department of Electrical Engineering, KU Leuven(KU Leuven电气工程系)
;
Department of Computer Science, KU Leuven(KU Leuven计算机科学系)
Sustainability assessment using multimodal AI agents
使用多模态AI代理进行可持续性评估
Zhihan Zhang, Alexander Metzger, Yuxuan Mei, Felix Hähnlein, Zachary Englhardt, Tingyu Cheng, Gregory D. Abowd, Shwetak Patel, Adriana Schulz, Vikram Iyer
机构
*
Paul G. Allen School of Computer Science & Engineering, University of Washington(保罗·G·艾伦计算机科学与工程学院,华盛顿大学)
;
Computer Science and Engineering, University of Notre Dame(计算机科学与工程,诺丁汉大学)
;
Electrical and Computer Engineering, Northeastern University(电气与计算机工程,东北大学)
机构
*
South China University of Technology(华南理工大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
Peking University(北京大学)
;
University of Electronic Science and Technology of China(电子科技大学)
MedVeriSeg: Teaching LISA-Like Medical Segmentation Models to Verify Query Validity Without Extra Training
MedVeriSeg: 教授LISA-like医学分割模型验证查询的有效性而无需额外训练
Qinyue Tong, Xiaozhen Wang, Ziqian Lu, Jun Liu, Yunlong Yu, Zheming Lu
机构
*
School of Aeronautics and Astronautics, Zhejiang University(浙江大学航空宇航学院)
;
Southern Medical University(南方医科大学)
;
School of Computer Science and Technology (School of Artificial Intelligence), Zhejiang Sci-Tech University(浙江科技学院计算机科学与技术学院(人工智能学院))
;
College of Information Science and Electronic Engineering, Zhejiang University(浙江大学信息科学与电子工程学院)
OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models
OpenMedReason: 医学视觉语言模型的科学推理监督
Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
机构
*
York University(约克大学)
;
Vector Institute(向量研究所)
;
University of British Columbia(不列颠哥伦比亚大学)
;
University of Toronto(多伦多大学)
;
Unity Health Toronto / St. Michael’s Hospital(多伦多联合健康/圣迈克尔医院)
;
University Health Network(大学健康网络)
;
Arc Institute(弧研究所)
;
Queen's University(女王大学)
Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends
具身基准构建的智能自动化:流程、具身、模拟器与趋势
Jinshan Lai, Jianwei Hu, Baoyang Jiang, Fengchun Zhang, Leyuan Wang, Haotian Li, Yida Wang, Tingxuan Huang, Xi Ren, Qiang Ma
机构
*
University of Electronic Science and Technology of China(电子科技大学)
;
Qiyuan Lab(启元实验室)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Tsinghua University(清华大学)
;
Beihang University(北京航空航天大学)
RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark
RAIL: 基于CHC框架重新思考大型音频语言模型中的听觉智能
Hongyu Jin, Siyi Wang, Yang Xiao, Jiaheng Dong, Shihong Tan, Kaiyuan peng, Georgiana Juravle, Shanquan Chen, Gongping Huang, Hong Jia, Eun-Jung Holden, James Bailey, Ting Dang
机构
*
School of Computing and Information Systems, The University of Melbourne(墨尔本大学计算与信息系统学院)
;
Faculty of Psychology and Educational Sciences, Alexandru Ioan Cuza University of Iași(亚历山德鲁伊万库扎大学心理学与教育科学学院)
;
School of Electronic Information, Wuhan University(武汉大学电子信息学院)
;
School of Public Health, The University of Hong Kong(香港大学公共卫生学院)
;
School of Computer Science, The University of Auckland(奥克兰大学计算机科学学院)
;
Department of Data Science and Artificial Intelligence, Monash University(莫纳什大学数据科学与人工智能系)
MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning
MODF-SIR:面向社交智能推理的多智能体全模态蒸馏框架
Shang Ma, Jisheng Dang, Wencan Zhang, Yifan Zhang, Bimei Wang, Hong Peng, Bin Hu, Qi Tian, Tat-Seng Chua
机构
*
School of Information Science and Engineering, Lanzhou University(兰州大学信息科学与工程学院)
;
School of Medical Technology, Beijing Institute of Technology(北京理工大学医学技术学院)
;
Cloud and AI BU, Huawei(华为云与AI业务部)
;
School of Computing, National University of Singapore(新加坡国立大学计算机学院)