TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
机构 * Korea University(韩国大学)
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments AAAI 2026 (Main Technical Track)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Korea University(韩国大学)
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments AAAI 2026 (Main Technical Track)
机构 * School of Computer Science, Northwestern Polytechnical University(西北工业大学计算机学院)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
机构 * Shandong University(山东大学)
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV
机构 * Columbia University(哥伦比亚大学)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.AI
Comments More videos can be found on our website:https://binghao-huang.github.io/touch_in_the_wild/
机构 * The Pennsylvania State University(宾夕法尼亚州立大学) ; Intel(英特尔) ; NVIDIA(英伟达)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
机构 * School of Software, Shandong University(山东大学软件学院) ; Shenzhen Loop Area Institute(深圳河套学院) ; School of Computing and Artificial Intelligence, Shandong University of Finance and Economics(山东财经大学计算机与人工智能学院)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025
机构 * Radar Technology Research Institute, School of Information and Electronics, Beijing Institute of Technology(雷达技术研究院,信息电子学院,北京理工大学) ; School of Computer Science and Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) ; Innovative Equipment Research Institute, Beijing Institute of Technology(创新装备研究院,北京理工大学) ; Key Laboratory of Electronic and Information Technology in Satellite Navigation (Beijing Institute of Technology), Ministry of Education(卫星导航电子信息技术重点实验室(北京理工大学),教育部) ; Beijing Racobit Electronic Information Technology Co., Ltd.(北京瑞科比特电子信息技术有限公司)
专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments Submitted to Pattern Recognition
机构 * Xi’an-Jiaotong Liverpool University(西交利物浦大学) ; University of Liverpool(利物浦大学) ; Duke Kunshan University(杜克昆山大学)
专题命中 多模态训练与对齐 :multi-modal(abstract);MLLM(abstract);分类 cs.AI
机构 * S-Lab, Nanyang Technological University(南洋理工大学S实验室) ; Shanghai AI Laboratory(上海人工智能实验室) ; Sensetime Research(商汤科技研究院) ; Zhejiang University of Technology(浙江工业大学) ; UCAS-Terminus AI Lab,University of Chinese Academy of Sciences(中国科学院大学Terminus AI实验室)
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments ICCV 2025
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.MM
Comments ACM Multimedia 2025 Accepted
Journal ref In Proceedings of the 33st ACM International Conference on Multimedia (MM '25), 2025
机构 * School of Information Engineering, Guangdong University of Technology(广东技术大学信息工程学院) ; TikTok, ByteDance Inc(字节跳动) ; School of Computer Science, Anhui University(安徽大学计算机科学学院) ; School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments For the first time, angle-based perception was introduced into the multi-modality image fusion task
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; Shanghai Jiao Tong University(上海交通大学)
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025 main
机构 * University of Southern California(美国南加州大学) ; Emory University(埃默里大学)
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV
机构 * Beijing Institute of Technology(北京理工大学) ; Basic Algorithm Center, PCG, Tencent(腾讯基本算法中心)
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV
机构 * Department of Biomedical Informatics, Stony Brook University, NY, USA(生物医学信息学系,石溪大学,纽约,美国) ; Department of Computer Science, Stony Brook University, NY, USA(计算机科学系,石溪大学,纽约,美国)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
机构 * University of Maryland, College Park(马里兰大学学院公园分校) ; CUHK MMLab(香港大学MMLab) ; ByteDance(字节跳动)
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Project page: https://hywang66.github.io/bridge/
机构 * College of Computer Science and Technology(计算机科学与技术学院) ; Jilin University(吉林大学) ; College of Data Science(数据科学学院) ; Taiyuan University of Technology(太原理工大学) ; Department of Computer Science(计算机科学系)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
机构 * University of Amsterdam(阿姆斯特丹大学) ; ETH Zürich(苏黎世联邦理工学院) ; University of Copenhagen(哥本哈根大学)
专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.AI
Comments ICML 2025 Assessing World Models Workshop; EMNLP 2025 Findings
机构 * School of Earth and Space Sciences, Peking University(地球与空间科学学院,北京大学) ; College of Urban and Environmental Sciences, Peking University(城市与环境科学学院,北京大学) ; State Key Laboratory of Multimodal Artificial Intelligence Systems Institute of Automation, CAS(多模态人工智能系统国家重点实验室,中国科学院自动化研究所) ; Intelligent Maintenance and Operations Systems Lab(智能维护与运营系统实验室)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments NeurIPS 2025
机构 * School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院) ; Jiangnan University(江南大学) ; Centre for Vision, Speech and Signal Processing(视觉、语音与信号处理中心) ; University of Surrey(Surrey大学)
专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments 23 pages, 17 figures
机构 * College of Electronics and Information Engineering, Sichuan University(四川大学电子信息工程学院) ; School of Mechanical Engineering and Automation, Fuzhou University(福州大学机械工程与自动化学院) ; Yunnan Key Laboratory of Software Engineering, Yunnan University(云南软件工程重点实验室)
专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
机构 * Yale University, USA(耶鲁大学) ; Nanyang Technological University, Singapore(南洋理工大学) ; Dalian University of Technology, China(大连理工大学)
专题命中 多模态训练与对齐 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025
机构 * School of Electronics and Communication Engineering(电子与通信工程学院) ; Guangdong Provincial Key Laboratory of Advanced IntelliSense Technology(广东省先进智能感知技术重点实验室) ; Department of Radiology(放射科)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
机构 * School of Biomedical Engineering, Division of Life Sciences ; Medicine, University of Science ; Technology of China (USTC), Hefei Anhui, 230026, China Center for Medical Imaging, Robotics, Analytic Computing \& Learning (MIRACLE), Suzhou Institute for Advance Research, USTC, 215123, China Stanford University, Palo Alto, California, 94025, United States Jiangsu Provincial Key Laboratory of Multimodal Digital Twin Technology, Suzhou Key Laboratory of Precision ; Intelligent Chemistry, USTC Anhui IFLYTEK CO., Ltd.
专题命中 多模态训练与对齐 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
Comments Accepted by MICCAI 2025
机构 * College of Computing and Data Science, Nanyang Technological University(computing and Data Science学院,南洋理工大学) ; School of Electrical and Electronic Engineering, Nanyang Technological University(Electrical and Electronic Engineering学院,南洋理工大学) ; School of Artificial Intelligence, Wuhan University(Artificial Intelligence学院,武汉大学) ; ByteDance(字节跳动)
专题命中 多模态训练与对齐 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Extended version of our conference paper arXiv:2410.09855
机构 * University of Science and Technology of China(中国科学技术大学) ; Peking University(北京大学) ; Kuaishou Technology(快手科技)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
机构 * University of Science, VNU-HCM(越南胡志明市科学大学) ; Thong Nhat Hospital(通纳特医院)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments ACM Multimedia 2025
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments Preprint,Underreview
机构 * National Engineering Laboratory for Integrated Aero-Space-Ground-Ocean Big Data Application Technology(集成空天地海大数据应用技术国家工程实验室) ; Northwestern Polytechnical University(西北工业大学) ; Huiying Medical Technology Company Ltd.(慧影医疗技术有限公司) ; The School of Computer Science(计算机学院) ; The University of Sydney(悉尼大学) ; Department of Computer Science and Engineering(计算机科学与工程系) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
专题命中 多模态训练与对齐 :multimodal(abstract);cross-modal(abstract);分类 cs.AI
Comments 21 pages, 25 figures