Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning
通过定制代理推理增强VLMs在少样本多模态时间序列分类中的能力
Lin Li, Jiawei Huang, Qihao Quan, Dan Li, Boxin Li, Xiao Zhang, Erli Meng, Wenjie Feng, Jian Lou, See-Kiong Ng
机构
*
Sun Yat-sen University(中山大学)
;
Xiaomi Corporation(小米公司)
;
University of Science and Technology of China(中国科学技术大学)
;
National University of Singapore(新加坡国立大学)
CMTA: Leveraging Cross-Modal Temporal Artifacts for Generalizable AI-Generated Video Detection
CMTA:利用跨模态时间特征进行通用的AI生成视频检测
Hang Wang, Chao Shen, Chenhao Lin, Minghui Yang, Lei Zhang, Cong Wang
机构
*
Xi’an Jiaotong University(西安交通大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
Guangdong OPPO Mobile Communications Co., Ltd.(广东OPPO移动通信有限公司)
;
City University of Hong Kong(香港城市大学)
CommentsThis paper has been withdrawn by the author. After further review, the author believes that the current version does not meet the desired standards and plans to revise the work before any potential resubmission
A multimodal and temporal foundation model for virtual patient representations at healthcare system scale
面向医疗系统规模的多模态与时间基础模型:虚拟患者表示
Andrew Zhang, Tong Ding, Sophia J. Wagner, Caiwei Tian, Ming Y. Lu, Rowland Pettit, Joshua E. Lewis, Alexandre Misrahi, Dandan Mo, Long Phi Le, Faisal Mahmood
机构
*
Department of Pathology, Mass General Brigham, Harvard Medical School(病理学系,马萨诸塞州总医院与哈佛医学院)
;
Cancer Program, Broad Institute of Harvard and MIT(癌症计划,哈佛-麻省理工Broad研究所)
;
Data Science Program, Dana-Farber Cancer Institute(数据科学计划,达纳-法伯癌症研究所)
;
Health Sciences and Technology, Harvard-MIT(健康科学与技术,哈佛-麻省理工)
;
Harvard John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛约翰·A·保罗森工程与应用科学学院,哈佛大学)
;
Department of Biomedical Informatics, Harvard Medical School(生物医学信息学系,哈佛医学院)
;
Electrical Engineering and Computer Science, Massachusetts Institute of Technology (MIT)(电气工程与计算机科学,麻省理工学院(MIT))
;
School of Computer and Communication Sciences, EPFL, Lausanne, Switzerland(计算机与通信科学学院,EPFL,瑞士洛桑)
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
通过视觉记忆机制扩展多模态大语言模型的长视频理解
Tao Chen, Kun Zhang, Qiong Wu, Xiao Chen, Chao Chang, Xiaoshuai Sun, Yiyi Zhou, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University(厦门大学多媒体可信感知与高效计算教育部重点实验室)
;
National University of Defense Technology(国防科技大学)
Cross-Modal Transferable Image-to-Video Attack on Video Quality Metrics
跨模态可迁移的图像到视频攻击视频质量度量
Georgii Gotin, Ekaterina Shumitskaya, Anastasia Antsiferova, Dmitriy Vatolin
机构
*
Lomonosov Moscow State University(罗蒙诺索夫莫斯科国立大学)
;
ISP RAS Research Center for Trusted Artificial Intelligence(俄罗斯科学院信息与系统研究所在可信人工智能研究中心)
;
MSU Institute for Artificial Intelligence(莫斯科大学人工智能研究所)
;
Laboratory of Innovative Technologies for Processing Video Content(视频内容处理创新技术实验室)
Multimodal Laryngoscopic Video Analysis for Assisted Diagnosis of Vocal Fold Paralysis
多模态喉镜视频分析用于辅助声带麻痹诊断
Yucong Zhang, Xin Zou, Jinshan Yang, Wenjun Chen, Juan Liu, Faya Liang, Ming Li
机构
*
School of Computer Science(计算机科学学院)
;
Suzhou Municipal Key Laboratory of Multimodal Intelligent Systems(多模态智能系统苏州市重点实验室)
;
Sun Yat-sen Memorial Hospital of Sun Yat-sen University(中山大学孙逸仙纪念医院)
;
School of Artificial Intelligent(人工智能学院)
;
Department of Otorhinolaryngology(耳鼻喉科部门)
C^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models Reasoning
C^2ROPE: 3D 大多模态模型推理中的因果连续旋转位置编码
Guanting Ye, Qiyan Zhao, Wenhao Yu, Xiaofeng Zhang, Jianmin Ji, Yanyong Zhang, Ka-Veng Yuen
机构
*
State Key Laboratory of Internet of Things for Smart City, University of Macau(物联网智能城市国家重点实验室,澳门大学)
;
Department of Automation, Shanghai Jiaotong University(上海交通大学自动化系)
;
Institute of Advanced Technology, University of Science and Technology of China(中国科学技术大学先进技术研究院)
;
School of Computer Science and Technology, USTC(中国科学技术大学计算机科学与技术学院)
;
School of Artificial Intelligence and Data Science, USTC(中国科学技术大学人工智能与数据科学学院)
Simulating the Real World: A Unified Survey of Multimodal Generative Models
模拟现实世界:多模态生成模型的统一综述
Yuqi Hu, Longguang Wang, Xian Liu, Ling-Hao Chen, Yuwei Guo, Yukai Shi, Ce Liu, Anyi Rao, Zeyu Wang, Hui Xiong
机构
*
Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)(人工智能前沿技术研究所,香港科学与技术大学(广州))
;
Department of Computer Science and Engineering, The Hong Kong University of Science and Technology Hong Kong SAR(计算机科学与工程系,香港科学与技术大学香港特别行政区)
;
MMLab, The Hong Kong University of Science and Technology(多模态实验室,香港科学与技术大学)
;
School of Electronics and Communication Engineering, Shenzhen Campus of Sun Yat-sen University(电子与通信工程学院,中山大学深圳校区)
;
The Chinese University of Hong Kong, Hong Kong, China(香港中文大学,香港,中国)
;
Tsinghua University, Guangdong, China(清华大学,广东,中国)
;
Bosch (China) Investment Co., Ltd., Shanghai, China(博世(中国)投资有限公司,上海,中国)