CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding
CausalEmbed: 在潜在空间中实现自动回归多向量生成用于视觉文档嵌入
Jiahao Huo, Yu Huang, Yibo Yan, Ye Pan, Kening Zheng, Wei-Chieh Huang, Yi Cao, Mingdong Ou, Philip S. Yu, Xuming Hu
机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
Alibaba Cloud Computing(阿里云计算)
;
University of Illinois Chicago(伊利诺伊大学芝加哥分校)
M3D-Net: Multi-Modal 3D Facial Feature Reconstruction Network for Deepfake Detection
M3D-Net:面向深度伪造检测的多模态3D面部特征重建网络
Haotian Wu, Yue Cheng, Shan Bian
机构
*
College of Mathematics and Informatics(数学与信息学院)
;
South China Agricultural University(华南农业大学)
;
Wushan Road, Tianhe District(天河南路)
;
Guangzhou(广州)
;
China(中国)
机构
*
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Columbia University(哥伦比亚大学)
;
Massachusetts Institute of Technology(麻省理工学院)
;
Harvard University(哈佛大学)
机构
*
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, School of Information Science and Technology, University of Science and Technology of China(脑启发智能感知与认知MoE实验室,信息科学与技术学院,中国科学技术大学)
;
iFlytek Research, iFlytek Co., Ltd., Hefei, China(科大讯飞研究院,科大讯飞股份有限公司,合肥,中国)
CommentsA benchmark for evaluating multimodal both voice and text LLM agents in dualcontrol settings. We introduce persona adaptive prompting and 12 new metrics to assess robustness safety efficiency and recovery in customer support scenarios
CAVERS: Multimodal SLAM Data from a Natural Karstic Cave with Ground Truth Motion Capture
CAVERS:从自然喀斯特洞穴获取多模态SLAM数据并采用地面真实运动捕捉
Giacomo Franchini, David Rodríguez-Martínez, Alfonso Martínez-Petersen, C. J. Pérez-del-Pulgar, Marcello Chiaberge
机构
*
Polytechnic of Turin Interdepartmental Centre for Service Robotics (PIC4SeR)(都灵理工大学服务机器人跨部门研究中心(PIC4SeR))
;
Systems Engineering and Automation Department, Universidad de Málaga(马德里大学系统工程与自动化系)
Time-RA: Towards Time Series Reasoning for Anomaly Diagnosis with LLM Feedback
Time-RA:面向时间序列推理的异常诊断方法:基于LLM反馈
Yiyuan Yang, Zichuan Liu, Lei Song, Kai Ying, Zhiguang Wang, Tom Bamford, Svitlana Vyetrenko, Jiang Bian, Qingsong Wen
机构
*
University of Oxford(牛津大学)
;
Nanjing University(南京大学)
;
MSRA(微软研究院)
;
SJTU(上海交通大学)
;
Abel AI
;
Outsampler
;
University of Strasbourg(斯特拉斯堡大学)
;
Squirrel Ai Learning(Squirrel AI学习)
机构
*
Nanjing University(南京大学)
;
M-A-P
;
Jiutian Research(九天研究)
;
National University of Singapore(新加坡国立大学)
;
Nanjing University of Science and Technology(南京理工大学)
An Edge-Cloud Collaborative Architecture for Proactive Elderly Care: Real-Time Risk Assessment and Three-Level Emergency Response
面向主动老年护理的边缘-云协作架构:实时风险评估与三级应急响应
Lijie Zhou, Luran Wang
机构
*
School of Computer Science University of Nottingham Ningbo China Ningbo, China(计算机科学学院 南阳大学宁波分校 中国 宁波)
;
School of Mathematical Sciences University of Nottingham Ningbo China Ningbo, China(数学科学学院 南阳大学宁波分校 中国 宁波)
Geoparsing: Diagram Parsing for Plane and Solid Geometry with a Unified Formal Language
几何解析:一种统一形式语言用于平面和立体几何的图表解析
Peijie Wang, Ming-Liang Zhang, Jun Cao, Chao Deng, Dekang Ran, Hongda Sun, Pi Bu, Xuan Zhang, Yingyao Wang, Jun Song, Bo Zheng, Fei Yin, Cheng-Lin Liu
机构
*
MAIS, Institute of Automation of Chinese Academy of Sciences(中国科学院自动化研究所MAIS)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Future Living Lab of Alibaba(阿里巴巴未来生活实验室)