Logit-Level Uncertainty Quantification in Vision-Language Models for Histopathology Image Analysis
视觉-语言模型在病理图像分析中的logit级不确定性量化
Betul Yurdem, Ferhat Ozgur Catak, Murat Kuzlu, Mehmet Kemal Gullu
机构
*
Department of Electrical and Electronics Engineering, Izmir Bakircay University(电气与电子工程系,伊兹密尔巴克伊大学)
;
Department of Electrical Engineering and Computer Science, University of Stavanger(电气工程与计算机科学系,斯塔万格大学)
;
Batten College of Engineering and Technology, Old Dominion University(工程与技术学院,老奥布良大学)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
Crab$^{+}$: 一种可扩展且统一的音频视觉场景理解模型,具有显式合作
Dongnuan Cai, Henghui Du, Chang Zhou, Xi Chen, Dan Guo, Hongyuan Zhang, Xuelong Li, Di Hu
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学耿丽人工智能学院)
;
Institute of Artificial Intelligence of China Telecom (TeleAI)(中国电信人工智能研究院)
;
AI Technology Center, Online Video Business Unit, Tencent PCG(腾讯PCG在线视频业务单元AI技术中心)
;
Hefei University of Technology(合肥工业大学)
;
The University of Hong Kong(香港大学)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
鲁棒的基于大语言模型的音频视觉语音识别与稀疏模态对齐和视觉单元引导的细化
Fei Su, Cancan Li, Juan Liu, Wei Ju, Hongbin Suo, Ming Li
机构
*
School of Computer Science, Wuhan University, China(武汉大学计算机学院)
;
School of Artificial Intelligence, Wuhan University, China(武汉大学人工智能学院)
;
School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)人工智能学院)
;
AI Center, OPPO, China(OPPO人工智能中心)
;
Digital Innovation Research Center, Duke Kunshan University, China(杜克大学昆山数字创新研究中心)
Sim2Sea: Sim-to-Real Policy Transfer for Maritime Vessel Navigation in Congested Waters
Sim2Sea: 用于拥挤水域船舶导航的仿真到现实策略迁移
Xinyu Cui, Xuanfa Jin, Xue Yan, Yongcheng Zeng, Luoyang Sun, Siying Wei, Ruizhi Zhang, Jian Zhao, Haifeng Zhang, Jun Wang
机构
*
Institute of Automation, CAS(中国科学院自动化研究所)
;
School of Artificial Intelligence, UCAS(中国科学院大学人工智能学院)
;
Zhongguancun Academy(中关村学院)
;
University College London(伦敦大学学院)
RIVER: A Real-Time Interaction Benchmark for Video LLMs
RIVER:面向视频大语言模型的实时交互基准
Yansong Shi, Qingsong Zhao, Tianxiang Jiang, Xiangyu Zeng, Yi Wang, Limin Wang
机构
*
School of Information Science And Technology, University of Science and Technology of China(信息科学与技术学院,中国科学技术大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
College of Computer Science and Artificial Intelligence, Fudan University(计算机科学与人工智能学院,复旦大学)
;
State Key Lab of Novel Software Technology, Nanjing University(新型软件技术国家重点实验室,南京大学)
Segment-to-Act: Label-Noise-Robust Action-Prompted Video Segmentation Towards Embodied Intelligence
基于动作的视频分割:面向具身智能的标签噪声鲁棒动作引导分割
Wenxin Li, Kunyu Peng, Di Wen, Ruiping Liu, Mengfei Duan, Kai Luo, Kailun Yang
机构
*
School of Artificial Intelligence and Robotics and the National Engineering Research Center of Robot Visual Perception and Control Technology, Hunan University, China(人工智能与机器人学院及机器人视觉感知与控制技术国家工程研究中心,湖南大学,中国)
;
Institute for Anthropomatics and Robotics, Karlsruhe Institute of Technology, Germany(人机学与机器人研究所,卡尔斯鲁厄理工学院,德国)
机构
*
School of Biomedical Engineering, Tsinghua University(清华大学生物医学工程学院)
;
School of Biomedical Engineering, Shanghai Jiao Tong University(上海交通大学生物医学工程学院)
;
Longwood Valley MedTech
机构
*
College of Command and Control Engineering, Army Engineering University of PLA(中国人民解放军陆军工程大学指挥控制工程学院)
;
School of Future Technology, Dalian University of Technology(大连理工大学未来技术学院)
;
School of Information and Communication Engineering, Dalian University of Technology(大连理工大学信息与通信工程学院)
InEdit-Bench: Benchmarking Intermediate Logical Pathways for Intelligent Image Editing Models
InEdit-Bench:智能图像编辑模型中间逻辑路径的基准测试
Zhiqiang Sheng, Xumeng Han, Zhiwei Zhang, Zenghui Xiong, Yifan Ding, Aoxiang Ping, Xiang Li, Tong Guo, Yao Mao
机构
*
State Key Laboratory of Optical Field Manipulation Science and Technology, Institute of Optics and Electronics, Chinese Academy of Sciences(光学场操控科学与技术国家重点实验室,光学电子研究所,中国科学院)
;
National Laboratory on Adaptive Optics(自适应光学国家实验室)
;
University of Chinese Academy of Sciences(中国科学院大学)
Factuality Matters: When Image Generation and Editing Meet Structured Visuals
事实性至关重要:当图像生成与编辑遇见结构化视觉
Le Zhuo, Songhao Han, Yuandong Pu, Boxiang Qiu, Sayak Paul, Yue Liao, Yihao Liu, Jie Shao, Xi Chen, Si Liu, Hongsheng Li
机构
*
CUHK MMLab(香港大学多模态实验室)
;
Beihang University(北京航空航天大学)
;
Krea AI(Krea人工智能)
;
Shanghai Jiao Tong University(上海交通大学)
;
Shanghai AI Lab(上海人工智能实验室)
;
Hugging Face
;
National University of Singapore(新加坡国立大学)
;
ByteDance(字节跳动)
;
The University of Hong Kong(香港大学)