When One Modality Is Not Enough: Multimodal Sex and Life-Stage Classification of Red Deer from Aerial RGB-Thermal Video
当单模态不够时:基于航拍RGB-热红外视频的马鹿多模态性别与生命阶段分类
Hugo Markoff, Christoph Praschl, Ivan Ludoški, Sara Beery, Michael Ørsted, David C. Schedl
机构
*
Aalborg University(奥尔堡大学)
;
University of Applied Sciences Upper Austria(上奥地利应用科学大学)
;
University of Novi Sad(诺威萨德大学)
;
Massachusetts Institute of Technology(麻省理工学院)
I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning
我在视频中寻找你:面向以人为中心的视频推理的身份条件查询
Shibo Gao, Chongxiao Wang, Chenglong Huang, Jie Ma, Haolin Shi, Fei Ding, Jing Li, Qiang Lyu, Yangyang Liu, Yang Liu, Jun Liu, Linlin Huang, Peipei Yang
机构
*
Beijing Jiaotong University(北京交通大学)
;
HUJING Digital Media & Entertainment Group(汇晶数字媒体与娱乐集团)
;
MAIS Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所MAIS)
机构
*
School of Artificial Intelligence, Beihang University(北京航空航天大学人工智能学院)
;
School of Humanities and Social Sciences, Beihang University(北京航空航天大学人文与社会科学高等研究院)
;
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
;
School of Informatics, Xiamen University(厦门大学信息学院)
CommentsDocMemo is a memory-guided framework for long-document reasoning that uses tri-level memory and dynamic Bayesian belief updating to overcome static retrieval limits and improve evidence tracking. 16 pages, 4 figures, 14 tables
Comments18 pages, 4 figures, 2 tables. Published in the Proceedings of ASCAAD 2025
Journal refProceedings of the 13th International Conference of the Arab Society for Computation in Architecture, Art and Design (ASCAAD 2025), Riyadh, Saudi Arabia, 2025
Walkable to Whom? Capturing Subjective Variability in Walkability Perception Using Multimodal Deep Learning
对谁而言是可步行的?使用多模态深度学习捕捉步行感知中的主观变异性
Moloud Damandeh, Meead Saberi
机构
*
School of Civil and Environmental Engineering, University of New South Wales (UNSW)(新南威尔士大学土木与环境工程学院)
;
Research Centre for Integrated Transport Innovation (rCITI)(综合交通创新研究中心)
GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models
GraphVerse:面向多模态大语言模型的综合性视觉图推理基准
Yuanfu Sun, Yuanhang Ren, Kang Li, Chuanhao Ji, Jiaxi Li, Jiajin Liu, Ninghao Liu, Qiaoyu Tan
机构
*
New York University(纽约大学)
;
Sensetime Research(商汤科技研究院)
;
Tsinghua University(清华大学)
;
New York University Shanghai(上海纽约大学)
;
University of Georgia(佐治亚大学)
;
The Hong Kong Polytechnic University(香港理工大学)
机构
*
Alibaba Group(阿里巴巴集团)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tsinghua University(清华大学)
;
University of Alberta(阿尔伯塔大学)
;
Zhejiang University(浙江大学)