Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Visual-Seeker:通过主动视觉推理实现视觉原生多模态智能搜索
Zhengbo Zhang, Changtao Miao, Jinbo Su, Zhaowen Zhou, Chunxia Zhang, Xukai Wang, Ruiqi Liu, Kaiyuan Zheng, Jiansheng Cai, Bo Zhang, Zhe Li, Shiming Xiang, Ying Yan
机构
*
School of Artificial Intelligence UCAS(中国科学院大学人工智能学院)
;
Institute of Automation CAS(中国科学院自动化研究所)
;
Ant Digital Technologies Ant Group(蚂蚁数字科技蚂蚁集团)
;
RUC(中国人民大学)
;
BIT(北京理工大学)
机构
*
Princeton University(普林斯顿大学)
;
Springer Heidelberg(斯普林格海德堡)
;
ABC Institute(ABC研究所)
;
Rupert-Karls-University Heidelberg(海德堡鲁珀特-卡尔大学)
;
Hangzhou Dianzi University(杭州电子科技大学)
;
Zhejiang University(浙江大学)
;
Children’s Hospital, Zhejiang University School of Medicine, National Clinical Research Center for Children and Adolescents’ Health and Diseases(浙江大学医学院儿童医院,国家儿童青少年健康与疾病临床研究中心)
EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective
EvoMemBench: 从自演化视角评估智能体记忆
Yuyao Wang, Zhongjian Zhang, Mo Chi, Kaichi Yu, Yuhan Li, Miao Peng, Bing Tong, Chen Zhang, Yan Zhou, Jia Li
机构
*
Hong Kong University of Science and Technology (Guangzhou)(香港理工大学(广州))
;
Createlink Technology(创-link科技)
;
Beijing University of Posts and Telecommunications(北京邮电大学)
;
Beijing Institute of Technology(北京理工大学)
CommentsWe propose SCOUT, a detector allocation framework that predicts each detector's accuracy and latency on a given input before running it, letting operators control the safety-utility trade-off with a single threshold and route to an LLM judge only when needed
机构
*
The Hong Kong Polytechnic University(香港理工大学)
;
Nanyang Technological University(南洋理工大学)
;
Tsinghua University(清华大学)
;
Royal Melbourne Institute of Technology(皇家墨尔本理工大学)
;
Beijing University of Aeronautics and Astronautics(北京航空航天大学)
Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy
虚拟言语治疗师:一种临床医生参与的AI言语治疗代理,用于个性化和监督式治疗
Shakeel Sheikh, Patrick Marmaroli, MD Sahidullah, Slim Ouni, Fabrice Hirsch, Goncalo Leal, Bjorn W Schuller
机构
*
The Kashmir Hub for Artficial Intelligence(喀布尔人工智能中心)
;
Microsoft / Vocametrix(微软 / Vocametrix)
;
IAI, TCG CREST(IAI,TCG CREST)
;
Université de Lorraine, CNRS, Inria, LORIA(洛林大学,CNRS,Inria,LORIA)
;
Laboratoire Praxiling, UMR5267, CNRS et Université Paul-Valéry Montpellier 3(Praxiling实验室,UMR5267,CNRS及蒙彼利埃Paul-Valéry大学)
;
Speechcare iStutter, Portuguese Catholic University(Speechcare iStutter,葡萄牙天主教大学)
;
CHI – Chair of Health Informatics, TUM University Hospital(健康信息学系,TUM大学医院)
;
GLAM – Group on Language, Audio, & Music, Imperial College London(语言、音频与音乐小组,伦敦帝国理工学院)
CommentsThis paper has been accepted by ICML 2026. If you find our project helpful, please consider giving it a star: https://github.com/dog-last/E-mem
机构
*
MoE Key Lab of BIPC, University of Science and Technology of China(中科院大学科学技术大学MoE关键实验室)
;
Shanghai Innovation Institute(上海创新研究院)
;
Shanghai AI Laboratory(上海人工智能实验室)
Medical world models: representing medical states, modelling clinical dynamics and guiding intervention policies
医疗世界模型:表示医疗状态、建模临床动态与指导干预策略
Ke Liu, Mengxuan Li, Yanyi Bao, Tianyun Zhang, Chong Chu, Jiajun Bu, Haishuai Wang
机构
*
College of Computer Science, Zhejiang University(浙江大学计算机科学与技术学院)
;
School of Medicine, Zhejiang University(浙江大学医学院)
;
Department of Biomedical Informatics, Harvard University(哈佛大学生物医学信息学系)
Steering Emotional Dynamics for Art Therapy: Controllable Narrative Script Generation through Hierarchically Guided LLM Agents
引导艺术治疗的情感动态:通过分层引导的LLM智能体实现可控叙事脚本生成
Suqing Wang, Qinghai Miao, Chao Guo, Yisheng Lv
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)