Any4D: Open-Prompt 4D Generation from Natural Language and Images
Any4D: 从自然语言和图像生成开放提示的4D生成
Hao Li, Qiao Sun
专题命中
复杂问题求解
:reasoning(abstract);分类 cs.AI
AI总结
本文提出Primitive Embodied World Models,通过限制视频生成时间范围,实现语言与视觉表示的细粒度对齐,降低学习复杂度,提升数据效率,并减少推理延迟,支持复杂任务的组合泛化。
CommentsThe authors identified issues in the 4D generation pipeline and evaluation that affect result validity. To ensure scientific accuracy, we will revise the methodology and experiments thoroughly before resubmitting. This version should not be cited or relied upon
Journal ref"On the Limits of LLM Reasoning: Evidence From Contamination, Translation, and Answer Modification in Multiple-Choice Benchmarks," in IEEE Access, vol. 14, pp. 9384-9393, 2026
Evidence-based diagnostic reasoning with multi-agent copilot for human pathology
基于多智能体助手的证据驱动诊断推理
Luca L. Weishaupt, Chengkuan Chen, Drew F. K. Williamson, Richard J. Chen, Guillaume Jaume, Tong Ding, Bowen Chen, Anurag Vaidya, Long Phi Le, Guillaume Jaume, Ming Y. Lu, Faisal Mahmood
机构
*
Health Sciences and Technology, Harvard-MIT(哈佛-MIT健康科学与技术)
;
Department of Pathology, Massachusetts General Hospital, Harvard Medical School(麻省总医院病理科,哈佛医学院)
;
Cancer Program, Broad Institute of Harvard and MIT(哈佛-MIT博德研究所癌症项目)
;
Harvard John A. Paulson School of Engineering and Applied Sciences, Harvard University(哈佛大学约翰·A·保尔森工程与应用科学学院)
;
Electrical Engineering and Computer Science, Massachusetts Institute of Technology (MIT)(麻省理工学院电气工程与计算机科学)
;
Harvard Data Science Initiative, Harvard University(哈佛大学数据科学计划)
MA-Bench: Towards Fine-grained Micro-Action Understanding
MA-Bench:迈向细粒度微动作理解
Kun Li, Jihao Gu, Fei Wang, Zhiliang Wu, Hehe Fan, Dan Guo
机构
*
CVLab, College of Information Technology, United Arab Emirates University(阿拉伯联合酋长国大学信息技术学院CVLab)
;
University College London(伦敦大学学院)
;
Hefei University of Technology(合肥工业大学)
;
Institute of Artificial Intelligence, Hefei Comprehensive National Science Center(合肥综合性国家科学中心人工智能研究院)
;
CCAI, Zhejiang University(浙江大学计算机辅助设计与图形学国家重点实验室)
CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models
CoVFT:面向多模态大语言模型的上下文感知视觉微调
Nan Zhou, Huiqun Wang, Yaoyan Zheng, Di Huang
机构
*
State Key Laboratory of Complex and Critical Software Environment, Beihang University(北京航空航天大学复杂关键软件环境国家重点实验室)
;
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
Reframing Human-Robot Interaction Through Extended Reality: Unlocking Safer, Smarter, and More Empathic Interactions with Virtual Robots and Foundation Models
通过扩展现实重新定义人机交互:利用虚拟机器人和基础模型实现更安全、更智能和更具同理心的交互
Yuchong Zhang, Yong Ma, Danica Kragic
机构
*
Division of Robotics, Perception, and Learning, KTH Royal Institute of Technology(KTH皇家理工学院机器人、感知与学习部)
;
Department of Clinical Medicine, University of Bergen(卑尔根大学临床医学系)
PepThink-R1: LLM for Interpretable Cyclic Peptide Optimization with CoT SFT and Reinforcement Learning
PepThink-R1:基于CoT SFT和强化学习的可解释环状肽优化LLM
Ruheng Wang, Hang Zhang, Trieu Nguyen, Shasha Feng, Hao-Wei Pang, Xiang Yu, Li Xiao, Peter Zhiping Zhang
机构
*
Merck & Co., Inc., Rahway, NJ, USA(默克公司,美国新泽西州拉威)
;
UT Southwestern Medical Center, Dallas, TX, USA(得克萨斯大学西南医学中心,美国得克萨斯州达拉斯)
;
University of Pittsburgh, Pittsburgh, PA, USA(匹兹堡大学,美国宾夕法尼亚州匹兹堡)
Comments10 pages, 14 figures, conference (FORGE '26: 2026 IEEE/ACM Third International Conference on AI Foundation Models and Software Engineering, Rio de Janeiro, Brazil, April 2026)