Stanford's Fei-Fei Li team open-sources OpenWAM world-action framework
A Stanford team including Fei-Fei Li, Jiajun Wu and Ehsan Adeli published OpenWAM, an open framework for composable world-action models. It starts from Alibaba's open Wan2.2-5B video model, continues pretraining on roughly 3.34 million robot and human-interaction videos with a causal, future-blind setup, then joins a 5B video expert and a 2B action expert via a shared Mixture-of-Transformers architecture with four configurable video-action interaction orders (VTA, ATV, Joint, Decoupled). The team also built LIBERO-Long-CF, a counterfactual dataset of about 32,000 clips and 4.1 million control records, and trained transferable local inverse- and forward-dynamics components; they report a frozen local IDM reaching 84.0% average success on four new LIBERO-90 tasks versus 47.0% for a full-context IDM and 21.5% for demonstration-only local IDM, plus 92.1% and 91.9% average success for VTA and Joint on real dual-arm Franka FR3 tasks. Training, evaluation and composition code and pretrained video-model weights have been open-sourced; the piece also notes Black Forest Labs' FLUX 3 Action world-action model.