{"id":"294affed-c8f5-420c-9811-9ad968290493","slug":"expo-ft","name":"EXPO-FT","description":"EXPO-FT is a Stanford research system (Dong, Hung, Gao, Sadigh, Finn) for sample-efficient reinforcement learning finetuning of pretrained vision-language-action policies. It pairs a supervised VLA base policy with a lightweight edit policy and Q-guided on-the-fly action selection, optionally with human-in-the-loop corrections. Paper-claimed results reach 30/30 on evaluated real-robot tasks in ~19 minutes of online interaction on average. Open-source codebase released with the paper.","brainType":"research-model","isOpen":true,"maturityStage":"research","architecture":"EXPO-style base VLA + bounded edit policy + RedQ-style Q ensemble; on-the-fly max-Q selection over base and edited action chunks; human-in-the-loop interventions during online RL","reviewStatus":"reviewed","sources":[{"url":"https://arxiv.org/abs/2605.25477","title":"EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models","sourceName":"arXiv","publishedAt":"2026-05-25"},{"url":"https://pd-perry.github.io/expo-ft/","title":"EXPO-FT project page (Stanford)","sourceName":"EXPO-FT authors (project site)","publishedAt":"2026-05-25"},{"url":"https://github.com/pd-perry/expo-ft/","title":"EXPO-FT open-source codebase","sourceName":"GitHub (pd-perry/expo-ft)","publishedAt":"2026-05-25"},{"url":"https://x.com/robotsdigest/status/2103188667342975009","title":"Robots Digest secondary summary of EXPO-FT blog","sourceName":"Robots Digest (@robotsdigest)","publishedAt":"2026-09-24"}],"keyFacts":[{"label":"Method (paper-claimed)","value":"Paper (arXiv 2605.25477, Perry Dong / Chelsea Finn et al., Stanford): EXPO-FT finetunes a pretrained VLA with online RL. Keeps the VLA as a supervised base policy, trains a lightweight bounded edit policy to maximize a learned Q-function, and uses an on-the-fly policy to pick the highest-Q action among base and edited candidates. Human-in-the-loop corrections allowed during training. Academic system, not a commercial product"},{"label":"Real-robot results (paper-claimed)","value":"Paper-claimed: perfect task performance (30/30 successes) across all evaluated real-robot manipulation tasks within an average of 19.1 minutes of online robot interaction data, outperforming prior RL-from-scratch and VLA finetuning baselines in the paper's tables. Evaluation suite covers 8 tasks (digest said 6; prefer paper count): flower insert, string-light routing (3 subtasks), egg flip, candy scoop, pool shot, cube pick. Academic result"},{"label":"pi0.5 backbone open-source claim (paper abs)","value":"arXiv abs 2605.25477 (Dong, Hung, Gao, Sadigh, Finn; Stanford; project https://pd-perry.github.io/expo-ft/) presents EXPO-FT for sample-efficient RL finetuning of pretrained VLAs, instantiating with pi0.5 as backbone, extending EXPO to action chunks and human-in-the-loop corrections. Paper claims 30/30 successes across evaluated real-robot tasks within an average of 19.1 minutes of online robot data and releases an open-source codebase. Treat results as paper-claimed, not independently audited here (arXiv)"},{"label":"Task suite egg flip pool shot (paper abs)","value":"The same abstract/paper highlight suite includes routing string lights and inserting the plug, striking a pool ball into a pocket, and inserting a flower into a wine bottle, emphasizing precision, dynamic actions, and varied initial states versus prior RL-from-scratch and VLA finetuning baselines. Treat task list as paper-stated (arXiv)"},{"label":"Actor learner training architecture","value":"EXPO-FT's project page describes a sample-efficient reinforcement-learning method that fine-tunes a pretrained vision-language-action policy with a learner and real-robot actor workflow (EXPO-FT)."},{"label":"Open research implementation","value":"The EXPO-FT repository publishes the research implementation for sample-efficient VLA fine-tuning and robot rollout experiments (EXPO-FT repository)."}],"aliases":["EXPO FT","EXPO-FT VLA","Sample-Efficient RL Finetuning for VLAs"],"collisionRisk":"low","reviewNote":null,"builtOnBrainId":null,"createdAt":"2026-09-24T21:00:27.398Z","updatedAt":"2026-10-09T22:00:40.522Z","jsonLd":{"@context":"https://schema.org","@type":"SoftwareApplication","@id":"https://registry.deploy.report/brains/expo-ft","url":"https://registry.deploy.report/brains/expo-ft","name":"EXPO-FT","alternateName":["EXPO FT","EXPO-FT VLA","Sample-Efficient RL Finetuning for VLAs"],"description":"EXPO-FT is a Stanford research system (Dong, Hung, Gao, Sadigh, Finn) for sample-efficient reinforcement learning finetuning of pretrained vision-language-action policies. It pairs a supervised VLA base policy with a lightweight edit policy and Q-guided on-the-fly action selection, optionally with human-in-the-loop corrections. Paper-claimed results reach 30/30 on evaluated real-robot tasks in ~19 minutes of online interaction on average. Open-source codebase released with the paper.","identifier":"294affed-c8f5-420c-9811-9ad968290493","applicationCategory":"research-model","publisher":{"@id":"https://deploy.report/#organization"}},"framework_metadata":{"framework_schema_version":"0.1.0","verification_status":"verified","maturity_stage":"research","lifecycle_state":null,"architectural_position":{"cohort":null,"sub_cohorts":[]},"within_cohort_verified_vs_claimed_pair":null,"cap_flags":[],"verification_depth":{"sources_count":4,"primary_source_types":["code-repository","preprint"]}}}