Brain
EXPO-FT is a Stanford research system (Dong, Hung, Gao, Sadigh, Finn) for sample-efficient reinforcement learning finetuning of pretrained vision-language-action policies. It pairs a supervised VLA base policy with a lightweight edit policy and Q-guided on-the-fly action selection, optionally with human-in-the-loop corrections. Paper-claimed results reach 30/30 on evaluated real-robot tasks in ~19 minutes of online interaction on average. Open-source codebase released with the paper.
Research model · Maturity: Research · Open source
Machine-readable surfaces
- Markdown mirror: /brains/expo-ft.md
- RSS feed: /brains/expo-ft/feed.xml
- JSON-LD: embedded in this page’s head
- REST API: /v1/brains/294affed-c8f5-420c-9811-9ad968290493
- Revision history: /brains/expo-ft/history
- Data documentation: /data
- Query this programmatically: Deploy MCP
Architecture
EXPO-style base VLA + bounded edit policy + RedQ-style Q ensemble; on-the-fly max-Q selection over base and edited action chunks; human-in-the-loop interventions during online RL
Key facts
- Method (paper-claimed)
- Paper (arXiv 2605.25477, Perry Dong / Chelsea Finn et al., Stanford): EXPO-FT finetunes a pretrained VLA with online RL. Keeps the VLA as a supervised base policy, trains a lightweight bounded edit policy to maximize a learned Q-function, and uses an on-the-fly policy to pick the highest-Q action among base and edited candidates. Human-in-the-loop corrections allowed during training. Academic system, not a commercial product
- Real-robot results (paper-claimed)
- Paper-claimed: perfect task performance (30/30 successes) across all evaluated real-robot manipulation tasks within an average of 19.1 minutes of online robot interaction data, outperforming prior RL-from-scratch and VLA finetuning baselines in the paper's tables. Evaluation suite covers 8 tasks (digest said 6; prefer paper count): flower insert, string-light routing (3 subtasks), egg flip, candy scoop, pool shot, cube pick. Academic result
- pi0.5 backbone open-source claim (paper abs)
- arXiv abs 2605.25477 (Dong, Hung, Gao, Sadigh, Finn; Stanford; project pd-perry.github.io) presents EXPO-FT for sample-efficient RL finetuning of pretrained VLAs, instantiating with pi0.5 as backbone, extending EXPO to action chunks and human-in-the-loop corrections. Paper claims 30/30 successes across evaluated real-robot tasks within an average of 19.1 minutes of online robot data and releases an open-source codebase. Treat results as paper-claimed, not independently audited here (arXiv)
- Task suite egg flip pool shot (paper abs)
- The same abstract/paper highlight suite includes routing string lights and inserting the plug, striking a pool ball into a pocket, and inserting a flower into a wine bottle, emphasizing precision, dynamic actions, and varied initial states versus prior RL-from-scratch and VLA finetuning baselines. Treat task list as paper-stated (arXiv)
- Actor learner training architecture
- EXPO-FT's project page describes a sample-efficient reinforcement-learning method that fine-tunes a pretrained vision-language-action policy with a learner and real-robot actor workflow (EXPO-FT).
- Open research implementation
- The EXPO-FT repository publishes the research implementation for sample-efficient VLA fine-tuning and robot rollout experiments (EXPO-FT repository).
Common questions
What is EXPO-FT?
Is EXPO-FT open source?
What type of AI is EXPO-FT?
What is EXPO-FT's maturity stage?
Which robots run on EXPO-FT?
Sources (4)
- https://arxiv.org/abs/2605.25477 · 2026-05-25
- https://pd-perry.github.io/expo-ft/ · 2026-05-25
- https://github.com/pd-perry/expo-ft/ · 2026-05-25
- https://x.com/robotsdigest/status/2103188667342975009 · 2026-09-24
Methodology: Verified · 4 sources (no primary) · last reviewed 2026-10-09
Verification posture
Verified
Low confidence
Review state
Stable
Last reviewed 2026-10-09
Maturity + lifecycle
Maturity stage: research
Sources by quality tier
- 2
- unclassified
- Unclassified source
- 1
- preprint
- Preprint
- 1
- code-repository
- Code repository
The framework is documented at /methodology. Corrections at /corrections. Reviewer: DEPLOY editorial team.
Methodology surface for EXPO-FT.Canonical ID 294affed-c8f5-420c-9811-9ad968290493