DEPLOYDatabase

Brain

EXPO-FT

EXPO-FT is a Stanford research system (Dong, Hung, Gao, Sadigh, Finn) for sample-efficient reinforcement learning finetuning of pretrained vision-language-action policies. It pairs a supervised VLA base policy with a lightweight edit policy and Q-guided on-the-fly action selection, optionally with human-in-the-loop corrections. Paper-claimed results reach 30/30 on evaluated real-robot tasks in ~19 minutes of online interaction on average. Open-source codebase released with the paper.

Research model · Maturity: Research · Open source


Machine-readable surfaces

Architecture

EXPO-style base VLA + bounded edit policy + RedQ-style Q ensemble; on-the-fly max-Q selection over base and edited action chunks; human-in-the-loop interventions during online RL

Key facts

Method (paper-claimed)
Paper (arXiv 2605.25477, Perry Dong / Chelsea Finn et al., Stanford): EXPO-FT finetunes a pretrained VLA with online RL. Keeps the VLA as a supervised base policy, trains a lightweight bounded edit policy to maximize a learned Q-function, and uses an on-the-fly policy to pick the highest-Q action among base and edited candidates. Human-in-the-loop corrections allowed during training. Academic system, not a commercial product
Real-robot results (paper-claimed)
Paper-claimed: perfect task performance (30/30 successes) across all evaluated real-robot manipulation tasks within an average of 19.1 minutes of online robot interaction data, outperforming prior RL-from-scratch and VLA finetuning baselines in the paper's tables. Evaluation suite covers 8 tasks (digest said 6; prefer paper count): flower insert, string-light routing (3 subtasks), egg flip, candy scoop, pool shot, cube pick. Academic result
pi0.5 backbone open-source claim (paper abs)
arXiv abs 2605.25477 (Dong, Hung, Gao, Sadigh, Finn; Stanford; project pd-perry.github.io) presents EXPO-FT for sample-efficient RL finetuning of pretrained VLAs, instantiating with pi0.5 as backbone, extending EXPO to action chunks and human-in-the-loop corrections. Paper claims 30/30 successes across evaluated real-robot tasks within an average of 19.1 minutes of online robot data and releases an open-source codebase. Treat results as paper-claimed, not independently audited here (arXiv)
Task suite egg flip pool shot (paper abs)
The same abstract/paper highlight suite includes routing string lights and inserting the plug, striking a pool ball into a pocket, and inserting a flower into a wine bottle, emphasizing precision, dynamic actions, and varied initial states versus prior RL-from-scratch and VLA finetuning baselines. Treat task list as paper-stated (arXiv)
Actor learner training architecture
EXPO-FT's project page describes a sample-efficient reinforcement-learning method that fine-tunes a pretrained vision-language-action policy with a learner and real-robot actor workflow (EXPO-FT).
Open research implementation
The EXPO-FT repository publishes the research implementation for sample-efficient VLA fine-tuning and robot rollout experiments (EXPO-FT repository).

Common questions

What is EXPO-FT?
EXPO-FT is a Stanford research system (Dong, Hung, Gao, Sadigh, Finn) for sample-efficient reinforcement learning finetuning of pretrained vision-language-action policies. It pairs a supervised VLA base policy with a lightweight edit policy and Q-guided on-the-fly action selection, optionally with human-in-the-loop corrections. Paper-claimed results reach 30/30 on evaluated real-robot tasks in ~19 minutes of online interaction on average. Open-source codebase released with the paper.
Is EXPO-FT open source?
Yes. EXPO-FT is recorded as open-source / open-weights on the DEPLOY registry, meaning model weights or source code are publicly available.
What type of AI is EXPO-FT?
EXPO-FT is a research model, built on a EXPO-style base VLA + bounded edit policy + RedQ-style Q ensemble; on-the-fly max-Q selection over base and edited action chunks; human-in-the-loop interventions during online RL architecture on the DEPLOY registry.
What is EXPO-FT's maturity stage?
EXPO-FT is at the research stage on the DEPLOY maturity ladder. Research stage means active development without commercial deployments on file.
Which robots run on EXPO-FT?
No robot models on the DEPLOY registry are recorded as running EXPO-FT. DEPLOY wires brain-to-model connections only when the wiring is verifiable from primary sources; absence may reflect pre-deployment or unverified manufacturer claims.

Sources (4)

  1. https://arxiv.org/abs/2605.25477 · 2026-05-25
  2. https://pd-perry.github.io/expo-ft/ · 2026-05-25
  3. https://github.com/pd-perry/expo-ft/ · 2026-05-25
  4. https://x.com/robotsdigest/status/2103188667342975009 · 2026-09-24
Methodology: Verified · 4 sources (no primary) · last reviewed 2026-10-09

Verification posture

Verified

Low confidence

Review state

Stable

Last reviewed 2026-10-09

Maturity + lifecycle

Maturity stage: research

Sources by quality tier

2
unclassified
Unclassified source
1
preprint
Preprint
1
code-repository
Code repository

The framework is documented at /methodology. Corrections at /corrections. Reviewer: DEPLOY editorial team.

Methodology surface for EXPO-FT.

Canonical ID 294affed-c8f5-420c-9811-9ad968290493