# EXPO-FT: robot AI research model

EXPO-FT is a [Stanford](/locations/stanford.md) research system (Dong, Hung, Gao, Sadigh, Finn) for sample-efficient [reinforcement learning](/glossary/reinforcement-learning.md) finetuning of pretrained [vision-language-action](/glossary/vision-language-action-model.md) policies. It pairs a supervised VLA base policy with a lightweight edit policy and Q-guided on-the-fly action selection, optionally with human-in-the-loop corrections. Paper-claimed results reach 30/30 on evaluated real-robot tasks in ~19 minutes of online interaction on average. Open-source codebase released with the paper.

- **Slug:** expo-ft
- **Type:** Research model
- **Maturity (DEPLOY ladder):** Research
- **Open source:** yes

## Architecture

EXPO-style base VLA + bounded edit policy + RedQ-style Q ensemble; on-the-fly max-Q selection over base and edited action chunks; human-in-the-loop interventions during online RL


## Key facts

- **Method (paper-claimed):** Paper (arXiv 2605.25477, Perry Dong / [Chelsea Finn](/people/chelsea-finn.md) et al., [Stanford](/locations/stanford.md)): EXPO-FT finetunes a pretrained [VLA](/glossary/vision-language-action-model.md) with online RL. Keeps the VLA as a supervised base policy, trains a lightweight bounded edit policy to maximize a learned Q-function, and uses an on-the-fly policy to pick the highest-Q action among base and edited candidates. Human-in-the-loop corrections allowed during training. Academic system, not a commercial product
- **Real-robot results (paper-claimed):** Paper-claimed: perfect task performance (30/30 successes) across all evaluated real-robot manipulation tasks within an average of 19.1 minutes of online robot interaction data, outperforming prior RL-from-scratch and [VLA](/glossary/vision-language-action-model.md) finetuning baselines in the paper's tables. Evaluation suite covers 8 tasks (digest said 6; prefer paper count): flower insert, string-light routing (3 subtasks), egg flip, candy scoop, pool shot, cube pick. Academic result
- **pi0.5 backbone open-source claim (paper abs):** arXiv abs 2605.25477 (Dong, Hung, Gao, Sadigh, Finn; [Stanford](/locations/stanford.md); project https://pd-perry.github.io/expo-ft/) presents EXPO-FT for sample-efficient RL finetuning of pretrained VLAs, instantiating with pi0.5 as backbone, extending EXPO to action chunks and human-in-the-loop corrections. Paper claims 30/30 successes across evaluated real-robot tasks within an average of 19.1 minutes of online robot data and releases an open-source codebase. Treat results as paper-claimed, not independently audited here (arXiv)
- **Task suite egg flip pool shot (paper abs):** The same abstract/paper highlight suite includes routing string lights and inserting the plug, striking a pool ball into a pocket, and inserting a flower into a wine bottle, emphasizing precision, dynamic actions, and varied initial states versus prior RL-from-scratch and [VLA](/glossary/vision-language-action-model.md) finetuning baselines. Treat task list as paper-stated (arXiv)
- **Actor learner training architecture:** EXPO-FT's project page describes a sample-efficient reinforcement-learning method that fine-tunes a pretrained [vision-language-action](/glossary/vision-language-action-model.md) policy with a learner and real-robot actor workflow (EXPO-FT).
- **Open research implementation:** The EXPO-FT repository publishes the research implementation for sample-efficient [VLA](/glossary/vision-language-action-model.md) fine-tuning and robot rollout experiments (EXPO-FT repository).


## Sources (4)

1. **EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models** · https://arxiv.org/abs/2605.25477 · 2026-05-25
2. **EXPO-FT project page (Stanford)** · https://pd-perry.github.io/expo-ft/ · 2026-05-25
3. **EXPO-FT open-source codebase** · https://github.com/pd-perry/expo-ft/ · 2026-05-25
4. **Robots Digest secondary summary of EXPO-FT blog** · https://x.com/robotsdigest/status/2103188667342975009 · 2026-09-24


## Common questions

### What is EXPO-FT?

EXPO-FT is a [Stanford](/locations/stanford.md) research system (Dong, Hung, Gao, Sadigh, Finn) for sample-efficient [reinforcement learning](/glossary/reinforcement-learning.md) finetuning of pretrained [vision-language-action](/glossary/vision-language-action-model.md) policies. It pairs a supervised VLA base policy with a lightweight edit policy and Q-guided on-the-fly action selection, optionally with human-in-the-loop corrections. Paper-claimed results reach 30/30 on evaluated real-robot tasks in ~19 minutes of online interaction on average. Open-source codebase released with the paper.

### Is EXPO-FT open source?

Yes. EXPO-FT is recorded as open-source / open-weights on the DEPLOY registry, meaning model weights or source code are publicly available.

### What type of AI is EXPO-FT?

EXPO-FT is a research model, built on a EXPO-style base [VLA](/glossary/vision-language-action-model.md) + bounded edit policy + RedQ-style Q ensemble; on-the-fly max-Q selection over base and edited action chunks; human-in-the-loop interventions during online RL architecture on the DEPLOY registry.

### What is EXPO-FT's maturity stage?

EXPO-FT is at the research stage on the DEPLOY maturity ladder. Research stage means active development without commercial deployments on file.

### Which robots run on EXPO-FT?

No robot models on the DEPLOY registry are recorded as running EXPO-FT. DEPLOY wires brain-to-model connections only when the wiring is verifiable from primary sources; absence may reflect pre-deployment or unverified manufacturer claims.


_API: GET /v1/brains/294affed-c8f5-420c-9811-9ad968290493 · canonical URL: /brains/expo-ft_
