{"id":"501f1b47-683d-400e-afc8-74beff0d8937","slug":"smolvla","name":"SmolVLA","description":"Hugging Face's compact 450M parameter VLA model for accessible robotics. Part of the LeRobot framework, designed for low-compute training and deployment on consumer hardware.","brainType":"foundation-model","isOpen":true,"maturityStage":"research","architecture":"Compact VLA, 450M params","reviewStatus":"reviewed","sources":[{"url":"https://huggingface.co/blog/smolvla","date":"2025-06-03","title":"SmolVLA: Efficient Vision-Language-Action Model trained on Lerobot Community Data","sourceName":"Hugging Face"},{"url":"https://huggingface.co/docs/lerobot/smolvla","title":"SmolVLA documentation","sourceName":"Hugging Face LeRobot"},{"url":"https://huggingface.co/lerobot/smolvla_base","title":"lerobot/smolvla_base","sourceName":"Hugging Face"}],"keyFacts":[{"label":"SmolVLA release (primary Hugging Face)","value":"Hugging Face's June 3 2025 post introduces SmolVLA as a compact 450M-parameter open-source vision-language-action model that runs on consumer hardware, pretrained on compatibly licensed community datasets under the lerobot tag. The base checkpoint is lerobot/smolvla_base. The post says the VLM backbone is SmolVLM2 and the action expert is a flow-matching transformer of about 100M parameters. Treat size and hardware claims as Hugging Face-stated ([Hugging Face](https://huggingface.co/blog/smolvla))","sourceUrl":"https://huggingface.co/blog/smolvla"},{"label":"Reported performance (primary Hugging Face)","value":"The same post says SmolVLA-450M outperforms much larger VLAs and ACT on LIBERO, Meta-World, SO100, and SO101, and that asynchronous inference is about 30% faster (9.7s vs 13.75s) with 2x task throughput (19 vs 9 cubes) at similar success. These are author-reported benchmark results, not an independent audit ([Hugging Face](https://huggingface.co/blog/smolvla))","sourceUrl":"https://huggingface.co/blog/smolvla"},{"label":"LeRobot deployment documentation","value":"The LeRobot documentation provides SmolVLA installation, policy configuration, dataset preparation, and training guidance for robot learning workflows ([LeRobot SmolVLA documentation](https://huggingface.co/docs/lerobot/smolvla))","sourceUrl":"https://huggingface.co/docs/lerobot/smolvla"},{"label":"SmolVLA base checkpoint","value":"Hugging Face publishes the lerobot/smolvla_base checkpoint as the base SmolVLA model for use with the LeRobot ecosystem ([SmolVLA base model](https://huggingface.co/lerobot/smolvla_base))","sourceUrl":"https://huggingface.co/lerobot/smolvla_base"},{"label":"Compact open VLA","value":"Hugging Face describes SmolVLA as a compact 450 million parameter open-source vision-language-action model for robotics that runs on consumer hardware ([SmolVLA](https://huggingface.co/blog/smolvla)).","sourceUrl":"https://huggingface.co/blog/smolvla"},{"label":"Community-data pretraining","value":"Hugging Face says SmolVLA is pretrained on publicly available community-shared robotics datasets under the LeRobot tag ([SmolVLA](https://huggingface.co/blog/smolvla)).","sourceUrl":"https://huggingface.co/blog/smolvla"},{"label":"Architecture","value":"Hugging Face says SmolVLA combines a SmolVLM2 vision-language backbone with a flow-matching transformer action expert ([SmolVLA](https://huggingface.co/blog/smolvla)).","sourceUrl":"https://huggingface.co/blog/smolvla"},{"label":"Asynchronous inference","value":"Hugging Face reports that SmolVLA's asynchronous inference setup delivers about 30 percent faster response and twice the task throughput in its reported evaluation ([SmolVLA](https://huggingface.co/blog/smolvla)).","sourceUrl":"https://huggingface.co/blog/smolvla"},{"label":"LeRobot integration","value":"Hugging Face documents SmolVLA through the LeRobot framework, including the SmolVLAPolicy class and a pretrained smolvla_base checkpoint ([SmolVLA documentation](https://huggingface.co/docs/lerobot/smolvla)).","sourceUrl":"https://huggingface.co/docs/lerobot/smolvla"}],"aliases":[],"collisionRisk":"low","reviewNote":null,"builtOnBrainId":null,"createdAt":"2026-08-09T14:35:27.762Z","updatedAt":"2026-09-29T22:15:01.854Z","jsonLd":{"@context":"https://schema.org","@type":"SoftwareApplication","@id":"https://registry.deploy.report/brains/smolvla","url":"https://registry.deploy.report/brains/smolvla","name":"SmolVLA","description":"Hugging Face's compact 450M parameter VLA model for accessible robotics. Part of the LeRobot framework, designed for low-compute training and deployment on consumer hardware.","identifier":"501f1b47-683d-400e-afc8-74beff0d8937","applicationCategory":"foundation-model","publisher":{"@id":"https://deploy.report/#organization"}},"framework_metadata":{"framework_schema_version":"0.1.0","verification_status":"verified","maturity_stage":"research","lifecycle_state":null,"architectural_position":{"cohort":null,"sub_cohorts":[]},"within_cohort_verified_vs_claimed_pair":null,"cap_flags":[],"verification_depth":{"sources_count":3,"primary_source_types":["model-repository"]}}}