{"id":"c984c50c-8305-467e-a9f4-1bb4b6d65e7c","name":"DeepInfra","slug":"deepinfra","description":"DeepInfra is an inference-focused GPU cloud that runs open-weight and open-source models (Llama, Qwen, DeepSeek, Kimi, GLM, Nemotron, Mistral, IBM Granite, and 100+ others) behind a pay-as-you-go API. Founded in September 2022 in Palo Alto by Nikola Borisov (CEO), Georgios Papoutsis, and Yessenzhar Kanapin, the team previously ran the infrastructure behind the imo messenger, which processed billions of messages a day for over 200 million monthly users. That operator DNA is the pitch: squeeze open-weight inference far below hyperscaler list price and hand developers a token meter instead of an instance bill.\n\nUnlike training-oriented neoclouds such as CoreWeave or Crusoe, DeepInfra explicitly positions on the inference side of the split: serverless per-token endpoints for the long tail of open models, plus a \"DeepCluster\" dedicated-GPU tier for customers who want SXM-connected A100, H100, H200, B200, and B300 boxes on weekly invoicing. Advertised dedicated pricing runs $0.89/hr for an A100 80GB, $2.20/hr for an H100 80GB, $2.69/hr for an H200, $3.69/hr for a B200, and $4.89/hr for a B300, with the company claiming roughly 70% savings versus public-cloud on-demand rates. Infrastructure sits in what the company describes as its own inference-optimized fleet inside secure US-based data centers; it has not publicly disclosed a GPU count, MW footprint, or colocation partners.\n\nDeepInfra raised a $107 million Series B in 2026 to expand the fleet, on top of an earlier Series A that included NVIDIA, A Capital, Felicis, 500 Global, SV Angel, Crescent Cove, and PEAK6. Public reference customers include Salesforce, Hugging Face, Abacus.AI, interface.ai, and Requesty. The company holds SOC 2 and ISO 27001 and markets a zero-retention policy, and is leaning into data-sovereignty framing for regulated buyers who want open-weight models rather than proprietary APIs.","status":"active","founded":2022,"hq":"Palo Alto, California, USA","hqLat":null,"hqLng":null,"hqGeocodedAt":null,"hqAddress":null,"hqPrecision":null,"website":"https://deepinfra.com","fundingTotal":"107000000","type":null,"roles":[],"reviewStatus":"reviewed","lifecycleState":"active","supersededByCompanyId":null,"sources":[{"url":"https://deepinfra.com/about","title":"About DeepInfra (founders, investors, HQ)","publisher":"DeepInfra"},{"url":"https://deepinfra.com/pricing","title":"DeepInfra pricing (per-token and dedicated GPU hourly rates)","publisher":"DeepInfra"},{"url":"https://deepinfra.com/","title":"DeepInfra homepage (customers, model catalog, positioning)","publisher":"DeepInfra"},{"url":"https://deepinfra.com/blog","title":"DeepInfra blog (Series B announcement, model launches)","publisher":"DeepInfra"}],"keyFacts":[{"label":"Headquarters","value":"Palo Alto, California","sourceUrl":"https://deepinfra.com/about"},{"label":"Founded","value":"September 2022 by Nikola Borisov, Georgios Papoutsis, and Yessenzhar Kanapin (ex-imo messenger infra team)","sourceUrl":"https://deepinfra.com/about"},{"label":"Positioning","value":"Inference-focused GPU cloud for open-weight models (Llama, Qwen, DeepSeek, Kimi, GLM, Nemotron, Mistral, Granite); per-token serverless plus DeepCluster dedicated tier","sourceUrl":"https://deepinfra.com/"},{"label":"Latest funding","value":"$107M Series B announced in 2026; earlier Series A backed by NVIDIA, A Capital, Felicis, 500 Global, SV Angel, Crescent Cove, PEAK6","sourceUrl":"https://deepinfra.com/about"},{"label":"GPU inventory","value":"Not publicly disclosed; company describes its own inference-optimized fleet in US-based data centers. No MW or campus-level footprint disclosed.","sourceUrl":"https://deepinfra.com/"},{"label":"Dedicated GPU pricing (SXM, on-demand)","value":"A100 80GB $0.89/hr, H100 80GB $2.20/hr, H200 141GB $2.69/hr, B200 180GB $3.69/hr, B300 270GB $4.89/hr; weekly invoicing","sourceUrl":"https://deepinfra.com/pricing"},{"label":"Cost claim vs public cloud","value":"DeepCluster dedicated hardware advertised at ~70% savings vs public-cloud equivalents","sourceUrl":"https://deepinfra.com/"},{"label":"Silicon partnerships","value":"NVIDIA (A100/H100/H200/B200/B300 across the fleet; NVIDIA also a strategic investor). No disclosed AMD, Cerebras, Intel Gaudi, or custom-silicon program.","sourceUrl":"https://deepinfra.com/pricing"},{"label":"Named customers","value":"Salesforce, Hugging Face, Abacus.AI, interface.ai, Requesty","sourceUrl":"https://deepinfra.com/"},{"label":"Model catalog","value":"100+ open-weight and open-source models: text generation, embeddings, ASR, reranking, TTS, text-to-image, text-to-video","sourceUrl":"https://deepinfra.com/"},{"label":"Compliance","value":"SOC 2 and ISO 27001; zero-retention policy on inference traffic","sourceUrl":"https://deepinfra.com/"},{"label":"Status","value":"Private; VC-backed","sourceUrl":"https://deepinfra.com/about"}],"aliases":[],"collisionRisk":"low","reviewNote":null,"atsProvider":null,"atsSlug":null,"factoryLocations":null,"manufacturingCapacity":null,"orgType":null,"burnSignal":null,"litigationEvents":null,"insurancePostureDisclosed":null,"unitEconomicsDisclosed":null,"tickerSymbol":null,"stockExchange":null,"secCik":null,"revenueUsd":null,"revenueBasis":null,"revenueAsOf":null,"employeeCount":30,"employeeCountBasis":"Approximate, based on visible team roster on deepinfra.com/about; company has not disclosed a precise headcount.","employeeCountAsOf":"2026-08-01T00:00:00.000Z","marketCapUsd":null,"marketCapAsOf":null,"marketCapSource":null,"createdAt":"2026-08-28T01:48:26.925Z","updatedAt":"2026-08-28T03:45:17.421Z","jsonLd":{"@context":"https://schema.org","@type":"Organization","@id":"https://registry.deploy.report/companies/deepinfra","url":"https://registry.deploy.report/companies/deepinfra","name":"DeepInfra","description":"DeepInfra is an inference-focused GPU cloud that runs open-weight and open-source models (Llama, Qwen, DeepSeek, Kimi, GLM, Nemotron, Mistral, IBM Granite, and 100+ others) behind a pay-as-you-go API. Founded in September 2022 in Palo Alto by Nikola Borisov (CEO), Georgios Papoutsis, and Yessenzhar Kanapin, the team previously ran the infrastructure behind the imo messenger, which processed billions of messages a day for over 200 million monthly users. That operator DNA is the pitch: squeeze open-weight inference far below hyperscaler list price and hand developers a token meter instead of an instance bill.\n\nUnlike training-oriented neoclouds such as CoreWeave or Crusoe, DeepInfra explicitly positions on the inference side of the split: serverless per-token endpoints for the long tail of open models, plus a \"DeepCluster\" dedicated-GPU tier for customers who want SXM-connected A100, H100, H200, B200, and B300 boxes on weekly invoicing. Advertised dedicated pricing runs $0.89/hr for an A100 80GB, $2.20/hr for an H100 80GB, $2.69/hr for an H200, $3.69/hr for a B200, and $4.89/hr for a B300, with the company claiming roughly 70% savings versus public-cloud on-demand rates. Infrastructure sits in what the company describes as its own inference-optimized fleet inside secure US-based data centers; it has not publicly disclosed a GPU count, MW footprint, or colocation partners.\n\nDeepInfra raised a $107 million Series B in 2026 to expand the fleet, on top of an earlier Series A that included NVIDIA, A Capital, Felicis, 500 Global, SV Angel, Crescent Cove, and PEAK6. Public reference customers include Salesforce, Hugging Face, Abacus.AI, interface.ai, and Requesty. The company holds SOC 2 and ISO 27001 and markets a zero-retention policy, and is leaning into data-sovereignty framing for regulated buyers who want open-weight models rather than proprietary APIs.","identifier":"c984c50c-8305-467e-a9f4-1bb4b6d65e7c","foundingDate":"2022","address":"Palo Alto, California, USA","publisher":{"@id":"https://deploy.report/#organization"}},"framework_metadata":{"framework_schema_version":"0.1.0","verification_status":"verified","maturity_stage":null,"lifecycle_state":null,"architectural_position":{"cohort":null,"sub_cohorts":[]},"within_cohort_verified_vs_claimed_pair":null,"cap_flags":[],"verification_depth":{"sources_count":4,"primary_source_types":[]}}}