Skip to content
ℹ️ Now part of NVIDIA: NVIDIA acquired Lepton AI and the product now ships as DGX Cloud Lepton, where lepton.ai redirects. See our recommended alternatives for current comparisons.

Lepton AI

Discontinued Last verified: September 2026

Cloud-native AI inference platform with per-minute compute billing. Discontinued as a standalone product; lepton.ai now redirects to NVIDIA DGX Cloud Lepton

4.3/5

What is Lepton AI?

Status update, checked 2 August 2026: lepton.ai no longer serves a standalone product site. It returns a permanent redirect to NVIDIA's DGX Cloud Lepton page, docs.lepton.ai no longer resolves, and the console now lives at dashboard.dgxc-lepton.nvidia.com. Lepton continues as NVIDIA's DGX Cloud Lepton GPU marketplace rather than as the independent inference service described below, and NVIDIA publishes no per-token or storage rates on that page. The prices in this review are therefore kept as a record of what Lepton published before the transition; they could not be re-verified against any current official page. Treat them as historical and get a quote from NVIDIA.

Lepton AI is a cloud-native platform for deploying, serving, and scaling AI models with a developer-first focus on simplicity and fast iteration, founded by former Alibaba AI researchers and acquired by NVIDIA in March 2025. Following the acquisition, Lepton was integrated into NVIDIA's DGX Cloud Lepton offering, which unifies Lepton's managed inference tooling with NVIDIA's global GPU supply across multiple cloud providers, giving developers access to NVIDIA Blackwell, H100, H200, and A100 GPUs through a single Lepton interface. Lepton's core product was a serverless-style inference platform where you could deploy any open-weight model from Hugging Face, a custom PyTorch model, or one of Lepton's pre-deployed endpoints, and get a production-ready API in minutes. Pricing was usage-based and billed per minute of actual compute, with no idle charges when your endpoint scaled to zero. For pre-deployed LLM inference, Lepton published aggressive per-token rates: Llama 3.2 3B at $0.03 per million tokens and Llama 3.1 8B at $0.07 per million tokens made it one of the cheapest inference options on the market. Storage for models, data, and logs was charged at $0.153 per GB per month. None of those rates can be bought today. Where Lepton differentiated itself was flexibility: you could self-deploy any custom model, use it for non-LLM workloads, and mix inference with batch training jobs on the same platform.

Lepton AI demo video

Watch Lepton AI's official demo to see Lepton AI in action before reading our full review.

Official video by Lepton AI via YouTube, embedded for reference. ToolChase does not host or claim this video.

⚡ Quick Verdict

Best for

Developers who want cheap pay-per-compute inference with flexibility to deploy custom models

Not ideal for

Teams that need the largest model catalog or the fastest LPU-style inference

Starting price

Discontinued. Historically from $0.03 per million tokens; NVIDIA publishes no rates for the successor

Free plan

None. The standalone Lepton AI site no longer exists

Key strength

Cheap pay-per-token pricing plus flexible custom model hosting on NVIDIA infrastructure

Limitation

Smaller catalog than Together AI or OpenRouter

Bottom line: Lepton scores 4.3/5 on the product as it was last reviewed, but it is no longer a standalone service: evaluate it as NVIDIA DGX Cloud Lepton and confirm current rates with NVIDIA, because none are published.

Pricing

Discontinued, historical record. Lepton is no longer sold as a standalone service and the rates below can no longer be purchased. Checked 2 August 2026: lepton.ai returns a 301 redirect to NVIDIA DGX Cloud Lepton, which publishes no per-token, per-GPU-hour or storage pricing, so these figures are Lepton's last published self-serve prices and could not be re-verified against an official page.

Pre-deployed LLM inference (per-token): Llama 3.2 3B at $0.03 per million tokens · Llama 3.1 8B at $0.07 per million tokens · Larger models priced proportionally. These were among the cheapest per-token rates on the market while Lepton sold them directly.

Custom model deployment, Pay-per-compute: Was billed by the minute for actual GPU usage, with no idle charges when your endpoint scaled to zero. Supported NVIDIA Blackwell, H100, H200, A100, and A10G GPUs.

Storage: Was $0.153 per GB per month for models, data, and logs stored on the Lepton platform.

DGX Cloud Lepton: Following NVIDIA's 2025 acquisition, Lepton is now sold as NVIDIA DGX Cloud Lepton, a marketplace that unifies GPU supply across many cloud providers. Pricing there is not published publicly, so budget by requesting a quote from NVIDIA or the underlying cloud partner.

Key Features

  • Cloud-native inference platform with per-minute billing
  • Pre-deployed Llama and other open LLMs, historically from $0.03 per million tokens
  • Custom model deployment from Hugging Face or PyTorch
  • Scale-to-zero for cost efficiency
  • NVIDIA DGX Cloud Lepton integration
  • Access to Blackwell, H100, H200, A100 GPUs
  • Supports LLM, vision, audio, and custom models
  • Python SDK and REST API

Pros & Cons

Pros

  • Very cheap per-token pricing for pre-deployed Llama models while it was sold
  • NVIDIA backing means reliable GPU supply and new hardware access
  • Flexible, serve any open-weight or custom model, not just a catalog
  • Scale-to-zero saves money for bursty workloads

Cons

  • Smaller ecosystem than Together AI or Replicate
  • Less polished docs than AWS Bedrock or Vertex AI
  • No standalone product any more, lepton.ai redirects to NVIDIA and pricing is no longer published
✅ Pricing shown is historical (product discontinued) · ⚠️ lepton.ai returns a 301 redirect to NVIDIA DGX Cloud Lepton · ✅ Independently reviewed · ✅ Scoring methodology

FAQ

Is Lepton still independent after the NVIDIA acquisition?

No. Lepton was acquired by NVIDIA in March 2025 and folded into the DGX Cloud Lepton product line. Checked on 2 August 2026, lepton.ai no longer hosts a standalone site at all: it permanently redirects to NVIDIA's DGX Cloud Lepton page, docs.lepton.ai has been withdrawn, and the console has moved to dashboard.dgxc-lepton.nvidia.com. The Lepton name survives as an NVIDIA product, a marketplace that connects developers to GPU capacity across many cloud providers, not as an independent vendor you can buy from directly.

How does Lepton pricing compare to Together AI?

These rates are historical and can no longer be purchased: Lepton is no longer sold as a standalone service. Lepton's pre-deployed Llama 3.2 3B at $0.03 per million tokens was cheaper than Together AI's equivalent tier, and Llama 3.1 8B at $0.07 per million was also competitive. For custom model hosting, Lepton's per-minute compute billing was simpler than Together's per-second pricing on dedicated endpoints.

Can I deploy my own fine-tuned model?

Not through a standalone Lepton account any more, because lepton.ai redirects to NVIDIA DGX Cloud Lepton. Lepton supported deploying any open-weight model from Hugging Face or a custom PyTorch model via their Python SDK: you packaged your model code, pushed it to Lepton's platform, and got a production endpoint with automatic scaling. Teams that have fine-tuned Llama and need managed hosting should price the equivalent with NVIDIA.

Does Lepton scale to zero?

That is how the standalone service worked. Custom deployments on Lepton could scale to zero when idle, meaning you only paid for compute when actual requests were being processed. It was a major cost advantage for bursty or low-traffic workloads compared to always-on dedicated endpoints on AWS SageMaker or Azure ML. Confirm whether NVIDIA DGX Cloud Lepton keeps the same behavior before budgeting for it.

What GPUs are available?

As part of NVIDIA DGX Cloud Lepton, the platform provides access to NVIDIA Blackwell, H200, H100, A100 80GB, and A10G GPUs across multiple cloud regions. For the latest GPUs (Blackwell), availability is tighter but Lepton has direct NVIDIA supply advantages over third-party providers.

Is Lepton good for non-LLM workloads?

Yes, this was one of Lepton's strengths compared to LLM-only providers. Lepton could host vision models, speech models, recommendation systems, and custom PyTorch pipelines. For teams with mixed AI workloads, Lepton simplified infrastructure by consolidating on one platform.

📋 Good to know

Setup

Sign-up now runs through NVIDIA DGX Cloud Lepton at dashboard.dgxc-lepton.nvidia.com; lepton.ai itself redirects to NVIDIA. From the console you deploy an endpoint or push a custom model via the SDK.

Privacy

Lepton did not train on your data, and enterprise agreements were available for regulated workloads. Confirm the equivalent terms with NVIDIA.

When to upgrade

No longer applicable, since Lepton is not sold as a standalone service. Historically you moved from pre-deployed endpoints to custom model hosting when you needed fine-tuned weights.

Learning curve

Moderate, Python SDK is straightforward but custom deployment involves more concepts than pure inference APIs.

📝 Report incorrect info about Lepton AI