Backblaze B2 — Goodput Calculator (Neocloud) · Guide
How to read the demo · what the numbers mean · the math behind them
← Back to demo

How this demo works

From plain English to the math. Starts simple, gets deep. Designed so anyone — a finance lead, an infrastructure VP, or an engineer — can follow the story and verify the numbers.

Estimate only — not a quote. Figures throughout this document are projections derived from the Configure inputs and Backblaze list pricing on the date shown. Actual results depend on your specific workload, cluster utilization, contracted rates, and implementation. This is for discussion purposes and does not constitute a sales quote, pricing commitment, service-level agreement, or warranty. The I/O wait and flash-replacement defaults are working defaults pending pilot validation — instrument your own cluster and override the Configure fields before relying on these numbers for budget decisions. Final terms require a signed agreement.
The 30-second version

Want the math? Jump to §6. Want to validate against your own workload? §7. Quick walkthrough of the demo? §4.

What's in here

  1. Start here — what's "goodput" and why should I care?
  2. The opportunity — who has this problem
  3. What this tool shows
  4. Demo Walkthrough
  5. Every field, explained — and what we left out
  6. The math — formulas and derivation
    1. Goodput Recovered
    2. Checkpoint Stall Eliminated
    3. Flash capacity reclaimed
    4. New-capacity revenue
    5. Tier Premium and Net ROI
    6. The bandwidth-capability cap
  7. Assumptions and what to validate before showing externally
  8. FAQ — questions you'll get from engineers and CFOs

1. Start here — what's "goodput" and why should I care?

The kitchen story

Imagine a very expensive chef. The chef costs $1,000 an hour. The chef can cook fast, but only if the ingredients show up on time. If the delivery truck is late, the chef just stands there with a knife. You still pay the chef.

That's what's happening to GPUs every day. An H100 GPU rents for about $3 to $4 per hour. A cluster of 256 H100s costs more than $1,000 per hour to run. AI companies pay that whether the GPUs are training a model or sitting there waiting for the next batch of data.

Throughput vs goodput

Two words that sound the same but mean different things:

A truck with a 200 mph top speed and zero packages delivered has high throughput and zero goodput. A truck driving 60 mph and delivering every package on schedule has lower throughput and perfect goodput. Goodput is what you actually get paid for.

For an AI training cluster, goodput is the percentage of GPU time spent actually computing, not waiting on storage. A "95% goodput" cluster means 5% of the time the GPUs are idle — waiting for the next shuffle of training data, or waiting for a checkpoint to finish flushing to disk.

Why a few percentage points matter so much

A 256-GPU H100 cluster at $4/hr per GPU is $1,024/hr. Over a month (720 hours) that's $737,000.

If your goodput is 95% (5% idle), you waste ~$37,000/mo on idle GPU time.
If you can get to 99% goodput, you save ~$30,000/mo on the same hardware bill.
At 512 GPUs the same delta becomes ~$60,000/mo.

That's the whole idea in one paragraph: storage that keeps GPUs busy converts idle GPU dollars into billable training dollars. The faster and steadier the storage, the more of the GPU bill is doing useful work.

When does goodput actually start to matter?

Two conditions have to be true at the same time. If either one isn't, the goodput math reduces to zero (and Standard B2 is the right answer).

ThresholdWhere it tips
Cluster size large enough that 1% idle = real money ~32+ GPUs (~$100/hr cluster cost). Below this, even significant goodput improvements are too small to justify a tier upgrade.
Working set large enough that storage actually stalls the cluster ~3+ PB hot working set, OR bursty checkpoints ≥ 100 GB on a 64+ GPU cluster. Below this, Standard B2's 50 Gbps already keeps the GPUs fed.

Below both thresholds — Fine-tune lab · 1 PB territory in the demo — Standard B2 is the right answer and the goodput tiles correctly read $0. The tool will recommend Standard when you click "Auto-select optimal tier" in this regime.

Above both thresholds — Mid-scale tenant · 5 PB and up — the goodput story dominates. The Net ROI math turns positive for an Overdrive tier sized to the cluster's demand.

2. The opportunity — who has this problem

The four customer segments this is built for

SegmentWhat they sellTheir painHow Overdrive helps
GPU cloud providers
"Neoclouds" renting GPU capacity by the hour
GPU capacity by the hour to AI builders Margin is GPU $/hr minus storage and ops. Every idle GPU minute is dead margin. Flash is expensive — paid for whether customers re-read it or not. Higher goodput → more billable GPU hours from the same hardware. Flash displacement → smaller flash footprint to amortize.
GPU orchestration platforms
Schedulers that pack training jobs across clusters
Software that schedules and packs training jobs across clusters Their value prop IS goodput — they promise "we make your GPUs busier." If the underlying storage stalls, the scheduler can't fix it. Overdrive is the storage layer that makes their orchestration claims defensible end-to-end.
AI infrastructure providers
Bare-metal training, fine-tune-as-a-service, inference platforms
Compute platforms measured in tokens/sec or training $/epoch SLA pressure from customers who measure everything. A 2-minute checkpoint stall × 4 jobs × 24 hours is hours of lost training per day. Checkpoint stall elimination — directly recovers customer-visible throughput.
Data providers for AI training
Web-scale crawlers, media aggregators, dataset publishers
Crawl, curate, and package petabytes of training data for downstream AI builders Petabytes of cold data re-read on every training cycle. Flash for this is economically impossible; slow object storage stalls every downstream training run. Overdrive serves PB-scale catalogs at training-cluster line rate, without the flash bill.

What they all share

The core idea: the conversation about AI infrastructure storage is moving away from storage line-item comparisons. The bigger question is whether storage can convert idle GPU dollars into billable GPU dollars and let you displace expensive flash. This demo models that math.

3. What this tool shows

The page has four big numbers across the top — the "hero tiles." Each is one piece of the goodput story.

TileWhat it answersWho cares most
Goodput Recovered / mo"How much of my GPU bill stops being wasted on I/O wait?"CRO, CFO — direct revenue impact
Checkpoint Stall Eliminated"How much GPU time stops being burned during checkpoint flushes?"Platform engineer, AI customers
Flash capacity reclaimed"How much expensive flash can I free — reclaimed capacity to onboard more customers — because B2 serves the hot working set?"Infrastructure VP, finance
New-capacity revenue / mo"What can I earn reselling that reclaimed flash capacity to the next customer at full price?"CRO — the growth story

Why these four, and not a simple storage cost comparison?

Storage is a small fraction of any AI training bill. Even cutting it 80% doesn't move the conversation. GPU spend is 10–100× the storage spend, which means a 1% goodput improvement is worth more than a 50% storage discount. This tool surfaces that math directly instead of burying it inside a per-TB rate sheet.

How to read the baseline

When the page loads on Standard B2 (50 Gbps), every goodput tile shows $0 with an explainer ("Standard tier — no goodput delta to claim"). That's intentional, not a bug. Standard B2 IS the baseline — there's no delta to compare against itself. Switch to an Overdrive tier (100, 200, 300, 400, or 800 Gbps) to see the math kick in and the dollar values change.

4. Demo Walkthrough

The fastest way to learn what the tool does is to click through it. The walkthrough below takes about five minutes.

  1. Open the page. Mid-scale tenant · 5 PB (128 GPUs, 70B model) is the default scenario. The hero tiles show roughly $170K/mo total monthly impact on its recommended 100 Gbps Overdrive tier.
  2. Click the Standard B2 tier button. Every Overdrive-dependent tile drops to $0. This is the baseline — Standard B2 is the comparison point and has no "Overdrive delta" by definition.
  3. Click "⚡ Auto-select optimal tier." The tool picks the feed floor — the smallest tier that fully streams the dataset on schedule (where Net ROI peaks) — and shows the reasoning in the live event stream below.
  4. Read the recommendation. The B2 tier strip marks the recommended tier as a combined ★⚡ — best value and keeps-training-fed land on the same tier (the feed floor, the smallest tier that fully streams the dataset). It's selected by default. A higher tier buys marginally faster checkpoints/rehydration for more premium; it's a manual click, not the default.
  5. Switch to Fine-tune lab · 1 PB (32 GPUs). The optimizer recommends Standard B2, not Overdrive. At this scale the math doesn't justify the Overdrive premium — and the tool says so honestly.
  6. Switch to Frontier multimodal · 60 PB (768 GPUs, 405B model). Flash capacity reclaimed jumps to about 14.4 PB on the recommended 800 Gbps tier — the capacity-reclaim story for the largest training accounts.
  7. Project the fleet. Set "Tenants like this (× N)" in Flash economics — the per-tenant monthly value multiplies into a total fleet opportunity (e.g. 20 mid-scale tenants → ~$2.6M/mo). Default N = 1 is the single mapped customer.
  8. Click "Run." The full 8-stage pipeline executes against the demo bucket. Hero tiles fill in progressively as each stage completes; the "Open in B2 console" link on each finished stage card surfaces the actual object in the bucket so you can verify nothing is simulated. Prefer to skip the wait? Click the small "skip" link next to the Run button — tiles populate from the math immediately.

Tuning the scenario to your workload

Open the Configure panel to adjust every input directly — cluster size, GPU $/hr, checkpoint cadence, hot working set %, flash $/TB-mo, and I/O wait %. The hero tiles and Tier Comparison panel update live as you change any field. The Tier Comparison panel breaks every tier down (goodput / stall / flash / premium / net ROI) so you can compare them side-by-side.

Switching to Full pipeline mode

The Run button executes all 8 pipeline stages end-to-end in sequence: Data Lake → Training Prep → Checkpoints → Model Registry → RAG Knowledge → KV Cache → Inference → Exhaust. The first 5 stages carry the goodput / stall / flash-displacement headline economics; the last 3 (RAG, KV Cache, Exhaust) demonstrate the full set of B2 patterns in a production AI pipeline. Individual stage cards also have their own Run / Re-run buttons if you want to demo a single pattern in isolation.

5. Every field, explained — and what we left out

Fields included

Workload profile

FieldDefaultWhy it's here
Fine-tune lab · 1 PB1 PB, 32 GPUs, 7BCredibility profile. Math correctly recommends Standard B2 here.
Mid-scale tenant · 5 PB (default)5 PB, 128 GPUs, 70BThe default scenario. Math recommends 100 Gbps Overdrive (★⚡ — fully feeds the demand at the lowest premium).
Growth tenant · 10 PB10 PB, 256 GPUs, 70BScaling-up tenant. Prefetch demand outruns Standard, so the math recommends 200 Gbps (★⚡); both calculators agree at this size.
Large tenant cluster · 25 PB25 PB, 512 GPUs, 70BWhere Overdrive economics carry strongly. Math recommends 400 Gbps (★⚡).
Frontier multimodal · 60 PB60 PB, 768 GPUs, 405BThe realistic training ceiling and the biggest capacity-reclaim story. Math recommends 800 Gbps (★⚡).

Flash cost model

FieldDefaultWhy it's here
Drive price ($/TB hardware)$490 (QLC) / $570 (TLC)Q1 2026 reference pricing for 30.72 TB enterprise drives at volume.
Amortization (months)36Standard 3-year refresh cycle most CFOs assume.
DC overhead (%)50Power + cooling + fabric + ops markup on raw drive cost. 30-60% is typical for high-density flash.
Sell price ($/TB·mo to customers)$100What a neocloud charges its customers for flash. Used in the New-capacity revenue math.
Tenants like this (× N)1How many fleet tenants match this profile. Multiplies the per-tenant monthly value into a total fleet opportunity. Default 1 = the single mapped customer. (Replaces the older "total flash fleet / cold %" framing — a cleaner, explicit multiplier.)

Goodput scenario

FieldDefaultWhy it's here
Cluster size (GPUs)128 (Mid-scale default)Drives cluster $/hr in every goodput formula.
GPU $/hr$4.00H100 on-demand rate. Reserved instances are lower; CFO will tune.
Model size70BAuto-fills checkpoint size (~13 bytes/param for full training state per MLCommons MLPerf Storage benchmark).
Checkpoint size (GB)800Hourly checkpoint of a 70B model. MLPerf anchor buttons next to the input auto-fill canonical sizes: 8B → 105 GB, 70B → 912 GB, 405B → 5.29 TB, 1T → 15 TB (MLCommons MLPerf Storage).
Checkpoint cadenceevery 1 hrFrontier pretraining typical. Faster cadence = more stall payoff for Overdrive.
Concurrent jobs4Multi-tenant clusters checkpoint each job independently. Multiplies the stall payoff linearly.
Hot working set (% of dataset)32%Fraction of the dataset that has to stay on flash for active shuffled reads.
Flash $/TB-mo (displacement)$100The flash cost basis Overdrive is replacing. Mirrors the sell-price field above.
Std / Overdrive I/O wait %4.0% / 0.8%Working defaults — should be validated against your cluster's telemetry before treating as authoritative.
Overdrive flash-replace %80%Fraction of hot working set Overdrive can serve at line rate (so flash isn't needed for it).
Epochs per run (Advanced)2Passes over the dataset per run. With run duration, sets the prefetch bandwidth demand behind the ⚡ "keeps training fed" tier (dataset × epochs × 8 / run-time).
Run duration (days) (Advanced)14Length of one training run — the other lever in the prefetch-demand math for the ⚡ tier. Matches the GenAI tool's basis so the two tools agree on the keeps-fed tier.

Fields intentionally not exposed as inputs

FieldWhy it's not a knob
"Training runs / month" Real training I/O isn't sequential per run — it's continuous shuffled reads plus checkpoint flushes. Multiplying per-run load-time savings by runs/month would overstate Net ROI into implausible territory, so the model treats I/O as continuous and leaves this knob out.
"S3 egress fees" / "S3 storage rate" The model leads with the GPU-dollar conversation; comparing storage line items is a distractor at this stage of the pitch.
"Cold tier vs warm tier vs hot tier" sliders Too much detail for a 5-minute opener. Rolled into a single "Hot working set %" input. Engineers who want the breakdown can read the stage flash-freed fractions in the source.
"$/TB·mo for B2 Overdrive" as an input Hardcoded per tier — these are published Backblaze rates, not tunable inputs. The Standard rate is $6.95/TB·mo; Overdrive tiers are listed in §6.5.
"Per-GPU bandwidth need" as a separate knob Not needed — the bandwidth demand is the prefetch demand (dataset × epochs × 8 / run-time), computed from the editable Epochs per run and Run duration inputs in Advanced. Heavier I/O (shorter runs, more epochs) raises the demand directly; the Tier Comparison Capability column shows whether each tier keeps up.
A separate "Calculator" panel Used to have its own inputs duplicating Configure. Removed for clarity — one set of inputs drives one set of outputs.

6. The math — formulas and derivation

Every formula below is what the tool computes. Each is presented with a worked example using the Large tenant cluster · 25 PB profile (512 GPUs, 70B, 400 Gbps — the tier the math recommends) so the chain of reasoning is reproducible by hand.

The premise behind the math. Public sources establish that storage performance materially affects GPU economics: The math below converts these established storage-side levers into per-tier dollar impact for the operator's economic model. Defaults are positioned mid-range conservative; the Configure panel lets you tune every input to your fleet's measured telemetry.

6.1 — Goodput Recovered / mo

prefetch_demand C     = dataset_TB × 1000 × epochs × 8 / (run_days × 86400)   // Gbps, sustained
streamed_fraction     = cold (uncached) % of the dataset, streamed from B2 each pass
starvation(g)         = streamed_fraction × max(0, 1 − g / C)                 // GPU idle fraction
training_hrs_per_mo   = 720 × cluster_utilization                            // default 80% → 576 hrs
goodput_$_per_mo      = cluster_$_per_hr × (starvation(standard) − starvation(tier)) × training_hrs_per_mo

What it measures — keeping the GPUs fed

The dollar value of GPU time recovered by keeping the cluster fed with sharded training data. The cold/uncached portion of the dataset is streamed live from B2 during the run. Prefetching and sharding hide latency, not a bandwidth shortfall — so when a tier can't sustain the prefetch demand, the prefetch buffer drains and the GPUs starve waiting on cold reads. This is the read path — independent of checkpoint mode (async checkpointing does nothing to it). It's the core reason Overdrive matters for datasets too big to cache. Grounded in NVIDIA DGX SuperPOD storage targets (~4 GB/s/GPU read) and the tf.data / PyTorch fact that prefetch is latency-hiding only — verified 2026-06-06 (NVIDIA, tf.data).

Two read-path levers in this number. "Goodput recovered" bundles (1) steady-state feed (above) and (2) run-start rehydration — the GPU idle while cold shards are pulled into the hot tier at each run start (cold_GB = training_TB × 1000 × rehydrate%; hours saved = (cold_GB×8/Standard − cold_GB×8/tier)/3600 × runs/mo, capped at 30% of monthly training hours; × cluster $/hr). Both improve with bandwidth and are GPU time the operator's Overdrive recovers for the tenant — the same way the GenAI tenant tool counts them. Editable via Cold % rehydrated / run and Training runs / month in Advanced. (Checkpoint stall — the write path — is a separate tile.)

Worked example (Large tenant cluster · 25 PB on 400 Gbps Overdrive, 70% streamed)

prefetch_demand C   = 25,000 × 1000 × 2 × 8 / (14 × 86400) ≈ 331 Gbps
starvation(50)      = 0.70 × max(0, 1 − 50/331)  = 0.594   // Standard idles ~59%
starvation(400)     = 0.70 × max(0, 1 − 400/331) = 0       // 400 fully feeds the demand
cluster_$/hr        = 512 × $4 = $2,048
training_hrs_per_mo = 720 × 0.80 = 576
goodput             = 2,048 × (0.594 − 0) × 576
                    ≈ $700,900 / mo
The capability column shows min(1, tier_Gbps / prefetch_demand) — what fraction of the cluster's read demand a tier can serve, measured against the prefetch demand (dataset × epochs / run-time), the same basis as the ⚡ keeps-fed tier. A tier at or above the demand reads 100%; below it the GPUs starve by the shortfall.
Streamed (uncached) % is the swing input — and it's a modeled default (70%), not a measurement. Datasets that fit in local cache approach $0 goodput here (nothing is re-streamed); PB-scale video / multimodal that can't be cached sit near 100%. The brief shows Year-1 sensitivity across 50 / 70 / 90% streamed. Pilot-measure the true cold fraction before relying on this for budget.

6.2 — Checkpoint Stall Eliminated

Why checkpoint stall is a real economic lever, not a model artifact. Public sources confirm checkpoint overhead is one of the largest underclaimed sources of GPU idle time in production training: The default checkpoint sizes used in the math are from the MLCommons MLPerf Storage benchmark: 8B → 105 GB, 70B → 912 GB, 405B → 5.29 TB, 1T → 15 TB. The Configure panel exposes these as model-size anchor buttons.
stall_std_min          = (ckpt_GB × 8) / standard_gbps / 60
stall_od_min           = (ckpt_GB × 8) / tier_gbps     / 60
delta_min              = max(0, stall_std_min − stall_od_min)
ckpts_per_hr           = 60 / cadence_min
training_hrs_per_mo    = 720 × cluster_utilization              // default 80% → 576 hrs
hours_recovered_per_mo = (delta_min / 60) × ckpts_per_hr × training_hrs_per_mo × concurrent_jobs
stall_$_per_mo         = hours_recovered_per_mo × cluster_$_per_hr × ckpt_stall_multiplier
Synchronous vs async checkpointing. This lever assumes synchronous checkpointing — GPUs fully block while the checkpoint flushes — which is the optimistic upper bound. Many large shops checkpoint async / local-first (PyTorch async DCP, CheckFreq, Nebula): GPUs snapshot to local NVMe in seconds and resume while a background process uploads to B2, so GPU idle decouples from B2 bandwidth. PyTorch's own Distributed Checkpoint cuts checkpoint stall ~92% (18.5s → 1.5s on a Llama-3 405B run; verified 2026-06-06, PyTorch DCP). The Checkpoint write toggle in Advanced sets ckpt_stall_multiplier — 1 for synchronous, the editable async capture % (default 20%, residual upload-keep-up + crash recovery) for async. Unlike goodput, this lever is checkpoint-mode-dependent; lead with goodput + flash for async tenants.

What it measures

Every training run flushes checkpoints periodically. Under synchronous checkpointing the cluster stalls — every GPU waits on the write. Standard B2 at 50 Gbps takes about 2.1 minutes to flush an 800 GB checkpoint. Overdrive at 400 Gbps does it in ~16 seconds. That ~1.87 minute delta, repeated hourly across 4 concurrent training jobs over the 576 effective training hours, adds up (scaled by the checkpoint-write multiplier above).

Worked example (Large tenant cluster · 25 PB on 400 Gbps)

stall_std           = (800 × 8) / 50  / 60 = 2.13 min
stall_od            = (800 × 8) / 400 / 60 = 0.27 min
delta               = 1.87 min
ckpts/hr            = 60 / 60 = 1
training_hrs_per_mo = 720 × 0.80 = 576
hours/mo            = (1.87/60) × 1 × 576 × 4 jobs
                    = 71.7 hours
stall_$/mo          = 71.7 × $2,048
                    ≈ $146,801 / mo
Multi-tenant cluster modeling: the tool multiplies stall recovery by concurrent_jobs. In a multi-tenant cluster each job checkpoints on its own cadence, so the payoff scales linearly with the number of concurrent flushers. Single-tenant clusters should set concurrent_jobs = 1 in the Configure panel for an unmultiplied result.

6.3 — Flash capacity reclaimed

Hot/cold split is the lever VAST and WekaIO have monetized for years. Their published whitepapers and customer references repeatedly cite hot working sets in the 20–40% range of total training corpus depending on epoch size, shuffle pattern, and dataset scale. Cold tail data (60–80%) is the displaceable portion the Overdrive flash-replace % math operates on. The default 32% hot working set is mid-range conservative against those public ranges. Should be validated against your fleet's actual hot/cold telemetry — instrument with object-access timestamps to derive your real number.
dataset_PB     = pipeline_training_TB / 1,000     (decimal PB, matches spec §4)
hot_PB         = dataset_PB × hot_working_set% / 100
displaced_PB   = hot_PB × overdrive_replace% / 100 × tier_capability
displaced_TB   = displaced_PB × 1,000
flash_$        = displaced_TB × flash_$_per_TB_per_mo
b2_overdrive_$ = displaced_TB × tier_storage_rate
net_savings    = max(0, flash_$ − b2_overdrive_$)

What it measures

The amount of flash a customer no longer needs to provision because Overdrive serves the hot working set at line rate — reclaimed capacity the neocloud can resell. For the 60 PB Frontier multimodal tenant with an 18 PB hot tier (30%), Overdrive 800 serving 80% of that hot tier at 100% capability on a 768-GPU cluster reclaims about 14.4 PB. At $100/TB·mo flash and $29/TB·mo B2 Overdrive 800, that's ~$1.02M/mo of net storage savings.

Worked example (Large tenant cluster · 25 PB on 400 Gbps)

dataset       = 25,000 TB / 1,000 = 25 PB
hot           = 25 × 0.32        = 8 PB
displaced     = 8 × 0.80 × 0.78  = 5.0 PB (5,000 TB)
flash_$       = 5,000 × $100     = $500,000
b2_od_$       = 5,000 × $21      = $105,000
net           = $500,000 − $105,000
              ≈ $395,000 / mo

6.4 — New-capacity revenue / mo

freed_TB        = Σ STAGE_FLASH_FREED_TB[stage] for stage in completed_stages
margin_per_TB   = flash_sell_price − b2_storage_rate
revenue_$_per_mo = max(0, freed_TB × margin_per_TB)

What it measures

The neocloud's resale opportunity. As each demo stage offloads its data tier to B2, that flash capacity is freed up. The neocloud can resell that flash to a new GPU customer at the going rate (~$100/TB-mo) while paying B2 only the storage rate (~$6.95–$19/TB-mo depending on tier). The margin is recurring revenue.

Per-stage flash-freed assumptions

StageFlash freedWhat stays hot
Data Lake90% of training corpus~10% re-training cache
Checkpoints75% of checkpoint historyActive + last 2-3 checkpoints
Model Registry24% of model tierCurrently deployed model
RAG Knowledge7.5% of model tierVector index + embeddings
KV Cache8.25% of model tierHot KV blocks for inference
Exhaust70% of logs72-hour log window

These per-stage fractions are educated estimates. Validate against actual customer deployment patterns before externalizing as authoritative numbers.

6.5 — Tier Premium and Net ROI (Tier Comparison table)

storage_TB         = training_TB + checkpts_TB + models_TB + logs_TB
standard_$_per_mo  = storage_TB × $6.95
tier_$_per_mo      = max(tier_min_monthly_commit,
                         storage_TB × tier_storage_rate + tier_network_fee)
tier_premium       = max(0, tier_$_per_mo − standard_$_per_mo)

net_ROI_per_mo     = goodput_$ + stall_$ + net_flash_$ − tier_premium

This is what the "Auto-select optimal tier" button uses to pick the best tier. The winner is whichever Overdrive tier maximizes net ROI for the configured workload. If no Overdrive tier produces positive ROI (small workloads where the tier premium exceeds the gains), it recommends Standard B2.

Two recommendations: ★ best value and ⚡ keeps training fed

The tier strip marks two tiers, because the cheapest-ROI tier and the tier that fully feeds the cluster aren't always the same:

⚡ never marks a tier below ★ — if the best-value tier already covers the feed floor, the two combine into a single ★⚡ badge ("best value & keeps training fed"). ⚡ only appears as a separate, higher tier when the best-value pick would under-feed the cluster, so it always reads as the performance up-tier.

Why this matters across the two tools: both calculators recommend the feed floor — the smallest tier that fully streams the dataset on schedule — using the same prefetch formula, so they recommend the same tier for the same workload (5 PB → 100, 10 PB → 200, 25 PB → 400, 60 PB → 800 in both). The operator and tenant frame the value differently (the operator as capacity reclaimed for resale → growth, the tenant as their own spend), but the recommended throughput tier is identical.

6.6 — The bandwidth-capability cap

prefetch_demand_gbps = dataset_TB × 1000 × epochs × 8 / (run_days × 86400)
tier_capability      = min(1, tier_gbps / prefetch_demand_gbps)

An Overdrive tier that can't sustain the prefetch demand can't fully feed the cluster. The capability cap is the fraction of that demand a tier serves — measured against the prefetch demand (dataset × epochs / run-time), the same basis as the ⚡ keeps-fed tier and the feed-driven goodput — so the Capability column, the recommendation, and the goodput all agree. It prorates Goodput Recovered and Flash Displaced.

Examples (25 PB profile, prefetch demand ≈ 331 Gbps):

Same basis as the GenAI tool. Both calculators derive bandwidth demand from the training pipeline — working_set × epochs × 8 / run_duration_seconds, the sustained throughput needed to stream the dataset on schedule. So both compute the same capability, recommend the same feed-floor tier, and produce the same goodput for the same workload. Grounded in NVIDIA DGX SuperPOD feed targets (~4 GB/s/GPU read) + the tf.data/PyTorch fact that prefetch is latency-hiding only — verified 2026-06-06.

7. Assumptions to validate against your workload

The numbers in this tool are conservative defaults pending pilot validation. Before relying on any of these numbers for a real budget decision, validate the inputs below against your own telemetry. Every Configure field can be tuned to match your environment.
AssumptionDefaultNotesHow to validate
Cluster utilization %80%Production-reservation typical Fraction of the month the cluster is actively training. 720 hr/mo × 80% = 576 effective training hours. Production reservations often run higher (90%+) on reserved capacity; research clusters can run lower. The math multiplies goodput recovered + stall recovery by this — overstating utilization overstates monthly impact linearly.
Standard I/O wait %4.0%Working default Pull cluster utilization telemetry from a representative training run. GPU SM efficiency in nvidia-smi or DCGM gives the inverse — the fraction of time GPUs were doing useful compute.
Overdrive I/O wait %0.8%Working default Same approach in an environment with sustained 200+ Gbps storage throughput. If you don't yet have one, treat this as the modeled outcome and re-measure after a pilot.
Prefetch demanddataset × epochs × 8 / run-timeFrom editable inputs Shorter runs or more epochs raise it; the Capability column shows each tier's coverage. NVIDIA DGX SuperPOD targets ~4 GB/s/GPU read at the upper bound.
Flash $/TB-mo (sell price basis)$100Mid-market reference Use your own flash sell price or internal chargeback rate. Common range is $80–$150.
Overdrive flash-replace %80%Modeled outcome This is a workload claim, not a pricing input. Validate by piloting Overdrive against a representative hot working set.
Stage flash-freed fractions90 / 75 / 40 / 30 / 55 / 70%Educated estimates from typical AI pipeline patterns Compare against your retention policy. Numbers vary by training cadence and how much old data stays warm.
QLC flash hardware cost$490/TBPublic reference (2026 Q1, 30.72 TB drives in volume) Use your own quote. Sustained-volume agreements often come in lower.

Where the I/O wait baselines come from

The 4.0% (Standard) and 0.8% (Overdrive) defaults sit in the middle of public ranges reported by storage and ML infrastructure vendors. They are not measurements of your specific workload — they describe what well-tuned vs un-tuned training-storage stacks typically look like in the industry.

SourceWhat it reports
AWS S3 Mountpoint & S3 Express One Zone performance documentationCold-stream I/O wait on training workloads in the high-single-digit range without local NVMe staging. Throughput improves dramatically with prefetch + parallel range reads (the pattern this tool models).
MLPerf Storage benchmark (MLCommons)Uses I/O wait as the proxy for "good vs bad storage." The 2–8% envelope separates well-tuned from un-tuned object-storage stacks across published submissions.
VAST Data & WekaIO public whitepapers + production telemetryCite measured I/O wait reductions on un-tuned object-storage backed training in the 3–10% band depending on shard pattern.
Hyperscaler + academic telemetry (AWS reports, Azure reports, Meta SPEC paper, Google Pathways)I/O wait can swing 1–15% depending on dataset size, shard pattern, dataloader concurrency, and checkpoint cadence. Azure specifically reports checkpoint overhead averaging 12% of total training time, up to 43% in some large-model scenarios.
What we don't have: Backblaze does not yet publish a customer-pilot data point of measured baseline I/O wait. The defaults are positioned in the middle of the public envelope as conservative starting points, but the only way to validate the number for your specific workload is to instrument your training run. The Configure panel is where you override the defaults with measured data when available.

8. FAQ

Where do the 4% / 0.8% I/O wait numbers come from?

They are working defaults pending validation against telemetry from a representative training run. The "Where the I/O wait baselines come from" subsection in §7 above cites the public source ranges. To produce numbers you can budget against, take a measurement of your own cluster (GPU SM efficiency from nvidia-smi or DCGM gives the inverse of I/O wait %) and update the Configure fields accordingly.

Why do the ★ best-value and ⚡ keeps-fed tiers sometimes differ?

The recommended ★⚡ tier is the feed floor: the smallest tier that fully streams your dataset on schedule (prefetch demand). Best-value and keeps-fed are the same tier — feeding below it starves the GPUs, and paying above it buys feed you already have — so every profile shows a combined ★⚡ (5 PB → 100, 10 PB → 200, 25 PB → 400, 60 PB → 800). A higher tier is always available as a manual click for faster checkpoints/rehydration. See §6.5.

Does the GenAI tool recommend the same tier as this one?

Yes. Both calculators recommend the feed floor — the smallest tier that fully streams the dataset on schedule — using the same prefetch formula, so they recommend the same tier for the same workload (5 PB → 100, 10 PB → 200, 25 PB → 400, 60 PB → 800 in both). The operator and tenant frame the value differently (the operator as flash capacity reclaimed for resale → growth, the tenant as their own spend), but the recommended throughput tier is identical.

Why isn't there an object-storage vs object-storage comparison?

The headline math here measures GPU-productivity impact, not storage line-item savings. A 5% improvement on a multi-hundred-thousand-dollar monthly GPU bill is worth more than even a large percentage discount on a storage line item. The tool is intentionally focused on the GPU-economic story.

What if my workload's training I/O is heavier?

The bandwidth demand is the prefetch demanddataset × epochs × 8 / run-time — derived from the editable Epochs per run and Run duration inputs. Heavier I/O shows up as shorter runs or more epochs, which raise the demand directly; raise the streamed (uncached) % too if more of the dataset must come from B2 each pass. The Tier Comparison "Capability" column shows whether a given tier can keep up with the resulting demand.

The demo only uploads ~200 MB. Why are you projecting petabyte numbers?

The live stage runs demonstrate the patterns end-to-end — multipart uploads, parallel range reads, mid-run checkpointing, model registry, inference cold-start — using small data so the demo finishes in under a minute. The dollar projections run against the configured workload (the workload profile preset). Each completed stage card explicitly says "↓ Projects to X PB freed" to flag the scale shift.

Why does the page open with Configure expanded?

So you can see — and change — every input that drives the headline numbers. Nothing is hidden behind a magic button. If you want a cleaner view, click the Configure summary to collapse it; the hero tiles stay populated.

Can you explain the Capability column in the Tier Comparison table?

It's the fraction of the cluster's read demand the tier can actually serve. Yellow (under 100%) = the tier is undersized for this cluster; goodput and flash benefits are prorated. Green (100%) = the tier saturates the workload; full benefit. The optimizer prefers the smallest 100%-capable tier because anything larger adds tier premium without adding benefit.

Why doesn't the Goodput tile change when I switch between Overdrive tiers?

It does — but it only changes when the bandwidth capability changes. Once a tier is big enough to saturate the cluster's read demand (cluster GPUs × 1 Gbps), every larger tier delivers the same goodput recovery at 100% capability; smaller tiers show lower goodput because they're undersized and prorated. This is the model behaving honestly — once a tier feeds the cluster, more bandwidth doesn't recover more goodput.

What does "Standard rates — recommend Overdrive to recover goodput" mean? Is Standard bad?

No. For small workloads Standard is the right answer (the Fine-tune lab · 1 PB profile is built around exactly that case). The note is just saying Standard is the baseline by definition, so there's no delta to recover when comparing it to itself. At Overdrive tiers the math measures the delta from the Standard baseline.

What does the Clean up button do? Is it safe?

Cleanup deletes only objects under the configured DEMO_PREFIX (default goodput-demo/). The confirmation prompt names the exact bucket and prefix being wiped. Other prefixes in the same bucket are untouched. Even so — use a dedicated demo bucket if you're running this against your own credentials.


Live demo: https://genai.backblazedemos.xyz/goodput/
Author: Kevin Lott · klott@backblaze.com