Media endurance · write-amplification study

How much do our disks actually wear?

A distributed home-hosted cloud fragments every file, scrubs daily, and self-heals on churn. A fair question from a hardware reviewer: does all that activity burn out the storage media before their natural end of life? Here is our honest model — and an interactive simulator so you can turn the dials yourself.

Three findings, up front

Write-wear is not the limiting factor.

For a memories workload (write-once, read-many), the dominant failure mode is random hardware mortality, not worn-out cells. Three results drive that conclusion:

1 · Reads don't wear flash
The daily integrity scrub (hash verification) and client access are pure reads. Their contribution to flash wear is ≈ 0, so they leave the write budget entirely.
2 · Cruise writes are churn-dominated
Initial fill is a one-shot cost. In steady state, rewrites come almost entirely from repairing fragments after a node leaves permanently. A node that switches off at night and back on in the morning writes nothing — a repair timeout absorbs transient absences. The rule of thumb: writes/day ≈ permanent-churn-rate × capacity × write-amplification.
3 · An HDD doesn't wear on write at all
Magnetic platters have no program/erase cycles. For an HDD the wear vector moves to a mechanical one — head-parking cycles (SMART attribute 193, ceiling ≈ 600,000) — which is a firmware/config problem, not a physics limit.
Turn the dials

Interactive endurance model

An analytical model (order-of-magnitude, not an event-level simulation). Pick a medium, set the workload, and watch the wear over your horizon. For an HDD, watch what happens when you push head-parking from a healthy 24/day toward an aggressive 2,000/day.

Usable capacity1 TB
Fragment size4 MB
Write amplification (WAF)2.0×
RS 8+4 erasure coding ≈ 1.5×; with replicated keys/index ≈ 2×.
Permanent churn0.20 %/day
Definitive node departures only, not nightly power cycles.
Horizon5 years
Rated endurance (TBW)400 TB
Good SSD 1 TB: 300–600 TB. Cheap USB key: 30–50 TB.
Initial fill (one-shot)
262,000
fragments = writes, once, spread over the fill phase
Steady state (recurring)
4.0 GB/day
≈ 1,000 fragments / day
Flash wear over 5 years
3.6 %
7.3 TB written / 400 TB rated
The design consequence

The controller lives in the network, not the disk.

An SSD's on-board controller does wear-leveling (useless on an HDD), bad-block remapping (an HDD already does this), ECC and caching. The valuable intelligence can also live above the disk — and in blob it already does.

Bad-block management
Reed-Solomon 8+4 — the same idea as remapping dead sectors, but at the scale of whole disks. A lost fragment is reconstructed from the survivors.
Silent-corruption detection
Daily hash scrub — ECC, but at the file level rather than the sector.
Integrity & placement
Cross-node dispersion (≤1 fragment/node) + proof-of-retention — distributed integrity, continuously auditable.
Why this matters for sourcing
A node's disk can be simple and cheap because the network is intelligent. A dumb, reliable disk + erasure coding + scrubbing + dispersion is collectively more resilient than a single expensive SSD in isolation (one SSD dies → all lost; one blob disk dies → rebuilt). The one thing software cannot fix is an HDD's mechanical power draw — which is why our node design still measures energy at the wall.
Where this is weak

Limits of the model

Order of magnitude
Analytical, not event-level. It captures steady-state averages, not the burst structure of correlated node failures.
Redundancy as one factor
Amplification is folded into a single WAF multiplier; the exact replication-vs-dispersion split for keys and index is still being characterized.
Single-disk view
This models one disk. The fleet aggregate (N clients → number of nodes, aggregate writes) is future work.
Nameplate figures
TBW and load-cycle ceilings are typical vendor values; we intend to confirm them against measured SMART data on real reconditioned hardware.

If you have run endurance studies on reconditioned flash or on distributed erasure-coded stores, we would genuinely value a critique of this model.

Tell us where this is naive.

The PoC is real and the model is deliberately honest about its limits. Fifteen minutes of a researcher's time — a critique, a pointer, a collaboration — would go a long way.