Skip to download

hesela.dev

Storage glossary · datasets · llms.txt · hesela.com

Synthetic trace / no device measurements

Deduplication restore locality: a six-request trace

Six synthetic cold-cache LRU scenarios isolate repeated container reads from unused fetched bytes. No device timings or real backup records.

Download CSV6 scenarios / 13 columns

JSON and provenance / Model code / Research explanation

Read accounting

Process each six-request trace in order from an empty LRU container cache. A miss fetches one full 65536-byte container; a hit refreshes recency. Each request emits 4096 bytes. Count misses and multiply by container size, then divide by total emitted bytes. Three cache capacities are evaluated for each of two placements.

Same six distinct 4 KiB chunks and three 64 KiB containers in both placements. Only the mapping from logical order to containers changes. 1 KiB = 1,024 bytes.
LayoutTraceCache slotsReadsFetched KiBRestored KiBByte ratio
clusteredAABBCC13192248x
clusteredAABBCC23192248x
clusteredAABBCC33192248x
interleavedABCABC163842416x
interleavedABCABC263842416x
interleavedABCABC33192248x

Reproduce

The downloadable module needs only Node.js. This command prints the identical CSV; no network, device access, random seed or external dependency is needed.

node --input-type=module -e 'import { toCsv } from "./dedup-restore.mjs"; process.stdout.write(toCsv())'

JSON includes all rows, input units and SHA-256 hashes of the model and CSV. Generation time records artifact creation, not an observation period. With two slots, interleaving causes six misses instead of three; with three slots, both placements fetch 192 KiB to output 24 KiB. Fetched-byte ratios are not elapsed-time multipliers.

Limits

Column definitions

layout
Synthetic placement: clustered or interleaved, not a measured system.
trace
Container IDs for requests to six distinct chunks in fixed logical order.
cacheContainers
LRU cache capacity in whole containers; cache starts empty.
cacheBytes
Capacity in bytes, excluding cache metadata.
requests
Requested equal-size chunks; six in each example.
uniqueContainers
Distinct containers referenced, not total reads.
chunkBytes
Bytes restored per request; illustrative fixed-size input.
containerBytes
Bytes fetched per miss; full-container read assumption.
containerReads
Cache misses, each causing one whole-container fetch.
cacheHits
Requests served from the modeled container cache.
restoredBytes
Total logical output bytes; denominator of read amplification.
fetchedBytes
Total fetched container bytes; includes unused data and rereads.
readAmplification
Fetched bytes / restored bytes, dimensionless; not a time ratio.

Sources and license

CC0-1.0 for Hesela's original model and rows. Cited publications retain their own rights.

  1. Mark Lillibridge, Kave Eshghi and Deepavali Bhagwat. Improving Restore Speed for Backup Systems that Use Inline Chunk-Based Deduplication. USENIX FAST 2013, pp. 183-197; sections 2-4. (accessed 2026-10-11)
  2. Min Fu, Dan Feng, Yu Hua, Xubin He, Zuoning Chen, Wen Xia, Fangting Huang and Qing Liu. Accelerating Restore and Garbage Collection in Deduplication-based Backup Systems via Exploiting Historical Information. USENIX ATC 2014, pp. 181-192; sections 3-6. (accessed 2026-10-11)

Definitions: deduplication and chunk fragmentation.