Storage

Compaction and collection

Merging small segments and reclaiming superseded blocks with resumable background work.

Writes append. A table written in many small batches ends up as many small segments, and every update or delete leaves the blocks it superseded on disk until something reclaims them. Two background jobs handle both, and both are stepped and resumable — a browser tab can be backgrounded, throttled, or closed mid-job.

Compaction

Merges small segments into larger ones, which is what keeps a scan from paying per-segment overhead after a long run of small writes.

await db.compactTable("orders");

It runs automatically by default: a scan or a run of commits that finds a table past the threshold plans a fold and drives it to publication in small steps, yielding between them, and keeps going while the table is still due. A table is due when it has many small segments, many delete, update, or upsert segments, or deltas that hold as many rows as the data they change (once they pass a few thousand rows) or more patched rows than a query keeps in a quarter of its memory budget. A single such delta is enough. That rule is for tables refreshed wholesale: each full-table upsert adds another copy of every row, and folding it right away, with no further write needed, keeps a scan at about two reads per row instead of more than thirty. A query that finds a delta too large for its budget also queues the fold. A step yields to the event loop or to this database's next commit, whichever comes first, so a loop that awaits one statement after another without pausing still lets the fold advance. A fold reads and decodes its source blocks in groups sized by its memory budget, so a step costs the loop a turn per group rather than a turn per block. A finished fold publishes as a writer: it takes a turn of its own in the database's writer queue, and a statement that gets its turn first lands the fold inside that turn before its own work, so the fold never races a write for the manifest.

Setting autoCompact: false transfers that responsibility to the caller. Small-write segments then accumulate, so scans become progressively slower and writes eventually meet the fixed segment-history ceiling unless the application drives compactTable or compactTableStep. To drive it yourself with a slice per idle callback. Use a worker to keep encoding off the main thread:

const db = new MinnowDatabase(store, { autoCompact: false });

requestIdleCallback(async function step() {
  const progress = await db.compactTableStep("orders", { maxBlocks: 64 });
  if (progress.result === null) requestIdleCallback(step);
});

Each step processes at most maxBlocks output blocks and checkpoints; progress.result stays null until the job publishes.

Deletes and updates before compaction

Compaction is not what makes a mutated table readable at speed. A query applies the table's deltas over its appended data directly, so a table that has been deleted from or updated answers from the same scan every other table gets, plus the cost of the deltas themselves. What compaction adds is returning the table to a plain append — the deltas stop being re-read on every query, and the storage they occupy is freed.

Upserts replay through the same streamed scan. Repeated writes to one key keep only the patches needed for its latest column values. Replay reads historical key blocks in bounded groups and keeps a bitmap with one bit per historical insert, base, or upsert row plus about eight bytes per patched row, built once per commit and shared by the queries after it. Temporary patch data is released after each scan window. Compaction still reduces the history that must be read and the bitmap's size.

A delta of any size stays readable. When a query's patches would take more than a quarter of its execution budget, the query keeps only the bitmap and replays the patches one range of rows at a time as its scan reaches each range. When even the replay's temporary key lookup would not fit, it splits the touched keys into groups and replays each group in turn. Both paths return exactly the rows the in-memory replay returns, in the same order — they cost time, not correctness — and both queue the fold that removes the delta. A reload alone preserves the stored mutation history.

Partitioned folds

A folded table is a run of level-one partitions of at most partitionRows rows (16,384 by default) in the order the rows were written. On a keyed table, a fold rewrites only the partitions its deltas touch — the ones holding a key some delete, update, or upsert names — plus the last partition while new rows are still joining it. Every other partition is left exactly as it is, so absorbing a run of point updates costs the partitions they land in, not the table: thirty-two updates rewrite at most thirty-two partitions, however many rows the table holds. Rows keep their hidden row IDs and their order across every fold.

An append-only table uses the same partition shape through the rechunk path. Its old partitions never need rewriting; only a partial tail joins the next batch. Partition descriptors use stable fractional logical orders between their unchanged neighbours, so an oversized partition written by an older Minnow version is split the next time a fold includes it instead of being left large because adjacent commit versions have no spare integer between them.

Smaller partitions make folds cheaper and give a scan more blocks to walk; the default suits tables from thousands to millions of rows. Set it per call, or for every fold the database plans — background ones included — when you know better:

await db.compactTable("orders", { partitionRows: 4096 });
const db = new MinnowDatabase(store, { compaction: { partitionRows: 4096 } });

Merging a keyed table plans the merge in memory proportional to the distinct keys its deltas touch. The table's size, how many times those keys were rewritten, and how wide the rows are do not add to it, so eighty refreshes of the same seven thousand rows cost seven thousand keys. Every output block is measured against a memory budget before it is written. The default budget is 32 MiB.

Automatic compaction needs no setting. It sizes each fold to the budget: a fold that does not fit is cut to fewer level-zero segments and planned again, and the table's next fold starts from the size that fit. When even the smallest fold the table allows does not fit, automatic compaction gives that one fold the memory it needs — about what the writes that produced it already held — rather than leaving the table to slow every read. compactTable without a memoryBudgetBytes option fits its fold the same way. A compactTable given an explicit budget, and every compactTableStep, plans exactly what its options ask for and reports CompactionMemoryBudgetError when that needs more:

await db.compactTable("orders", { memoryBudgetBytes: 256 * 1024 * 1024 });

The order rows arrive in does not matter. A fold's job record holds its sources, output windows, and partitions — not which source row feeds each output cell. That mapping, the replay, is a pure function of the immutable sources: the engine that plans a fold keeps it in memory as compact runs, and any engine that resumes the fold — another tab, the same tab after a reload — recomputes it and checks it against a checksum the record carries before writing a block. A refresh written in an unrelated order therefore folds as quickly as one in the table's own order, and the record every step rewrites stays small. Jobs written by Minnow 0.12 and earlier keep their stored ranges and resume as planned.

Responsiveness

Background work never holds the thread for long. Folds, garbage collection, index and full-text builds, OPFS checkpoints, and live-query updates run in slices of a few milliseconds and hand the event loop a turn between them, so a query or write that arrives in the middle waits for one slice, not for the job. Long statements are sliced the same way: a scan yields between batches, a large write yields between its passes over the rows, and the first lookup through a new index yields while it hashes each block's keys. One clock paces all of it, so a statement made of many short steps, or a run of statements awaited back to back, yields as often as one long loop.

On a laptop, the longest single block measured during background work is about 35 ms, on the memory and OPFS stores alike, for every shape in scripts/stall-survey.mts: wide tables refreshed in any order, hundreds of thousands of rows of point updates or deletes, index and full-text builds, and live aggregates under heavy writes. Writes are sliced the same way. A large one checks its keys and adds them to their memberships behind a mask that keeps them invisible, checks its index changes, checksums its blocks, and on OPFS encodes its log frame and writes all but its last piece, a slice at a time, before it commits; committing then drops the mask, at no cost per key. Upserting keys that already exist adds nothing to publish. Writing 500,000 rows in one call and then upserting all of them holds the thread for at most about 21 ms on the memory store and 22 ms on OPFS. A query on OPFS does not wait for that commit either: reader leases are logged between its slices, and index lookups read without the leader's queue. A single write of 2,000,000 rows into a keyed table with two indexes, and an upsert of all of them, succeed on every store; the longest blocks are then the JavaScript engine's own garbage collection of millions of rows, up to about 200 ms in Node. On an OPFS database of a million keyed rows, the longest block while loading it 50,000 rows at a time is under 80 ms, creating a UNIQUE index over all of it never holds the thread for more than about 30 ms, and opening it again holds it for under 40 ms at a time.

An index change larger than one stored chunk — a write of more than 65,536 distinct values into one index — is folded into the index's base right after its commit instead of after more commits. Until it is, an OPFS checkpoint can outgrow its 256 MiB slot; such a checkpoint stops as soon as it passes the slot and is tried again once the fold has pruned the change, while the log keeps everything durable.

Automatic compaction backs off a table whose attempt does not help, rather than retrying it on every query. A fold that fails waits a delay that doubles from a quarter of a second per consecutive failure up to a minute, then runs again whether or not anything was written. A fold compaction declines because the table's layout or keys put it out of reach also waits for the table to reach twice its segment count, capped at the largest level-zero prefix one fold consumes. A threshold crossed while another fold is running is queued. Writes are sampled during a burst and its final tail is checked after a short quiet period, so the last writes cannot leave a table permanently above the fold threshold.

Jobs are records in the store, so listCompactionJobs() finds one a previous session left behind and resumeCompactionJob(jobId) picks it up. A job whose owner lease has expired — a tab that closed mid-fold — is taken over by whoever resumes it, background maintenance included: the dead owner is aborted and the job continues under a fresh one. cancelCompactionJob(jobId) stops one cleanly — compaction is visible-data-neutral by construction, so cancelling it can never lose data. The store atomically admits only one non-terminal compaction for a table, so two tabs racing the plan step converge on the same durable job instead of creating parallel roots. Planning reads stay outside the writer turn; job registration takes a short turn alongside DDL. A resumer reconciles replaced sources only after confirming that the durable job advanced.

Background compaction reselects active work after each yield. If another tab abandons a fold after a schema change, the next step plans against the current schema. A job abandoned between selection and execution reports a CompactionJobConflictError with its ID and changed revision. Explicitly resuming an already aborted job still reports its recorded failure; unexpected I/O and corruption remain errors. Resuming a ready job preserves its publication readiness. Automatic folds retry schema conflicts within the write retry limit and stop cleanly when another tab drops the table or cancels the job. Unexpected I/O failures still reach the caller or background diagnostic hook, including failures that coincide with another tab publishing or cancelling the job, or arrive after a job transition was persisted. A failure while recording the original error is reported separately. Before opening a write snapshot, the writer assists one pending compaction claim at each entry point. Newly arriving jobs cannot keep that entry point draining indefinitely. A claim that uses the writer's turn cancels its redundant queue wait; its actual storage work keeps the turn until completion. Backpressure can still drive the steps needed to keep the table below its hard limit. IndexedDB validates visible segment owners in groups of at most 128 reads within the publishing transaction. Every owner is still checked, and multiple segments owned by one transaction count separately toward the level-zero ceiling. A corrupt or missing owner aborts publication atomically. Segment discovery reads at most 1,024 records from the selected table's index. A complete smaller partition avoids one cursor turn per segment; a full batch falls back to the exact cursor path. Block visibility is still checked through fresh point reads in the same publishing transaction.

Level-zero growth is guarded on the write path in two tiers. Past 512 level-zero segments — twice the prefix one background fold absorbs — a write drives one fold step before its own commit and then proceeds, so a loop that adds a segment per statement faster than the fold retires them lends the fold its own turn: each statement is delayed by one step until the fold publishes, and the backlog stays within one fold's steps of that threshold. A table may retain at most 4,096 level-zero segments; a write at that ceiling takes the same step and is refused with CompactionBacklogError only if the table is still at the ceiling afterwards. A step that fails or cannot help backs the table off as a background fold would, so a table that cannot be folded costs one attempt per backoff period, not one per write.

Garbage collection

Reclaims blocks and segments no live version references any more, prunes old manifests, and removes terminal transaction records after nothing still needs them. Committed records wait for their manifest and segments; aborted records wait for their pending blocks and segments:

await db.collectGarbage();

A record is only collectable when no manifest, no open reader lease, no segment, and no non-terminal job still roots it. That is what makes collection safe while a report is being read: an open snapshot scope pins its version, and collection skips everything that version needs.

Isolation does not eliminate storage contention. IndexedDB collection validates retained records and deletes candidates in one read/write transaction; foreground operations can wait for it. Validation uses bounded batches and reuses ownership information within that transaction, but a candidate-count limit is not a wall-clock deadline. See IndexedDB collection cost.

The completed pass also removes expired reader leases, abandoned query-spill owners and pages, old manifest summaries below the earliest readable checkpoint, and finished maintenance-job records beyond the small diagnostic tail. Retired block intervals live in a separate paged provenance index until collection removes each payload and its provenance atomically. Summary pruning therefore cannot lose a large cleanup's durable discovery cursor. These are record bounds, not just block-byte bounds: normal commits, cancelled folds, and crashed readers do not leave one permanent metadata record per event. Segment discovery reads bounded listSegmentPage() windows directly, so reclaiming old segment and transaction records does not depend on keeping an unbounded compaction-job history. Segment discovery reads at most 64 records per page, batches owner lookups, and persists the last visited ID once per page. A page never exceeds its remaining candidate budget. A retained segment record protects every block it references, including retired history. Collection removes the segment before its blocks, or removes both in the same atomic step. A block nominated before its segment is retained for a later pass, keeping recovery valid between cleanup steps. New discovery cycles visit segments before retired blocks to avoid spending an extra cycle on blocks still protected by segment metadata. Existing durable discovery cursors remain resumable. Concurrent background collectors retry a job conflict even if another collector has already finished and removed that job. This normal race does not report a worker error or mark maintenance as failed; other storage failures still do.

Only one planned/running collection job may exist. Its durable cursor advances across the ID-ordered retired-provenance and segment pages, so a large obsolete history is read once across bounded steps and reopen—not rescanned from its first ID on every step.

It runs automatically by default. autoCollect is independent of autoCompact: disabling folds must not also disable reclamation of manifests, expired leases, abandoned transactions, spill, or job records. A pass follows every background fold, runs every 64 commits, and runs a minute after open or the last commit so a reopened, read-only tab still shrinks. A large backlog schedules another bounded pass until its durable discovery cursor finishes instead of stopping at one internal batch. After a complete discovery cycle, collection waits for the next commit, compaction, or idle trigger; it does not repeatedly scan retained history merely because the last cycle reclaimed something. A continuation keeps its destination while paging retained manifests, so a long history cannot restart completed segment discovery and starve later cleanup phases. Background passes keep the 64 most recent versions readable for up to a minute, so query({ version }) on a version you were just handed still finds it. An explicit collectGarbage() reclaims everything no lease or job pins; pass retainRecentVersions to keep a window. It first reconciles compaction jobs the way a background pass does — an owner past its lease is aborted and its job resumed or cancelled — so an explicit call also repairs a fold a dead owner left behind.

Collection failures are visible and bounded. db.maintenanceStatus() reports pending commit debt, whether a pass is running or queued, retry time, consecutive failures, and the last error. Automatic compaction checks and jobs, lazy full-text and secondary index builds, and posting tail folds also report failures through onBackgroundError. Without a handler, the direct engine writes the error and operation context to console.error; in a worker these reports reach onWorkerError. A throwing diagnostic callback is logged and contained. Scan fallbacks remain available when an optional index build fails. maintenanceStatus().backgroundFailureCount counts all reported background failures; backgroundErrors keeps the last 32 in reporting order, with sequence, context, name, message and time. Each text field is capped at 1,024 characters. The snapshot is defensive and stays available after close. Observers still receive the original error. Collection-specific counters and lastError describe collection only. Foreground failures reject their owning call; a successful fallback still reports the unexpected accelerator failure. Failures retry on a capped timer even when no more writes arrive. Debt resets after any pass that reclaims something. If repeated failures let it reach the safety ceiling, the next write first assists collection and, if that pass still reclaims nothing, throws MaintenanceBacklogError; already committed data is unchanged.

Active writers are durable but not immortal. A one-stage explicit scope starts with a renewable reader lease on its pre-scope manifest, so compaction and collection cannot reclaim that snapshot while its only staged batch is still process-local. When staged artifacts exceed one local batch and make the journal durable, ownership passes without a gap to the transaction record. The live scope renews whichever owner is current, including while an awaited write() callback or SQL transaction is idle. A deadline that has passed cannot be resurrected. Collection atomically races record renewal against expiry-abort, then reclaims the aborted journal and its staged blocks; an expired pre-journal lease is ordinary reader-lease garbage. Artifacts left by a killed tab therefore need neither a later write nor manual recovery.

Posting accelerators have their own growth fuse. Full-text and scalar secondary indexes fold their per-commit delta tail back into a base in the background. If a persistent storage failure prevents rebuilding, Minnow first marks the accelerator invalid (so every reader scans), then removes its base and deltas before the tail can grow without limit. Query results stay exact and a later relevant query rebuilds the accelerator after storage recovers. A UNIQUE index keeps its separate fail-closed membership set throughout; losing acceleration never weakens enforcement.

The same stepped shape applies:

let progress = await db.collectGarbageStep({ maxItems: 128 });
while (progress.result === null) {
  progress = await db.resumeGarbageCollectionJob(progress.jobId, { maxItems: 128 });
}

Public maintenance steps accept at most 1,024 planning or work items per call. Larger cleanups use the cursor and another call, keeping adapter transactions, WAL frames, and response arrays bounded.

Scheduling

With autoCompact and autoCollect on, nothing needs scheduling: a scan or a run of commits starts a fold, a fold or a quiet minute starts a collection pass, and both run in small steps that yield between them. To drive them yourself instead, turn both off and use the stepped calls:

  • After a bulk load, compact the tables it touched.
  • On an idle callback, take one step of whatever is outstanding.
  • At startup, resume jobs a previous session left and call cleanupQuerySpill() if automatic collection is off.

Leaving them undone costs storage and scan speed. It never changes an answer — a database that is not compacted or collected still reads correctly over more blocks and bytes — but new growth is eventually refused at the fixed segment, history, or maintenance-debt safety ceiling. Run or restore maintenance rather than treating those fuses as an operating target.

A loop of single-row UPDATEs that never awaits anything else between statements is the hardest shape for background folds, because each statement adds a level-zero segment and gives maintenance one turn. Grouped block reads and the write-path step above keep it bounded: on a folded 20,000-row table, 6,000 such statements hold between about 1.4 and 3.6 ms each with the level-zero count never past 512, where before 0.7.6 latency climbed to 18 ms and the loop was refused at the ceiling by its 4,200th statement. A bulk update through updateBatch() or UPDATE … FROM a VALUES source is still cheaper: it lands as one segment.

Background work can lose its build ownership or job revision to maintenance in another tab. Those typed refusals also reach the diagnostic hook; they do not imply that an acknowledged application write was lost. Index readers use the correct scan fallback while the accelerator is rebuilt. An I/O or corruption failure remains a failure and is reported through the same hook.

On this page