Reference

Architecture

How Minnow works and why it's built this way.

Minnow is built for how browsers actually behave: tabs close without warning, storage is slow, and the same app may be open in five tabs at once. Everything below follows from that.

The short version: data lives in compressed column blocks inside IndexedDB or OPFS. Each read sees one stable version, and queries work through the data in batches.

The ideas that guide it

Everything else is built around these ideas.

  • The durable store is the source of truth. BroadcastChannel, Web Locks, and page lifecycle events can all fail silently. Minnow uses them to speed things up, never for correctness. Correct behavior always rests on a committed storage write — an IndexedDB transaction, or a checksummed entry in the OPFS store's command log.
  • A committed write is the visibility boundary. Durable adapters acknowledge only after their atomic storage commit. Their default strict mode also requests or performs the final flush; relaxed keeps atomicity but explicitly accepts that power loss can remove an acknowledged suffix. Origin eviction and deliberate site-data clearing remain separate risks, so data that must survive those events still needs an independent copy.
  • Where the engine runs is your choice. Put it in a worker and hold an async proxy in the page, or construct it on the main thread. The API is async everywhere, so your code — and any adapter written against it — looks the same either way.
  • Published blocks never change. Once written, a block stays as it is. That makes retries, stable reads, and cleanup much easier to reason about.
  • Nothing special to deploy. No COOP/COEP headers, no SharedArrayBuffer, no WASM file to host. npm install and go.
  • The engine is designed for browsers. Minnow contains its own parser, planner, optimizer, and executor instead of wrapping SQLite or DuckDB. Its SQL dialect is a documented embedded subset, tracked in the PostgreSQL compatibility page.

Why columns, and why blocks never change

Browser storage charges a lot per operation and little per byte — IndexedDB per transaction, OPFS per file open. So Minnow stores a few large values instead of many small ones:

  • Data is packed into compressed columnar blocks — one column for a group of rows, roughly a megabyte before compression.
  • No row is ever its own storage entry, and no table is one giant entry.
  • Each block carries separate checksums over its stored payload, logical uncompressed payload, and header plus statistics. Recovery can reject damaged compressed bytes without decompressing them, while queries can trust min/max statistics and skip blocks that cannot match a filter. A declared scalar or composite secondary index adds a durable candidate map when block statistics are not selective enough.

Columns beat rows here because most reads touch a few columns across many rows: filter, aggregate, scan. Similar values sit together, so they compress well, and queries fetch only the columns they use.

gzip is adaptive per column and per physical output. Inputs below 4 KiB stay raw; larger probes keep gzip only when it compresses at least 1.2 to 1, and an unhelpful column is re-probed every 32 blocks. The configured codec is therefore a preference: each self-describing block records the codec it actually used. This keeps repetitive strings compressed without making random numeric data pay a compression and decompression pass for single-digit percentage savings. Compression also stops at the format's 64 MiB stored-payload ceiling before joining output chunks. A legal wide value that gzip expands falls back to raw without first allocating a second oversized compressed result.

Keeping published blocks unchanged pays for everything else:

  • A crash before publication can strand unused blocks; incomplete blocks cannot become visible. Acknowledged durability depends on the adapter and durability mode, and detected corruption stops recovery rather than guessing.
  • Retrying an unpublished block write preserves its immutable bytes. Retrying an application mutation with an unknown outcome requires reconciliation by durable application ID first.
  • Another tab can keep reading its leased old version; if it falls far enough behind, fixed retained-history ceilings backpressure writers rather than reclaiming data under the reader.
  • A snapshot reuses the published block bytes verbatim, so export never rewrites or recompresses table data.

Locked block envelope

Block format 2 is little-endian. Every block starts with this fixed 44-byte header, followed by canonical UTF-8 JSON metadata and then the stored payload:

OffsetWidthField
04ASCII magic BRDB
44envelope CRC-32 over bytes [8, 44 + metadata byte length)
82format version (2)
102header length (44)
121logical type ID
131physical encoding ID
141compression ID
151mandatory flags; currently zero
164row count
204null count
244metadata byte length
284uncompressed physical payload byte length
324stored payload byte length
364CRC-32 of the uncompressed physical payload
404CRC-32 of the stored payload

Readers verify the envelope before trusting a header field or statistic. A storage verifier can then check all stored bytes without decompressing them; a query additionally checks the logical checksum after decompression. Metadata is capped at 1 KiB, either payload at 64 MiB, and one block at 1,048,576 rows. The row ceiling keeps even bitmap-dense blocks safe to decode into JavaScript values. Unknown flags and identifiers are rejected, and raw blocks require their stored and logical checksums and stored/uncompressed lengths to match. CRC-32 detects accidental corruption; it is not a cryptographic authenticity check.

Logical and physical type IDs are 1 boolean, 2 number, 3 string, and 4 datetime; the two fields must match in format 2. Compression ID 0 is raw and 2 is gzip. ID 1 is permanently reserved for an abandoned pre-lock prototype and can never be reassigned. Metadata is exactly {} or {"zoneMap":{"min":number,"max":number}} in that key order with no insignificant whitespace; unknown fields and alternate JSON spellings are rejected.

All physical columns begin with a validity bitmap. Booleans add a value bitmap. Numbers use little-endian IEEE-754 float64 values, datetimes use float64 Unix milliseconds, and strings use rowCount + 1 little-endian uint32 offsets followed by strictly well-formed UTF-8. Minnow's physical rewrite operations validate and preserve these same canonical bitmap, null-slot, padding, numeric, UTF-8, offset, metadata, checksum, and size rules without materializing row objects. The public low-level decodeColumn applies that validation too; only the already-verified block reader uses its internal no-second-pass materializer. Frozen vectors cover every physical type, and the native database fixture locks the block and snapshot containers together.

Writes append small delta segments: an update stores just the key and the changed columns, a delete stores a key marker. Background compaction folds deltas into larger read-friendly segments later. Nothing is ever edited in place.

A query reads a table with deltas by scanning the appended data and applying the deltas over it: deleted keys mask rows out, updated keys patch the cells they changed, and the row groups a delta cannot reach are skipped from their statistics alone. The cost is the size of the deltas, not the size of the table — deleting one row of a million does not make the next query re-read the million.

How a write commits

Readers see the database through a manifest version. Its record is a bounded summary; exact block membership lives in an ordered provenance record per block, with the version that added it and, after retirement, the version that removed it. Each commit writes only changed provenance plus the next summary, so publishing costs the size of the change rather than cloning the live block-ID set. A commit publishes version N + 1, in strict order:

1. encode and compress the new blocks
2. write the blocks               (nothing points to them yet)
3. open a short metadata transaction
4. check the manifest is still at version N
5. publish summary N + 1 and the changed membership intervals

Data first, version last. The version change is all or nothing. A crash anywhere in between leaves unused blocks for the garbage collector — never a version pointing at half-written data.

Step 4 is the defensive check. Writers take turns before step 1 — a queue per store in each context, and the Web Lock minnowdb-write:<store> across tabs — so among cooperating connections nothing publishes between a writer's snapshot and its commit. If a writer that did not take a turn published first, the commit fails cleanly and a plain write retries against the new version. Conflicts surface as typed errors, never as silent interleaving.

A simple write — an insert, a point update, a delete, with no triggers staging rows of their own — folds steps 2 to 5 into one storage transaction where the store offers it (every shipped store does): the blocks, the segment, the transaction record, and the manifest flip land together or not at all. Durable storage commits are what a write costs in a browser, so this matters: a point update or delete is one, an insert two (its row ids are reserved first), and the first read after a commit re-pins the engine's shared reader lease in place rather than creating a new one and removing the old one later.

Encoding is bounded but parallel across independent columns. Blocks within each column keep their original order, and the staged metadata stays in schema order, so native compressors can overlap without making the committed layout depend on completion timing.

Sharing the database across tabs

Minnow assumes several tabs, unaware of each other, some frozen or already gone.

  • One writer at a time. Tabs take turns through a Web Lock before a write reads anything, so two tabs never prepare conflicting writes. Underneath, the store still serializes the manifest flip — IndexedDB in a transaction, the OPFS store through the browser's own exclusive lock on its log — so a writer that did not take a turn cannot interleave either: only one publishes, the other retries. A tab that stops inside its turn holds the others until the browser releases its lock; the engine reports the wait and never bypasses it.
  • Notifications are hints. BroadcastChannel announces that a new version exists, but every tab reconciles against the durable store. A missed message costs a little latency, never a stale result. Live queries are built on this.
  • Dead tabs are handled by leases. A long-running read holds a lease — a stored record with an expiry, renewed while the tab is alive. If the tab vanishes, the lease expires and whatever it pinned becomes reclaimable. Correctness depends on the durable lease record, not an in-memory heartbeat or a lifecycle notification.

How queries run

  • Data moves through the query runner in batches of 2,048 rows. Numbers and dates use Float64Array; strings share a small dictionary; null markers are packed into bits. This avoids creating millions of short-lived row objects while a query runs.
  • Grouping reuses the small number assigned to each repeated string. For several grouping columns, Minnow picks a compact lookup based on how many combinations are likely. Either way, it avoids repeatedly encoding and hashing the same strings for every row.
  • Queries accept a memory budget (executionMemoryBudgetBytes). Memory is reserved before it is allocated. Under a budget, sorts and grouped aggregations spill to durable temp pages instead of blowing up the tab; past it, you get a typed QueryMemoryBudgetError.
  • Spill pages are lease-protected, so a query abandoned by a dead tab gets cleaned up.
  • Compiled plans are cached separately, by statement text, so re-issuing a statement doesn't re-parse or re-plan it. Plan, statement, pattern, full-text-term, collation, constraint, and catalog-state caches all have fixed entry or byte bounds; oversized accepted text executes but does not become permanent process state.
  • Everything else repeated is cached in one byte-bounded buffer pool (bufferPoolBytes, default 64 MiB): decoded blocks by immutable block id, assembled column vectors and zone descriptions by the same, and computed results — whole-block results, the columnar forms of derived and windowed sources, and whole statement results — by exact visible-segment version key. A commit changes those keys, so computed entries stop matching, but unchanged blocks stay decoded: the next statement pays re-assembly, not re-fetch and re-decompression. Passing memoize: false to a query bypasses the computed-result entries and measures execution, which is what the benchmarks do.

The budget models query allocations rather than measuring the JavaScript heap. See Limits for the operations it does not cover.

Background work

Compaction and garbage collection run for a long time inside a tab that can die at any moment. So:

  • Every job is saved. Minnow saves the plan, then advances it in small steps with checkpoints. A dead tab loses at most one step, not the whole job.
  • Cancellation and publication cannot both win. Whichever finishes first decides the outcome.
  • Cleanup keeps anything still in use. The current version, open readers, live transactions, and running compactions protect what they need. Each deletion checks again before it happens. Finished transactions and jobs are removed once nothing refers to them, along with expired reader records and abandoned temporary query files. Ordinary use does not leave one permanent metadata record for every operation.
  • Compaction has a write budget. A rewrite that would cost more than a set multiple of the data it consolidates is refused up front, not discovered in the quota bill.

The shape of the system

Page (UI thread)       async proxy: requests, cancellation, results
        |
Engine worker          catalog, snapshots, transactions, planning, live queries
        |
Storage                IndexedDB or OPFS, block codecs, manifests, leases, GC
  • Each tab owns its own engine worker when it uses the worker client.
  • Buffers cross the boundary as transferable ArrayBuffers — no shared memory, which is why no cross-origin isolation is needed. Query results and live change events cross as one array per column (typed arrays transferred, strings as one flat text), and the proxy rebuilds the row objects on the page side; row objects are the API, never the wire.
  • The worker protocol is versioned RPC; each handle answers a fixed list of methods, nothing else.
  • Inside @minnowdb/core, the layers are separate modules — block-format, storage, transactions, and the engine — each importable on its own.
  • Within the engine, internal controllers own catalog mutations, writer admission, query preparation/execution, and maintenance scheduling. They share the database’s existing state and inject storage and lifecycle operations; physical execution kernels remain in the database. These are internal boundaries, with no new public entry points.
  • IndexedDB and the memory/OPFS record engine share commit freshness checks, level-zero admission planning, and accelerator retention rules. Each adapter still owns its native transaction, block visibility checks, and durability boundary.

Crash testing

Fault injection has been required since the first storage code. Tests kill the engine before and after every block write and manifest commit, and these must hold no matter where the crash lands:

  1. A visible manifest only references complete, checksum-valid blocks.
  2. A stale manifest check cannot publish anything.
  3. Retrying a block write cannot change published bytes.
  4. Unpublished data never affects a read and is always safe to collect.
  5. Worker execution keeps engine work off the UI thread; bounded pages and cancellation limit transfer and retained work. Running the engine directly on the page can still block rendering.

The same hook is public: FaultInjectingBlockStore in @minnowdb/core/testing wraps any store, so your own tests can crash the engine on purpose too.

Limits

  • The memory budget is not a hard heap limit. Some selected columns are built before accounting starts, and some joins cannot spill. The budget prevents several large working sets but does not cap the JavaScript heap.
  • Performance depends on the browser and workload. Run the benchmarks to measure the queries and storage adapters on your target device.
  • The block format is explicitly versioned. Version 2 freezes its 44-byte envelope and raw physical layouts in permanent byte vectors. A future format gets a new version and reader; existing bytes are never reinterpreted.
  • Indexes use tuple postings, not a general B-tree. Scalar and composite keys prune leftmost equality/IN prefixes plus a following range. UNIQUE has independent atomic membership. Exact non-null indexes can also deliver ORDER BY and covering reads for keyless append-only tables; other shapes keep the ordinary scan and bounded sort.

On this page