Storage

Writing a storage adapter

The shared storage interface, the rules an adapter follows, its test kit, and reusable tools for building one.

Minnow's engine holds one BlockStore and uses it for all saved data. Implement that interface for a new storage system and you have a working database with no engine changes. The three included adapters follow the same rules. A React Native store, an object store such as R2, the Node filesystem, or an encrypted wrapper can do the same.

import type { BlockStore } from "@minnowdb/core/storage/contracts";

class MyBlockStore implements BlockStore {
  // ...
}

const db = new MinnowDatabase(new MyBlockStore());

What the interface covers

BlockStore combines seven smaller interfaces, each responsible for one part of storage:

InterfaceOwns
BlockPayloadStoreImmutable byte blobs — the columnar data itself.
CatalogStoreTable records, row-id and auto-increment counters, unique-key lookups.
TransactionStoreManifests, segments, transaction records, and the atomic commit.
LeaseStoreReader pins that protect versions from garbage collection.
MaintenanceStoreResumable compaction and garbage-collection job records.
FtsIndexStoreFull-text and secondary postings bases, deltas, and candidate reads.
TempSpillStoreQuery spill pages and their owner leases.

The TSDoc in @minnowdb/core/storage/contracts explains the exact order of operations, the error types, and which changes must happen together. The guide below gives the friendlier overview.

The rules

A store holds records and byte blocks, not rows. The engine hands you compressed blocks that never change after publication, plus small structured records: manifests saying which blocks are live at each version, segments mapping blocks to tables, transactions, leases, jobs, counters, keys. Because published blocks never change and a version is just a set of block ids, readers hold versions open while writers commit — the concurrency story is the data model, not locks you invent.

Each operation is all or nothing. A local adapter method that resolves has happened; one that rejects has not. commitTransaction is the heart of the interface: it checks the expected version, publishes the new version, finishes its segments, updates keys and full-text data, and marks the transaction complete in one saved step. A crash cannot leave half of that change visible.

OPFS follower RPC has one explicit transport-loss exception. OpfsUncertainOutcomeError means a leader vanished after receiving the named mutation, its reply was lost, and the log the next leader recovered can neither answer the re-sent request nor prove it never ran: the request was first sent before the recovered ledger's coverage begins, the leader died in the instant between the mutation's frame and the frame that records its return value, or the return value was too large for the ledger to keep (a staging call's answer is the whole transaction record). Every other loss is resolved from the log — a re-send is answered with the value the mutation already returned, or run once because the log proves it never ran — and a graceful handover never raises the error at all. Preserve the class and its method field across any wrapper or worker boundary; inspect or reopen the database instead of blindly retrying. An ordinary Error still means no effect and must never be used to hide an uncertain outcome.

Artifact journaling has the same required boundary. stageTransactionArtifacts saves new blocks and segments and appends their IDs to the active transaction in one operation. rollbackTransactionArtifacts compare-and-swaps the journal back to the exact saved set and deletes only the proven difference in that same operation. The adapter validates that the duplicate-free retained and removed sets are disjoint and exactly partition the current journal before changing anything. There is no sequential fallback: a crash between a record update and byte deletion would otherwise create either an uncollectable orphan or a journal pointing at missing data.

Keep each staging call within the exported MAX_TRANSACTION_STAGE_BLOCKS, MAX_TRANSACTION_STAGE_SEGMENTS, and MAX_TRANSACTION_STAGE_BYTES limits, checked before opening a write transaction. The engine splits larger writes into resumable batches; the per-call ceiling prevents one caller from creating an unbounded storage transaction. A transaction's journal itself has no fixed length: only the aggregate staged-artifact quotas bound it, so an adapter must not rewrite or re-verify the whole journal on every staging call — append to it, and validate what this call adds. A generic updateTransaction that carries pending ids must extend the journal in order; the IndexedDB adapter refuses one that reorders or drops journaled ids, and the engine never sends one.

Transaction liveness is part of that boundary. Persist ownerId and expiresAt with the transaction itself. renewTransaction may extend only a matching, still-active owner whose old deadline is strictly after the supplied cutoff, and it must not change the transaction's data revision. abortTransactionIfExpired atomically checks owner, status, and deadline before moving the record to aborted. Renewal and abort must have one winner; an expired owner can never be resurrected. The collector treats only the successfully aborted journal as reclaimable provenance.

Schema serialization is part of commit too. getCatalogProbe() atomically returns the current manifest version, broad catalog epoch, and structural schema epoch. beginTransaction() copies that structural epoch into schemaEpochGuard; commitTransaction() and writeTransaction() must compare it in the same transaction that publishes the manifest and throw SchemaConflictError before any mutation when it differs. Advance the structural epoch for table/view/column/default, trigger, and write-semantic index changes. Do not advance it for ordinary manifests or accelerator build progress such as building/ready/invalid stamps: those do not change row validation, and treating them as schema DDL would make unrelated writers conflict.

Catalog cleanup follows the same rule. When updateTable removes a column or secondary index, that operation also deletes all of its postings base chunks, staged generations, and commit deltas. A dropped UNIQUE index also removes its membership namespace. removeFtsColumn provides the postings cleanup directly. A dropped column or index must not leave storage growing invisibly for the rest of a long-running database's life.

Publishing a UNIQUE build is one catalog operation: updateTable checks expectedManifestVersion, publishes the ready index, and installs uniqueKeySeed together. The seed namespace must belong to that ready index, duplicate seed tokens reject the whole update, and a writer that changes the table without covering every enforced unique namespace rejects with UniqueIndexCoverageError. That rejection is how a writer prepared before the DDL is restarted instead of slipping through unenforced.

Postings output is always chunked. beginFtsBaseBuild, writeFtsBaseBuildChunk, finishFtsBaseBuild, and abortFtsBaseBuild stage one small chunk at a time and publish the complete generation atomically. Starting a replacement reclaims an abandoned generation. Common append/update/delete histories are scanned through bounded row windows. An exact UNIQUE build over an uncompacted upsert history is the deliberate exception: it materializes the current projected key columns, never the historical rows, before writing the same bounded chunks. These methods are required: an adapter must not store postings as one database-sized value during CREATE INDEX.

readFtsPostings returns the canonical term-then-row-ID merge of that same snapshot-bounded base and delta tail. Ordered and covering secondary-index scans use it; returning a partial generation would change row order or omit rows, so hasBase, coversVersion, and deltaChunkCount must match readFtsCandidates exactly.

runGarbageCollectionStep owns record reclamation as well as blob reclamation. It processes manifests, segments, blocks, then terminal transactions; a committed transaction can disappear only after its manifest is pruned and no segment or unfinished compaction names it, while an aborted one waits for every pending artifact to disappear. Adapters must apply that whole step as one operation. removePrunedManifestRecords then deletes only markers that no readable version needs. Block discovery does not depend on those bounded summaries: listRetiredManifestBlockPage walks the independent, ID-ordered provenance intervals for blocks retired through a version. Payload and provenance are removed atomically only after no readable version, transaction, or maintenance job roots them, so a small pass can resume after the summary is gone without stranding bytes. A file-backed adapter may delete immutable block payloads after logging the resolved step, as the toolkit example does.

Maintenance discovery must also stay bounded. listManifestPage, listSegmentPage, listTransactionPage, and listCompactionJobPage use exclusive cursors so the engine can sweep old metadata without materializing the database's complete history. Return records sorted by the cursor field and reject a non-positive or storage-unsafe page limit, matching the interface contract. listManifestBlockPage and listRetiredManifestBlockPage page exact block membership and retired provenance independently of summary pruning. Do not add a whole-history fallback for an old unpaged API: the v1 contract begins with these cursor boundaries.

Global growth ceilings are atomic admission rules. The exported MAX_* storage constants bound live owners, staged journals and accelerator builds, catalog table/view records, manifest summaries, segment records, temporary spill, retained terminal metadata, pinned retired history, and total obsolete payload bytes. Sweep safely expired owners first, then either save the record and its accounting together or throw StorageResourceLimitError before either changes. Catalog accounting includes a transaction-owned pending table. Charge catalog and segment records by their exact canonical wire bytes; charge an unpruned manifest for the exact bytes of its eventual 24-byte UTC tombstone too, so pruning cannot deadlock at the byte ceiling. Treat a negative, overflowing, over-limit, or checksum-invalid durable ledger as corruption, not as an empty ledger.

A stable identity is what lets writers take turns. liveQueryChannelName names the storage underneath the adapter object — the shipped adapters use minnowdb-live:<kind>:<database> — and the engine keys its writer queue and the cross-tab Web Lock on it. Give a durable adapter one, and two engines over the same storage take turns whether they share the object, the context, or only the browser profile. Without it, only engines sharing the same store object take turns: db.writeCoordination reports instance, and a second connection in another context is an uncoordinated writer that storage compare-and-swap alone keeps out. See writer turns.

Control records have fixed structural bounds too. Validate them before I/O with the exported helpers and constants: database names stop at 256 characters; storage IDs and catalog names at 1,024; and posting terms at 65,536. Tables stop at 1,024 columns, indexes, or constraints, 256 triggers, and 4,096 enum values; one table record stops at 1,048,576 characters and 65,536 walked entries. Persist only well-formed Unicode, finite or safe numeric counters, canonical unsigned 64-bit row IDs, and exact 24-character UTC timestamps. Required v1 fields are not defaults: for example, a level-two segment has a stable partitionOrdinal, and a persisted compaction checkpoint carries its full rewrite plan, cursors, source levels, and memory accounting. Missing pre-v1 fields are corruption, not a reason to infer state and continue.

Conflicts use the exported error classes. Version-check failures throw SchemaConflictError, WriteConflictError, TransactionRecordConflictError, UniqueKeyConflictError, UniqueIndexCoverageError, or the matching exported class directly. The engine uses those classes to decide when to retry, and the worker client recreates them on the main thread.

Platform failures pass through. A quota refusal escapes as the environment's own QuotaExceededError, unwrapped, with prior data intact and the same write succeeding once space frees.

Nothing is shared. Copy bytes and records in both directions; the engine may reuse the buffers it hands you and mutate the records you return.

An optional shortcut still has to be all or nothing. writeTransaction, moveLease, putTempRunPages, and the framed snapshot session methods are optional so an adapter that can do something in one atomic step may say so. getCatalogProbe and beginTransaction are required: schema-dependent writes need one coherent catalog proof, and CREATE TABLE … AS SELECT needs a transaction-owned table reservation that sequential calls cannot emulate safely. Callers trust a present fused method completely and fall back to the required atomic primitives where that fallback exists — never implement one as the sequential calls in a trench coat. Framed snapshot support is different: the export and import session families are each complete capabilities, and an absent family means that direction is unavailable rather than unbounded.

getCatalogProbe returns the manifest version, catalog epoch, and structural schema epoch from one coherent read. The engine may reuse cached catalog pages only while the first two remain unchanged; writers retain the schema epoch until commit. Current and historical query preparation use listTableSegmentPage, getTransactions, and version-scoped manifest-membership reads; no query shortcut may return every segment, transaction, or block ID in a retained history.

Each snapshot direction is one complete capability: begin/read/close frame export, and begin/renew/append/finish/cancel frame import. Do not expose a partial family. Export must pin one version and page all six frame kinds; import must compare same-sequence replay bytes and publish only after the footer totals and rolling checksum match. See Snapshots.

Tabs are plural and mortal. Several connections may open one database; readers keep a stable version while writers publish new versions. Physical I/O may queue. A connection can die between any two operations without corrupting anything. How you achieve that is yours: the IndexedDB adapter leans on storage transactions, the OPFS adapter on a write-ahead log behind a browser-arbitrated leader. (An adapter for a single-process environment — Node, React Native — has an easier version of this problem, but follows the same rules.)

The shared test kit

The kit is the runnable half of this page. It works with any test framework and checks the rules above: all-or-nothing operations, exact conflict classes, safe copies, ordering, and data after a reopen:

import { blockStoreConformanceCases } from "@minnowdb/core/testing";

for (const conformanceCase of blockStoreConformanceCases()) {
  it(conformanceCase.name, () =>
    conformanceCase.run({
      create: () => MyBlockStore.open({ name: crypto.randomUUID() }),
      reopen: async (store) => {
        store.close();
        return MyBlockStore.open(/* the same database */);
      },
    }),
  );
}

Provide reopen so the kit can verify that saved data survives. The kit is a floor, not a ceiling: Minnow's own adapters additionally run fault sweeps that interrupt every storage operation, concurrency soaks, and quota injection, and the kit itself runs against all three shipped adapters in CI so it cannot drift from what they do. FaultInjectingBlockStore from the same entry point wraps any store to fail at named points when you want to test your own crash handling.

Two architectures that work

Direct — map each record family onto your storage system's own transaction support, the way the IndexedDB adapter maps them onto object stores and commitTransaction onto one read-write transaction. This works well when the storage system can update several keys in one all-or-nothing transaction.

Log-structured, single writer — the OPFS adapter's shape, and the natural one for storage systems that only offer files or blobs: keep every record in memory, append each change as a checksummed frame to a write-ahead log, fold into checkpoints, pack blobs into extent files, and let one connection own the writes. Recovery is checkpoint-plus-tail; a torn tail frame reads as "not written". This maps directly onto the Node filesystem (fs handles in place of sync access handles), React Native storage, or an object store like R2 — where the log grows by conditional puts and the single-writer election uses the store's own conditional-write support. Multi-client coordination is the part you own; everything above the persistence layer behaves identically — and ships as reusable pieces in the toolkit below.

Either way, snapshots come almost free — one committed version as a portable file — and are also the migration path between your store and the shipped ones.

The adapter toolkit

@minnowdb/core/storage/toolkit is the library the shipped adapters are assembled from, published so a new adapter can start from working parts instead of a blank interface. It is deliberately not required by the interface: the engine never imports it, the test kit never requires it, and an adapter that stores records another way is equally valid. It exists because the hardest parts of an adapter are often the same.

ExportWhat it is
RecordCoreThe in-memory record engine behind the memory and OPFS adapters: one synchronous method per operation, with validation, conflict errors, safe copies, and dump()/load() for checkpoints.
WalWriter, replayWalFramesChecksummed write-ahead-log frames over one held file handle. Replay ignores a truncated final frame, which is exactly what a crash can leave, and rejects complete corrupt frames. A payload too large for one frame is written as continuation frames and joined on replay.
ExtentPoolPacked append-only files for bulk bytes, addressed by Placement { extent, offset, length, checksum }, with sealing, full-payload recovery checks, live-byte accounting, fragmentation detection, and a read-handle cache.
encodeRecordJson and friendsThe JSON codec (bigints included) and the checksummed, versioned envelopes for checkpoints and immutable chunks.
SyncFileHandleThe small file interface used by these tools: positioned read, write, truncate, and flush. A browser FileSystemSyncAccessHandle already has this shape; a Node file descriptor can be wrapped to match it.
readFully, writeFullyComplete a positioned transfer across permitted short reads or writes. They reject invalid or overflowing ranges before I/O, throw on EOF, zero progress, or an invalid byte count, and avoid copying on both the one-call fast path and retries.

Wrapping RecordCore obligates you to two things. One writer at a time — its methods validate and then mutate synchronously, so calls must never interleave (a promise-chain queue, a leader, or a process-wide lock all work). Durability is yours — the core is memory: log each change, save checkpoints with dump() (serialize the result immediately because its arrays still refer to the live state), and replay the log on open. For any single-log design, flush every payload the checkpoint references, then write and flush the checkpoint before resetting the log — never the other way around. Use readFully and writeFully for every sync-handle transfer; the platform is allowed to return fewer bytes than requested.

A shared test checks that these parts work together: toolkit-example.test.ts builds a complete, persistent, log-structured BlockStore from these exports alone: record handling from RecordCore, one log, blocks in packed files, and a checkpoint at open. It passes the same test kit as the included adapters. Reading it top to bottom is the fastest way to see where your storage system fits.

For running your adapter in plain Node tests, MemoryOpfs from @minnowdb/core/testing is an in-memory origin-private file system with the behaviour that matters in storage tests: exclusive sync-access locks that throw the real DOMExceptions, handles that resolve by path as browsers do (a handle to a removed directory throws NotFoundError), injectable write faults, and deterministic short-transfer limits for quota, crash, and partial-I/O suites.

What the fused methods buy

The required surface includes getCatalogProbe, beginTransaction, stageTransactionArtifacts, and rollbackTransactionArtifacts. writeTransaction is the optional fusion that matters most for writes: it begins (or continues) a transaction, stages its blocks and segments, and commits, all in one storage transaction, so a point update or delete costs one durable commit instead of three and an insert two (its row ids are reserved by beginTransaction first). moveLease re-pins the engine's shared reader lease to each new version in place, which saves a lease create and a later remove on the first read after every commit. Implement the required probe and transaction begin first, then writeTransaction and moveLease; add putTempRunPages when your storage system charges per call, and framed snapshot sessions when your store should be exportable. Snapshot support means the complete export/import session family, not one database-sized record helper. stageTransactionArtifacts and rollbackTransactionArtifacts are correctness boundaries every adapter must implement, not performance options.

On this page