Writing a storage adapter
The shared storage interface, the rules an adapter follows, its test kit, and reusable tools for building one.
Minnow's engine holds one BlockStore and uses it for all saved data. Implement that interface for
a new storage system and you have a working database with no engine changes. The three included
adapters follow the same rules. A React Native store, an object store such as R2, the Node
filesystem, or an encrypted wrapper can do the same.
import type { BlockStore } from "@minnowdb/core/storage/contracts";
class MyBlockStore implements BlockStore {
// ...
}
const db = new MinnowDatabase(new MyBlockStore());What the interface covers
BlockStore combines seven smaller interfaces, each responsible for one part of storage:
| Interface | Owns |
|---|---|
BlockPayloadStore | Immutable byte blobs — the columnar data itself. |
CatalogStore | Table records, row-id and auto-increment counters, unique-key lookups. |
TransactionStore | Manifests, segments, transaction records, and the atomic commit. |
LeaseStore | Reader pins that protect versions from garbage collection. |
MaintenanceStore | Resumable compaction and garbage-collection job records. |
FtsIndexStore | Full-text and secondary postings bases, deltas, and candidate reads. |
TempSpillStore | Query spill pages and their owner leases. |
The TSDoc in @minnowdb/core/storage/contracts explains the exact order of operations, the error
types, and which changes must happen together. The guide below gives the friendlier overview.
The rules
A store holds records and byte blocks, not rows. The engine hands you compressed blocks that never change after publication, plus small structured records: manifests saying which blocks are live at each version, segments mapping blocks to tables, transactions, leases, jobs, counters, keys. Because published blocks never change and a version is just a set of block ids, readers hold versions open while writers commit — the concurrency story is the data model, not locks you invent.
Each operation is all or nothing. A local adapter method that resolves has happened; one that
rejects has not. commitTransaction is the heart of the interface: it checks the expected version, publishes
the new version, finishes its segments, updates keys and full-text data, and marks the transaction
complete in one saved step. A crash cannot leave half of that change visible.
OPFS follower RPC has one explicit transport-loss exception. OpfsUncertainOutcomeError means a
leader vanished after receiving the named mutation, its reply was lost, and the log the next
leader recovered can neither answer the re-sent request nor prove it never ran: the request was
first sent before the recovered ledger's coverage begins, the leader died in the instant
between the mutation's frame and the frame that records its return value, or the return value
was too large for the ledger to keep (a staging call's answer is the whole transaction record). Every other loss is
resolved from the log — a re-send is answered with the value the mutation already returned, or
run once because the log proves it never ran — and a graceful handover never raises the error
at all. Preserve the class and its method field across any wrapper or worker boundary;
inspect or reopen the database instead of blindly retrying. An ordinary Error still means no
effect and must never be used to hide an uncertain outcome.
Artifact journaling has the same required boundary. stageTransactionArtifacts saves new blocks
and segments and appends their IDs to the active transaction in one operation.
rollbackTransactionArtifacts compare-and-swaps the journal back to the exact saved set and
deletes only the proven difference in that same operation. The adapter validates that the
duplicate-free retained and removed sets are disjoint and exactly partition the current journal
before changing anything. There is no sequential
fallback: a crash between a record update and byte deletion would otherwise create either an
uncollectable orphan or a journal pointing at missing data.
Keep each staging call within the exported MAX_TRANSACTION_STAGE_BLOCKS,
MAX_TRANSACTION_STAGE_SEGMENTS, and MAX_TRANSACTION_STAGE_BYTES limits, checked before opening
a write transaction. The engine splits larger writes into resumable batches; the per-call ceiling
prevents one caller from creating an unbounded storage transaction. A transaction's journal itself
has no fixed length: only the aggregate staged-artifact quotas bound it, so an adapter must not
rewrite or re-verify the whole journal on every staging call — append to it, and validate what
this call adds. A generic updateTransaction that carries pending ids must extend the journal
in order; the IndexedDB adapter refuses one that reorders or drops journaled ids, and the
engine never sends one.
Transaction liveness is part of that boundary. Persist ownerId and expiresAt with the
transaction itself. renewTransaction may extend only a matching, still-active owner whose old
deadline is strictly after the supplied cutoff, and it must not change the transaction's data
revision. abortTransactionIfExpired atomically checks owner, status, and deadline before moving
the record to aborted. Renewal and abort must have one winner; an expired owner can never be
resurrected. The collector treats only the successfully aborted journal as reclaimable provenance.
Schema serialization is part of commit too. getCatalogProbe() atomically returns the current
manifest version, broad catalog epoch, and structural schema epoch. beginTransaction() copies
that structural epoch into schemaEpochGuard; commitTransaction() and writeTransaction() must
compare it in the same transaction that publishes the manifest and throw SchemaConflictError
before any mutation when it differs. Advance the structural epoch for table/view/column/default,
trigger, and write-semantic index changes. Do not advance it for ordinary manifests or
accelerator build progress such as building/ready/invalid stamps: those do not change row
validation, and treating them as schema DDL would make unrelated writers conflict.
Catalog cleanup follows the same rule. When updateTable removes a column or secondary index,
that operation also deletes all of its postings base chunks, staged generations, and commit
deltas. A dropped UNIQUE index also removes its membership namespace. removeFtsColumn provides
the postings cleanup directly. A dropped column or index must not leave storage growing invisibly
for the rest of a long-running database's life.
Publishing a UNIQUE build is one catalog operation: updateTable checks
expectedManifestVersion, publishes the ready index, and installs uniqueKeySeed together. The
seed namespace must belong to that ready index, duplicate seed tokens reject the whole update, and
a writer that changes the table without covering every enforced unique namespace rejects with
UniqueIndexCoverageError. That rejection is how a writer prepared before the DDL is restarted
instead of slipping through unenforced.
Postings output is always chunked. beginFtsBaseBuild, writeFtsBaseBuildChunk,
finishFtsBaseBuild, and abortFtsBaseBuild stage one small chunk at a time and publish the
complete generation atomically. Starting a replacement reclaims an abandoned generation. Common
append/update/delete histories are scanned through bounded row windows. An exact UNIQUE build
over an uncompacted upsert history is the deliberate exception: it materializes the current
projected key columns, never the historical rows, before writing the same bounded chunks. These
methods are required: an adapter must not store postings as one database-sized value during
CREATE INDEX.
readFtsPostings returns the canonical term-then-row-ID merge of that same snapshot-bounded base
and delta tail. Ordered and covering secondary-index scans use it; returning a partial generation
would change row order or omit rows, so hasBase, coversVersion, and deltaChunkCount must match
readFtsCandidates exactly.
runGarbageCollectionStep owns record reclamation as well as blob reclamation. It processes
manifests, segments, blocks, then terminal transactions; a committed transaction can disappear
only after its manifest is pruned and no segment or unfinished compaction names it, while an
aborted one waits for every pending artifact to disappear. Adapters must apply that whole step
as one operation. removePrunedManifestRecords then deletes only markers that no readable version
needs. Block discovery does not depend on those bounded summaries:
listRetiredManifestBlockPage walks the independent, ID-ordered provenance intervals for blocks
retired through a version. Payload and provenance are removed atomically only after no readable
version, transaction, or maintenance job roots them, so a small pass can resume after the summary
is gone without stranding bytes. A file-backed adapter may delete immutable block payloads after
logging the resolved step, as the toolkit example does.
Maintenance discovery must also stay bounded. listManifestPage, listSegmentPage,
listTransactionPage, and listCompactionJobPage use exclusive cursors so the engine can sweep
old metadata without materializing the database's complete history. Return records sorted by the
cursor field and reject a non-positive or storage-unsafe page limit, matching the interface
contract. listManifestBlockPage and listRetiredManifestBlockPage page exact block membership
and retired provenance independently of summary pruning. Do not add a whole-history fallback for
an old unpaged API: the v1 contract begins with these cursor boundaries.
Global growth ceilings are atomic admission rules. The exported MAX_* storage constants
bound live owners, staged journals and accelerator builds, catalog table/view records, manifest
summaries, segment records, temporary spill, retained terminal metadata, pinned retired history,
and total obsolete payload bytes. Sweep safely expired owners first, then either save the record
and its accounting together or throw StorageResourceLimitError before either changes. Catalog
accounting includes a transaction-owned pending table. Charge catalog and segment records by
their exact canonical wire bytes; charge an unpruned manifest for the exact bytes of its eventual
24-byte UTC tombstone too, so pruning cannot deadlock at the byte ceiling. Treat a negative,
overflowing, over-limit, or checksum-invalid durable ledger as corruption, not as an empty ledger.
A stable identity is what lets writers take turns. liveQueryChannelName names the storage
underneath the adapter object — the shipped adapters use minnowdb-live:<kind>:<database> — and
the engine keys its writer queue and the cross-tab Web Lock on it. Give a durable adapter one,
and two engines over the same storage take turns whether they share the object, the context, or
only the browser profile. Without it, only engines sharing the same store object take turns:
db.writeCoordination reports instance, and a second connection in another context is an
uncoordinated writer that storage compare-and-swap alone keeps out. See
writer turns.
Control records have fixed structural bounds too. Validate them before I/O with the exported
helpers and constants: database names stop at 256 characters; storage IDs and catalog names at
1,024; and posting terms at 65,536. Tables stop at 1,024 columns, indexes, or
constraints, 256 triggers, and 4,096 enum values; one table record stops at 1,048,576 characters
and 65,536 walked entries. Persist only well-formed Unicode, finite or safe numeric counters,
canonical unsigned 64-bit row IDs, and exact 24-character UTC timestamps. Required v1 fields are
not defaults: for example, a level-two segment has a stable partitionOrdinal, and a persisted
compaction checkpoint carries its full rewrite plan, cursors, source levels, and memory accounting.
Missing pre-v1 fields are corruption, not a reason to infer state and continue.
Conflicts use the exported error classes. Version-check failures throw
SchemaConflictError, WriteConflictError, TransactionRecordConflictError, UniqueKeyConflictError,
UniqueIndexCoverageError, or the matching exported class directly. The engine uses those classes
to decide when to retry, and the worker client recreates them on the main thread.
Platform failures pass through. A quota refusal escapes as the environment's own
QuotaExceededError, unwrapped, with prior data intact and the same write succeeding once
space frees.
Nothing is shared. Copy bytes and records in both directions; the engine may reuse the buffers it hands you and mutate the records you return.
An optional shortcut still has to be all or nothing. writeTransaction, moveLease,
putTempRunPages, and the framed snapshot session methods are optional so an adapter that can do
something in one atomic step may say so. getCatalogProbe and beginTransaction are required:
schema-dependent writes need one coherent catalog proof, and CREATE TABLE … AS SELECT needs a
transaction-owned table reservation that sequential calls cannot emulate safely.
Callers trust a present fused method completely and fall back to the required atomic primitives
where that fallback exists — never implement one as the sequential calls in a trench coat. Framed
snapshot support is different: the export and import session families are each complete
capabilities, and an absent family means that direction is unavailable rather than unbounded.
getCatalogProbe returns the manifest version, catalog epoch, and structural schema epoch from one
coherent read. The engine may reuse cached catalog pages only while the first two remain unchanged;
writers retain the schema epoch until commit. Current and historical query
preparation use listTableSegmentPage, getTransactions, and version-scoped manifest-membership
reads; no query shortcut may return every segment, transaction, or block ID in a retained history.
Each snapshot direction is one complete capability: begin/read/close frame export, and begin/renew/append/finish/cancel frame import. Do not expose a partial family. Export must pin one version and page all six frame kinds; import must compare same-sequence replay bytes and publish only after the footer totals and rolling checksum match. See Snapshots.
Tabs are plural and mortal. Several connections may open one database; readers keep a stable version while writers publish new versions. Physical I/O may queue. A connection can die between any two operations without corrupting anything. How you achieve that is yours: the IndexedDB adapter leans on storage transactions, the OPFS adapter on a write-ahead log behind a browser-arbitrated leader. (An adapter for a single-process environment — Node, React Native — has an easier version of this problem, but follows the same rules.)
The shared test kit
The kit is the runnable half of this page. It works with any test framework and checks the rules above: all-or-nothing operations, exact conflict classes, safe copies, ordering, and data after a reopen:
import { blockStoreConformanceCases } from "@minnowdb/core/testing";
for (const conformanceCase of blockStoreConformanceCases()) {
it(conformanceCase.name, () =>
conformanceCase.run({
create: () => MyBlockStore.open({ name: crypto.randomUUID() }),
reopen: async (store) => {
store.close();
return MyBlockStore.open(/* the same database */);
},
}),
);
}Provide reopen so the kit can verify that saved data survives. The kit is a
floor, not a ceiling: Minnow's own adapters additionally run fault sweeps that interrupt every
storage operation, concurrency soaks, and quota injection, and the kit itself runs against all
three shipped adapters in CI so it cannot drift from what they do. FaultInjectingBlockStore
from the same entry point wraps any store to fail at named points when you want to test your
own crash handling.
Two architectures that work
Direct — map each record family onto your storage system's own transaction support, the
way the IndexedDB adapter maps them onto object stores and commitTransaction onto one
read-write transaction. This works well when the storage system can update several keys in one
all-or-nothing transaction.
Log-structured, single writer — the OPFS adapter's shape, and the natural one for
storage systems that only offer files or blobs: keep every record in memory, append each change
as a checksummed frame to a write-ahead log, fold into checkpoints, pack blobs into extent
files, and let one connection own the writes. Recovery is checkpoint-plus-tail; a torn tail
frame reads as "not written". This maps directly onto the Node filesystem (fs handles in
place of sync access handles), React Native storage, or an object store like R2 — where the
log grows by conditional puts and the single-writer election uses the store's own
conditional-write support. Multi-client coordination is the part you own; everything above the
persistence layer behaves identically — and ships as reusable pieces in the toolkit below.
Either way, snapshots come almost free — one committed version as a portable file — and are also the migration path between your store and the shipped ones.
The adapter toolkit
@minnowdb/core/storage/toolkit is the library the shipped adapters are assembled from,
published so a new adapter can start from working parts instead of a blank interface. It is
deliberately not required by the interface: the engine never imports it, the test kit never
requires it, and an adapter that stores records another way is equally valid. It exists because
the hardest parts of an adapter are often the same.
| Export | What it is |
|---|---|
RecordCore | The in-memory record engine behind the memory and OPFS adapters: one synchronous method per operation, with validation, conflict errors, safe copies, and dump()/load() for checkpoints. |
WalWriter, replayWalFrames | Checksummed write-ahead-log frames over one held file handle. Replay ignores a truncated final frame, which is exactly what a crash can leave, and rejects complete corrupt frames. A payload too large for one frame is written as continuation frames and joined on replay. |
ExtentPool | Packed append-only files for bulk bytes, addressed by Placement { extent, offset, length, checksum }, with sealing, full-payload recovery checks, live-byte accounting, fragmentation detection, and a read-handle cache. |
encodeRecordJson and friends | The JSON codec (bigints included) and the checksummed, versioned envelopes for checkpoints and immutable chunks. |
SyncFileHandle | The small file interface used by these tools: positioned read, write, truncate, and flush. A browser FileSystemSyncAccessHandle already has this shape; a Node file descriptor can be wrapped to match it. |
readFully, writeFully | Complete a positioned transfer across permitted short reads or writes. They reject invalid or overflowing ranges before I/O, throw on EOF, zero progress, or an invalid byte count, and avoid copying on both the one-call fast path and retries. |
Wrapping RecordCore obligates you to two things. One writer at a time — its methods
validate and then mutate synchronously, so calls must never interleave (a promise-chain queue,
a leader, or a process-wide lock all work). Durability is yours — the core is memory:
log each change, save checkpoints with dump() (serialize the result immediately because its
arrays still refer to the live state), and replay the log on open. For any single-log design,
flush every payload the checkpoint references, then write and flush the checkpoint before
resetting the log — never the other way around. Use readFully and writeFully for every
sync-handle transfer; the platform is allowed to return fewer bytes than requested.
A shared test checks that these parts work together:
toolkit-example.test.ts
builds a complete, persistent, log-structured BlockStore from these exports alone: record
handling from RecordCore, one log, blocks in packed files, and a checkpoint at open. It passes
the same test kit as the included adapters. Reading it top to bottom is the fastest way
to see where your storage system fits.
For running your adapter in plain Node tests, MemoryOpfs from @minnowdb/core/testing is an
in-memory origin-private file system with the behaviour that matters in storage tests:
exclusive sync-access locks that throw the real DOMExceptions, handles that resolve by path as
browsers do (a handle to a removed directory throws NotFoundError), injectable write faults,
and deterministic short-transfer limits for quota, crash, and partial-I/O suites.
What the fused methods buy
The required surface includes getCatalogProbe, beginTransaction,
stageTransactionArtifacts, and rollbackTransactionArtifacts. writeTransaction is the
optional fusion that matters most for writes: it
begins (or continues) a transaction, stages its blocks and segments, and commits, all in one
storage transaction, so a point update or delete costs one durable commit instead of three and
an insert two (its row ids are reserved by beginTransaction first). moveLease re-pins the
engine's shared reader lease to each new version in place, which saves a lease create and a
later remove on the first read after every commit. Implement the required probe and transaction
begin first, then writeTransaction and moveLease; add
putTempRunPages when your storage system charges per call, and framed snapshot sessions when your
store should be exportable. Snapshot support means the complete export/import session family, not
one database-sized record helper. stageTransactionArtifacts and
rollbackTransactionArtifacts are correctness boundaries every adapter must implement, not
performance options.