# How to store data

> A plain-language, drill-down guide to where your app should keep its data: a file, git, SQLite, Postgres, or a hosted service. The five questions that decide it, what each option really trades away, and a checklist to hand your agent.

Source: https://iambenschmidt.com/shape/how-to-store-data/ · 2026-10-07


<section class="sh-layer" id="start" data-nav="What this is, and why it bites">

Every app you build has to keep what it knows somewhere. Deciding *where* sounds like a detail. It is actually one of the few early choices you cannot cheaply undo.

"Save the results somewhere" is not one instruction. It is four or five different engineering decisions wearing a trench coat, and your agent will make all of them for you, quietly, in the first few seconds, then move on as if nothing happened. When the choice is wrong you do not find out now. You find out later, as lost work, a corrupted file, or an app that falls over the week it finally gets busy.

Here is the reassuring part: you do not need to understand databases to make this call well. You need to understand the *shape* of it just enough to ask the two or three questions that force your agent to choose on purpose, out loud, instead of silently. That is the whole skill, and it is the whole point of this page.


Anyone directing an agent to build something real, whether or not you write the code yourself.
Enough fluency to name the trade-offs, plus a checklist you hand your agent before it picks.


**How to read this.** Everything below starts as a one-line verdict you can skim in a couple of minutes. Open only the sections you care about for the depth. Dotted terms like this one give you a plain definition on hover or tap.

</section>









<section class="sh-layer" id="shape">

## The shape of the choice

A flat file, a git repo, a local database, a hosted database, a key-value service. Each is a real answer, and each quietly trades away something the others keep. You do not pick by memorizing them. You pick by answering a handful of questions about *your* data, and letting the answers point at the simplest option that fits. Those questions are next, and they are the part worth actually holding in your head.

</section>

<section class="sh-layer" id="questions">

## The 5 questions that decide it

Each maps to a place your agent will otherwise guess silently. These are the "what should I ask" set.


A single save is rarely a single action underneath. The simplest approach, truncate then write, means there is a moment where the old data is gone and the new data is not there yet. A crash in that gap loses everything. The fix is to write the change atomically, or to use a store that keeps a replayable log.








One writer is easy. Two writers racing to change the same thing is where flat files quietly corrupt. Real databases serialize this for you with a lock or a transaction; files and git do not.








Fetching one item by its name (key-value) is something almost anything can do. "Everyone who signed up last week and has not logged in since" is a query, and the moment you need those, a file forces you to load everything and loop by hand. That is the line where a database stops being optional.








Your agent will happily "add backups" by writing a second copy next to the first, on the same disk, corrupted by the same crash. A backup only counts once you have restored from it and watched the system come back. Until then it is a hope.








The honest answer is usually "one machine, and I am over-worrying." But the moment data lives on more than one machine you inherit a genuine law: you cannot have perfect consistency, constant availability, and survival of a network split all at once. You pick which bends.

If nobody chooses, the system chooses for you, usually at the worst possible moment.







</section>

<section class="sh-layer" id="options">

## The options

Verdict first. Open one for what it really is, what it trades, and the gotcha your agent will not mention.

A rule of thumb before you drill in: start as far *left* on the ladder as your five answers allow, and move right only when a real need pushes you there. Most over-engineering is reaching right by habit.


**What it is.** You read the whole thing into memory, change it, write it back. No server, no setup, you can open it in a text editor. For a config file or a one-off script it is the right call.

#### What it trades away
Concurrency and durability the moment the app is real. Two writers clobber each other; a crash mid-write can leave a half-file; there is no query beyond "load it all and loop."

**Verdict.** Perfect for small, single-writer, human-readable data. The first thing to outgrow, and it outgrows quietly.








**What it is.** Still files, but with a complete, auditable history and the ability to branch and merge. Wonderful for content, configuration, and anything a human edits deliberately.

#### What it trades away
It is not built for programs writing constantly or at once. Two processes committing at the same time is a conflict, not a transaction.

**Verdict.** A versioning layer, not a database. If your data reads like a document, git is underrated. If it reads like live state, it is a trap.







**What it is.** A full SQL database with transactions and real queries, that lives in one file and needs no server. Durability and safe concurrent reads with almost none of the operational cost of a server.

#### What it trades away
Heavy concurrent *writes* from many processes at once, and running across multiple machines. For most apps that never happens, which is why "just use SQLite" is good advice more often than people admit.

**Verdict.** When in doubt, start here. Real database guarantees, and you can move to Postgres later if you genuinely outgrow it.







**What it is.** A separate program your app talks to over a connection. Built for many clients writing at once (high concurrency), strong guarantees, and serious querying. The workhorse once an app has real concurrent users.

#### What it trades away
Simplicity and operational cost. Now there is a server to run, secure, back up, and connect to. Taking that on before you need it is the most common over-engineering your agent will cheerfully do.

**Verdict.** Right when you genuinely have concurrent writers or heavy queries. Wrong when you reached for it out of habit.







**What it is.** A company runs the database, backups, and scaling for you, behind an API or a connection string. You trade money and control for not being the one paged at 3am.

#### What it trades away
Control and portability. Your data and often your data model now live in someone else's shape; moving off later is real work. Every call also crosses the network, so "fast" has a floor you do not set.

**Verdict.** Great leverage for a small team that should not run infrastructure. Just know what the exit costs before the data is shaped around it.






</section>

<section class="sh-layer" id="deep">

## Go deeper: the fundamentals underneath

The concepts the verdicts lean on, in plain terms, each with a prompt that makes your agent apply it to *your* project.


A file is bytes; everything else is your code's job. A database bundles the hard parts: a change either fully happens or not at all (a transaction), two writers do not corrupt each other (concurrency), and you can query across the data without loading it all. You are not paying for storage, you are paying for those guarantees.












Atomicity (all-or-nothing), Consistency (rules never break), Isolation (concurrent writes do not see each other half-done), Durability (once it says saved, it survives a crash). When people say "a real database," this is mostly what they mean.











The law behind "Scale" above. The network *will* split sometimes; when it does, you either refuse writes to stay consistent, or accept them and reconcile later to stay available. Pick on purpose, because the default is chosen for you under failure.

<details class="sh-more"><summary>Go deeper: the three promises, and why you only get two</summary>

##### Consistency
Every reader sees the latest write, everywhere, immediately. Ask two different servers the same question at the same moment and you get the same answer.

*Think of it like a single shared whiteboard. There is one board, so everyone is always looking at exactly the same thing.*

##### Availability
Every request gets an answer, right now, without waiting. The system never says "come back later."

*Think of it like a shop that is always open. You might occasionally get slightly out-of-date information, but the door is never locked.*

##### Partition tolerance
The system keeps working even when the network between its machines breaks and they can no longer talk to each other.

*Think of it like two branches of that shop whose phone line just went down. Each one keeps serving customers on its own, without checking with the other.*

##### Why you only keep two
On a single machine there is no network to split, so none of this bites. The moment your data lives on more than one machine, the network *will* occasionally break. When it does, a machine that receives a write has to choose: accept it, and now the two sides disagree, so you have given up consistency; or refuse it until the line is back, and now you have given up availability. You cannot do both. And across machines you do not really get to drop partition tolerance, because the split happens whether you planned for it or not. So the real choice is only ever this: under failure, does consistency bend, or does availability?

<p class="sh-ex"><b>Real example: an ATM network</b>Cut off from the bank, an ATM can refuse every withdrawal to stay perfectly consistent, or let you take out up to $200 to stay available and reconcile the balance later, accepting a small risk of overdraft. Most banks choose availability with a cap. That is the whole theorem, playing out in a parking lot.</p>

</details>











Before changing the data, the database records "I am about to do X" in an append-only log. If it crashes, on restart it replays the log and finishes or discards cleanly. It is why a database can promise durability and a naive file-overwrite cannot.










</section>

<section class="sh-layer" id="examples">

## Worked examples: the same choice, three apps

The questions only click once you run them on something real. Here is the reasoning out loud for three apps people actually build. The answer comes out different every time, and the simplest option wins more often than you would expect.


One person, one device, a few hundred items at most. Walk the five questions and listen to how fast they collapse.

Durability matters a little: you would hate to lose today's list to a crash, but a careful atomic save covers that. Concurrency is a non-issue, because there is exactly one writer, you, and you are not tapping the same checkbox from two phones at once. Query is trivial: you load "my lists" and show them, you are never asking "all todos across all users where due date is Tuesday." Recovery is mostly handled for you, since the phone's own backup already copies the file to the cloud. Scale is one device, probably forever.

Add it up and the honest answer is a single file on the device. The only thing that nudges you up a rung is wanting instant search across thousands of items, or the same list on your laptop too, and then it is on-device SQLite, *still* no server.

**The pick.** A flat file. Move to on-device SQLite only if you add real search or multi-device sync. Reach for a hosted database here and you have built a cathedral for a garden shed.



Now almost every answer flips. Concurrency is real: a background job writes fresh numbers every few minutes while several teammates load the page at once, so you have concurrent readers and a writer touching the same data. Query is the entire point: "revenue by week," "top ten customers this month," all of which is querying across records, the one thing files are worst at. Durability and recovery stop being shrugs, because this is shared business data and losing it, or failing to restore it, is an actual incident. Scale is likely just one machine, but it has to be dependably up during the workday.

Add it up and you want a real database server. A flat file would corrupt under the concurrent writes and choke on the queries by the second week.

**The pick.** Postgres, or a hosted database if nobody on the team wants to run a server. Files and git are off the table from the start.



This one sits in between, and it is the most instructive. One primary user and modest data, which sounds like a file. But you will absolutely want to query it, "everyone I have not talked to in ninety days" is the whole reason the thing exists, and that pushes you past a flat file to SQLite. Then the second pressure shows up: you want it on your phone *and* your laptop, and the instant two devices need to see the same data, a local file on one of them stops working.

Notice the deciding question here is not durability or scale. It is "do I need this in two places," and only you know the answer. That single yes-or-no moves you a whole rung.

**The pick.** SQLite if it honestly lives on one device. A small hosted service the moment you want it in two. The multi-device wish is the entire decision.






</section>


Before you choose how to store this data, answer out loud, for THIS project:
  1. Durability  - if you crash mid-write, is the data whole or corrupt?
  2. Concurrency - how many writers at once, and what happens on a collision?
  3. Query       - fetch-by-key only, or query across records?
  4. Recovery    - what is the restore procedure, and have we tested it?
  5. Scale       - one machine forever, or will this cross machines?
Then recommend the SIMPLEST option that satisfies all five, and name what it trades away.



---
This is the plain-markdown twin of https://iambenschmidt.com/shape/how-to-store-data/. The whole site offers one: append index.md to any page URL. Index for agents: https://iambenschmidt.com/llms.txt
