Pick a depth. Each prompt opens in your AI pre-loaded with the lesson. Click a row to preview the prompt.
A generator that only produces one size is a generator that only works on one machine. The default dataset here is about 1.4GB across seven million rows, which is the right size for the lessons — big enough that a missing index costs seconds, small enough to reload over lunch. But it will not fit inside a free Neon or Atlas tier, and on a small laptop the large variant will take an uncomfortable amount of time. Making the size a single environment variable costs almost nothing and means nobody is blocked. The other property worth building in is resumability: a load of seven million rows will occasionally be interrupted, and a script that starts again from zero every time turns a two-minute annoyance into a twenty-minute one. Batching plus a resume check gives you both speed and the ability to pick up where it stopped.
The three scales, the generator for both engines, and the two properties that make it survive contact with a real machine: batching and resumability.
# ── The three scales ────────────────────────────────────────────
#
# SCALE users orders events approx size use when
# --------------------------------------------------------------------
# small 20,000 200,000 500,000 ~140 MB free tier,
# small laptop
# default 200,000 2,000,000 5,000,000 ~1.4 GB THE DEFAULT —
# every number
# quoted in this
# course
# large 2,000,000 20,000,000 50,000,000 ~14 GB module 11's
# partitioning,
# module 7's BRIN
#
# Pick with the env var from task 1:
$ SCALE=small node seed.js
$ SCALE=default node seed.js # or just: node seed.js
$ SCALE=large node seed.js
# ── What the numbers in later modules assume ───────────────────
# Every quoted figure — "1402ms", "keys=59921", "32210 buffers" — comes
# from SCALE=default. On small they will be proportionally lower; on
# large, higher. The RATIOS are what matter and they hold at every
# scale, which is the point of the exercise.
# ── Timing, roughly, on a 2023 laptop ──────────────────────────
# small ~15 seconds
# default ~3 minutes
# large ~35 minutes (run it before lunch, not before a lesson)
$ SCALE=default node seed.js
postgres users 200000/200000
postgres orders 2000000/2000000
postgres events 5000000/5000000
postgres ANALYZE done
mongodb users 200000/200000
mongodb orders 2000000/2000000
mongodb events 5000000/5000000
seeded in 174s