Bring up both databases — your own or the compose file — load the dataset at a scale you choose, and produce the evidence that the lab is sound. The deliverable is not 'it ran'; it is a verification report a reader could use to confirm your environment will produce the numbers the rest of the course expects, plus a committed baseline to measure everything else against.
Do the connection check before the seed, not after. A load against a half-working connection produces a partially populated dataset that passes a casual glance and then makes every measurement in the course quietly wrong. If you are on a free managed tier, start with SCALE=small and confirm the whole pipeline works end to end before deciding whether the default will fit — the storage limits on those tiers are lower than people expect.
$ node check.js
postgres ok PostgreSQL 16.4 on aarch64-unknown-linux-gnu
mongodb ok MongoDB 7.0.14
$ SCALE=default node seed.js
postgres users 200000/200000
postgres orders 2000000/2000000
postgres events 5000000/5000000
postgres ANALYZE done
mongodb users 200000/200000
mongodb orders 2000000/2000000
mongodb events 5000000/5000000
seeded in 174s
$ mongosh "$MONGO_URL" --quiet verify.js
users 200000 / 200000
orders 2000000 / 2000000
events 5000000 / 5000000
orders/user cv 4.18
data 612MB uncompressed, 288MB on disk, cache 256MB
lab verified — every later measurement will be comparable
$ ./baseline.sh > baseline-20260913.txt
## postgres
orders by user+status median 1.402s
plan:
Seq Scan on orders (actual time=204.1..1402.1 rows=23 loops=1)
Filter: ((user_id = 8812) AND (status = 'refunded'::text))
Rows Removed by Filter: 1999977
Buffers: shared hit=32210
scale=default · MacBook Pro M2, 16GB · docker compose