Not another dashboard course. The code you write to emit telemetry — span kinds, cardinality, and the browser half nobody instruments.
Your API p99 is eighty milliseconds and a user waited three seconds. Your dashboard is green and nobody can tell you which service was slow. This course is about the code that closes that gap. You start from what you were asked during your last three incidents, then write the instrumentation that would have answered it: span kinds, events and status, and the traceparent header — including the step everyone forgets, extracting a context and never making it current, which silently splits every trace in two. Then the parts most services never instrument: trace context through Kafka, SQS and Celery, and the queue time that falls between two traces and is usually the real latency. Cardinality, where four reasonable labels multiply into twenty million series and the ingester is OOM-killed during the incident that label was added to debug. Then the browser, where the missing seconds usually are: Web Vitals with attribution that names the element, error capture, source maps, and one waterfall running from the click to the database. It ends with the Collector and a migration off ad-hoc logging. Python, Node and TypeScript, every technique verifiable on localhost with no vendor account.
Built by Lakshya Kumar
Paste this into any AI chat. Fill in the bracketed parts with your context — you'll get back a straight answer on whether this belongs on your plate.
We grant free access case-by-case — students, career-switchers, builders on a tight budget. Sign in to send us a note.
Sign in to applyFinished the tasks? Take the prompt to your AI and get tested on it. We copy the prompt and open the app — just paste it in.
Span kind decides how backends compute service graphs. Events record moments. Status decides what counts as an error.
One 55-character header is what turns per-service spans into one distributed trace. Read it byte by byte, then carry it correctly.
A message is a carrier. Put the context in it, link rather than parent, and stop losing the half of your system that runs later.
Six instrument types, one aggregation pipeline you can reshape without touching code, and exemplars that jump to a trace.
Series count is the product of every label's distinct values. Multiplication is why one innocuous label ends a quarter's budget.
Inject trace_id into every log record and the pivot between 'which request' and 'what the code decided' becomes one click.
Lab tools measure one fast laptop. Real users are on cheap Androids on bad networks, and only they can tell you the truth.
A minified stack trace is useless and a swallowed rejection is invisible. Capture everything, then symbolicate it.
One trace from the click to the database. The browser becomes the root span, and CORS is the thing that stops you.
Receivers, processors, exporters and the tail sampler — the one place you can change telemetry behaviour without a deploy.
A staged plan for a real codebase: audit, pick boundaries, strangle service by service, prove parity, delete the old path.
I'm taking "Instrument Your Stack: Traces from Browser to Database" — a telemetry instrumentation course in the Engineering track. Twelve modules: signal design, manual spans (kinds/events/status), context propagation, queues and async boundaries, the OTel metrics API, cardinality, logs joined to traces, browser Web Vitals as real-user telemetry, browser error capture and source maps, stitching browser spans to server traces, the Collector (pipelines, OTTL, tail sampling), and migrating a codebase off ad-hoc logging. My context: 1. My stack and languages: [describe] 2. What telemetry exists today: [nothing / logs only / a vendor APM / partial OTel] 3. Whether I have a front end I control: [yes/no, framework] 4. Whether anything crosses a queue: [describe] 5. My monthly telemetry spend, if I know it: [number] 6. The last incident I could not diagnose quickly: [describe] Given that, answer: - Which module should I start with, given what already exists? Skipping ahead is fine if the earlier ground is covered. - Which parts of this course do NOT apply to my situation, and why? - For my last undiagnosable incident, which specific signal and attribute would have answered it? - What is the smallest change that would most improve my ability to diagnose the next one?
Short and worth reading in full before module 3. The processing-model section is the part people skip.