Skip to content
New this month 24 fresh C++, C#, F#, JavaScript, TypeScript and Lua courses just landed. Browse new releases Use code WELCOME10 for 10% off your first order · 14-day refund

Production Readiness: Observability, Incidents and On-Call

Metrics, traces, alerts and postmortems: what it takes to own a service people depend on.

12 students

TE Created by Tom Eriksen

  • Last updated August 2026
  • English
  • 21h of material
  • 157 lessons

What you will learn

8 concrete outcomes

Every bullet below is something you will have built, shipped or be able to explain by the time you finish the last lesson.

  • Instrument structured logs with identifiers that survive service boundaries
  • Choose metrics with affordable cardinality around the four key signals
  • Trace a request across services and a queue with sensible sampling
  • Define service level objectives and manage an error budget honestly
  • Alert on symptoms with burn rates rather than on every cause
  • Run an incident with a commander, a scribe and clear communication
  • Write a blameless postmortem and design an on-call rotation people can sustain
  • Run a readiness review before a service carries real traffic

Course curriculum

8 modules · 157 lessons · 21h of material

20 lessons running 2h 34m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.

Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.

20 lessons running 2h 40m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.

Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.

20 lessons running 2h 44m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.

Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.

20 lessons running 2h 46m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.

Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.

19 lessons running 2h 38m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.

Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.

19 lessons running 2h 36m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.

Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.

19 lessons running 2h 34m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.

Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.

20 lessons running 2h 28m in total. Each lesson ships with the finished source files and a short written recap, so you can follow along in your own editor and skim the module again later.

Lesson-by-lesson titles, code downloads and exercises live inside the course library you get access to straight after checkout.

8 modules · 157 lessons

21h total length

Requirements

Short list, and deliberately so. If you meet these you can start today.

  • Experience running services in production
  • Comfortable with containers, deployment pipelines and a cloud environment
  • Some exposure to being on-call, even informally
  • Docker and a local Kubernetes cluster for the lab environment

About this course

There is a moment when a service stops being a project and becomes something people depend on. Everything changes at that point, and very little of what changes is taught anywhere. This course is twenty-one hours on exactly that transition.

Observability comes first, built rather than described. Structured logging with correlation identifiers that survive a service boundary. Metrics with cardinality you can afford, instrumented at the four signals that matter. Distributed tracing through a request that touches three services and a queue, with sampling decisions that keep the cost sane.

Alerting is treated as a design problem with a human on the other end. Service level objectives and error budgets, alerting on symptoms rather than causes, burn rate alerts that fire early enough to matter and rarely enough to be trusted, and a ruthless audit of alerts that wake people up and teach them nothing.

The second half is the human system. Running an incident with a defined commander and scribe, communicating during an outage to people who are not engineers, writing a blameless postmortem that produces changes rather than promises, sustainable on-call rotations and handover, chaos and load exercises, and the readiness review a service should pass before it is allowed to carry real traffic. A complete review checklist ships with the course.

Frequently asked questions

Still unsure about something? Write to misteryjj100@gmail.com and a human answers, usually the same working day.

OpenTelemetry for instrumentation with Prometheus, Grafana and a tracing backend running locally, so the skills transfer to whatever your company already pays for.

Especially so. A dedicated module covers doing this work with three engineers rather than thirty, including which practices to adopt first.

No. A deliberately faulty sample system runs locally and reproduces every incident scenario, including the ones that only appear under load.

Checkout is handled on our provider's secure payment page. The moment your payment clears we email your personal access link and access code to the address you used at checkout, and the same link appears in your account library. There is nothing to install and nothing to wait for.

Email misteryjj100@gmail.com within 14 days of your purchase, quote your order number, and we refund the full amount to your original payment method. No form to fill in and no questions about how much of the course you watched.

What students say

No reviews yet

Nobody has reviewed this course yet, so there is no score to show. Buy it, work through it, and your review could be the one that helps the next developer decide.

Your instructor

TE

Tom Eriksen

Platform engineer and DevOps instructor

  • 219 students taught
  • 13 courses published
  • 4.8 instructor rating
  • Docker
  • Kubernetes
  • CI/CD
  • Testing and tooling

Tom has built delivery pipelines for teams of three and teams of three hundred, and openly prefers the smaller ones. He teaches Docker, Kubernetes and CI/CD outward from a single commit, so nothing ever appears by magic, and he treats the test suite as part of the pipeline rather than a separate subject. Every course includes a deliberate outage that you have to diagnose from logs and metrics before you can continue. He is a CNCF contributor and a reluctant but formidable YAML expert.

$179 USD

One-time payment · lifetime access

The WisdomCharms dispatch

One useful email a week. No fluff, no spam.

New course releases, discount codes before anyone else, and a short, practical breakdown of one technique — a prompt pattern, a C++ idiom, a TypeScript trick — that you can use the same day.

  • Subscriber-only launch pricing
  • Unsubscribe in one click
  • We never sell your address

By subscribing you agree to our Privacy Policy. Questions? Write to misteryjj100@gmail.com.