Rzemiosło · The Craft
What is The CraftLevelsGet startedChapters Search Download PLEN
Level 3 · Codex ~4 min read

In one sentence: What happens after launch, when real users and external systems hit the service — and how to harden it.

This chapter in plain terms

This chapter is about what happens to an app after launch — when real users and an unreliable outside world hit it. The motto: the worst failures don’t shout, they quietly hang, log people out, or drain your budget.

A few defences worth knowing by name:

One bad request must not take down the whole server. A global “error catcher” is enough so that one request’s error breaks that one, not the service for everyone at once.

Treat other people’s services as unreliable. Every connection to an external API — An agreed way for two programs to talk to each other — one asks, the other answers in a set format. Through an API your app connects to outside services (payments, maps, AI). Treat an API key like a password. gets a timeout and retry with growing pauses (backoff) — otherwise one hung connection hangs forever. And check the content of the response, not just the status: “200 + an error page” is failure disguised as success.

Every paid resource behind a hard cap. An Endpoint — A single “address” in an API you send a request to for a specific thing (e.g. the list of orders). Apps talk through endpoints — one endpoint = one function you expose. calling a paid API (AI, maps, mail) with no limit is an open tab — set an amount per user, with a clear message and a reset date.

Know that prod is alive (observability). Logs with levels, a simple “health check” and one external ping — so you learn about a failure from the system, not from a customer.

Watch cyclical jobs by their effect, not the “green report”. A nightly job fails quietly — nobody’s watching at 4 a.m., and an outage looks just like a calm night. Exit code “0” and a log saying “26 scrapers, 0 errors” can be true while the data has sat frozen for three days. So the Watchdog — An independent monitor that checks by itself whether data is fresh / the app is alive, and alarms when something stalls. For cyclical jobs measure the effect (age of the freshest record), not the job’s exit code — a green report can lie. measures the age of the freshest record per source (when something actually last landed), not whether the job ran. And remember: the scheduler’s settings (catch-up after a missed start, “don’t run on battery”) are production too — they live outside the repo, so write them into the runbook.

Example of a silent trap: a hosting server can block mail port 465. A send “set to the secure port” then hangs until the timeout instead of throwing an error — which is why you use port 587 (STARTTLS) and hard time limits, so a failure is loud.

Read more: rate limiting · exponential backoff (retrying) · observability.

14 — Operational resilience and external dependencies

Commandments III, V, VI at the runtime layer: prod lives in an unreliable world. The network drops, the provider blocks ports, the Scraping — Automatically gathering data from websites with a program instead of copying by hand. Powerful for acquiring data, but it needs manners: respect others’ rules (robots.txt), don’t overload the server. dies halfway, a paid API — An agreed way for two programs to talk to each other — one asks, the other answers in a set format. Through an API your app connects to outside services (payments, maps, AI). Treat an API key like a password. costs on every request. The doctrine in chapters 03–05 covers code and deployment; this chapter is about what happens after — when real users and independent systems hit a running service.

These lessons were paid for dearly: nearly every one is a prod incident, not theory. The common denominator — the worst bugs don’t shout, they quietly hang, log out, or drain the budget.

1. One bad request must not take down the process

An uncaught throw/rejection in an async handler can kill the entire server (Node/Express: rejection → process exit → pm2 — A manager that keeps the (Node) app running all the time — restarts it after a crash. Without it, after a crash or server restart the app just sits idle. pm2 keeps it “alive”. restart loop). Build a net at the process level:

  • Global catchers: process.on('unhandledRejection') and 'uncaughtException') — they log and shut down in a controlled way, leaving no zombie process.
  • An async-route wrapper that passes the error to next(err) instead of losing it in an uncaught promise.
  • 500 middleware that doesn’t leak the stack trace to the user (→ 09).
  • Defense on edge data: an OAuth account with no password, a null field where the code assumes a string — these are real inputs once you let real users in (often surfaces after a database swap, → 05).

Rule: a throw in one request degrades that request, not the service. Test it for regressions (→ 03).

2. Long jobs: resumable and detached from the agent session

A “collect everything → save once” scraper/ETL loses 100% of its work on every crash. Write in batches:

  • Checkpoint completed units (a file/table of URLs/IDs) — a restart resumes from where it left off, not from scratch.
  • Save every N, not at the end — a crash costs the last batch, not the whole run.
  • Match the runner to the job’s lifetime. The agent’s background is durable within the session — it survives turns and reports back when it finishes, so use it (→ 16) — but it dies with the session: tearing down the session host kills the process halfway (it’s not anti-bot, it’s a vanishing runner). Anything that must outlive the conversation goes to a detached process or the OS scheduler (→ § 7), never to the session background.
  • End-to-end Idempotency — A property of an operation you can run many times with the same result — no duplication. Key for scripts and events: a re-run doesn’t break data. Like an “ON” switch. (→ 04): a resumed job doesn’t duplicate already-saved data.

3. Treat external sources as hostile

Other people’s APIs and pages drop connections, rate-limit (429), return 200 with an error page. Assume they’re unreliable:

  • A timeout on every call — without it the socket hangs forever (a silent freeze, not an error).
  • Retry with exponential backoff (e.g. 10/20/30 s), with an upper bound on attempts.
  • Rotate session and User-Agent under heavy I/O — a fresh Session (new TCP/cookies) per batch, a UA from a pool of real browsers (some hosts drop you after a few hundred requests from one session).
  • Validate the response, not the status — a “200 + error page” is a failure; check the shape of the data.

4. The provider’s infrastructure imposes limits — verify end-to-end

Hosting has its own network rules that break “working” code only in prod:

  • Ports can be blocked. E.g. outbound SMTP 465 may be closed → use 587 + STARTTLS. secure:true on a blocked port hangs every send until timeout — set hard connection/greeting timeouts so the error is loud.
  • Mail deliverability isn’t “I sent it.” A verified sender domain (SPF/DKIM), a real test send, GDPR — The EU’s data-protection law — how you may collect, keep and delete users’ data. It concerns every app with people’s data. Better to write in consents and retention from the start than pay fines later.-compliant opt-in (→ 09). An email that “went out” but landed in spam/nowhere is a bug.
  • Check this on the provider’s prod/staging, not locally — your local network doesn’t have these blocks (→ 03).

5. Every paid resource behind a hard quota

An Endpoint — A single “address” in an API you send a request to for a specific thing (e.g. the list of orders). Apps talk through endpoints — one endpoint = one function you expose. calling a paid API (LLM, geocoding, email) with no limit is an open tab and an abuse vector:

  • A quota per user (monthly/daily window) with a clear message and a reset date up front.
  • Differentiate per tier (free vs premium), enforce it server-side.
  • Rate-limit + security headers (helmet/limiter) as a permanent part of the stack (→ 08).
  • Tie it to the law and the terms of service: a limit and how you communicate it are also protection against abuse (→ 09).

6. Know prod is healthy — observability

The defenses above keep prod from dying; observability tells you it’s alive — after the deploy window closes, before a user emails you. “Verify, don’t declare” (→ 03) applied to a running system:

  • Structured logs with levels. error / warn / info (not print everywhere). Errors carry context (request id, user, what failed) — but never secrets or full PII. You grep logs at 2 a.m.; make them greppable.
  • A health endpoint. /healthz returning 200 + a cheap check (DB reachable, version) — for the load balancer and an uptime monitor. An external uptime ping tells you it’s down before the customer does.
  • Alert on what hurts, not on noise: spikes in 5xx, failed payments/emails, a job that didn’t run, the selector-drift / “200 + error page” case (→ § 3). One actionable alert beats a hundred dashboards.
  • A few real metrics over vanity: error rate, p95 latency, queue depth, daily cost of paid APIs (→ § 5). Measure before you optimize (→ 13).
  • Errors → a place you’ll see them (a log drain / error tracker), not just stdout that scrolls away.

Start tiny: logs with levels + a health check + one uptime ping covers most of the value. Grow only when a real incident shows the gap (capture the lesson in the runbook → 05).

7. Recurring jobs: watch the result, not the runner

A nightly job is the part of prod that fails silently by design — nobody is watching at 4 a.m., and a failure looks exactly like a quiet night. Three rules, each paid for with an incident:

  • The Watchdog — An independent monitor that checks by itself whether data is fresh / the app is alive, and alarms when something stalls. For cyclical jobs measure the effect (age of the freshest record), not the job’s exit code — a green report can lie. measures the data, not the job. Exit code 0, a log file that grew, “26 scrapers, 0 errors” — every one of those can be true while the data is frozen. Assert on the age of the freshest record per source (MAX(updated_at) grouped by shop/feed/tenant) instead. That survives a job that ran and fetched nothing, an ingest that failed after a successful fetch, and a source added next month — none of which a log-scraping check survives. In the reference project a green nightly report ran for three days beside 44% of the catalogue frozen: the report wasn’t lying, it simply knew nothing about the second data path.
  • A green report only covers what it knows about. When part of the pipeline runs elsewhere — another machine, another schedule, a manual step — the main report’s silence about it is not good news. Either every path ends in its own evidence, or one watchdog sees them all.
  • Scheduler config is production, and it lives outside the repo. That incident wasn’t a code bug: the scheduled task had “run as soon as possible after a missed start” switched off and a “don’t start on battery” condition, so a sleeping laptop meant a skipped night that was never retried. No test, review or git log could have caught it. Write the settings that matter into the runbook (retry-on-miss, power/network conditions, the account it runs as, the working directory → 05) and re-check them after every OS, hardware or account change.

Anti-patterns

  • 🚫 No global error catcher — one bad request restarts the service for everyone.
  • 🚫 A scraper with no checkpoints run from the session background — crash = run from scratch, vanishing runner = perpetual failure.
  • 🚫 An external call with no timeout and backoff — a silent freeze or a ban after a string of 429s.
  • 🚫 “I sent the email” ≠ I delivered it — no verification of port/domain/deliverability.
  • 🚫 A paid API with no quota — a surprise bill and an open abuse vector.
  • 🚫 Sessions in process memory — every deploy logs everyone out (→ 05).
  • 🚫 No logs/alerts — you find out prod is down from the user, not the system; “it works” with no way to know.
  • 🚫 A recurring job watched by its own exit code — it reports success for a week while the data stands still.
  • 🚫 Scheduler settings nobody wrote down — the run is “automated” right up to the night it silently isn’t.
  • 🚫 Secrets/PII in logs — the log drain becomes the breach.

For new projects

Add to Day 0 (→ 07) before the first real user shows up: a global error handler + a 500 with no leak, timeouts and backoff on every external I/O, quotas on paid endpoints, a durable session store in a separate file. It’s cheaper now than as a 2 a.m. incident. Every prod incident → an entry in the runbook with a date and a numbered lesson (→ 05, 06).

The canonical doctrine is written in English.