Skip to main content
By the end of this guide, you have a Slack channel and an on-call email defined in code, retries and degraded thresholds that absorb blips, run-based escalation with one reminder, a location threshold on a parallel monitor, and a maintenance window that mutes a planned deploy.
A Checkly API check detail page for Books catalog, passing again, with run results showing failed runs from N. Virginia and Ireland, each retried twice in the same location before the final result
To follow along without your own app, clone the sample project. It monitors the API and homepage of the Danube demo shop.
To run this guide from your terminal or your coding agent, run npx checkly init in your project first. It installs the Checkly CLI and Checkly Skills for your agent. Then paste the prompt below into Claude Code, Cursor, Codex, or any agent that supports skills. It builds the same setup as this guide, proves it with npx checkly test --record, and stops for your confirmation before npx checkly deploy.
Prompt
The steps below are what the agent does, in the open.

Step 1: Route alerts by who needs them

An alert is only useful if it lands with someone who can act on it. Split your channels by urgency. The team channel hears everything, including slow responses. The on-call inbox only hears about failures and their recovery.
__checks__/alert-channels.ts
SlackAppAlertChannel needs the Checkly Slack app installed in your workspace first. Invite the app to the channel if it is private. Attach both channels as project defaults, so every check gets them without anyone remembering to add them.
checkly.config.ts

Step 2: Absorb blips before they count

Most noise comes from single failed requests: a DNS lookup that times out, a dropped connection, a cold container. The retry strategy in the config above handles those. A failed run is retried twice, 30 seconds apart, in the same region. Only if all three attempts fail does the run count as failed. Retrying in the same region confirms the problem where it happened, instead of hiding it behind a pass from somewhere else. Slow is not the same as down. Give each check two thresholds. Above degradedResponseTime the result is degraded: the Slack channel hears about it, and the on-call inbox does not, because it has sendDegraded: false. Only above maxResponseTime does the check fail.
__checks__/api/books.check.ts
Set the degraded threshold from what you measure, not from what you hope. This endpoint answers in about 20 milliseconds from N. Virginia and 290 from Ireland, so 1 second leaves room for normal variance and still catches a real slowdown.

Step 3: Escalate on runs, and let regions vote

Retries confirm that one run failed. Escalation decides when that is worth a notification. The same check file sets a run-based escalation: two failed runs in a row before anyone is alerted, then one reminder 10 minutes later if the check is still failing. When the check recovers, any pending reminder is cancelled. The threshold is a trade between noise and speed. On a check that runs every minute, two failed runs plus their retries took about three minutes in the test below. On a check that runs every 10 minutes, it is twenty, so lower the threshold or raise the frequency for anything that pages. For a check that runs from several locations at once, count locations instead of runs. The homepage runs in parallel from three regions and alerts only when half of them fail.
__checks__/web/homepage.check.ts
One failing region is recorded against that location and alerts nobody. Two failing regions is a real outage and alerts on the first run. Use an odd number of locations so a 50% threshold never lands on a tie. Monitor from around the globe covers how to choose them.

Step 4: Mute a planned deploy

A deploy that restarts the API is not an incident. Put a tag on what the deploy touches and schedule a maintenance window for that tag.
__checks__/api/group.ts
__checks__/maintenance.check.ts
Checks and groups with a matching tag skip their scheduled runs for the length of the window. Pick a tag only the affected checks carry. A broad tag like api also pauses every other team’s checks that use it. Test everything, then deploy:
Terminal
Terminal
Terminal

Verify it works

Break the check on purpose. In __checks__/api/books.check.ts, change statusCode().equals(200) to statusCode().equals(201) and run npx checkly deploy. This is what happened when the sample was deployed that way:
  1. The first run, from Ireland, failed three times in a row, 30 seconds apart. That was one failed run and no alert.
  2. The next run, from N. Virginia, did the same. Two failed runs met the threshold, and the alert went out about three minutes after the deploy.
The team channel gets the failure with the assertion that broke and the request that was sent:
Slack channel ops-alerts with two messages from the Checkly app: Books catalog failed in N. Virginia with the error Expected 200 to be equal to 201, then Books catalog recovered in N. Virginia 19 minutes later
The on-call inbox gets the same failure:
Checkly failure alert email for Books catalog with the check link, 20 ms response time, location N. Virginia, group Shop API, tag shop, and the error Expected 200 to be equal to 201
Change the assertion back to equals(200) and deploy again. The first passing run sends a recovery to both channels. Open the check in the web app and filter run results by Has retries to see each failed run with its two retries.

Next

Communicate availability with status pages: once alerts reach your team reliably, tell your users what is going on with the same checks.

Reference