Back to blog

Posts in Incidents

When something breaks, two clocks start: the one until you notice, and the one until your client does. These posts are about keeping the first one short, reading error patterns, routing alerts to someone who is actually awake, and putting something useful on a status page before your inbox fills up.

Maintenance window configuration with a scheduled time range and paused alerting across several client monitors
Incidents

Maintenance Windows Done Right: Planned Downtime Without the Alert Avalanche

How a maintenance window pauses alerting during planned work and shows the downtime cleanly as maintenance instead of an outage in your status history.

Florian Zaskoku8 min read
A monitoring alert routed simultaneously to Slack, Microsoft Teams, and a webhook
Incidents

Setting Up Monitoring Alerts in Slack, Teams & via Webhook

Setup per channel: Slack, Microsoft Teams, webhook, plus the webhook payload structure (JSON) for your own pipelines. With examples.

Florian Zaskoku11 min read
An escalation chain where an unacknowledged alert automatically moves to the next responsible person
Incidents

Setting Up Escalation Policies & On-Call Routing for Small Teams

Escalation chains without PagerDuty complexity: how small agencies without a 24/7 NOC set up reliable on-call routing, step by step.

Florian Zaskoku11 min read
A correctly sent email landing in the spam folder instead of the inbox
Incidents

Client Emails Landing in Spam? What Causes It and How to Catch It Early

Why client emails land in spam: blacklisting, missing DKIM/SPF, poor IP reputation, and how continuous monitoring catches it early.

Florian Zaskoku11 min read
Browser security warning over an expired SSL certificate, scaring visitors away
Incidents

An Expired SSL Certificate: The Reputation Disaster No Client Forgives

What happens when an SSL certificate expires, browser warning, loss of trust, and how automatic expiry alerts prevent the reputation disaster.

Florian Zaskoku10 min read
Overview of HTTP 5xx error codes 500, 502, 503 and 504 with cause and meaning
Incidents

HTTP Error Codes Explained: 500, 502, 503, 504 and What They Mean for Your Clients

An error reference for the key HTTP 5xx codes: cause, fix, and how monitoring reports the error before your client calls.

Florian Zaskoku11 min read
Server running but the site won't load, the DNS resolution between user and server is broken
Incidents

"Site Unreachable" While the Server Runs: When DNS Is the Problem

Why a site looks "down" while the server runs fine, DNS hijacking, expired delegation, wrong records, and how to catch it early.

Florian Zaskoku10 min read
A user checking the server status of a large online service on a status page during an outage
Incidents

When Big Services Go Down (YouTube, WhatsApp & Co.), and What They Teach About Status Communication

Why millions check the server status during an outage, and what the giants' professional status communication gets right that any agency can adapt at small scale.

Florian Zaskoku10 min read
Timeline of status updates during an incident, from the first notice to the all-clear
Incidents

Communicating Incidents & Maintenance Windows: Status Updates That Build Trust

Templates for incident and maintenance communication that reassure the client instead of alarming them, from the first update to the all-clear.

Florian Zaskoku11 min read