Implementation and support
Health checks need business dependency signals
A server can be up while the workflow is not ready. Useful health checks separate liveness from readiness and include the business dependencies that decide whether work can actually move.
Green uptime can hide stopped work
A web application may return a healthy response while invoices are not sending, imports are delayed or a required provider is unavailable. That kind of failure frustrates teams because every basic monitor says the system is up. The workflow is still stuck.
A better health design distinguishes “the process is alive” from “the business workflow is ready.” The first answer helps infrastructure decide whether to restart a process. The second helps operations decide whether users can rely on the system right now.
Separate liveness from readiness
Kubernetes documentation describes probes that can restart unhealthy containers or stop sending traffic to containers that are not ready. That distinction is useful even outside Kubernetes. A liveness check should answer whether the application process can continue running. A readiness check should answer whether it should receive work.
Do not overload one endpoint with every possible dependency. If a transient reporting provider fails, restarting the main application may make the situation worse. If the process is wedged, merely removing it from traffic may not recover it. Separate signals allow different responses.
Include dependencies that affect the task
Readiness should include the dependencies required for the next unit of work. For an intake workflow, that may be the database, file storage and queue. For a scheduled reporting workflow, it may be source freshness and transformation completion. For a notification workflow, it may be provider availability and retry backlog.
Each dependency needs a threshold. A queue with ten seconds of delay may be normal. A queue with three hours of delay may mean the team should stop promising same-day processing. Business thresholds turn raw measurements into useful status.
Show degraded states honestly
Not every problem is total outage. A system may accept new requests but delay attachments, show reports from yesterday or hold outgoing messages. A single red or green badge hides the operational choice.
Use degraded states that explain impact: intake open but notifications delayed, reports available through yesterday, exports paused pending provider recovery. Give support a short explanation and the next owner. Users do not need the stack trace; they need to know whether they can continue the task.
Test checks by breaking dependencies
A health check is only believable if the team has watched it fail correctly. Disable the provider in a test environment, delay the queue, break a credential, make the database read-only and confirm the status changes and alerts match the expected business impact.
The takeaway: health checks should not merely prove that code can answer an HTTP request. They should help the team decide whether work can move, whether traffic should continue and who needs to act when it cannot.