Useful observability lets an engineer move from “users cannot complete a task” to an explanation of where and why that task failed. Metrics, logs and traces become valuable when they describe the same system and can be investigated together.
This guide outlines an observability design for a small distributed backend. The running example is an appointment request that crosses an API gateway, a booking service and a database. It is a proposed investigation workflow, not a benchmark or an account of an employer's internal systems.
Begin with questions, not dashboards
Before instrumenting every method, list the questions an on-call engineer needs to answer. Are users completing bookings? Which route became slower? Did the delay begin after a deployment? Is the application waiting for a database connection or executing an expensive query? These questions determine which measurements and contextual fields are useful.
Start with the externally meaningful request boundary. Count eligible requests and successful outcomes, measure duration and record the service version. Distinguish expected validation failures from unexpected service failures. If a user selects an unavailable time slot, the application may be behaving correctly even though the HTTP response is not a success code.
Choose stable names for services and environments. “booking-api” should mean the same thing in a trace, an alert and a deployment record. A naming scheme that changes across teams creates investigative joins that humans have to perform manually during an incident. Keep a small dictionary of resource attributes and review additions before they spread.
Give metrics, traces and logs different jobs
OpenTelemetry's signals documentation describes how telemetry represents system activity. For this workflow, use metrics to identify changes over time, traces to follow the path of a request and logs to capture discrete contextual events. OpenTelemetry instrumentation and collection still need an appropriate storage and analysis backend.
| Signal | Useful question | Example |
|---|---|---|
| Metric | How widespread is the problem? | Latency distribution for the booking route |
| Trace | Where did this request spend time? | Connection acquisition dominates a request span |
| Log | What discrete event explains the wait? | A pool timeout with service version and trace identifier |
A log line saying “database failed” does not identify which operation failed or whether the error affected users. A trace showing a long database span does not establish how frequently that happens. A latency chart cannot explain an individual transaction by itself. Design the investigation path so each signal answers a question the previous signal raised.
Carry context across service boundaries
A request trace is only as complete as its context propagation. Verify that supported instrumentation passes context across HTTP clients, message producers and consumers. An asynchronous job may need a linked trace relationship rather than a simplistic parent-child story, especially when one batch combines work from many requests.
Include trace and span identifiers in structured application logs when the logging integration supports them. Add service identity, deployment environment and version as resource context. Keep business identifiers separate from unbounded metric labels. The following is an illustrative log shape; adapt field names to your logging and telemetry conventions.
{
"severity": "ERROR",
"service": "booking-api",
"environment": "staging",
"event": "database_pool_timeout",
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
"span_id": "00f067aa0ba902b7",
"deployment_version": "example-release"
}
Avoid logging access tokens, session cookies or complete request bodies by default. Redaction should happen before sensitive content leaves the process or collection boundary. An observability platform should not become an accidental second database of private user input. Test redaction with representative fields rather than trusting a policy document.
Turn user expectations into a measurable objective
A service-level indicator (SLI) measures an aspect of service behavior. A service-level objective (SLO) sets a target for that indicator over a defined window. For example, a team might choose a hypothetical objective that 99.9% of eligible booking requests complete successfully over 30 days. The definition must specify what counts as eligible and how success is measured.
The resulting error budget is the allowed bad fraction, in this example 0.1%. A burn rate compares the observed bad fraction with that allowance. If 1% of eligible requests are failing, the instantaneous ratio is 0.01 divided by 0.001, or 10. That means the service is consuming its error allowance at ten times the sustainable rate under the same measurement assumptions.
Google's SRE workbook chapter on SLO alerting explains the motivation for multiwindow burn-rate alerts. A short window helps detect a current problem, while a longer window can confirm that it is sustained. Thresholds and windows should reflect the service's traffic and response expectations; copying numbers without understanding them can create noisy or insensitive alerts.
Work through a concrete incident hypothesis
Suppose the booking success rate falls after a deployment while average CPU remains normal. First compare eligible request volume and failure categories. This establishes whether the apparent change comes from service errors, a traffic-mix change or an instrumentation mistake. Mark the deployment time on the same timeline.
Next inspect representative failing traces. Imagine that most time is spent acquiring a database connection, with short query execution after acquisition. That evidence points toward pool pressure or connection handling rather than an expensive query plan. Check connection usage, waiting requests and the change in instance count before increasing database capacity.
Finally, correlate logs for affected traces and compare the new version with the previous one. A code path may hold connections longer than expected. The immediate mitigation might be a rollback, but the conclusion should remain a hypothesis until the relevant behavior is reproduced and corrected. Observability should narrow uncertainty rather than turn the first correlation into a confident root-cause claim.
Bound telemetry volume and cardinality
Metric dimensions multiply. If a metric includes 20 routes, five methods and six status classes, it can already describe hundreds of combinations before adding service instances or regions. A user identifier can expand that space dramatically. Use normalized route templates such as /bookings/{id} rather than raw paths when defining bounded request dimensions.
Tracing also needs a sampling strategy. Sampling every request may be affordable in a small test environment and expensive at production volume. Whatever strategy is chosen, document what evidence can be missed. A sampled trace set is not an exact count of all failed requests; retain a suitable unsampled aggregate metric for that purpose.
Give telemetry a retention policy and an operational budget. Store detailed data where it supports an investigation and aggregate long-term trends where full detail is unnecessary. Connect this decision to unit-cost analysis, while ensuring the cost reduction does not remove the evidence needed to diagnose failures.
Roll out one complete investigation path
Instrument one important user journey from entry point to persistence. Verify that an engineer can start from an alert, open the relevant service view, inspect a trace and find matching logs. Trigger a controlled failure in staging and have someone unfamiliar with the implementation follow the path.
- Confirm that resource names and deployment versions match across signals.
- Check that errors reach the backend even when export is temporarily delayed.
- Ensure telemetry failure does not block application requests indefinitely.
- Validate that alert messages include ownership and a practical response guide.
- Review whether the observed data answers the original operational questions.
Expand instrumentation after that path works. A small set of coherent signals often provides more operational value than a large collection of disconnected dashboards. The test is whether the team can explain a user-visible problem accurately and choose a measured response.
