Dmitriy Kononov.
Let’s talkContact

News and analysis

Observability starts with a decision the team needs to make

Connecting OpenTelemetry, service indicators and SLOs to business workflows, operational ownership and infrastructure priorities.

InfrastructurePublished:

Suppose staff report that orders take too long to reach the next processing stage. CPU usage looks moderate, the server is reachable and HTTP error charts are quiet. The system appears operational while the business outcome is delayed. This hypothetical example shows why technical dashboards should follow an agreed workflow definition.

Observability helps when the team can locate the delay and choose an action: fix a query, adjust a worker, limit incoming load or address a dependency on an external provider. The evidence needs to connect the relevant parts of the process. Collecting more indicators does not prove that the necessary answer is available.

Define a successful user action first

A successful event could mean retrieving an up-to-date order status, completing an approval or saving a document. Agree what counts as success, which duration matters and which operations enter the calculation. An HTTP 200 response, for example, cannot establish that the returned information is current.

Google SRE presents SLOs as reliability targets used to guide priorities. An SLI measures a selected property of the service. A percentage copied from another organisation's example needs scrutiny against user expectations, traffic patterns and support costs before it becomes a useful target.

A small team can start with one important workflow. Measure it first and understand gaps in the evidence. Then agree a target, an evaluation period and the response to deterioration. Without a named owner and an action that owner can take, the target becomes another reporting obligation.

Connect a broad symptom to an individual operation

OpenTelemetry explains complementary observability signals, including metrics, traces and logs. In application diagnosis, a metric can indicate the scale of a problem, a trace can show the operation's path, and a log can explain a particular event. Check these roles against the system's actual workflow.

Imagine that saving an order touches the application, a database and an external API. An operation identifier helps connect events across components. If one part is not instrumented, the trace has a gap. Treat it as a diagnostic limitation rather than evidence that this component has no problems.

Observability does not require storing the contents of customer documents or conversations. Duration, operation type, outcome and an appropriate identifier are often more useful. Agree collection boundaries, access and retention before enabling broad capture. Unrestricted collection makes both operation and access control harder to manage.

An alert should lead to a response

Separate notifications that need immediate attention from evidence kept for investigation. High CPU usage may be useful context without requiring intervention every time it changes. Give greater priority to symptoms that disrupt an agreed user workflow and call for an operational response.

An alert should identify the affected process, measurement window, diagnostic entry point and responsible owner. Background operations need their own timeliness indicators, such as the age of waiting jobs. A quiet HTTP error chart says little about whether the queue is being processed on time. Overnight processes also need a defined completion condition.

Review false alarms. Their cause may be an inappropriate threshold, a duplicated signal or an evaluation window that is too short. Silencing every warning removes visibility; adding another delivery channel increases noise. A useful correction addresses the definition of the problem and the path to action.

Choose the first investment around an unanswered question

Selecting an observability platform depends on data volume, support capacity and the required depth of investigation. Before expanding collection, establish storage cost, configuration ownership and the questions that remain unanswered. Instrumenting the entire application at once can delay useful evidence for one critical path.

A strong first stage measures the chosen workflow, allows one operation to be followed across components and defines a response to failure. The findings can then guide decisions about release reliability, database queries and external integrations. This matters when engineering capacity is limited.

For infrastructure work, I suggest starting the discussion with these questions. An SLO does not replace customer commitments or guarantee the absence of failures. Its value appears when the team actually uses the measurement to choose its next action.

In High Ridge Hydroponics, I configured networks, servers and microcontrollers, developed the backend and connected equipment to a GUI for greenhouse climate and indicator management. That work supplies a concrete frame: a person needs to perform an action through the interface, while the engineer examines the path to the equipment. The public case contains no measured SLOs; targets and checks need separate definition.

My guide to moving from spreadsheets to a system continues the discussion of useful outcomes. It starts with coordination problems and missed transitions. A particular transition can then anchor monitoring that helps someone decide what to do.

Sources

Sources checked on 7 October 2026.