DevOpsArk Engineering

Engineering team, DevOpsArk

Articles written collectively by the DevOpsArk engineering team, the people who build the modules described on this site. Technical deep dives, architecture notes and the reasoning behind specific product decisions.

Distributed systemsObservabilityContainer securityContinuous delivery
16 articles

Written by DevOpsArk Engineering

7 min read

Blameless postmortems: the format that actually gets filled in

Most postmortem templates go unused after the first few sections. Why that happens, and what a blameless postmortem format looks like when people actually complete it.

7 min read

What is infrastructure drift, and why your Terraform state keeps lying to you

Infrastructure drift is the gap between what your Terraform state says and what is actually running. Why it happens, what breaks, and how to catch it before an incident.

7 min read

SIEM vs audit log vs observability: three different jobs, constantly confused

SIEM, audit logs and observability all involve data about your systems, but they answer different questions for different people. How to tell them apart.

6 min read

SSL certificate management: preventing the outage nobody planned for

Why certificate expiry still causes outages, how to find the certificates nobody documented, and how to automate renewal so the problem stops recurring.

9 min read

Kubernetes troubleshooting: a decision tree that works

The common Kubernetes failures, what each one actually means, and the order of checks that reaches the cause fastest.

6 min read

AI versus traditional automation: choosing the right one

A decision framework for when a deterministic script is the correct answer and when a reasoning agent earns its extra complexity and risk.

7 min read

AI log analysis: making millions of lines legible

How log pattern extraction works, why it finds errors that search never will, and how to use novelty and rate detection without generating a new source of noise.

8 min read

AI agents in DevOps: where they help and where they do not

A practical assessment of where AI agents genuinely improve infrastructure work, where they are oversold, and how to introduce them without creating a new class of incident.

8 min read

Kubernetes deployment strategies: rolling, blue/green and canary

How rolling updates, blue/green and canary deployments differ, what each actually protects you from, and how to choose per environment rather than per opinion.

7 min read

Logs, metrics and traces: which signal answers which question

What each telemetry type is genuinely good at, what it costs, and how to decide where a given piece of information belongs.

9 min read

How to automate Docker builds without hand-writing Dockerfiles

Multi-stage builds, layer caching that actually works, hardening defaults, and how to generate and maintain container definitions across a large service estate.

6 min read

Monitoring vs observability: a distinction worth keeping

The difference between monitoring and observability, why the distinction is more than marketing, and what each one is actually for.

8 min read

DevOps automation: what to automate, in what order

A practical sequence for automating a delivery path, which step gives the most back first, which automation tends to be regretted, and how to tell when a stage is genuinely done.

11 min read

Kubernetes monitoring: what to measure and why

A practical guide to Kubernetes monitoring: the cluster, node and workload signals that predict failure, the ones that only look useful, and how to alert on them without drowning.

8 min read

What is observability, and how is it different from monitoring?

A definition of observability that does more work than "the three pillars": what property you are actually trying to obtain, and how to tell whether you have it.

9 min read

Kubernetes architecture explained: the parts that matter operationally

What each Kubernetes component actually does, which ones you will meet during an incident, and the mental model that makes cluster behaviour predictable.

Questions about any of this?

The people who write here also take the demo calls.