Blog

Written by people who run this in production

31 articles across six clusters. Technical, specific, and willing to say when a popular practice is not worth the trouble.

Short answer

What does the DevOpsArk blog cover?

The DevOpsArk blog publishes 31 articles on DevOps, Kubernetes, observability, AI DevOps, DevSecOps and cloud, written by the engineers who build the platform.

All articles

Everything, newest first

7 min read

Blameless postmortems: the format that actually gets filled in

Most postmortem templates go unused after the first few sections. Why that happens, and what a blameless postmortem format looks like when people actually complete it.

7 min read

Compliance readiness vs compliance certification: what auditors actually ask for

"Compliance ready" and "certified" get used interchangeably, but auditors treat them very differently. What the difference is, and what audit evidence actually looks like.

7 min read

What is infrastructure drift, and why your Terraform state keeps lying to you

Infrastructure drift is the gap between what your Terraform state says and what is actually running. Why it happens, what breaks, and how to catch it before an incident.

7 min read

SIEM vs audit log vs observability: three different jobs, constantly confused

SIEM, audit logs and observability all involve data about your systems, but they answer different questions for different people. How to tell them apart.

7 min read

DevOps audit trails: why full event history is your cheapest incident tool

Most audit trails are built for a compliance reviewer and used, months later, by an on-call engineer at 2am. What a useful trail captures and how to make it serve both jobs.

6 min read

SSL certificate management: preventing the outage nobody planned for

Why certificate expiry still causes outages, how to find the certificates nobody documented, and how to automate renewal so the problem stops recurring.

9 min read

Kubernetes troubleshooting: a decision tree that works

The common Kubernetes failures, what each one actually means, and the order of checks that reaches the cause fastest.

8 min read

Multi-cloud DevOps: making one operating model work across providers

Why organisations end up multi-cloud, what it genuinely costs, and how to build one operating model across providers without pretending they are identical.

6 min read

AI versus traditional automation: choosing the right one

A decision framework for when a deterministic script is the correct answer and when a reasoning agent earns its extra complexity and risk.

7 min read

Secrets management: getting credentials out of your repositories

How to centralise credentials, replace long-lived secrets with short-lived ones, rotate without breaking consumers, and respond when a secret leaks.

7 min read

AI-powered incident response: compressing the first ten minutes

Where incident time actually goes, which parts an agent can take over safely, and how to structure incident response so the automation helps rather than adds noise.

7 min read

AI log analysis: making millions of lines legible

How log pattern extraction works, why it finds errors that search never will, and how to use novelty and rate detection without generating a new source of noise.

8 min read

Vulnerability management: turning a report into a work queue

How to make a twelve-thousand-row vulnerability report actionable: deduplication, exposure-based ranking, ownership and verified closure.

8 min read

DevOps platform or toolchain? An honest comparison

The real trade-off between assembling best-of-breed tools and adopting an integrated platform, including the costs of each that vendors on both sides tend not to mention.

8 min read

Alerting that people still read at 3am

How to build an alerting system with a high proportion of actionable pages: correlation, ownership routing, suppression, and deleting the rules that never produce a decision.

8 min read

AI agents in DevOps: where they help and where they do not

A practical assessment of where AI agents genuinely improve infrastructure work, where they are oversold, and how to introduce them without creating a new class of incident.

8 min read

Kubernetes deployment strategies: rolling, blue/green and canary

How rolling updates, blue/green and canary deployments differ, what each actually protects you from, and how to choose per environment rather than per opinion.

10 min read

Kubernetes security: the controls that matter most

A prioritised guide to securing Kubernetes (RBAC, pod security, network policy, secrets and supply chain) ordered by risk reduced rather than by chapter number.

7 min read

Logs, metrics and traces: which signal answers which question

What each telemetry type is genuinely good at, what it costs, and how to decide where a given piece of information belongs.

8 min read

Kubernetes cost optimization: where the money actually goes

Why Kubernetes clusters cost more than they should, how to attribute spend to teams, and the specific changes that produce the largest savings.

9 min read

Container security: the practices that actually reduce risk

Build-time hardening, runtime restriction and supply chain controls for containers, ordered by how much risk each one removes rather than by how often it is mentioned.

9 min read

How to automate Docker builds without hand-writing Dockerfiles

Multi-stage builds, layer caching that actually works, hardening defaults, and how to generate and maintain container definitions across a large service estate.

6 min read

Monitoring vs observability: a distinction worth keeping

The difference between monitoring and observability, why the distinction is more than marketing, and what each one is actually for.

10 min read

Multi-cluster Kubernetes management: patterns that hold up

Why organisations end up with many Kubernetes clusters, what actually gets hard at that point, and the operational patterns that survive contact with a growing fleet.

8 min read

DevOps automation: what to automate, in what order

A practical sequence for automating a delivery path, which step gives the most back first, which automation tends to be regretted, and how to tell when a stage is genuinely done.

8 min read

What is DevSecOps? Beyond "shift left"

What DevSecOps means in practice, why shifting left fails when the feedback is not actionable, and the practices that actually change security outcomes.

11 min read

Kubernetes monitoring: what to measure and why

A practical guide to Kubernetes monitoring: the cluster, node and workload signals that predict failure, the ones that only look useful, and how to alert on them without drowning.

8 min read

What is observability, and how is it different from monitoring?

A definition of observability that does more work than "the three pillars": what property you are actually trying to obtain, and how to tell whether you have it.

9 min read

What is agentic DevOps?

A precise definition of agentic DevOps, how it differs from scripted automation and from AIOps, and the conditions under which an agent is safe to give real access.

9 min read

Kubernetes architecture explained: the parts that matter operationally

What each Kubernetes component actually does, which ones you will meet during an incident, and the mental model that makes cluster behaviour predictable.

9 min read

What is DevOps? A definition that survives contact with practice

What DevOps actually means, where the definition came from, what the lifecycle looks like in practice, and the common misreadings that turn it into a job title instead of a way of working.

FAQ

Questions about the blog

Prefer a structured route through this?

The learning tracks arrange these articles into ordered paths per discipline.