Cloud and cost
Multi-cloud operations, cost optimisation, infrastructure drift and infrastructure hygiene.
What does the Cloud and cost cluster cover?
The Cloud and cost cluster on the DevOpsArk blog collects 5 articles on multi-cloud operations, cost optimisation, infrastructure drift and infrastructure hygiene.
Everything in Cloud and cost
What is infrastructure drift, and why your Terraform state keeps lying to you
Infrastructure drift is the gap between what your Terraform state says and what is actually running. Why it happens, what breaks, and how to catch it before an incident.
SSL certificate management: preventing the outage nobody planned for
Why certificate expiry still causes outages, how to find the certificates nobody documented, and how to automate renewal so the problem stops recurring.
Kubernetes troubleshooting: a decision tree that works
The common Kubernetes failures, what each one actually means, and the order of checks that reaches the cause fastest.
Multi-cloud DevOps: making one operating model work across providers
Why organisations end up multi-cloud, what it genuinely costs, and how to build one operating model across providers without pretending they are identical.
Kubernetes cost optimization: where the money actually goes
Why Kubernetes clusters cost more than they should, how to attribute spend to teams, and the specific changes that produce the largest savings.
What the platform does about this
Servers
Fleet inventory, patching and access
Security
Posture, policy and continuous verification
360 DITE
Delivery, infrastructure, testing and experience in one score
Kubernetes
Multi-cluster Kubernetes management
SSL Management
Certificates that renew before they expire
DNS Management
Records, zones and change safety
Monitoring
Infrastructure and application monitoring
Log Management
Centralised logs with structure and retention
Elsewhere on the blog
DevOps
Fundamentals, automation strategy, incident practice and the platform-versus-toolchain question.
Kubernetes
Architecture, monitoring, deployment strategy, multi-cluster operations and troubleshooting.
Observability
Metrics, logs, traces, alerting design and what observability actually means.
AI DevOps
Agentic DevOps, AI log analysis, incident response and where automation should stop.
DevSecOps
Container and Kubernetes security, vulnerability management, secrets handling, audit trails and compliance evidence.
Want a structured route through this?
Learning tracks arrange these articles into an ordered path with the glossary terms they depend on.