Your Terraform state describes the infrastructure you think you have. Drift is the gap between that and what is actually running, and it opens whenever a resource is changed outside the tool meant to manage it: a console edit, an emergency fix, another automation. Drift is close to inevitable, so the goal is detecting it quickly with scheduled plans, logging out-of-pipeline changes and codifying manual fixes.
What is infrastructure drift?
Infrastructure drift is the difference between the infrastructure declared in code (a Terraform state file, a CloudFormation stack or another infrastructure-as-code source of truth) and the infrastructure actually running. It happens when a resource is changed outside the managing tool, for example by a manual console edit, an emergency fix during an incident or a separate automation touching the same resource.
Why infrastructure drift happens
Drift happens whenever a change bypasses the infrastructure-as-code tool that is supposed to be the source of truth. The common causes:
- Manual console changes. Someone opens the AWS, Azure or GCP console, often during an incident and under time pressure, and makes a change that never finds its way back into the Terraform configuration.
- Emergency fixes. The fastest way through an active incident is rarely "write a PR, get it reviewed, run the pipeline". Engineers reach for the console or a direct CLI command, fully intending to backport the change later, and often do not.
- Other automation touching the same resources. A separate script, another team's pipeline or a third-party tool modifies a resource Terraform also manages, and Terraform never knows.
- Provider-side changes. Occasionally a cloud provider changes a default or a resource attribute without any action from the team.
None of these are exotic edge cases. They are ordinary parts of operating infrastructure day to day, which is why drift is closer to inevitable than rare.
Why Terraform state "lies" once drift happens
The Terraform state file is a snapshot of what Terraform believes about your infrastructure, based on the last time it applied a change. It does not continuously verify reality. It trusts its own record until something forces a re-check.
Once a resource has drifted, the state describes infrastructure that no longer exists in that form, and Terraform cannot know until someone runs a plan or apply and it queries the provider. Until then, every plan reasons about an assumption rather than a verified fact. That is how a drifted resource sits unnoticed for weeks and then produces a baffling result the next time someone touches that part of the configuration.
# Read-only drift check: refresh against the provider, change nothing.
# Exit code 2 means the live infrastructure differs from the configuration.
terraform plan -refresh-only -detailed-exitcode -input=falseWhat breaks when drift goes undetected
- An apply can revert a manual fix. If an emergency change was never backported into code, the next routine apply silently undoes it and reintroduces the original problem.
- Incident investigations take longer. An engineer reads the Terraform configuration, sees what should be running, and loses time before realising the real infrastructure has diverged.
- Security posture quietly degrades. A security group rule loosened temporarily during an incident and never reverted or codified can persist far longer than intended. It is exactly the change an audit trail should record and a drift check should flag.
- Trust in the source of truth erodes. After being burned a few times, engineers start doubting whether the code reflects reality at all, which undermines the point of infrastructure as code.
How teams detect and manage drift
- Scheduled drift detection. Run a read-only plan against live infrastructure on a schedule, even when nothing is being applied, so drift surfaces before it causes an incident.
- Treat manual changes as alert-worthy events, not silent ones. A change made outside the pipeline is exactly what an audit trail should capture. Undocumented changes are hard to trace precisely because they do not leave the record they should.
- Backport emergency fixes as a standard incident follow-up. Make "codify the manual fix" a required postmortem action item rather than optional cleanup.
- Restrict console access to production resources managed by infrastructure as code, so manual changes become a deliberate exception instead of the easy default under pressure.
Common mistakes
- Assuming Terraform will notice drift on its own without a scheduled detection step.
- Never backporting emergency manual fixes, so drift accumulates invisibly.
- Treating drift detection as a one-time setup task rather than an ongoing, scheduled practice.
- Not correlating drift with the audit trail, and so investigating the same problem twice.
How DevOpsArk approaches this
DevOpsArk manages clusters, servers and cloud accounts centrally across environments, and reports drift from a declared baseline as a specific difference rather than a vague warning. The audit trail records every DevOps event, including changes made outside the normal pipeline, with full history retained, so a drift report comes with a record of exactly when and where the manual change was made. 360 DITE adds automated infrastructure assessment and reliability analysis, which turns drift detection into an ongoing capability rather than a manual step that is easy to skip.
Key takeaways
- Drift is the gap between declared infrastructure and what is actually running.
- It comes from ordinary operations: console edits, emergency fixes and overlapping automation.
- Terraform state is a record of the last apply, not a live view of reality.
- Undetected drift can let a routine apply revert a manual fix and restart an incident.
- Schedule read-only plans, log out-of-pipeline changes and make codifying manual fixes a postmortem action.
Frequently asked questions
Any change made outside the infrastructure-as-code tool that manages a resource: manual console edits, emergency fixes, other automation touching the same resources, or provider-side changes.
Run terraform plan, or a read-only refresh-only plan, on a schedule against live infrastructure. It surfaces differences between the state file and the actual resource configuration without changing anything.
Yes. A routine terraform apply can silently revert a manual emergency fix that was never backported into code, reintroducing the original problem without warning.
Not entirely, because emergency fixes and manual changes will happen. It is more realistic to detect and correct drift quickly and consistently than to assume it will never occur.
An undocumented manual change is often the root cause of both drift and a gap in the audit trail. The same missing record that would explain what changed is what would have flagged the drift sooner.