Systems · Concept · 11 min read

Infrastructure as Code, and the File That Knows What Exists

Introductions draw the code and the infrastructure. There are three things, not two, and the one they leave out is where almost every bad day with these tools comes from.

Written by Marko Ristic, Editor Updated Sep 17, 2026
3Things in the picture, and introductions draw two
2Failure modes, and both are a disagreement with the state file
0Checks that compare the code to reality directly
1Habit that prevents most incidents: read the plan
Short answer

Infrastructure as code, usually shortened to IaC, means defining servers, networks and cloud services in files that an automation tool reads and applies, instead of building them by hand.

The files go in version control, so infrastructure gets the review, history and rollback software has had for decades. What the introductions leave out is that the tool has to remember what it created, and that memory lives in a state file.

  • Infrastructure is defined in files, versioned, and applied by automation
  • Declarative tools describe the end state; imperative ones describe the steps
  • The tool keeps state: a record of what it built and therefore owns
  • Drift is what happens when somebody changes something outside the tool
  • IaC does not remove the need to understand the platform underneath
On this page

How it worksHow infrastructure as code works

The practical answer to what is infrastructure as code is a workflow borrowed from software development. Infrastructure as code treats a network, a virtual machine or a database the way a development team treats application code, and the steps are the same in every tool.

1. Write. An engineer describes the infrastructure resources in configuration files: virtual machines, networks, load balancers, databases, storage and the permissions between them. The format is usually HCL, YAML or JSON. 2. Commit.

The files go into version control, normally Git, where every change to the infrastructure has an author, a date and a reviewer. 3. Plan. The IaC tool compares the configuration with what exists and lists what it would create, change or destroy.

4. Apply. The tool calls the cloud provider's API, whether that is AWS, Azure or Google Cloud, and provisions the resources. Nobody logs in to a console. 5. Repeat. The same configuration, with different variables, builds the development, test and production environments, so the environments match.

Infrastructure management then becomes a matter of changing files. To resize a server or add a network, a team edits the configuration, and the pipeline applies the difference to the running infrastructure.

This is why IaC is treated as a DevOps practice. The infrastructure configuration runs through the same CI/CD pipeline as the application, so an application deployment and the environment it needs are versioned, tested and released together. Operations teams stop provisioning by ticket, and development teams stop waiting for an environment.

The two stylesDeclarative and imperative IaC

There are two ways to write infrastructure as code, and the declarative vs imperative distinction sounds academic. It decides how the automation behaves when the infrastructure is already there.

Imperative code describes the steps: create this server, then attach that disk, then open that port. It reads like a script because it is one. The problem is running it twice. The second run tries to create a server that already exists, and unless every step checks first, it fails or duplicates.

Declarative code describes the end state: there should be a server of this size, with that disk, and that port open. The tool works out what to do to get there. If everything already matches, it does nothing. If one port is wrong, it changes that port and touches nothing else.

Declarative is the dominant style for provisioning, and the property that makes it work is idempotence: an idempotent definition applied twice leaves the same result. That is what allows a configuration to be run on a schedule, in a pipeline, by anybody, without anyone having to know whether it ran before.

The two styles are not competitors so much as layers. Terraform and its equivalents are provisioning tools that create the cloud resources. Ansible and its equivalents are configuration management tools that configure what runs on them. Plenty of teams use one of each and the boundary between them is a design decision worth making deliberately.

ToolStyleLayerTypical use
TerraformDeclarativeProvisioningMulti-cloud infrastructure
OpenTofuDeclarativeProvisioningThe community fork of Terraform
AWS CloudFormationDeclarativeProvisioningAWS only
Bicep, ARMDeclarativeProvisioningAzure only
PulumiDeclarativeProvisioningInfrastructure in a general language
AnsibleMostly declarativeConfigurationConfiguring existing machines
Puppet, ChefDeclarativeConfigurationContinuous configuration enforcement

The layer column is the one that matters when choosing. A team searching for ansible vs terraform is usually comparing two IaC tools that do different jobs, and the answer is often both.

StateThe state file, and why it decides everything

A declarative IaC tool has to answer a question every time it runs: what did I create last time? Nothing in the cloud provider says which of a thousand resources belong to your configuration. So the tool keeps its own record, and that record is the state.

Terraform state is the best known example. AWS CloudFormation keeps a similar record on the provider side and calls it a stack.

State does three jobs. It maps a resource in your code to the real object in the provider, so the tool knows which server is web_a. It stores what the settings were at the last apply, so a change can be detected. And it lets the tool know what to delete when you remove a block from the code.

Three consequences follow, and they account for most of the trouble people have.

State can disagree with reality. If somebody deletes a machine in the console, the state still says it exists. The next apply notices and recreates it, which is the tool doing its job and can still be a surprise.

State is a secret. It commonly contains connection strings, generated passwords and keys, in plain text, because those are attributes of the cloud resources it created. A state file in version control is a credential leak, and it is a common one.

State is shared, so it needs locking. Two people applying at once against one state file will corrupt it. Remote backends with locking exist for exactly this, and using a local state file for anything a second person will touch is a decision to have that incident eventually.

The habit that prevents most state incidents is trivial: run the plan, read it, and only then apply. A plan that proposes to destroy nine things you did not expect is the tool telling you the state is wrong, and it takes ten seconds to notice.

DriftDrift, and why it is not the enemy people think

Configuration drift is the gap between what the infrastructure code says and what actually exists. Somebody resizes a machine in the console during an incident, and the code no longer describes reality.

The reflex is to call this a failure of team discipline. It usually is not. It is what happens when a system is under pressure and the deployment pipeline takes twenty minutes, and the honest response is to expect it and to detect it rather than to forbid it.

Detection is cheap. Running a plan on a schedule and alerting when it is not empty tells you within a day that something changed outside the process. What you do next is a judgment call: bring the change into the code because it was right, or revert it because it was not. Both are fine; not knowing is not.

The dangerous case is drift in the other direction, where the code was changed but never applied. Then the repository describes an infrastructure that has never existed, and every reader of it is misled. A deployment pipeline that applies on merge is the usual fix, and it is worth more than any amount of policy about console access.

What it doesWhat it is genuinely good at

The IaC benefits that vendor pages list are speed, consistency and lower cost. The ones below are the benefits of infrastructure as code that hold up in a small or midsize environment.

Rebuilding. The strongest argument for IaC. Cloud infrastructure defined in configuration files can be recreated in a different region or account without archaeology, which is a disaster recovery property as much as a convenience.

Review. A change to a firewall rule arrives as a diff in version control that a second person reads before it happens, which is the same control software deployment has had for decades and infrastructure mostly has not.

Security review before deployment. Because the infrastructure configuration is text, a scanner can check it for an open storage bucket or a permissive security group before anything is provisioned. Security teams call this policy as code, and it moves the check from after deployment to before it.

Making environments match. Staging and production environments built from the same definition with different variables differ in the ways you chose rather than in the ways that accumulated.

A history that explains itself. The version control commit that opened the port has a date, an author and a reason attached, which is a different quality of answer from a cloud audit log entry.

Deleting things safely. Knowing which resources belong to a definition is what makes it possible to remove an environment completely, which nobody does confidently by hand.

What it does notWhat it does not do

IaC does not remove the need to understand the platform. A subnet defined in code is a subnet, and getting it wrong in a file is exactly as wrong as getting it wrong in a cloud console, only now the automation applies it to thirty accounts. The subnetting still has to be right.

It does not make changes safe. It makes them repeatable, which cuts both ways. A destructive change applied by a pipeline is faster and wider than the same mistake made by hand.

It does not eliminate the console. Reading is still done there, incidents are still handled there, and pretending otherwise is how drift becomes undiscussable rather than rare.

It is not free. There is a real cost in learning the IaC tools, building the deployment pipeline, managing state and maintaining modules, and the management overhead does not go away once the first deployment works.

That cost is worth paying at the point where the estate is big enough or rebuilt often enough. For six servers that never change, a documented build and a good backup can be the better answer.

PitfallsWhere implementations go wrong

Secrets in version control. Variables files with passwords, or a state file committed by accident. Both are common and both are a credential leak the moment the repository is shared. Use the cloud provider secret store and keep state in a remote backend.

No plan step. Applying without reading the plan is the single most common cause of an unexpected deletion. The plan is the tool telling you exactly what it is about to do.

Local state on a shared project. It works alone and corrupts the first time two people run at once, usually at the worst moment.

One enormous configuration. A single definition covering the entire infrastructure means every change plans against everything and one automation mistake touches everything. Splitting by lifecycle, so things that change together live together, is the usual fix.

Copying modules instead of referencing them. The fastest way to get twenty slightly different versions of the same network, which is the problem IaC was adopted to avoid. Modules belong in version control with a version number like any other dependency.

Code that was never applied. A merged change that no pipeline ran is worse than no code, because now the repository lies confidently.

Forgetting the IaC tools have versions too. Provider and tool upgrades change generated plans. Pinning versions and upgrading deliberately is the difference between a routine change and an interesting afternoon.

INTRODUCTIONS DRAW TWO BOXES. THERE ARE THREEThe middle one is where the bad days come from.The codein version controlRealitywhat the provider hasThe state filewhat the tool thinks it builtwhat everybody assumes is the only linemerged, never applieddriftNothing compares the code to reality directly. Both checks go through the state file.Which is why a lost or wrong state file proposes to destroy things you did not touch.
The dashed line is the relationship people think they are managing. The two solid ones are the relationships that actually break.

ComparisonIaC, configuration management, manual builds and images

CriterionInfrastructure as codeConfiguration managementManual buildsGolden images
Creates the machineYesNoYesYes
Configures what is on itPartlyYesYesBaked in
RepeatableYesYesNoYes
Reviewed before it happensYesYesNoAt build time
Detects driftYes, by planningYes, by runningNoNo
Rebuild in a new regionYesNeeds machines firstArchaeologyYes, if the pipeline exists
Setup costRealRealNoneReal

The last row is the honest one. Every column except manual has a setup cost, and the case for paying it is the fifth and sixth rows rather than any claim about speed.

Building one server by hand is faster than writing the code for it, every time. Building the ninth one, or rebuilding all of them somewhere else, is not.

FAQFrequently asked questions

What is infrastructure as code?

IaC means defining servers, networks and cloud services in files that an automation tool reads and applies, rather than building them by hand. The files are versioned, reviewed and applied by a pipeline, so infrastructure gets the controls software already has.

What is the difference between declarative and imperative IaC tools?

Declarative code describes the end state and lets the tool work out the steps, so running it twice changes nothing the second time. Imperative code describes the steps themselves, which means it has to check what already exists or it will fail on a second run.

What does idempotent mean here?

That applying the same definition twice produces the same result. It is what allows a definition to be applied on a schedule or by anybody without needing to know whether it has run before.

What is a Terraform state file?

The tool's record of what it created: which real resource corresponds to which block of code, what its settings were at the last apply, and therefore what to change or delete next time.

Why is the state file dangerous?

Because it commonly contains generated passwords, keys and connection strings in plain text, and because two people applying against it at once will corrupt it. Keep it in a remote backend with locking and never in a repository.

What is configuration drift?

The gap between what the code says and what actually exists, usually created by somebody changing something in the console. It is normal rather than shameful, and the fix is detecting it quickly rather than forbidding it.

How do I detect drift?

Run a plan on a schedule and alert when it is not empty. An empty plan means reality matches the code; anything else is a change made outside the process that somebody should decide about.

What is the difference between Terraform and Ansible?

Layer. Terraform provisions the infrastructure, creating the machine and the network. Ansible configures what runs on a machine that already exists. Many estates use one of each rather than choosing.

Does infrastructure as code make changes safer?

It makes them repeatable and reviewable, which is not the same thing. A destructive change applied by a pipeline reaches further and faster than the same mistake made by hand, which is why the plan step matters.

Is IaC worth it for a small environment?

Not always. The cost of the tools, the deployment pipeline, the state and the modules is real, and for a handful of servers that rarely change, a documented build and a tested backup can be the better answer. It pays off with scale or with frequent rebuilds.

Where should secrets live?

In the provider's secret manager, referenced by the code rather than contained in it. Variables files with passwords and committed state files are the two most common credential leaks in this area.

Why did my plan want to destroy everything?

Almost always a state problem: the wrong state file, a lost state file, or resources renamed in the code so the tool no longer recognizes them. Read the plan before applying and that discovery costs seconds instead of an outage.

Should one configuration cover the whole estate?

No. A single definition means every change plans against everything and one mistake reaches everything. Split by lifecycle, so resources that change together live together.

Is infrastructure as code part of DevOps?

Yes. IaC is how DevOps teams apply software development practices to infrastructure management: version control, review, automated testing and a pipeline. The application and the infrastructure it runs on are released the same way.

Which IaC tool should a small team start with?

The one native to the cloud in use if there is a single provider: AWS CloudFormation on AWS, Bicep on Azure. Terraform or OpenTofu when the infrastructure spans more than one provider. Add a configuration management tool such as Ansible only when the servers themselves need configuring.

Do I still need the console?

Yes, for reading and during incidents. The goal is that changes made there are noticed and reconciled, not that nobody ever makes one.

Read next · Addressing What Is a Subnet? The thing a definition still has to get right, because a wrong subnet in a file is applied to thirty accounts. Open this next15 min
Also worth reading
One packet a weekA short, illustrated explainer every Tuesday. No vendor pitches, unsubscribe in one click.