I’d judge IaC drift governance by verified fixes - not alert counts. Give each finding an owner, check its risk and keep a record linking approved code to the live result.
Firefly’s 2025 vendor research reports 68% IaC adoption, but only 6% full infrastructure coverage. That gap matters: scanning code-managed resources alone can leave parts of your cloud estate unchecked. These figures are not UK enterprise benchmarks.
In this review, I focus on four checks:
- Coverage: scan configuration, security, cost, state and resources outside IaC.
- Timing: combine event-driven checks with scheduled scans, then measure time from detection to verified closure.
- Policy: review runtime failures and expired exceptions; require approval before disruptive fixes.
- Accountability: name decision-makers, retain audit records and track missed scans, repeat drift and unresolved findings.
My starting point: measure what your controls cover and what they close. Set scan frequency and retention periods around your risks and duties - not a claimed universal standard.
::: @figure
{Enterprise IaC Drift Governance: Detection to Verified Closure}
:::
What Teams Scan and How Often
Configuration, Security, Cost and State Checks
Scan configuration, security, cost and state. Track configuration drift separately from policy compliance. An approved capacity change may differ from code without breaking policy. Meanwhile, unchanged infrastructure may fail a new control. Higher spend alone is not drift: check whether infrastructure changed or usage increased.[4][5][9]
Use this matrix as the baseline for assigning each drift type an owner, supporting evidence and a remediation path.
| Check category | Detection sources | Business impact | Primary owner | Supporting evidence | Remediation path |
|---|---|---|---|---|---|
| Configuration: regions, instance types, autoscaling and Kubernetes settings | IaC plans, provider APIs and configuration inventory | Outages, inconsistent environments and failed deployments | Platform or infrastructure team | Approved code, plan output, state version and deployment record | Reconcile through a reviewed pull request; import, update or recreate as needed |
| IAM and access | Cloud audit logs, provider APIs and policy checks | Unauthorised access, data exposure and audit findings | Security and platform teams | Identity, timestamp, API event, policy result and incident record | Revoke or narrow access, restore the declared setting and investigate the initiating identity |
| Network exposure | Firewall rules, routes, load-balancer settings and configuration rules | Service disruption or unintended public access | Network and security teams | Before-and-after configuration, flow logs and change event | Restore approved rules or approve a documented exception |
| Encryption, backups and logging | Key-management configuration, backup settings, logging controls and IaC policy scans | Data-protection failures, recovery risk and regulatory exposure | Security, platform and service owners | Key identifiers, control results, backup evidence and approval history | Enable the required control, rotate or replace keys, and record the change |
| Sizing and spend | Cloud billing data, resource metadata, utilisation metrics and IaC plans | Unplanned spend and capacity or performance problems | FinOps, platform and service owners | Cost allocation tags, usage data, budget alerts and change history | Resize, stop, delete or formally approve the higher-cost configuration |
| Tags and ownership | Resource inventory, tagging policies and billing reports | Unallocated spend and unclear accountability | FinOps and application teams | Tag-policy result, owner mapping and service catalogue entry | Add or correct tags through code or an approved controlled action |
| Unmanaged resources | Cloud inventory compared with IaC repositories and state | Security blind spots and orphaned costs | Platform or cloud governance team | Inventory export, repository search, state listing and creation event | Import into IaC, remove the resource or document an approved exception |
| State drift | State refresh, provider read operations and deployment history | Incorrect plans, accidental replacement and unreliable recovery | IaC platform team | State serial or version, lock records, refresh output and pipeline logs | Repair or re-import state under controlled review. Do not edit state manually outside controlled review. |
Resource-level checks depend on a complete inventory. Measure the whole estate, not just IaC state. Track enrolled accounts and subscriptions, production resources linked to owners and repositories, unmanaged assets and control coverage by family.
For each finding, record the resource ID, account or subscription, observed and expected values, detection time, approved commit, state version and deployment record. Include the audit actor and timestamp too.[6][7][8]
Coverage determines what you scan. Cadence determines how soon you find drift.
Continuous, Event-Driven and Scheduled Scans
Choose scan cadence based on how critical each control is and how often changes occur. These methods work together: cloud audit events can trigger targeted checks, while scheduled inventory reconciliation finds unmanaged resources and gaps in event capture. Continuous monitoring covers only the resources and attributes its integrations observe.[7][8]
| Approach | Coverage | Detection latency | Run cost | Evidence quality | Best for |
|---|---|---|---|---|---|
| Continuous monitoring | Supported resources and attributes | Low where observation is complete | Costs for ingestion, evaluation and triage throughout monitoring | Detailed timeline if logs are retained | Critical IAM, internet exposure, encryption and production control-plane changes |
| Event-driven checks | Changes captured by audit events | Near real time for captured changes | Depends on event volume and routing | Actor, timestamp and API context | IAM, public ingress and key-management changes |
| Scheduled estate scans | Inventoried resources and configured checks | Determined by schedule | Recurring API and scan costs | Point-in-time differences; actor context may be missing | Full production estates, lower-risk environments and resources missed by event monitoring |
| On-demand investigation | Selected resource or workspace | Starts when requested | Operator time and targeted scans | Detailed snapshot with linked logs | Incidents, suspected compromise and failed deployments |
Match scan frequency to the capacity of the team that must triage and close each finding.
Once enabled, HCP Terraform health assessments run at roughly 24-hour intervals. Treat that as product behaviour, not a benchmark.[10] Set schedules using asset criticality, change volume, acceptable detection delay, API limits and triage capacity.
Keep post-deployment checks even when CI/CD policy checks pass. Pipeline checks miss later console edits, compromised credentials and new unmanaged assets.[9]
How to Prevent IaC Drift: Audit-Ready Checklist
Remediation Timing and Policy Failures
After detecting drift, teams need to triage it, approve a fix and verify closure safely. Track how long each step takes.
Risk-Based Fixes and Approvals
Measure from detection to verified closure, with separate timings for detection-to-triage, triage-to-approval and approval-to-closure. Report median and percentile times by severity. Orca Security’s 2024 data, reported by Tamnoon, found that resolving a security alert took an average of 145 hours; 60% of organisations took longer than four days[18]. These figures cover security alerts, not IaC drift benchmarks.
Match the fix to the intent behind the change. Revert changes made without authorisation through a reviewed plan. Adopt legitimate changes through code. Import unmanaged resources only after checking ownership, identity, dependencies and configuration.
Terraform’s plan -refresh-only checks differences without changing deployed resources. An approved refresh-only apply updates state to match live values; it does not change the resource. Escalate uncertain findings instead of changing production automatically[11][12][13].
Aim for same-day review when findings involve public exposure, access without authorisation or weakened encryption. Schedule low-risk tag corrections only when they leave security, billing allocation and audit evidence unaffected.
Disruptive fixes need a service owner, a peer-reviewed plan, approval suited to the risk, dependency checks and rollback or recovery safeguards. Before execution, check whether the fix will replace resources, interrupt traffic or revoke access. Retain emergency-change evidence and verify the live result before closing the finding.
Runtime Policy Gaps and Expired Exceptions
Deployment approval does not prove continued compliance. Runtime enforcement must continue after pipeline approval[17]. Missing owners, incomplete inventories, unrestricted manual changes and exceptions without expiry are implementation risks - not measured failure rates.
Policy-as-code should set the scope, severity, owner and response for permissions, encryption, regions, sizing and mandatory tags. Use preventive, detective, corrective and compensating controls to block changes, detect problems, restore the required state or temporarily reduce exposure until the permanent fix is approved[14][15].
Every exception needs a named risk owner, justification, affected resources, approval, compensating controls, review date, expiry date and closure criteria. Expiry should trigger re-approval, escalation or safe enforcement - not silent renewal. Keep the original clock running: an expired exception remains an unresolved finding[16].
Avoid blanket ignore_changes rules. Limit exclusions to specific attributes and document who controls them. Track expired exceptions, repeat drift, owner coverage and independently verified closures to test whether enforcement works, rather than just counting policy checks[16].
Each closure should link the live fix to the approved change, exception record and audit evidence. Governance owners must be able to trace that evidence trail through to closure.
Audit Records and Governance Ownership
Closing a drift finding is only part of the job. Audit records provide proof that the issue was addressed - and that the evidence stands up to scrutiny.
Customers, internal audit and regulators need different evidence. Customers want to know that controls exist, run consistently and can be evidenced for their services. Internal audit needs coverage data, failed runs and independently checked closures. Regulatory duties vary by firm. Under SYSC 15A, FCA-regulated firms may need current, written records of important business services and their dependencies, with each version kept for at least six years.[20][21]
Tracing Findings from Approved Code to Closure
Create a searchable trail: approved pull request → deployment ID → live state at detection → policy result → ticket or incident → remediation or approved exception → closure scan.
Keep the policy identifier, version and effective date, alongside the expected and observed configuration, severity, approvals and linked exception decisions. Record timestamps for policy evaluation, alert generation, triage, approval, remediation deployment and closure verification. Use UTC. Identify the human approver, committing user, pipeline service account, cloud principal, remediator and risk owner separately, rather than grouping them under a generic administrator identity.[19]
A scan counts only when it completes successfully. Missing, failed or timed-out scans are coverage gaps, as are inaccessible accounts. Keep the scan scope, resource count, execution status, errors and evidence location.
Protect records against changes, deletion or disabling without permission, and restrict access to authorised personnel. Retention must support closure verification, exception control and audit review. Set periods by record type and obligation: NCSC recommends at least six months for the most important security logs, while ICO guidance calls for documented schedules, assigned responsibilities and regular review. Document who can read, export, amend or delete evidence. Minimise personal data in logs and tickets, and use immutable or write-once storage where appropriate.[19][22][24]
Team Roles and Accountability
Name both the evidence-retention owner and the accountable owner. Make clear who runs each control and who has the right to make decisions. Internal audit should review the evidence, not own its retention. Retention should sit with governance, risk, compliance or platform governance, with input from records-management and privacy teams.
Use the matrix to assign ownership of the evidence - not just responsibility for fixing the issue.
| Control type | Evidence retention owner | Accountable owner |
|---|---|---|
| Platform guardrails | Platform governance | Head of platform |
| Security configuration | Security governance | CISO or delegated security owner |
| Cost and tagging | FinOps governance | FinOps lead |
| Resource configuration | Service governance | Service owner |
| Operational resilience dependency | Risk and compliance function | Senior business-service owner |
Smaller firms can combine roles, but decision rights must remain clear. Document combined roles, require a second-person review for high-risk exceptions and keep all approvals in the ticketing system. Name the escalation owner for missed deadlines or expired exceptions. FCA expectations support explicit board and senior accountability for operational resilience.[23]
Conclusion: Priorities for Scaling Drift Governance
Prioritise traceable configuration ownership, scans across the estate, continuous or more frequent monitoring of high-risk changes, and risk-based remediation. These priorities are informed by evidence, not representative adoption data for 2026. Resources and attributes outside monitoring can still hide drift.[25][26][27] Track progress through coverage, age and closure metrics.
Working Practices and Governance Metrics
Measure control, not activity. Define each metric’s scope, population and reporting period. Track scan coverage, drift age, recurrence, median and percentile remediation times by severity, unresolved critical findings and their age, exception lifetime, and evidence completeness. Report production and non-production separately, and distinguish missing scans from clean results. Fewer findings mean little if coverage has fallen.[26][30]
Route fixes through reviewed pull requests. Use automatic correction only for pre-approved, low-risk changes that can be reversed.[10][25][29] When drift keeps returning, investigate the root cause - permissions, manual changes, module defaults or unclear ownership - instead of repeatedly closing tickets.
Track scan reliability, verification failures and manual hours spent on triage, reconciliation and evidence preparation. Report attributable tooling, cloud, engineering and incident costs in £ against a defined baseline and reporting period. These measures help teams evaluate results; they do not establish guaranteed savings.[28][30]
External Support for Control Automation
When internal capacity is limited, implementation support can speed up automation. Hokstad Consulting can help automate governance workflows, evidence capture and remediation. Internal teams retain responsibility for risk decisions, approvals and evidence ownership.
FAQs
How do we prioritise drift governance with a small team?
Start with a phased, risk-based approach: automate drift detection for your most critical production assets first. Use role-based access control to restrict write access to senior engineers, and require formal break-glass procedures for all emergency changes.
Set policy-as-code guardrails in non-blocking audit mode first, so they don’t disrupt workflows. As confidence grows, gradually extend continuous remediation to lower-risk components in batches of 5 to 10 resources.
How can we verify drift fixes independently?
After remediation, run a refresh-only or plan check to compare live infrastructure with the desired state in Git/IaC. Keep immutable audit evidence: a timestamped drift report, the affected resources and the remediation outcome.
Use CI/CD drift detection with non-zero exit codes, post-fix runtime drift scans, or both. Before scaling up automated fixes, run dry-run simulations and validate rollback paths. [1][2][3]
How do we set realistic drift remediation targets?
Roll out changes in phases, based on risk. Automate low-risk remediation, but require approval or manual intervention for high-risk assets. Keep changes moving without disrupting day-to-day services.
Scan production every 4–6 hours and development daily. Begin with non-blocking audits to gather data before switching to hard enforcement. Align targets with DORA metrics to measure recovery speed, and schedule freeze windows to protect critical business hours.