Can We Bring Operations Back? The Question Behind Every OT Backup Strategy

Most organizations running OT environments claim to have backups. That is the easy answer.

The harder question is, can they bring operations back?

Can it restore the right system, the right version, in the right order, without putting operations at risk? Can it prove this before an incident, not after one?

In an OT environment, this distinction matters. OT does not run only on servers and files. It runs on SCADA systems, HMIs, PLC logic, DCS configurations, engineering workstations, historian data, network configurations etc. If these are missing, outdated, corrupted or scattered, the plant may not come back as expected.

A backup that cannot restore operations is not resilience. It is stored uncertainty.

The challenge is familiar. OT environments are older, more sensitive and more constrained than ordinary IT environments. Some systems cannot support modern backup agents. Some cannot be restarted casually. Some sit inside segmented or air-gapped networks. A restored system may boot, but still fail to communicate with field devices. That is why OT backup and recovery cannot be copied from an IT template. It must be built around operational consequences.

The starting point is not technology. It is clarity.

What must be recoverable to bring operations back safely? The answer is rarely limited to databases or server images. It may include PLC and DCS logic, SCADA configurations, historian records, relay settings, engineering workstation images, network switch configurations, firmware versions, license files etc.

This is where many organizations may quietly fail. There may be backups, but not of what actually matters.

That is not a recovery strategy. That is hope. And hope is not control.

Ransomware has changed the entire backup conversation. Attackers understand recovery. They know that if backups are destroyed, the organization loses its bargaining power.

A strong OT backup programme begins with asset criticality. A safety system, SCADA server, historian, HMI, engineering workstation, domain controller, and network switch do not carry the same consequence. They should not have the same recovery priority. Recovery Time Objective and Recovery Point Objective must be based on operational impact, not generic policy language.

Some systems need redundancy. Some need full-image backups. Some need configuration exports. Some need offline copies. Some need rapid restoration. Some can wait. The maturity lies in knowing the difference.

A mature OT backup strategy must also move beyond the old comfort of one backup copy sitting somewhere on the network. The principle should be simple: multiple copies, different media, one protected offsite or offline copy, one immutable or air-gapped copy and recovery verification with no critical errors. In OT, a backup copy is not enough. The copy must be reachable by the recovery team, unreachable by the attacker, and proven usable before the crisis. This is not excessive caution. It is basic survival planning.

Version control is equally important. Teams should know what changed, who changed it, when it changed and which version is approved. Without this, recovery may restore an outdated or unsafe state.

A backup report may say “successful.” That only proves something was copied. It does not prove the plant can recover.

Testing in OT is difficult. Production cannot be disturbed casually. But avoiding tests does not reduce risk. It hides it. Where live restoration is not practical, a digital twin can provide a safer way to rehearse recovery. The goal is not to create a perfect replica of the plant. The goal is to test whether the restored system can boot, communicate, load the right configuration and perform its intended function before a real incident. Every recovery rehearsal should record what failed, what was corrected and how the runbook must improve. The recovery runbook must be precise. It should define the restore order, required tools, responsible teams, validation checks, escalation path, approval points and fallback options etc.

A restore that has never been tested is not a capability. It is an assumption.

Governance is what turns backup into resilience. IT may provide the platform, but OT must define what is critical, when backups are safe, how systems should be restored, and how recovery must be validated. Cybersecurity must protect the recovery path. Leadership must ask for evidence.

Not comfort. Not verbal assurance. Evidence.

Which systems are covered? Which are not? When was the last restore tested? Which backups are immutable? Which copies are offline? Who can delete them? Who approves restoration? What failed during the last drill? What was corrected? What is the fallback if old hardware is unavailable?

These are not server-room questions. They are governance questions.

The shift is simple. Stop asking only, “Did the backup job complete?” Start asking, “Can we bring operations back safely?”

Backup completion is a technical status. Recovery confidence is an operational capability.

Tools matter. But tools do not define criticality, ownership, restore order, or operational acceptance. Governance does.

When a SCADA server fails, an HMI goes dark, a historian is encrypted or an engineering workstation is compromised, no plant manager will ask how many backup jobs completed last night.

The question will be sharper.

Can we bring operations back?

The answer should not depend on memory, luck or one person’s laptop.

It should depend on a tested, governed and proven recovery capability.

Backups are only the beginning.

The real test is recovery.

Author Details

Rohini Haridas

She is an Associate Consultant in OT security with a Ph.D*. in Electrical Engineering from Malaviya National Institute of Technology, Jaipur. Her research focused on modelling Load Redistribution (LR) attacks from the attacker’s perspective, including stealth, resource constraints, and uncertainty in attacker knowledge. She previously served as a certified cybersecurity trainer at the National Power Training Institute (NPTI) and brings over 13 years of experience in teaching and research.

Leave a Comment

Your email address will not be published. Required fields are marked *