The backup job has reported success every night for the last four hundred nights. Nobody in the organisation has ever restored from it.
That sentence describes a large share of the environments we see across the region, and it stays true until the night it stops being survivable. The call comes at some inconvenient hour. The finance ERP is throwing file errors, and within twenty minutes it is clear this is not a disk problem. Someone opens the backup console and finds exactly what they hoped to find: green ticks, all the way back. Then someone notices the repository is a share on a server joined to the same domain, and that the last fourteen days of restore points are gone.
At that point the organisation discovers something uncomfortable. Every control it bought (firewall, endpoint agent, mail gateway, SIEM, awareness training) operates before the incident. Once the incident has happened, one control is still in play, and it is the one that was never rehearsed.
So here is the question worth sitting with. When did your organisation last restore a full production system, end to end, onto clean infrastructure, with the person who normally does it deliberately unavailable and a stopwatch running? Not a file restore. Not a check that the job completed. A service, back to a working state, timed.
If the honest answer is “we haven’t”, the rest of this is about you.
The gap between the recovery plan and the recovery
Almost every organisation of any size has a business continuity document with an RTO and an RPO in it. Four hours and fifteen minutes turn up again and again, far more often than any measurement in those organisations could justify.
Ask where the four hours came from. Usually a template, a consultant’s draft, or a prior employer’s policy. It was not derived from a measured restore of the actual data set, on the actual infrastructure, over the actual network. It is an aspiration written in the grammar of a commitment, signed off by people who assumed it was a finding rather than a wish.
The audit cycle reinforces this. Auditors ask whether a policy exists, whether backups are configured, whether retention is defined. All of those can be answered truthfully by an organisation that would take nine days to recover from a serious ransomware event. The evidence requested is the existence of the control, not its performance.
That is changing. Regulators across our territories are moving from “do you have a plan” to “show us when you last tested it”. Central bank and financial sector guidance in India, Sri Lanka, Bangladesh, Singapore, the UAE and Saudi Arabia increasingly speaks of operational resilience and tested recovery rather than documented intent. Data protection regimes, including India’s DPDP Act, Sri Lanka’s PDPA, Singapore’s PDPA, the UAE PDPL with the DIFC and ADGM regimes, Saudi Arabia’s PDPL and the NCA’s essential cybersecurity controls, and equivalents in Qatar, Oman and Bangladesh, attach obligations to availability and integrity, not only confidentiality. Losing customer data permanently is a data protection failure, not just an outage.
There is a second gap underneath the first. Backup is almost always owned as a secondary duty. Very few organisations in this region have a dedicated backup administrator; the task sits with a virtualisation engineer, a systems administrator, or an MSP’s level-two queue, and is measured by the absence of red in a console. Nobody is paid to be professionally paranoid about whether a restore would work, so nobody is.
The result is a capability staffed as housekeeping, quietly holding up the continuity of the business.
The failure modes nobody rehearses
None of these are exotic, and most organisations have several at once.
- The backup that has never been restored end to end
A completed job proves data was read and written. It does not prove the restore point is consistent, that the application inside it will start, that the database will pass its own integrity check, or that the licence keys, certificates and service accounts needed to bring the service back are captured at all.
The classic version is the line-of-business application whose data is backed up perfectly and whose configuration lives in a directory nobody included in the job. The restore works. The service does not. The team spends two days rebuilding something that was never in scope while the business waits.
- The backup infrastructure sits inside the blast radius
Ransomware operators do not start with your data. They start with your recovery. Obtain domain credentials, locate the backup server, expire the retention policy, remove restore points, wipe the repository, and only then encrypt production. They target backup infrastructure first because they know what forces a payment decision.
If the backup server is domain-joined, its console authenticates against the same directory as everything else, the repository is reachable from a production subnet, and its administrator uses an account that also administers the hypervisor, then the backup is not a separate copy of the data. It is the same failure domain with a different file path.
- The documented recovery time and the actual recovery time
Assume the backups are intact. Now walk the arithmetic. Several terabytes moving across a link shared with production, into hosts that may themselves need rebuilding because they are considered compromised, in a sequence where domain controllers must come back before anything that authenticates against them, each step taken by someone doing it for the first time.
Four hours is plausible for restoring one virtual machine onto healthy infrastructure. It is not plausible for restoring a business when the infrastructure underneath it is evidence. Organisations that measure this honestly usually find their real window runs to days.
The regional dimension matters. Grid variability, connectivity that is excellent in one office and thin in a branch two hundred kilometres away, and the cost of moving volume between sites all lengthen the real number in ways a template RTO never anticipated.
- Microsoft 365 and SaaS data that nobody is backing up
This is the most common misconception we meet in the field, and capable people hold it sincerely: because the data is in the cloud, the provider must be backing it up.
The provider protects its own infrastructure, uptime and availability. It is explicit that protecting your business data against deletion, malicious action inside your tenant, account compromise and departing employees remains your responsibility. Native retention, versioning and deleted-item recovery exist, but they are bounded, reversible by whoever holds tenant admin, and were never designed as a recovery capability against an attacker who already holds those credentials.
The exposure looks like this. A compromised global administrator account, or a departure with elevated rights, removes mailboxes and clears the recycle bin. Thirty days later there is nothing to recover from, and the organisation’s contracts, board papers and entire collaboration history live in Exchange Online, OneDrive, SharePoint and Teams, outside the scope of every backup job it runs.
- Retention set for the auditor rather than the recovery
Retention is often chosen to satisfy a compliance requirement, then applied uniformly to backup as if the two were the same problem.
They are not. Retention for compliance is about proving what existed. Retention for recovery is about having a restore point from before the attacker got in. Dwell time frequently exceeds a short backup retention window, which produces the worst outcome available: a complete set of restore points, every one of which contains the intrusion. The organisation restores, and reinfects itself.
- The DR site that is a second rack in the same building
“We have a DR site” deserves a follow-up question, and the answer is usually revealing. Often it is a second rack in the same data hall, or a facility in the same city, on the same power grid, sometimes reached over the same carrier, and always administered by the same credentials.
That protects against hardware failure. It does not protect against fire, flood, a district power event, a national connectivity disruption, or an attacker with domain rights, the scenarios DR sites exist for. And where residency requirements apply, an organisation may have correctly kept data inside the border and incorrectly concluded that inside the border means inside the same building.
- The runbook that assumes one person is available
Every environment has one. The engineer who built the cluster, knows the start-up order, remembers where the encryption passphrase is stored, and can talk anyone through a restore.
The runbook was written by that person, in their shorthand, silently assuming their context. It says “restore the VM and bring up the app” where the real procedure has eleven steps. It does not say where the backup console credentials are held if the password manager was on the encrypted volume. It does not account for that person being on leave or gone.
Recovery capability that lives in one head is not a capability. It is a dependency.
Why buying more storage does not fix this
The usual response to backup anxiety is to buy capacity. More disk, a bigger appliance, another bucket. It produces a number that goes up, and it addresses almost none of the above.
Capacity does not make a restore point verified. It does not remove the repository from the domain the attacker controls. It does not shorten the dependency chain that determines your real recovery time. It does not back up your Microsoft 365 tenant, and it does not give you a restore point from before dwell time began. A larger repository in the same rack is still in the same rack.
Worse, capacity bought on per-terabyte licensing creates pressure in exactly the wrong direction. When retention costs more every month, the rational local decision is to shorten it, exclude the large-but-boring data sets, and skip the second copy. Each decision is defensible in isolation and corrosive in aggregate. Nobody makes an explicit decision to reduce recoverability. It erodes, one renewal at a time.
The other common response is another point tool: one product for virtual machines, another for endpoints, another for SaaS, another for databases, each with its own console and restore procedure. Where nobody owns backup full time, that fragmentation is itself a recovery risk.
The fix is not more storage or more tools. It is treating recovery as a rehearsed, verified, isolated capability, and building the platform underneath it accordingly.
How BDRShield and TrueNAS approach this differently
EGUARDIAN distributes both BDRShield and TrueNAS across our territories, and the pairing is deliberate. One makes recovery a verified outcome rather than a hoped-for one. The other makes the storage that recovery depends on resistant to the same attack that made recovery necessary.
1. BDRShield: coverage that matches the environment you actually run
BDRShield, formerly BDRSuite, is a single platform covering the workloads mid-market and enterprise environments here actually run, rather than a subset assuming a homogeneous estate.
That includes physical Windows and Linux servers; Windows, Linux and Mac endpoints; virtual machines on VMware, Hyper-V, KVM, Proxmox VE and oVirt; cloud workloads on AWS and Azure; applications and databases including Microsoft Exchange, SharePoint, SQL Server, MySQL and PostgreSQL; file and NAS data; and Microsoft 365 and Google Workspace.
Breadth matters not for procurement convenience but because recovery risk concentrates in whatever is left out. One console, one retention model and one restore procedure is far easier to rehearse than four, and rehearsal is the whole point.
2. Immutability, air gap, and getting the backup out of the blast radius
BDRShield supports immutable backups with retention lock, so restore points cannot be modified or deleted before their retention expires, including by an administrator account, which is exactly what an attacker will have. It also supports air-gapped and offsite copies, so a copy exists somewhere the production network cannot reach.
This is the architectural answer to the domain-joined backup server: immutability changes the question from “can the attacker reach the repository” to “does reaching it help them”.
Backup targets are deliberately open: local storage, network shares, S3-compatible object storage including AWS S3, Azure Blob, Google Cloud, Wasabi, MinIO and Backblaze, and BDRShield’s own cloud. For organisations under residency requirements, that flexibility is what allows an immutable copy to sit on infrastructure satisfying the regulator rather than only the architecture diagram.
3.Verified restores rather than green ticks
This addresses the failure mode at the top of this article directly. BDRShield performs automated backup verification aimed at guaranteed recoverability, testing the restore point itself rather than only the job that produced it. The definition of a successful backup moves from “the job finished” to “the thing inside it has been shown to come back”.
For recovery itself, BDRShield provides instant recovery for virtual machines and physical servers, and granular restore of files, emails and application items. Its VembuHIVE backup file system produces instantly mountable restore points, so a service can be brought up from the backup copy rather than waiting for a full data transfer to finish before anything starts. That distinction is where most of the gap between a documented RTO and a real one is won or lost.
The platform also supports pre-restore malware scanning and restoration into isolated environments, so a restore point can be validated before reintroduction, addressing the reinfection problem directly. Anomaly detection on backup activity adds a further signal, because unusual change rates are often the earliest visible evidence that something is encrypting production.
4. SaaS treated as a first-class workload
Microsoft 365 protection covers Exchange Online, OneDrive, SharePoint and Teams, including group and shared mailboxes and in-place archives, with retention from days to years and a choice of storage target. Google Workspace is covered as well.
That closes the gap described earlier, in the same console and retention model as everything else, so the SaaS restore gets rehearsed alongside the server restore rather than being a separate procedure nobody has practised.
5. Built for service providers as well as internal teams
For MSPs, BDRShield provides multi-tenant centralised management, an MSP portal for customer licensing, billing and invoicing, self-service consoles for clients, white-label branding and PSA/RMM integration, with cloud storage that starts small and is allocated across tenants.
This matters regionally because much of the mid-market’s recovery capability in Sri Lanka, Bangladesh, India and the Gulf is delivered by service providers rather than in-house teams. Operability at tenant scale is what puts tested recovery within reach of organisations that will never hire a backup administrator.
6. TrueNAS: a storage foundation that resists the same attack
Backup software needs somewhere to land, and the properties of that landing place determine much of what the backup is worth. TrueNAS is open storage built on OpenZFS, and the relevant characteristics are structural rather than bolted on.
ZFS provides end-to-end checksums, automatic corruption detection and self-healing, along with RAID-Z protection, addressing the quieter failure mode where a backup is intact as a file and corrupt as data. TrueNAS supports ransomware-proof snapshots and snapshot-based local and remote replication, so a second copy at a second site is native to the platform rather than a separate product.
It is unified storage: file, block and object in one system, with SMB, NFS, iSCSI, Fibre Channel, NVMe-oF and S3-compatible services. The same platform serves as an immutable backup target for BDRShield and as general-purpose enterprise storage for virtualisation, file services or archive. Encryption covers data at rest and in flight, with FIPS 140 validation and KMIP support.
The line runs from the free, open-source TrueNAS Community Edition, formerly TrueNAS SCALE, through TrueNAS Enterprise with dual-controller high availability and automated failover, across appliance families including the NVMe-focused F-Series, high-capacity M-Series, power-efficient H-Series, single-controller R-Series and the Mini series for edge and branch sites.
7. The commercial shape is part of the architecture
TrueNAS does not charge per-capacity licence fees. That sounds like a procurement detail and is actually a resilience property. When retention does not get more expensive every time it gets longer, and a second copy at a second site carries no licensing penalty, the pressure to quietly shorten retention or drop the offsite copy disappears. Those decisions get made on the recovery scenario rather than the renewal quote.
What this looks like in practice
A workable design for a mid-market organisation with a virtualisation estate, a Microsoft 365 tenant and no backup administrator:
- BDRShield protects the full estate from one console (virtual machines, physical servers, endpoints, databases and the Microsoft 365 tenant) under a single retention model.
- The primary backup target is a TrueNAS system on-site, sized for the working retention window.
- Restore points on that target are immutable with retention lock, so credentials alone cannot shorten or delete them.
- TrueNAS replicates snapshots to a second TrueNAS system at a genuinely separate site (different building, different power feed, different failure domain), and an additional immutable copy goes to object storage in a jurisdiction that satisfies your residency requirement.
- The backup domain is isolated: separate credentials, separate administrative accounts, no shared directory trust with production.
- Automated verification runs continuously against the restore points themselves, so they are known-good rather than assumed-good.
- Recovery is rehearsed on a schedule. A production service is restored into an isolated environment, scanned, brought to a working state, and timed. That measured number replaces the one in the policy document.
- The runbook is rewritten from each rehearsal by someone who was not its author, and the next rehearsal is run by someone else.
Step seven is what changes the organisation’s actual risk position. Everything before it is preparation.
Mapping recovery failure modes to capabilities
| Recovery failure mode | How this pairing addresses it |
|---|---|
| Backups complete but have never been restored end to end | Automated verification of the restore point itself, plus scheduled rehearsals in isolation |
| Attacker holds domain credentials and deletes restore points or encrypts the backup share | Immutable backups with retention lock, air-gapped and offsite copies, an isolated backup domain, ZFS snapshots |
| Documented RTO has never been measured and the real one is far longer | Instant recovery for VMs and physical servers, instantly mountable restore points, and rehearsals that produce a measured number |
| Microsoft 365 and SaaS data assumed to be backed up by the provider | Native protection for Exchange Online, OneDrive, SharePoint, Teams and Google Workspace under one retention model |
| Retention set for the auditor, leaving no clean restore point before the intrusion | Retention from days to years, immutable long-term copies, and pre-restore scanning |
| DR site is a second rack in the same building or city | Snapshot replication to a second TrueNAS system in a separate failure domain, plus an immutable copy in object storage |
| Retention and second copies quietly reduced by per-terabyte licensing | Storage without per-capacity licence fees, so retention and replication are decided on recovery grounds alone |
Who benefits most
Hospitals and healthcare groups where clinical systems cannot be down for days, patient records carry statutory retention obligations, and the IT team is small relative to the criticality of the estate it runs.
Manufacturers running production-line systems, MES and ERP where downtime carries a calculable hourly cost, and whose plant sites sit on connectivity and power a head-office DR plan never modelled.
Banks, insurers and financial services firms under central bank expectations that increasingly ask for evidence of tested recovery, and under data protection regimes across India, Sri Lanka, Bangladesh, Singapore, the UAE and the Gulf.
Government departments and public sector bodies with in-country residency requirements, long retention obligations and mixed estates accumulated over a decade or more.
MSPs and system integrators delivering verified, multi-tenant backup and recovery as a service across many customers without growing a specialist team to match.
The reframe
Backup is usually budgeted as infrastructure and reviewed as a line item. It is more accurate to treat it as the last business control standing, the only one that still functions after prevention, detection and response have been overtaken by events.
Every other security investment reduces the probability of an incident. Recovery determines its consequence. And unlike almost every other control in the stack, its effectiveness cannot be inferred from a dashboard. It is only knowable by rehearsal. An organisation that has restored a production service under timed conditions in the last twelve months knows what its resilience is worth. An organisation that has not is holding an opinion.
This is a solvable problem with known engineering. Immutability, isolation, verification and replication are available now, at prices mid-market organisations in this region can carry. What is usually missing is not technology or budget. It is the decision to test.
Talk to our experts at EGUARDIAN. If you are carrying a documented recovery time nobody has measured, let us walk through your workloads, your constraints and what a tested, immutable, genuinely offsite recovery capability would look like for your organisation. Reach out to us at hello@eguardian.com.