Disaster Recovery Solution

Automated, Orchestrated Disaster Recovery

EasyStack Disaster Recovery (DR) Solution delivers automated, orchestrated protection for workloads running on ECF (Cloud Infrastructure Platform), ECNF (Cloud Native Infrastructure), and EAF (AI Infrastructure Platform). Built on the EOS Digital Native Engine and leveraging proven agent-based replication technology, the solution provides continuous data protection, predefined DR plans, one-click failover, and automated failback.

Illustration of disaster recovery: a production server stack replicating to a protected DR site shield, linked by a failover and failback loop
Summary

One DR control plane for VMs, containers, bare metal and AI

As enterprises migrate from VMware to open private cloud platforms and deploy AI workloads, disaster recovery requirements grow more complex. Workloads span VMs, Kubernetes containers, bare metal, and AI inference services. Data must remain within national borders. RTO and RPO expectations are tightening. Traditional DR solutions — based on storage replication or manual runbooks — are too slow, too expensive, and too fragile for modern hybrid workloads.

EasyStack DR Solution addresses these challenges. It integrates natively with ECF for VM and volume protection, ECNF for Kubernetes and container workloads, and EAF for AI model repositories and inference services. Using agent-based block-level replication, it continuously protects workloads from production sites to DR sites — whether on-premises, in a second EasyStack region, or in a sovereign cloud environment.

DR plans define the orchestration of recovery: which workloads to recover, in what order, with what network mappings and resource specifications. Failover can be triggered manually or automatically. Failback returns workloads to production with zero data loss or downtime.

With EasyStack DR, enterprises can:

RPO & RTO

Achieve RPO of minutes to zero and RTO of minutes.

Workload coverage

Protect virtual machines, containers, and AI inference services.

Failover orchestration

Automate failover orchestration with predefined DR plans.

DR testing

Test DR strategies without disrupting production.

Sovereignty & compliance

Maintain data sovereignty and compliance across regions.

No vendor lock-in

Reuse existing ECF / ECNF / EAF infrastructure without vendor lock-in.

Why Modern DR

Why modern disaster recovery is a strategic imperative

Downtime cost, regulation, VMware replacement and AI workloads are converging at the same time.

The cost of downtime

For financial services, government, energy, and healthcare organizations, even minutes of downtime can result in millions of dollars in losses, regulatory penalties, and reputational damage. Industry analysts consistently report that the majority of organizations experience at least one significant IT outage per year, and the average cost per hour of downtime continues to rise.

Regulatory and sovereignty requirements

Regulations such as Indonesia's PDP Law, Thailand's PDPA, and Europe's GDPR require data to be protected and often kept within national borders. Disaster recovery sites must comply with the same sovereignty requirements as production sites. EasyStack DR enables sovereign, on-premises DR across regions within the same jurisdiction.

VMware alternative drives DR modernization

As enterprises replace VMware with open platforms like EasyStack ECF, their existing VMware-centric DR solutions — such as SRM — may no longer be applicable. EasyStack DR provides a vendor-neutral, platform-agnostic alternative that protects KVM, OpenStack, container, and AI workloads from a single control plane.

AI workloads require new DR approaches

AI inference services, model repositories, and Agentic workflows are stateful, latency-sensitive, and often business-critical. Traditional backup-and-restore is insufficient. EasyStack DR supports continuous replication and automated failover for EAF workloads, including model service endpoints and inference configurations.

Challenges

Seven challenges in enterprise disaster recovery

Legacy DR tools were not designed for heterogeneous, sovereign, AI-era workloads.

Challenge

Description

Heterogeneous workloads

VMs, containers, bare metal, and AI services require different protection strategies.

Slow recovery

Manual runbooks and storage-based replication often lead to RTO of hours or days.

High cost

Dedicated DR hardware and standby infrastructure are expensive and underutilized.

Complex testing

DR drills disrupt production and are difficult to schedule and validate.

Sovereignty

DR sites must comply with the same data residency regulations as production.

Orchestration

Recovery order, network mapping, and dependencies must be automated, not manual.

Agent overhead

Traditional agents consume resources and require maintenance on every protected host.

Solution Architecture

Built on the EOS Digital Native Engine, orchestrated by DR plans

A DR control plane, replication agents, cloud agents, DR site storage and a DR plan repository.

EOS Digital Native Engine foundation

EOS — Digital Native Engine

EasyStack DR is built on the EOS Digital Native Engine, the same microservices-based, fully symmetric distributed cloud operating system that powers ECF, ECNF, and EAF. EOS provides:

  • Unified lifecycle management for over 150 microservices
  • OTA smooth upgrades with zero business impact
  • High availability and fault self-healing
  • Open APIs for integration

DR architecture components

DR Controller

Central control plane for replication management, DR plan orchestration, and failover/failback. Provides web UI and RESTful API.

Replication Agent

Installed on protected hosts (or provided as an appliance). Performs block-level changed block tracking (CBT) and sends incremental changes to the DR site.

Cloud Agent

Auxiliary VM on the target DR site that receives replicated data and writes it to target storage.

DR Site Storage

Target storage for replicated data: local disks, EasyStack volume disks, NFS, or S3-compatible object storage.

DR Plan Repository

Stores predefined DR plans containing machine specifications, network mappings, boot order, and orchestration instructions.

Replication architecture & DR plan orchestration

Agent-based block-level replication

EasyStack DR uses agent-based block-level replication:

  • Initial replication: Full synchronization of protected workloads to the DR site.
  • Incremental replication: Continuous changed block tracking (CBT) sends only modified blocks.
  • Replication intervals: Configurable from seconds to hours, enabling RPO from near-zero to minutes.
  • Bandwidth control: QoS settings prevent replication traffic from impacting production performance.
  • Deduplication and compression: Reduce storage and network consumption at the DR site.

DR plan orchestration

DR plans are the heart of recovery automation. Each plan defines:

  • Machine specifications: vCPU, RAM, disk, and flavor for failover VMs.
  • Network mappings: Source-to-target network and security group mappings.
  • Boot order and dependencies: The sequence in which workloads are recovered.
  • Pre/post scripts: Custom actions before or after failover.
  • Recovery point selection: The point-in-time snapshot to use for recovery.

Plans can be created in Basic mode (graphical interface) or Expert mode (JSON-based for complex scenarios). Plans can be generated automatically from protected machines or from a CSV file for large-scale environments.

Core Capabilities

From continuous protection to automated failback

Continuous data protection, predefined DR plans, one-click failover, non-disruptive testing, multi-tenant DR and compliance.

Capability

What EasyStack DR delivers

Continuous Data Protection
EasyStack DR continuously replicates protected workloads from the production site to the DR site. Replication is agent-based and operates at the block level, making it hypervisor-agnostic and workload-agnostic. Supported source workloads include:
  • ECF: KVM-based VMs, volumes, and hyperconverged clusters
  • ECNF: Kubernetes persistent volumes, container images, and application state
  • EAF: Model repositories, inference configurations, AI Workspaces, and API gateway settings
  • Bare metal: Physical servers via agent-based replication
Predefined DR Plans

DR plans eliminate manual runbooks. Administrators define recovery scenarios in advance — specifying which workloads to recover, in what order, and with what network and resource mappings. When disaster strikes, recovery is a single click.

One-Click Failover
Failover can be triggered manually or automatically based on health checks. The DR Controller orchestrates:
  • Provisioning target VMs on the DR site according to the DR plan
  • Attaching replicated data volumes
  • Applying network and security mappings
  • Starting workloads in the correct order
  • Validating recovery success
RTO is measured in minutes, not hours.
Non-Disruptive DR Testing

EasyStack DR supports non-disruptive DR testing. Administrators can run failover scenarios in an isolated environment without affecting production. Test results validate the DR plan, network mappings, and recovery time, ensuring readiness before a real disaster.

Automated Failback
After the production site is restored, failback returns workloads to production with zero data loss or downtime. Failback supports:
  • Incremental synchronization of changes made during failover
  • Automated VM creation on the source site
  • Validation before cutover
  • Minimal business impact
Multi-Tenant DR
For cloud service providers and large enterprises, EasyStack DR supports multi-tenant DR with:
  • Tenant-isolated replication and recovery
  • Role-based access control (RBAC)
  • Per-tenant DR plans and storage quotas
  • Audit logs retained for ≥180 days
Security and Compliance
  • Data never leaves the sovereign domain (both production and DR sites within jurisdiction)
  • Encrypted data transmission (TLS)
  • One-way hashed API keys
  • Audit logs for all DR operations
  • Compliance with Indonesia PDP Law, Thailand PDPA, GDPR, and financial regulatory requirements

RPO and RTO targets

Configuration

RPO

RTO

Continuous replication (seconds interval)

Near-zero to seconds

Minutes

Scheduled replication (15 min – 3 hours)

15 minutes – 3 hours

Minutes

Synchronous replication (for critical workloads)

Zero

Minutes

DR for ECF / ECNF / EAF

One control plane, every workload type

Region-to-region, hybrid, VMware-to-ECF, Kubernetes and AI inference DR — plus a unified view across all three platforms.

Diagram of heterogeneous workloads — virtual machines, Kubernetes containers, AI services and bare metal — protected by a single disaster recovery control plane

ECF Disaster Recovery

Protected workloads: KVM-based VMs, volumes, hyperconverged clusters, and software-defined storage.

DR scenarios

  • Region-to-region DR: Production in Region A, DR site in Region B (both EasyStack ECF).
  • Hybrid DR: Production on-premises ECF, DR site in a sovereign cloud or secondary data center.
  • VMware-to-ECF DR: Protect VMware VMs and fail over to ECF in case of disaster or migration.

Key capabilities

  • Continuous block-level replication of VM disks
  • DR plans with VM flavor, network, and security group mappings
  • One-click failover with RTO of minutes
  • Automated failback to production ECF
Diagram of heterogeneous workloads — virtual machines, Kubernetes containers, AI services and bare metal — protected by a single disaster recovery control plane

ECNF Disaster Recovery

Protected workloads: Kubernetes persistent volumes, container images, application state, and cluster configurations.

DR scenarios

  • Kubernetes cluster DR: Protect ECNF clusters and fail over to a DR ECNF cluster.
  • Application-level DR: Protect persistent volumes and application state for stateful workloads.
  • Cross-cluster recovery: Recover applications to a different ECNF cluster with different node specifications.

Key capabilities

  • Persistent volume replication
  • Kubernetes API integration for workload recovery
  • DR plans with namespace and service mappings
  • Automated failback with incremental synchronization
Diagram of heterogeneous workloads — virtual machines, Kubernetes containers, AI services and bare metal — protected by a single disaster recovery control plane

EAF Disaster Recovery

Protected workloads: Model repositories, inference configurations, AI Workspaces, API gateway settings, and token metering data.

DR scenarios

  • AI inference DR: Protect inference services and fail over to a DR EAF cluster.
  • Model repository DR: Replicate model files and version metadata to the DR site.
  • Agentic workflow DR: Protect multi-Agent collaboration configurations and sandbox environments.

Key capabilities

  • Model repository replication with version consistency
  • Inference endpoint failover with network remapping
  • API key and quota configuration recovery
  • Token metering data replication for billing continuity
Diagram of heterogeneous workloads — virtual machines, Kubernetes containers, AI services and bare metal — protected by a single disaster recovery control plane

Unified DR Across ECF / ECNF / EAF

EasyStack DR provides a single control plane for all workload types. Administrators can:

  • Create DR plans that span VMs, containers, and AI services
  • Define cross-workload dependencies and boot order
  • Monitor replication status from one dashboard
  • Trigger coordinated failover for the entire application stack
Deployment Scenarios

Seven ways enterprises deploy EasyStack DR

From region-to-region and sovereign cloud to DRaaS, government, financial, edge and remote sites.

Scenario

Description

Region-to-region DR

Production in Region A, DR site in Region B (both EasyStack).

On-premises to sovereign cloud

Production on-premises, DR site in a sovereign cloud environment.

VMware Alternative DR

Protect VMware VMs and fail over to ECF.

AI infrastructure DR

Protect EAF model repositories and inference services.

Cloud hosting provider DRaaS

Offer multi-tenant DR as a service to customers.

Government and financial DR

Sovereign DR with audit compliance and data residency.

Edge and remote site DR

Protect remote workloads with centralized DR orchestration.

Business Value & ROI

Measurable value across recovery, cost and compliance

Metric

Value

RTO

Minutes (vs. hours or days with traditional DR)

RPO

Near-zero to seconds (vs. hours with tape or daily backup)

DR infrastructure cost

Significantly reduced by reusing EasyStack infrastructure; no dedicated DR hardware

DR testing frequency

Unlimited, non-disruptive testing vs. annual manual drills

Operational efficiency

Automated orchestration reduces manual recovery effort by >80%

Compliance readiness

Audit logs, data residency, multi-tenancy

Business continuity

99.99% availability across production and DR sites

Cost Savings

EasyStack DR eliminates the need for dedicated DR hardware and standby infrastructure. The DR site can run on the same EasyStack platform as production, with resources provisioned on demand during failover. This reduces capital expenditure and operational overhead.

Risk Mitigation

Automated failover and predefined DR plans reduce the risk of human error during a crisis. Non-disruptive testing ensures DR plans are validated and ready. Compliance with sovereignty requirements avoids regulatory penalties.

Agility

Self-service DR testing and orchestrated failover enable IT teams to respond quickly to business continuity requirements. Multi-tenant DR enables cloud providers to offer DR as a value-added service.

Conclusion

Disaster recovery is not optional

Disaster recovery is not optional—it is a strategic requirement for private cloud, sovereign cloud, and AI infrastructure. EasyStack DR Solution provides automated, orchestrated protection for ECF, ECNF, and EAF workloads from a single control plane.

In an era where business continuity is non-negotiable, EasyStack DR is the reliable foundation for enterprise disaster recovery.

For enterprises replacing VMware, building sovereign clouds, or deploying AI workloads, EasyStack DR delivers:

  • Continuous block-level replication with RPO from near-zero to minutes
  • Predefined DR plans with automated orchestration and one-click failover
  • Non-disruptive DR testing and automated failback
  • Multi-tenant DR with audit and compliance
  • Sovereign data protection across regions
  • Integration with ECF, ECNF, and EAF

Ready to automate your disaster recovery?

Talk to our solution architects about region-to-region and sovereign DR, VMware-to-ECF protection paths and AI workload recovery — built on the EOS Digital Native Engine.