Who does the work in managed disaster recovery

When a disaster happens, a backup is the start of the job, not the end of it, because getting your data back and getting your business running again are two different pieces of work.

—

5 minutes
Woman looking at a computer screen

If you run your business on servers somebody else manages, backups are either already part of what you pay for, or you’ve recently found out that they aren’t. When a disaster happens, a backup is the start of the job, not the end of it, because getting your data back and getting your business running again are two different pieces of work.

Disaster recovery is the second piece. It covers standing up a replacement server, making sure your traffic can reach it, deciding what comes back in what order, and proving all of that works before the day you need it. Most of it isn’t something you buy, it’s something somebody does, continuously, for as long as you own it.

What running disaster recovery yourself involves

Running disaster recovery yourself means owning a list of tasks, and almost none of them happen during the outage.

Before anything breaks, someone on your team:

  • Provisions the recovery environment. This is a standby copy of your server that sits ready in case the live one fails, sized against what it would be replacing and kept in step as your systems change
  • Configures the network path so traffic can reach that standby copy, then tests the route and confirms it works
  • Writes the recovery sequence, so what comes back, in what order, and which systems depend on which
  • Sets a testing schedule, runs the tests, and files the results
  • Revisits all of it as your environment changes and people leave, or the sequence on the page stops matching the servers you run

Nothing looks wrong if the network path is missing, which is what makes it the one to check first. A standby server can start up perfectly and still be unreachable.

Then, during the outage, the same people:

  • Decide the outage is bad enough to declare a disaster
  • Find the credentials and start the failover, which is the switch that makes the standby copy your live server
  • Check the network path is behaving
  • Bring services back in the right order and confirm the data is usable
  • Redirect DNS so your customers reach the new server
  • Keep those customers informed while the rest of it is happening
  • Open a ticket with the recovery platform vendor if that platform is the thing misbehaving, then wait in their queue

What support covers, and what stays yours

Support moves part of that list off your team, and the rest stays with you no matter who you buy from.

What support should coverWhat stays yours whoever you buy from
Provisioning the standby copy, from the day you buy itDeciding when to declare a disaster
Configuring the network path and testing the routeStarting the failover and the failback
Monitoring the copying and raising alerts when it stopsCleanup inside your own applications
Running scheduled test failovers and filing dated evidenceMaintaining your recovery documentation
Owning the vendor relationship with the recovery platformHolding your credentials and encryption keys
Senior engineers on the line through failover and failbackControlling your DNS

A provider offering to take anything from the right-hand column off you should be able to say how they would know your business is having a disaster, or what a correct state looks like inside one of your applications.

You might expect a supported service to press the button for you, but most don’t. If your provider starts the failover, your recovery starts when somebody on their side picks up your ticket. If you start it, failover begins when you decide.

Ask to see the route test results for the network path. Anyone who configures it will have them.

The table also doesn’t tell you what happens during a live event. Being talked through a recovery on the phone and having somebody run it for you are different services, and both get called support.

Somebody has to write your recovery documentation

Somebody has to write your recovery documentation outlining what to bring back, and in what order. Without it, the first hour of an outage is spent working that out while the business waits.

That documentation is the runbook, and it also names which systems depend on which, who is allowed to declare a disaster, and what has to be true before anyone starts a failover. It is important to verify and update this runbook anytime a change is made to your servers or applications that run on them, and at minimal yearly.

You cannot buy an application or product to do this for you. Since the runbook documenting your recovery process describes your environment and not the platform’s,it means somebody has to sit down with your systems and their dependencies and write it, then keep it current as things move. It’s the item most likely to be nobody’s actual job, so ask any provider whether building and maintaining it is part of what you’d be paying for.

When paying for support is the right call

Paying for disaster recovery support is the right call when the people who would plan for, periodically test a simulation, and run your recovery either don’t exist or have something better to focus on during an outage. Some businesses have nobody who would be confident starting a failover. Some businesses just want the extra assurance of experts on hand during a critical outage. 

If you have determined to explore disaster recovery services, taking time today to document the order your systems come back on and who’s allowed to declare a disaster ensures your systems recover smoothly. Reach out to one of our Nexcess engineers for help on how to build a recovery strategy tailored to your infrastructure. We’ll help you set up, test, and manage it every step of the way.