MCC Mag / published record

Cellular Failover Test Plan for Critical Sites

A cellular backup is useful only when essential services still work after the primary connection fails. A router showing a cellular signal does not prove…

Evergreen critical-site building linked to remote services by an orange cellular backup route above a broken primary connection, with a test timing scale below.

A cellular backup is useful only when essential services still work after the primary connection fails. A router showing a cellular signal does not prove that an application is reachable, an alarm leaves the site, or an operator can connect. Test the service path from primary loss through backup operation and return to normal routing.

Set an acceptable interruption for each service before the exercise. NIST’s contingency planning guide calls for alternate telecommunications that support essential functions within an organisation-defined time period and asks whether primary and alternate services share a point of failure. Turn those questions into observations that a tester can record.

Define what must work

List a small set of essential services and a transaction that proves each one works: a completed application request, delivered telemetry message, acknowledged alarm, or authorised remote session. Name the service owner, normal route, expected cellular route, and maximum acceptable interruption. A successful ping to the gateway proves little about the application behind it.

Record dependencies that may fail despite a healthy radio link: local power, DNS, authentication, VPN termination, cloud endpoints, and inbound access tied to the primary circuit’s address. Mark traffic that must stay off cellular under site policy. If an owner cannot define a tolerable interruption or a reliable check, settle that before a live outage.

Write down the site, gateway configuration, primary circuit, cellular connection, selected endpoints, test window, operator and stop authority. Specify whether the exercise removes the primary path or forces a route change. A forced switch checks the backup path but cannot establish that the gateway would detect a real primary failure.

Rehearse the decision points

Bring the network operator, service owners and incident coordinator through the sequence on paper. Decide who starts the test, watches monitoring, checks transactions, restores the primary path and communicates when the usual channel depends on the affected connection. Discuss a cellular link that comes up while an application remains unavailable, and a primary circuit that returns while traffic is still on cellular. Agree which result calls for observation and which calls for immediate rollback.

FEMA’s exercise planning guidance recommends defined objectives, advance communication, communications checks and an after-action improvement record. Use the rehearsal to assign those tasks rather than discovering missing roles during the outage.

Complete a go/no-go check: confirm cellular availability, save the approved configuration and rollback steps, align clocks to a known time zone, arrange monitoring outside the site, and make sure service owners can perform their checks. Tell affected people the window, expected behaviour, reporting channel and stop condition. Follow the site’s operating procedure for safety-sensitive services; defer a disruption that cannot be run safely.

Capture a baseline and remove the primary path

Before changing anything, record the active route, relevant public address, tunnel state, DNS response, application results and alert state. Run the agreed transactions from inside and outside the site where each view matters. Save timestamps and actual results, rather than a single “working” label.

  1. Start deliberately. Record the action and time used to remove the primary path. Choose an approved, reversible method that matches the failure under test, such as isolating the circuit at the agreed boundary. Note whether the gateway still sees a physical link; that may affect its health check.
  2. Observe detection and routing. Record when primary health first reports a problem, when cellular becomes active and when a new connection passes traffic. Check the forwarding path against policy: intended networks should use cellular and excluded traffic should remain blocked. An interface status alone is insufficient.
  3. Run service transactions. Repeat each selected check. Record the last success before failure, the first failure, the first sustained success and any intermittent failures. Try a fresh session as well as an existing one, because a path or address change can affect them differently.
  4. Check names and access. Resolve selected names from a client on the backup path, then use the resulting service. Test approved remote administration and site-to-site access from the locations that depend on them. If inbound access was never designed for cellular, record that limitation against the plan.
  5. Confirm alerts. Check that the primary-loss alert arrived at its intended destination and that a test notification can leave the site. Record receipt and acknowledgement times. Report exercise status over the agreed separate communications channel.

Measure interruption at the service boundary. The time from circuit isolation to router state change may be shorter than the gap between successful application transactions. Keep both measurements. If a service exceeds its agreed limit, preserve logs and observations before changing settings; simultaneous adjustments obscure the cause.

While the site runs on cellular, observe the same transactions long enough to detect repeated disconnects or slow responses. Record signal quality, interface errors and data use if the gateway exposes them, alongside application results. These measurements help distinguish a radio problem from a routing or service problem. Run only the traffic agreed for the exercise; a successful light check does not establish how the backup will behave under a heavier operating load.

Verify security on cellular

Check that the backup path follows approved access rules. Confirm required VPN encryption and authentication, authorised remote identities, and firewall or segmentation rules. Test an intended flow and a selected prohibited flow. Check whether traffic still passes through the expected security service. An application reachable through an exposed management interface is a finding, not a pass.

Record the visible source address and tunnel state when they affect allowlists or audit trails. An address change may explain a failed connection; a tunnel reconnect may explain a gap after the cellular interface appears ready. The result should establish whether the usable path meets the site’s security design.

Restore service and test failback

Reverse the loss action at the agreed time. Record when the primary circuit reports healthy and whether traffic returns automatically or needs operator action, according to policy. Watch for route oscillation, stalled sessions, and delayed or duplicated work. Repeat the service, DNS, remote access, alert and security checks used during failover.

Confirm that transactions complete through the intended primary route and that cellular returns to its expected standby state. If policy requires a stable-primary interval before switching back, measure it. Tell stakeholders when normal routing is verified; identify any service still using backup or operating in a degraded state.

If failback needs manual action, record the command, operator and reason. Check that the action did not leave a temporary route, firewall exception or monitoring override in place. Save the final configuration identifier so the next tester knows which settings produced the observed result.

Keep a repeatable evidence record

Use one test identifier and time zone throughout. A compact record should contain:

  • Scope: site, owners, transactions, interruption limits, window and stop authority.
  • Baseline: configuration version, active path, addresses, tunnels, DNS, application checks and monitoring state.
  • Timeline: loss action, detection, cellular activation, failed and recovered transactions, primary restoration and confirmed failback.
  • Findings: affected service, user impact, duration, intermittent behaviour, unexpected routes or alerts, and supporting evidence.
  • Follow-up: corrective action, owner, closure check and retest result.

Label each check passed, failed or not tested. An unavailable remote tester or omitted endpoint cannot support a pass. Attach relevant timestamped logs or screenshots and note clock offsets. FEMA’s communications evaluation guidance asks evaluators to identify demonstrated systems and retain test results and corrective actions. Its emergency-preparedness setting differs, but that evidence practice applies here.

Review material gaps while the sequence is fresh. Compare measured interruptions with agreed limits, assign an owner to each correction, and state the transaction that will prove it worked. Record configuration changes before a retest so the two runs can be compared.