Skip to main content
runbooks/enterprise-networking-security/palo-alto-ha-pair-split-brain.md
CRITICAL SEVERITYEnterprise Networking & SecurityFirewalls

Palo Alto HA Pair Split Brain: Diagnose and Recover Safely

Severity
CRITICAL
Target Time
30 minutes to stabilise HA control communication
DomainEnterprise Networking & Security
Verified
Overview

Confirm Palo Alto Networks HA split brain by proving both peers are active, then restore HA1 control communication without suspending, rebooting or forcing either firewall before traffic ownership and out-of-band access are understood.

Share

01 // Diagnose

Symptom

Incident signalWhat responders observe

An active/passive pair reports both firewalls active, or an active/active pair reports both peers active-primary. Duplicate forwarding, asymmetric traffic, duplicate IP or MAC behaviour, session loss and route instability can follow. Treat a single peer-down alarm as insufficient: establish the state from both management planes.

Detection Signature

Detection evidenceMetrics, logs, and confirmation commands
  1. Use independent management or console access to both peers and record model, PAN-OS release, HA mode, local state, peer state and current traffic ownership.

  2. Run the read-only HA commands below on both peers and compare control-link counters, transitions and flap evidence.

  3. Check HA1, HA1 backup and management heartbeat paths end to end, including switch ports, VLANs, routing and packet loss.

  4. Correlate the transition time with system-resource pressure, upgrades, interface changes and network maintenance.

  5. Do not infer split brain solely from HA2/session-synchronisation failure; the defining evidence is simultaneous active ownership caused by lost HA control communication.

Root Cause Analysis

Causal chainWhy the incident occurred
  1. HA1 failed with no functioning HA1 backup, so each peer stopped receiving control heartbeats and elected itself active.

  2. The transport carrying HA1 or HA1 backup stopped passing heartbeat traffic because of a switch, router, VLAN, interface or path failure.

  3. Management-plane resource pressure delayed heartbeat processing even though the physical link remained up.

  4. Heartbeat Backup was unavailable or not enabled, removing the management-port fallback for heartbeat and hello messages.

  5. Less commonly, a PAN-OS defect or mismatched HA configuration caused repeated state transitions; confirm the exact release and support advisories before software action.

02 // Contain & Prevent

Blast Radius

  • Traffic traversing the pair can be duplicated, black-holed or routed asymmetrically while both peers claim active ownership.

  • Stateful sessions, VPNs, dynamic routing adjacencies, NAT and neighbouring switches may retain conflicting state.

  • A premature suspend, reboot or forced-functional action can remove the peer still carrying valid traffic and widen the outage.

Prevention Measures

Prevent recurrenceControls and architectural guardrails
  • Configure HA1 backup on a genuinely independent path and evaluate Heartbeat Backup on the management ports using current Palo Alto Networks guidance.

  • Monitor HA link state, heartbeat loss, transition count and flap statistics from both peers.

  • Keep peer models, content and PAN-OS versions compatible, and validate HA state after every network or software change.

  • Exercise a controlled HA failure test with documented traffic-owner checks and out-of-band access.

  • Maintain the related firewalls operational reference with the incident record.

03 // Fix & Intervention

Pre-Flight Checks

Change gateChecks required before intervention
  1. Establish out-of-band access to both peers and a separate path to the affected network.

  2. Identify which peer currently carries valid traffic by checking sessions, interface counters, routing adjacencies and upstream/downstream observations.

  3. Freeze commits, upgrades, reboots, suspend/functional requests and automated failover actions.

  4. Capture configuration and operational state from both peers under the incident record.

  5. Obtain network and firewall owner approval before changing an HA transport, election or peer state.

Execution CommandsCOMMANDS

# Run on both peers; these commands are read-only.
show high-availability all
show high-availability state
show high-availability control-link statistics
show high-availability transitions
show high-availability flap-statistics
show system resources

04 // Verify & Recover

Verification Steps

Recovery proofEvidence required before closure
  1. Confirm exactly one active peer in active/passive mode, or the designed primary/secondary roles in active/active mode, for at least 15 consecutive minutes.

  2. Verify HA1 and the intended backup path are up and heartbeat/control-link error counters are no longer increasing.

  3. Confirm configuration and session synchronisation return to the environment baseline.

  4. Check routing neighbours, VPNs, NAT, session establishment and application probes from both sides of the firewall.

  5. Keep enhanced monitoring through the next peak traffic period before closing the incident.

Rollback Protocol

Safe reversal path
  1. If an approved HA transport or election-setting change worsens state, revert only that change from the captured configuration through the normal commit process.

  2. Preserve the restored physical path while the rollback commits.

  3. Do not roll back by blindly suspending, rebooting or forcing a peer functional

  4. return to the last documented single-owner topology under vendor support direction.

Escalation

Conditions requiring additional ownership
  • Both peers remain active after HA1 communication is restored.

  • Traffic ownership cannot be established safely or out-of-band access is unavailable.

  • HA transitions continue, management resources are exhausted, or the pair shows a suspected PAN-OS defect.

  • Any action would require a peer suspend, reboot, forced state, software change or production commit without a tested plan.

  • Escalate to the network incident commander and Palo Alto Networks Support with tech-support files from both peers.

Authoritative Sources

Palo Alto HA Pair Split Brain: Diagnose and Recover Safely - Incident Runbook | KBY Technologies