Skip to main content

Business Continuity & Disaster Recovery

Business Continuity Planning (BCP) and Disaster Recovery (DR) ensure that a bank can continue critical operations and recover IT systems following a disruption โ€” whether from a technology failure, natural disaster, cyber incident, or pandemic.

๐Ÿ“Œ Key Conceptsโ€‹

TermDefinition
BCPBusiness Continuity Plan โ€” procedures to maintain critical business functions during a disruption
DRDisaster Recovery โ€” IT-focused plan to restore systems and data after a failure
RTORecovery Time Objective โ€” maximum acceptable time to restore a system or process after disruption
RPORecovery Point Objective โ€” maximum acceptable data loss measured in time (e.g. last backup was 4 hours ago)
BIABusiness Impact Analysis โ€” identifies critical processes, dependencies, and acceptable downtime
Crisis Management Team (CMT)Senior leadership group activated during a major incident to coordinate response
Hot SiteFully equipped backup data centre ready to take over immediately
Warm SiteBackup site with infrastructure ready but requiring some setup time
Cold SiteBackup facility with space and power but no pre-installed equipment

๐Ÿ—๏ธ BCP Frameworkโ€‹

Tier Classification

TierRTORPOExamples
Tier 1 (Critical)Institution-definedInstitution-definedCore banking and critical payment services
Tier 2 (Important)Institution-definedInstitution-definedImportant customer and operational services
Tier 3 (Normal)Institution-definedInstitution-definedSupporting and internal services

๐Ÿ› ๏ธ BCP Activation Workflowโ€‹

Minor Disruption (Technology / Process Failure)

  1. Incident reported to IT Service Desk and Business Continuity Officer (BCO)
  2. Incident assessed against documented severity and BCP trigger criteria
  3. If triggered: BCP team activated; workaround procedures invoked (manual processing, backup channels)
  4. Stakeholders notified: management, affected business units, customer communications if required
  5. Recovery actions tracked; system restored within RTO
  6. Post-incident review conducted; lessons captured

Major Disruption (Site Loss / Disaster)

  1. Crisis Management Team (CMT) convened by Group CEO or delegated authority
  2. Incident severity assessed; BCP invoked at department / bank-wide level
  3. Staff relocated to alternate site (hot/warm site) or WFH arrangements activated
  4. DR initiated for affected IT systems: failover to backup data centre
  5. Critical business functions restored in priority order (Tier 1 first)
  6. MAS notified within required timeframe (see regulatory requirements)
  7. Ongoing status updates to CMT; customer and media communications managed
  8. Gradual return to primary site once declared safe and stable

๐Ÿงฎ RTO / RPO Examplesโ€‹

Core Banking System

RTO = 4 hours
RPO = 15 minutes (continuous replication to DR site)

If primary data centre fails at 10:00 AM:
- DR site takeover initiated: 10:05 AM
- System available at DR site: 2:00 PM (within 4-hour RTO)
- Data loss: transactions from 9:45 AM onwards may need manual reprocessing

FAST Payment Processing

RTO = 2 hours
RPO = 0 (synchronous replication; no data loss)

Failover to backup payment gateway activates automatically
Manual fallback: hold payments in queue; release when system restored

๐Ÿงช Testing Requirementsโ€‹

Test TypeFrequencyScope
Tabletop ExerciseRisk-based scheduleCMT and department BCOs walk through scenarios
Component DR TestApproved test scheduleIndividual system recovery tested in isolation
Critical-System Recovery TestAt least annually where requiredRecovery capability and established RTO validated
Staff Alternate-Site TestApproved test scheduleStaff validate alternate-site or remote-working arrangements
Call Tree / Communication TestApproved test scheduleEmergency contact lists verified and tested

๐Ÿ“‹ Regulatory Requirementsโ€‹

  • MAS Technology Risk Management Guidelines: DR requirements for systems supporting critical banking services
  • MAS Business Continuity Management Guidelines: BCP framework, testing frequency, CMT structure
  • Qualifying system malfunctions and IT security incidents must be notified to MAS within the applicable timeframe, including the one-hour notification requirement where the relevant notice applies
  • Outsourcing arrangements must include appropriate continuity, recovery, oversight and exit controls
  • Test evidence, deficiencies and remediation actions must be documented and available for supervisory review
  • The board and senior management oversee resilience according to the institution's governance framework