Skip to content
Critical IT Services Restart Framework

Restart Framework

A structured, dependency-driven approach to restarting critical IT services after a destructive cyber incident. Vendor-neutral, sector-agnostic, and free to adopt as your organization's own template under a Creative Commons non-commercial license.

Why a restart framework

why a restart framework

Technical restoration does not equal trust.

After a destructive cyber incident, such as a wiper attack or enterprise-wide ransomware, one of the most dangerous moments is the rush to get back online. Pressure to restore whatever the loudest part of the business needs, rather than what is technically ready, reintroduces compromised credentials and persistence into an environment that was just cleaned. A booted server is not a clean server. Restoration is not remediation.

The Restart Framework governs that gap: the high-consequence period between containment and eradication on one side, and safe business-as-usual operations on the other. It replaces ad hoc, urgency-driven restoration with disciplined sequencing based on hard technical dependencies.

The urgency trap

Restoring by business loudness instead of technical readiness brings compromised credentials and persistence back in.

The integrity gap

A restored application does not prove a clean environment or verified data integrity.

Security as a prerequisite

EDR telemetry, Zero Trust isolation and identity validation must come before network reconnection, not after it.

Dependency-driven sequencing

Recovery follows hard technical dependencies (identity, network, data), not departmental urgency or convenience.

Scope

what it governs, and what it leaves to other disciplines

This framework governs

  • Restart sequencing and service prioritization, mapped by dependency
  • Minimum viable conditions before any service resumes
  • Dependency validation: identity, network, infrastructure, data, security controls and third parties
  • Identity, PKI and privileged access recovery
  • Trust and integrity re-establishment before return to production
  • Executive governance, sign-off and veto authority
  • Degraded and compensating operating modes during recovery

Left to other disciplines

  • Incident Response and Digital Forensics
  • Threat Hunting
  • Disaster Recovery planning and Business Continuity Management
  • Crisis Communications
  • Regulatory reporting obligations (e.g. DORA, NIS2)

The framework complements these disciplines. It does not replace them.

Principles

eleven principles

The framework rests on eleven principles drawn from frontline incident response experience and established cyber resilience standards.

  1. 1

    Validated trust in human and technical systems

    Credentials, code, configuration and access rights are validated before they are reintroduced.

  2. 2

    Structured stakeholder communication

    Deliberate, calibrated information flow between technology leaders, the board, regulators and third parties.

  3. 3

    Safety before speed

    Stability and non-reinfection outweigh arbitrary timelines and commercial pressure.

  4. 4

    Minimum viable operations first

    Re-establish core survival infrastructure and a minimum viable business before full feature sets.

  5. 5

    Trust before restoration

    Trust requires communicating root cause, current state and remedies, backed by verified system integrity.

  6. 6

    Dependency-driven recovery

    Services recover by hard technical prerequisite, such as DNS and identity, not by business urgency.

  7. 7

    Security controls as recovery prerequisites

    Telemetry, segmentation, EDR and monitoring are operational before production platforms are revived.

  8. 8

    Defined authority and accountability

    No layer progresses without governance sign-off. The CISO and the Executive Recovery Governance Committee hold veto power.

  9. 9

    Progressive, not big-bang

    Iterative, throttled restoration. Mass simultaneous boots create instability and reinfection vectors.

  10. 10

    Continuous monitoring during recovery

    Heightened logging, telemetry and active scanning cover every revived asset.

  11. 11

    Repeatability and auditability

    The full recovery trail is documented and governed, supporting compliance with DORA and NIS2.

Lifecycle

the four-phase restart lifecycle

Phase I

Assess

Determine exploit type, vector and blast radius. Sever hybrid sync links. Establish out-of-band communication channels.

Phase II

Remediate

Purge threat-actor footprints. Restore infrastructure to known-clean baselines using a hybrid brownfield / greenfield approach.

Phase III

Assure

Verify data and system integrity end to end. Complete Active Directory schema audits and finalize signed executive attestations.

Phase IV

Reconnect

Re-establish connectivity through low-value synthetic test transactions, throttled network adjacency and heightened SOC telemetry.

Each phase requires explicit risk acceptance and sign-off before the next begins. No phase is skipped under pressure. The phase structure is adapted from the CMORG Reconnection Framework (UK financial sector, v3.0, July 2025).

Sequence

six recovery layers

Critical IT services are grouped into six layers, sequenced by hard technical dependency. Each layer gates the next.

  1. 1

    Emergency Comm, Identity Resumption & Continuity

    Illustrative RTO: ~0.5h per service

    Out-of-band collaboration and communication, emergency and break-glass access, crisis response archive, internal user communication.

    Gate before next layer: All four services operational and free of indicators of compromise.

  2. 2

    Core IT Infrastructure

    Illustrative RTO: 0.5–2h

    Data center, network (WAN/LAN, perimeter, DNS/DHCP/NTP, cloud network, secure remote access, Zero Trust segmentation) and compute (hypervisor, storage, app and web servers, enterprise monitoring).

    Gate before next layer: Clean baseline reached; hypervisor restored from a clean image; enterprise monitoring live.

  3. 3

    Information Security Stack

    Illustrative RTO: 2–4h

    EDR, SIEM, DLP, identity and access management, cryptography and PKI, systems and device management.

    Gate before next layer: Most controls return in parallel. DLP waits until IAM reaches a baseline operational state.

  4. 4

    Data Vault & Recovery Orchestration

    Illustrative RTO: ~4h

    An air-gapped, immutable environment used to stage, scan and validate clean backups before anything returns to production.

    Gate before next layer: Recovery environment network-isolated and confirmed clean before any promotion.

  5. 5

    Mission-Critical Apps & Data

    Illustrative RTO: 4–8h

    Application recovery plan execution, integrations and middleware, data pipelines, application-level service desk.

    Gate before next layer: Layers 2–4 operational; application runbooks and clean backups available.

  6. 6

    Technology Operations

    Illustrative RTO: 2–8h

    Executive and major-event support, service desk, onsite support, manufacturing and supply chain support, file shares, productivity software, campus and wireless LAN.

    Gate before next layer: Layer 2 operational; IAM in place for access-dependent services.

RTOs are illustrative planning targets for a mature organization, not guarantees. Actual recovery time depends on incident severity, backup posture and preparedness. Real destructive events have taken weeks to fully reconnect.

Recovery model

brownfield or greenfield

The framework assumes a hybrid recovery model rather than a full rebuild by default.

Brownfield (preferred default)

Remediate and restore on top of existing infrastructure. It preserves forensic evidence, is faster, and avoids the false confidence a rebuild can create when the true scope of compromise is not yet known.

Use when

  • Only part of the environment is compromised
  • Clean, recent backups exist
  • Identity systems are intact or repairable
  • Forensic data is actionable
  • Time to recovery is critical to continuity

Greenfield (reserved for severe cases)

A full rebuild, reserved for situations where trust in the existing estate cannot be re-established.

Use when

  • Environment fully lost, with no usable backups
  • Suspected deep compromise (firmware or rootkit level)
  • Identity systems fully subverted
  • A legal or regulatory mandate requires it

In practice, most organizations run both at once: hardened components such as identity, PKI and the isolated recovery vault are rebuilt greenfield, while the wider estate is remediated brownfield.

Governance

governance: sign-off, veto and degraded modes

Sign-off and veto controls

Layer gate enforcers

The CISO and the Executive Recovery Governance Committee hold veto power at every layer threshold.

Bilateral attestation

Incident response / forensics and GRC leads both sign off before any production interconnect.

Evidentiary baselines

Verified clean backup hashes, credential rotation and cleared SOC telemetry are required.

Degraded operating modes

Controlled exposure

Essential services run in temporary, restricted modes while wider recovery continues.

Compensating safeguards

Manual approval queues, read-only database states and restricted user cohorts.

Heightened telemetry

Aggressive EDR policies and real-time egress monitoring during mid-recovery operations.

Adopt it

adopt it as your template

This is a starting point, not a fixed prescription. Map your own services, set your own targets, and rehearse before you need it.

  1. 1

    Map your services

    Place your own critical IT services into the six recovery layers.

  2. 2

    Set your own RTOs

    Calibrate the illustrative targets to your environment and backup posture.

  3. 3

    Assign gate owners

    Name the CISO and GRC owners accountable for each layer's sign-off.

  4. 4

    Rehearse it

    Run tabletop exercises so the sequencing is trusted, not theoretical.

  5. 5

    Align to regulation

    Map the governance evidence to obligations such as DORA and NIS2.

Downloads

downloads

Both resources are published by the Cyber Resilience Manifesto under the Creative Commons Attribution-NonCommercial 4.0 license (CC BY-NC 4.0): free to share and adapt for non-commercial use, with attribution.

Access Request

Provide your contact details to access the Restart Framework resources.

Submit your contact details once to unlock the handbook and the spreadsheet. The same access applies to both files.

Privacy Notice

We use your information to provide access to these materials and to send professional updates related to cyber resilience initiatives. See our Privacy Policy.

Included after request

Download the Handbook

The full framework: scope and boundaries, definitions, the eleven principles, the four-phase lifecycle, the brownfield vs. greenfield decision criteria, the six-layer restart sequence matrix with minimum activation criteria, and the governance and sign-off controls.

Critical-IT-Services-Restart-Framework-Handbook.pdf

Download the Spreadsheet

The restart sequence matrix in working form. Every critical IT service with its recovery layer ID, domain and sub-domain, service definition, illustrative RTO, sequencing type (series or parallel) and minimum requirements prior to activation. Adapt it to your own estate.

Critical-IT-Services-Restart-Framework-Spreadsheet.xlsx

Download access

Submit the form to unlock the files below.

Download the HandbookLocked
Download the SpreadsheetLocked

Credits

authors and contributors

Authors

Francesco Chiarini, Patrick Lechner, Shwetha Babu Prasad

Key contributors

Saketh Varma Namburi, Alex Sharpe, Jordan Schoenherr

© 2026 Cyber Resilience Manifesto. Licensed under CC BY-NC 4.0. License terms

Audience

who it is for

Cyber resilience leaders, CISOs, incident response and disaster recovery teams, infrastructure and operations leaders, business continuity professionals, technology risk leaders, recovery governance committees and executive decision makers.

Related

related

References

references

  1. 1. CMORG (Cross Market Operational Resilience Group), CMORG Reconnection Framework, v3.0, July 2025 (TLP:CLEAR). Link
  2. 2. Renfrow, H., “IT Network Restoration After Ransomware: Why Brownfield Beats Greenfield”, Cyber Technology Insights, 29 May 2025. Link