• +44(0)7855748256
  • bolaogun9@gmail.com
  • London

The SIGNAL Framework

A Practitioner’s Reference for Warning Lifecycle Management in Platform Engineering

Version: 1.0
Author: Bola Ogunlana
Classification: Original framework, Cockpit to Cloud series
Origin: Derived from post-mortem analysis of the 1995 Sampoong Department Store collapse, mapped to persistent warning failure modes in cloud and platform engineering operations


What SIGNAL Is

SIGNAL is a 6-principle operational framework for managing persistent warnings in platform engineering, DevSecOps, and cloud operations environments.

It addresses a specific, named failure mode: warning decay, the process by which legitimate, accurate alerts progressively lose their authority through repeated dismissal, until the organisation is structurally blind to the risk they represent.

SIGNAL does not help you find more warnings. It is for teams that are already receiving the warnings and failing to act on them appropriately.

SIGNAL stands for:

LetterPrinciple
SStamp every dismissal with a timestamped decision log
IIdentify the act: triage or dismissal (forced classification)
GGrant engineers formal escalation authority
NNail pass/fail criteria before the review begins
AAge unactioned findings upward in severity, not downward
LLog the commercial rationale for every deferral

The 6 principles are an interdependent system. Implementing 3 of them produces a weaker version of the problem with better documentation. All 6 are required for the framework to function as designed.


The Problem SIGNAL Solves: Warning Decay

Definition

Warning decay is the gradual erosion of an alert’s authority within an organisation. A warning that is raised and dismissed repeatedly loses its ability to compel action, regardless of whether the underlying risk has changed.

The decay curve follows a consistent pattern:

  1. First occurrence: The warning receives full attention. A review is conducted.
  2. Second occurrence: The warning is acknowledged. Someone recalls the previous review and defers.
  3. Fifth occurrence: The warning is familiar. The team assumes it will resolve as it did before.
  4. Tenth occurrence: The warning is background noise. Nobody remembers who first raised it or why.
  5. Nth occurrence: The warning fires on the day the risk materialises. The team is surprised.

This is not a failure of detection. The system found the problem every time. It is a failure of organisational memory and accountability.

How It Manifests in Platform Engineering

Warning decay appears in 4 distinct operational contexts:

Security posture drift: CSPM findings that sit unactioned across sprints. Each sprint cycle, the finding is reviewed and deferred. Over time, the finding’s age becomes evidence of its benignness, rather than evidence of escalating risk. The team has trained itself to read “28 days old” as “safe” rather than “overdue.”

IaC configuration drift: Terraform plan deviations that are acknowledged but not remediated. The finding is logged, the team is aware, and nothing changes. The drift widens. The finding becomes permanent background.

AI-generated monitoring findings: AI monitoring tools generate findings at a rate no team can fully action. Teams apply volume-based triage: findings that survive long enough without causing an incident are reclassified as false positives, regardless of whether they are.

Vulnerability backlogs: CVEs triaged to “accepted risk” status on the basis that exploitation hasn’t occurred yet. The acceptance was reasonable at the time of triage. Nobody has reviewed whether it remains reasonable 90 days later.

Why Experience Makes It Worse

This is the counter-intuitive finding at the centre of the framework, and the one most likely to generate pushback from senior practitioners.

Experienced engineers pattern-match faster. A senior SRE who has seen a particular alert category fire 40 times without incident will triage it in seconds. That speed is the problem. The triage is not an assessment of the current alert; it is a pattern-match against historical data. It is accurate until the day the underlying conditions change, at which point the historical pattern actively misleads.

Junior engineers, who haven’t seen the alert before, escalate more. They lack the confidence to dismiss. This is regularly treated as a problem (unnecessary escalations, noise) when it is, in some respects, a protective mechanism.

SIGNAL doesn’t argue for slow triage. It argues for triage that is legible and accountable, so the pattern-match can be challenged when conditions change.


Theoretical Foundations

SIGNAL is grounded in 3 bodies of established work:

Normalisation of Deviance (Diane Vaughan, 1996): Vaughan’s post-mortem of the Challenger disaster introduced the concept that organisations can gradually come to treat deviant conditions as normal through repeated exposure without consequence. The O-ring erosion that destroyed Challenger had been observed and documented in 7 previous launches. Each time, managers decided the risk was acceptable. Over time, erosion became a standard launch condition rather than a warning sign. SIGNAL applies Vaughan’s mechanism directly to cloud operations: repeated exposure to an alert without incident normalises the alert as acceptable.

Swiss Cheese Model of Accident Causation (James Reason, 1990): Reason’s model holds that system failures occur when the holes in multiple independent defensive layers align simultaneously. SIGNAL is designed as one layer in the defensive stack, specifically the organisational layer that prevents legitimate warnings from being filtered out before they reach decision-makers with the authority and context to act.

Safety-II and Resilience Engineering (Erik Hollnagel, 2012): Safety-I frameworks focus on preventing failure. Safety-II frameworks focus on understanding why things usually go right, so that capability can be preserved under novel conditions. SIGNAL incorporates this by making triage decisions visible and legible: the goal is not to prevent all dismissals, but to ensure every dismissal is a genuine assessment that can be revisited when conditions change.

The Sampoong Empirical Case: On 29 June 1995, the Sampoong Department Store in Seoul collapsed in under 20 seconds, killing 502 people. Engineers had documented structural warnings for 3 months. Management received those warnings, conducted inspections, and chose to continue trading. The collapse didn’t happen because the warning system failed; it happened because the organisational response to the warning system failed. SIGNAL names the 6 specific failure modes in Sampoong’s response and provides a structural countermeasure for each.


The 6 Principles in Detail

S — Stamp every dismissal

What it means:

Every deferred or dismissed alert must carry a human-authored decision record containing 4 fields:

  • Who reviewed it (a named individual, not a team)
  • When they reviewed it (timestamp to the day)
  • Why they deferred it (a sentence of stated reasoning, not a category tag)
  • When it must be reviewed again (a specific date or trigger condition)

Why it matters:

Automatic ticket timestamps record when an alert was closed. They don’t record why. The absence of documented reasoning makes every deferral invisible in retrospect. When an alert has been deferred 8 times, the current team has no way to assess whether each deferral was reasoned or reflexive.

Sampoong’s managers inspected the building multiple times. None of those inspections produced written criteria or accountability trails. When the final inspection concluded the building was safe, there was no record of what prior inspectors had seen, concluded, or committed to re-evaluate.

What failure looks like without it:

A post-incident review cannot reconstruct who made the call to defer a pre-existing finding, when they made it, or what they assessed at the time. The deferral was not a decision. It was a default.

Implementation:

Add mandatory fields to your alert management workflow (Jira, ServiceNow, PagerDuty, or equivalent): Reviewer, Review date, Deferral reasoning, Next mandatory review. Make these fields required for ticket closure. A one-line reasoning field is sufficient; the requirement is that a human states it, not that it is comprehensive.

Measure:

Percentage of deferred critical/high findings with complete decision records. Target: 100%. Current baseline for most teams: near zero.


I — Identify the act: triage or dismissal

What it means:

Every alert closure must be classified as one of 2 explicitly named acts:

  • Triage: I have assessed this risk. I understand its current probability and impact. I am consciously choosing to prioritise other work. I am accountable for this finding until the next review date.
  • Dismissal: I am classifying this alert as noise, false positive, or otherwise invalid. I am removing it from active monitoring. I accept that if this classification is wrong, the finding will not resurface.

These are not synonyms. They have different downstream consequences and require different review processes.

Why it matters:

Most alert workflows have a single closure state. “Resolved,” “closed,” “accepted risk.” There is no semantic distinction between an engineer who carefully assessed a finding and accepted the risk and an engineer who closed the ticket to reduce queue length. Both look identical in the audit trail.

The Sampoong facilities managers believed they were triaging. An independent observer would call it dismissal. The difference between the two is whether the underlying conditions were genuinely assessed against stated criteria.

What failure looks like without it:

Your triage rate is unknowable. You cannot distinguish a team that is operating with good risk judgement from a team that is clearing queues. You find out which one you had in the post-incident review.

Implementation:

Two closure states in your ticketing workflow: Triaged (with mandatory deferral record per the S principle) and Dismissed (with mandatory classification: false positive, accepted risk, or out of scope). Track the ratio. A team with a high dismissal rate and a low triage rate warrants investigation; they are likely clearing queues rather than managing risk.

Measure:

Triage-to-dismissal ratio by team, by alert category, and over time. A sudden shift in the ratio (e.g., dismissal rate spikes in the week before a major feature delivery) is a leading indicator of queue-clearing behaviour.


G — Grant engineers formal escalation authority

What it means:

Every engineer on the platform or security team must have access to a named, formal escalation path for alerts they believe are being dismissed rather than triaged. This path must:

  • Be documented and published to the team
  • Lead to a named individual with closure authority
  • Carry no career risk for the engineer who uses it
  • Result in a documented decision from the recipient

Why it matters:

The Sampoong structural engineers had no authority to close the building. They could file reports and raise verbal concerns. The closure decision sat entirely with commercial management, who had different incentive structures and different information.

Every platform team replicates this structure by default. Security engineers identify findings. Product and platform leads decide what gets fixed and when. When a security engineer believes a finding is being dismissed rather than triaged, they have no formal channel to challenge that call.

Informal escalation (“I mentioned it in standup”) is not an escalation path. It is verbal noise that produces no documented decision and carries no accountability.

What failure looks like without it:

Engineers who believe a risk is being dismissed stop raising it after 1 or 2 attempts. They are not wrong to stop; they have learned that raising it produces no outcome. The risk accumulates. The engineers know. The decision-makers don’t.

Implementation:

Document the escalation path in your team’s security runbook. Format: “If you believe a critical or high-severity finding is being deferred without adequate assessment, you may escalate to [named individual]. This escalation will result in a documented decision within [timeframe]. Using this path is a professional obligation; it is not insubordination.” Publish it. Name the recipient. Review annually.

Measure:

Number of escalations per quarter. Zero escalations in a team with a high alert volume is a warning sign, not evidence of good practice. It means the path isn’t being used, either because it doesn’t exist in practice or because using it carries implicit risk.


N — Nail criteria before the review

What it means:

Before any team member reviews a persistent alert to determine whether to escalate or defer, they must commit in writing to the specific, observable condition that would change their conclusion.

Format: “We will escalate this finding if [observable condition] occurs or if [threshold] is exceeded.”

This commitment is recorded before the review begins, not after.

Why it matters:

Any review conducted without pre-committed criteria defaults to confirmation bias. The reviewer assesses the current state of the alert against their expectation of what it should look like, shaped by their prior conclusion that it is acceptable. If the prior conclusion was deferral, the review will almost always produce deferral again, because the reviewer is not looking for evidence of escalation; they are looking for evidence that their prior call was correct.

The Sampoong facilities managers walked the fifth floor with no defined threshold for what “unsafe” looked like. Their inspection confirmed what they needed it to confirm. The cracks were visible. The criteria for “too many cracks” did not exist.

What failure looks like without it:

Your reviews are not reviews. They are audits of the prior decision. They will produce the same conclusion the prior decision produced, until the risk manifests at a severity that cannot be denied.

Implementation:

Add a mandatory pre-review field to your alert triage workflow: “Escalation trigger: this finding will be escalated if [condition].” This field is completed before the reviewer examines the current state of the finding. It can be as simple as: “We will escalate if the affected resource count exceeds 10” or “We will escalate if this appears in combination with [finding type B].” The specificity matters less than the act of pre-commitment.

Measure:

Percentage of reviews that include a documented escalation trigger. Track whether reviews that include a pre-committed trigger produce different outcomes (escalation rate, time-to-remediation) from those that don’t. In practice, teams with pre-committed triggers escalate more and faster.


A — Age unactioned findings upward

What it means:

Unactioned critical and high-severity findings must escalate in severity and visibility over time, not expire into the backlog. Specifically:

  • A critical finding unactioned for 14 days escalates to mandatory named-owner review
  • A critical finding unactioned for 30 days escalates to senior management visibility
  • A high-severity finding follows the same path at 30 and 60 days respectively

These thresholds are a starting point. Teams should calibrate based on their environment and compliance requirements.

Why it matters:

Most AI monitoring and CSPM tools are configured with a default that works against safe operations: findings that aren’t actioned simply age in the backlog. The older a finding gets, the more buried it becomes under newer findings. A critical finding raised in January that hasn’t been remediated by March is harder to see in May, not easier.

Sampoong’s cracks didn’t shrink because nobody fixed them. They widened. The finding didn’t become less dangerous with age; it became more dangerous. Your unactioned findings follow the same physics.

The AI monitoring context makes this acute. AI tools generate findings at a volume that exceeds team capacity by design. The economic outcome is that old findings are systematically de-prioritised in favour of new ones, because new findings feel more urgent. This is backwards. A finding that has been present for 90 days and unactioned is a finding that has survived 3 monthly sprint cycles without triggering remediation. Its persistence is evidence of organisational failure to respond, not evidence of low risk.

What failure looks like without it:

Your oldest findings are your most dangerous. They are also your least visible. A 6-month-old critical finding buried in a CSPM backlog has been through 6 rounds of implicit acceptance without a single documented triage decision. It represents the maximum expression of warning decay in your environment.

Implementation:

Audit the age-escalation configuration of every monitoring tool in your stack. For each tool, answer: what happens to a critical finding unactioned for 30 days? If the answer is “nothing,” the tool’s default configuration is working against you. Most enterprise CSPM, SIEM, and vulnerability management tools support age-based severity escalation rules. Enable them. Set the thresholds. Assign the escalation recipients.

Measure:

Mean age of critical findings at time of remediation. P90 age. Number of findings older than 30/60/90 days by severity tier. These metrics are the clearest leading indicator of warning decay in a given environment.


L — Log the commercial rationale

What it means:

When a critical or high-severity finding is deferred for business reasons, the deferral must include a documented commercial rationale in the following format:

“We are accepting the risk of [specific named outcome] because [specific business reason]. This decision stands until [specific date or named trigger condition].”

This document must be authored by a named individual with sufficient authority to accept the stated risk. It must be stored with the finding record. It must be reviewed on the stated review date.

Why it matters:

Commercial pressure is a legitimate input to risk management decisions. The problem is not that commercial considerations influence deferral. The problem is that commercial considerations influence deferral invisibly, without accountability, and without a structured review mechanism.

The Sampoong managers made an economically rational decision to keep the store open. Lost revenue, customer disruption, and reputational cost were concrete and immediate. The cost of collapse was probable but not yet certain. That calculation is not inherently unreasonable; it is the calculation that risk management exists to make explicit and accountable.

What made the Sampoong decision fatal was not the calculation itself. It was that the calculation was never written down, never reviewed, and never subjected to external challenge. The decision was invisible until it was catastrophic.

Every sprint planning session contains a version of this decision. A critical finding is deferred because a feature deadline takes priority. That is a legitimate business call. It should be made by someone with the authority to make it, documented with the specific risk being accepted, and reviewed on a stated date. If the person making the call won’t write those 3 sentences, they are not managing risk. They are hoping.

What failure looks like without it:

Post-incident reviews cannot reconstruct who accepted which risk, when, or why. The deferral looks like negligence in retrospect because there is no documentation of the reasoning. The individual who made the call has no protection and no recourse. The organisation has no learning artefact.

Implementation:

Build a Risk Acceptance Template into your alert management workflow. 3 mandatory fields: named risk, business reason, review date or trigger. Make this mandatory for any closure of a critical finding that is not full remediation. Store it with the finding. Route it to the named reviewer on the review date. If the form isn’t completed, the finding cannot be deferred.

Measure:

Percentage of critical finding deferrals with complete risk acceptance documentation. Mean time between risk acceptance and scheduled review. Number of risk acceptances that were actually reviewed on their stated date versus allowed to lapse.


How the Principles Interlock

SIGNAL is not a checklist. The 6 principles are designed as a system, and each one depends on the others to function correctly.

Failure modeWhich principles address it
Decisions made invisibly, no audit trailS (stamp), L (log)
Triage and dismissal look identicalI (identify)
Engineers silenced by hierarchyG (grant)
Reviews confirm prior decisionsN (nail)
Old findings buried by new onesA (age)
Business pressure undocumentedL (log)
Named accountability absentS (stamp), G (grant), L (log)

The most common partial implementation failure: teams adopt S and I (documentation and classification) without A (age escalation). This produces a well-documented backlog of ageing findings that are triaged on schedule but never remediated. The documentation creates the illusion of managed risk. The findings accumulate.

The second most common failure: teams adopt A and L (escalation and logging) without G (engineer authority). This creates a process where findings escalate on schedule to managers who make deferral decisions without the technical context the escalating engineers hold. The escalation produces paperwork, not action.

All 6 principles must operate together.


Implementation Roadmap

SIGNAL is designed to be adopted in 3 phases. Each phase delivers standalone value while building toward full implementation.

Phase 1: Visibility (Weeks 1–4)

Implement S and I. Stamp every dismissal. Classify every closure.

Goal: make the current state legible. You are not trying to change behaviour yet. You are trying to understand what behaviour is actually occurring.

Deliverables:

  • Decision log fields added to alert management workflow
  • Closure classification states (triage/dismissal) in ticketing system
  • Baseline metrics: triage rate, dismissal rate, mean finding age by severity

Phase 2: Accountability (Weeks 5–8)

Implement G, N, and L. Grant escalation authority. Require pre-review criteria. Document commercial rationale.

Goal: make deferral a named human decision rather than a process outcome.

Deliverables:

  • Escalation path documented and published
  • Pre-review commitment field in triage workflow
  • Risk acceptance template in alert workflow
  • Named owner on every critical finding

Phase 3: Lifecycle Management (Weeks 9–12)

Implement A. Configure age escalation across the monitoring stack.

Goal: ensure the passage of time increases pressure to remediate, not decreases it.

Deliverables:

  • Age escalation rules configured in all CSPM, SIEM, and vulnerability management tools
  • Escalation recipients named and notified
  • Age-based metrics in security posture dashboard
  • Quarterly SIGNAL review cadence established

What Good Looks Like

A team operating SIGNAL correctly exhibits the following measurable behaviours:

  1. Every critical finding older than 14 days has a named owner and a documented decision record
  2. The triage-to-dismissal ratio is stable and known. Sudden shifts in the ratio trigger a management review
  3. Engineer escalations occur at least quarterly. Zero escalations in a team with high alert volume is a red flag
  4. Age escalation fires reliably. Findings do not silently survive past their escalation thresholds
  5. Risk acceptance documents exist for every deferred critical finding, with stated review dates that are actually met
  6. Post-incident reviews can reconstruct the full decision history of any pre-existing finding, including who reviewed it, what they assessed, and what criteria they committed to

Common Objections

“This creates too much process overhead for the team.”

The overhead is proportional to the number of decisions the team is already making. SIGNAL doesn’t add decisions; it makes existing decisions legible. If the volume of decisions is unmanageable, that is a resourcing problem that SIGNAL surfaces but didn’t create.

“We already have SLAs on critical findings.”

SLAs measure time-to-close, not quality of closure. A finding closed as “accepted risk” in 24 hours meets the SLA and is invisible in the backlog. SIGNAL addresses closure quality, not closure speed.

“Experienced engineers can triage these alerts without a formal process.”

This is precisely the objection the framework anticipates. Experienced engineers pattern-match against historical data. When the underlying conditions change, historical pattern-matching produces incorrect conclusions at high speed. The framework doesn’t slow down triage. It makes the reasoning behind fast triage legible, so it can be challenged when conditions change.

“We don’t have the tooling to implement age escalation.”

Every enterprise-grade CSPM, SIEM, and vulnerability management tool in current use supports age-based escalation rules. The question is whether those rules have been configured. In most environments, they haven’t, because the default configuration optimises for finding throughput, not warning integrity.


Framework Provenance

The SIGNAL Framework was developed by Bola Ogunlana as part of the Cockpit to Cloud series, which maps historical failures in safety-critical industries to operational failure modes in cloud and platform engineering.

The Sampoong empirical case (Seoul, 29 June 1995) provided the 6 specific failure modes that the framework addresses. Vaughan’s normalisation of deviance provided the theoretical mechanism. Reason’s Swiss Cheese Model provided the defensive layer context.

Version history and supporting templates: blog.ogunlana.net

Leave a Reply

Your email address will not be published. Required fields are marked *