What the night before Challenger launched teaches DevSecOps engineers about organisational silence

On the evening of 27 January 1986, Roger Boisjoly spread his data across a table in a teleconference room at Morton Thiokol and made his case. The O-rings that sealed the joints of Challenger’s solid rocket boosters had never been tested below 53°F. The forecast for the next morning: 29°F. The correlation between cold and O-ring damage was unmistakable in his own data. He had written a formal memo six months earlier saying the same thing in precise engineering language.
Morton Thiokol’s senior vice president, Jerry Mason, looked at his engineering VP, Bob Lund, and said: “Take off your engineering hat and put on your management hat.”
Lund changed his vote. The launch was approved. The engineers who opposed it refused to sign the new recommendation. Seventy-three seconds after launch on 28 January 1986, Challenger disintegrated over the Atlantic. Seven people died on live television.
The O-rings failed exactly where Boisjoly said they would, at the temperature he warned about, in the way his July 1985 memo had described in writing.
The data was right. The decision was wrong. And the process used to make that decision looked, from the outside, entirely procedural.
That last part is the bit most post-incident analyses miss. Security engineers and SREs encounter this pattern weekly, not in the form of a dramatic teleconference, but in risk acceptance forms, change advisory board reviews, and “business context” conversations that reshape technical findings before they reach a decision-maker. The mechanism is quieter. The outcome, eventually, isn’t.
The warning was already on the table
On 31 July 1985, Boisjoly submitted an internal memo to Morton Thiokol management. It said the O-ring erosion problem posed a risk of “loss of human life” and that the situation had changed “drastically” from an accepted design limitation to a genuine flight-critical hazard. He attached photographs of erosion damage from previous flights. He used words like “unacceptable” and “catastrophe.”
The memo went into a file. Subsequent launches proceeded without incident. By January 1986, the warning had been rendered organisationally routine.
This is how documented technical risk gets neutralised without anyone making a conscious decision to dismiss it. Each launch that didn’t result in catastrophe reduced the apparent urgency of Boisjoly’s concern by one increment. The warning didn’t go away. It just stopped being heard as urgent. Call it warning decay: a formally documented concern that loses organisational weight through the accumulation of uneventful outcomes.
DevSecOps teams reproduce this mechanism at scale. Your CVE backlog is a warning decay engine. A critical finding that hasn’t triggered a breach in 90 days starts to feel less critical, not because anything about the risk has changed, but because the absence of consequences has been implicitly treated as evidence of acceptable risk. The “accepted” status in Jira represents accumulated organisational familiarity, often not a fresh risk decision.
One thing teams consistently discover when they audit old risk acceptances: many were signed by people who never read the original technical finding. The risk rating had been softened somewhere between the engineer’s assessment and the form submission. Nobody lied. The form just didn’t ask the right questions.
The action: Apply a half-life test to every accepted risk in your backlog. If a finding has sat at “accepted” for more than 60 days without documented re-evaluation, treat it as a new finding and require a fresh sign-off from someone who has read the original technical evidence. Not a summary. The original.
The management hat instruction
Mason didn’t tell Lund to ignore the data. He told him to change the frame he was using to interpret it.
When you’re wearing a “management hat,” the O-ring concern becomes one input among several: the launch window, the Christa McAuliffe teacher-in-space programme and its live television audience, the cost of a scrubbed launch, Morton Thiokol’s contract with NASA. Weighted this way, a technically unquantified risk at a temperature outside the tested range might genuinely look manageable relative to those competing pressures. The reframing feels rational. That’s what makes it dangerous.
There was a second mechanism at work that night. NASA had shifted the burden of proof during the teleconference. The working assumption had been inverted: rather than requiring evidence that the launch was safe, NASA wanted evidence that it was definitely not safe before agreeing to a delay. That is a very different standard, and it was applied in real time, under schedule pressure, to data that was inherently probabilistic.
Security teams live inside this inversion. “Can you prove it will be exploited?” is the management-hat question. The engineering-hat answer is: “We can’t prove it won’t be, and here’s the evidence for why that matters.” These are structurally different conversations, and conflating them is where technical findings get turned into ambiguous risk positions.
The Management Hat Test is a single check: does business context change the physics of the problem? Delivery pressure doesn’t change whether a vulnerability exists. A launch schedule doesn’t change O-ring behaviour at 29°F. If the answer is no, the technical assessment stands. The conversation about whether to accept the risk is a separate, legitimate conversation. But it should stay separate.
The action: When producing a technical risk assessment under deadline pressure, write two explicit documents. First: what the evidence shows, independent of business context. Second: a business impact analysis that addresses what accepting or mitigating the risk would cost. Never merge them into a single “risk recommendation.” The separation means the audit trail shows that someone had the technical evidence and made a conscious choice, not that the evidence was unclear.
Risk acceptance as organisational theatre
Morton Thiokol had a flight readiness review process. There were teleconferences, forms, and sign-offs at multiple levels. The Rogers Commission’s investigation found that this process was followed, and that a flawed decision was made anyway: the substantive technical objection was neutralised through an instruction that appeared in no form and left no paper trail.
The engineers who opposed the launch refused to sign the revised recommendation. Joe Kilminster, a Thiokol manager, signed it instead. The paperwork looked complete.
Most DevSecOps risk acceptance processes have the same structural vulnerability. They capture the risk description, the business justification, the compensating controls, and the sign-off authority. What they typically don’t capture is whether the technical expert who raised the finding revised their assessment after being asked to “consider the business context.” Whether the risk rating changed between the engineer’s initial finding and the submitted form. Whether any alternatives to acceptance were genuinely evaluated, or whether the form was filled in to document a decision already made elsewhere.
A risk acceptance form signed under these conditions isn’t a risk decision. It’s a procedural cover for one. The distinction matters when the thing that was “accepted” eventually materialises.
The Risk Acceptance Framework keeps three things deliberately separate: the technical finding (what the evidence shows, independent of organisational pressure), the risk position (who is accountable, what they are specifically accepting, and under what conditions that acceptance expires), and the escalation path (what happens if the risk materialises, and who gets notified before that point). Most existing processes collapse all three into a single form. That collapse is where the management hat instruction hides.
The action: Add one field to every risk acceptance form in your process: “Did the originating engineer revise their technical assessment after receiving business context?” If yes, require the original assessment as an attachment alongside the revised one. This costs nothing to add. What it does is make the management hat instruction visible in the audit trail rather than absorbed into a final risk rating that no longer reflects the engineering position.
The escalation trap
After the teleconference decision was made, Boisjoly and fellow engineer Arnie Thompson sat in silence. There was nowhere else to go. The organisational structure that existed to protect technical integrity had been the mechanism through which it was overridden. Alan McDonald, Morton Thiokol’s representative in Florida, was surprised when he saw the recommendation to launch. He appealed directly to NASA management not to proceed. He was also overruled.
This is the escalation trap: when the same authority structure that makes technical decisions also controls the escalation path for technical dissent, the dissent has nowhere to go. Boisjoly’s July 1985 memo had entered that same chain. Everything went to the same people.
In cloud and security organisations, the equivalent structure is common and rarely examined. The SRE flags a production risk. The engineering manager escalates to the delivery director. The delivery director is accountable for the release date. The person controlling the escalation path has a direct financial interest in the outcome. Most hierarchies are built exactly this way, and the result is the same: technical dissent that arrives at the person most motivated to resolve it in favour of the schedule.
The CISO-CTO reporting line creates an equivalent structural condition. A CISO who reports to a CTO accountable for delivery velocity is, structurally, in a similar position to a Morton Thiokol engineer in the same management chain as the programme manager asking for a green light. The architecture is the problem. Individual competence and professional integrity can’t overcome a reporting structure that routes the security authority through the delivery management chain. A CISO with a direct line to the board or CEO has a qualitatively different escalation path. Post-Challenger, NASA created exactly this: an independent safety office with direct access to the NASA administrator, outside the programme management chain.
The Escalation Mapping Process addresses this by documenting, before any significant incident or release, where the independent technical escalation path actually sits. Not “raise it with your manager”: if your manager is the person making the contested decision, here is the named individual with genuine authority and independence who should receive the escalation. In regulated industries, that person often exists by mandate: a Chief Risk Officer, an independent safety officer, a board-level audit committee. In most tech companies, the equivalent role doesn’t exist, and nobody has noticed.
The action: Before your next major release or significant infrastructure change, document in your runbook: who is the independent technical authority if the engineering lead and the delivery lead disagree? Name a person. Not a process. Not a committee. One person, with a reporting line outside the delivery chain. If you can’t name someone, that’s your finding.
The three frameworks as a system
The Management Hat Test, the Risk Acceptance Framework, and the Escalation Mapping Process aren’t three separate tools. They work as one system, and each depends on the others.
The Management Hat Test catches the moment of reframing. But catching it doesn’t help without a mechanism to act on it: that’s the Risk Acceptance Framework, which keeps the technical finding separate from the business position in the documented record. And both of those are insufficient without the Escalation Mapping Process: if the local authority overrides the technical finding regardless, there has to be somewhere else for the concern to go before the decision is final.
Boisjoly had none of these. He had correct data, a formal warning in writing, a structured review process, and no mechanism that could stop what happened. The data alone wasn’t enough. The process alone wasn’t enough. The organisational architecture was the constraint, and individual technical competence couldn’t overcome it on the night.
Most DevSecOps teams have some version of these three things. The question is whether the versions they have are designed to preserve independent technical authority, or designed to look like they do while allowing the management hat instruction to operate unobstructed in practice.
Running a tabletop exercise to find out is uncomfortable. It’s also considerably less uncomfortable than discovering the answer under production conditions.
The action: Pick a real risk your team accepted in the last six months. Walk backwards through the process. When was the original technical assessment made? Did the risk rating change between the engineer’s finding and the submitted form? Who had input? Was there an escalation path that didn’t pass through the delivery authority? If you can’t answer those questions from the documentation, the process has gaps. Find them before the physics does.
The pertinent question to ponder over
Boisjoly was right about the O-rings. About the temperature. About the failure mode. His technical judgment wasn’t disputed after the fact, and it wasn’t really disputed on the night either. The organisation found a way to set it aside using an instruction that sounded procedurally reasonable in the moment and left almost no paper trail.
He testified fully to the Rogers Commission. He was sidelined by Morton Thiokol afterwards and took early retirement. He spent the rest of his career lecturing on engineering ethics at universities across North America. He said the night of 27 January 1986 was “the worst night of my life.” He died in 2012.
The question every engineer in a security or reliability role has to sit with is simple: if you had data that clearly showed a system would fail under known conditions, and someone with authority asked you to put on your management hat and reconsider, what would you do?
The harder follow-on: does your organisation have a mechanism that would let you hold your position and have it matter?
Most don’t. Most have processes that look like those mechanisms. Most engineers learn, over time, which fights are worth having, and which aren’t. That learning is the organisation calibrating them to its tolerance for technical dissent. Once you’ve been calibrated, it’s difficult to see it clearly.
Boisjoly never got calibrated. That’s probably why he’s remembered, and why the seven people who died are remembered alongside him.
The Management Hat Test, Risk Acceptance Framework, and Escalation Mapping Process are documented together as an interdependent governance system at blog.ogunlana.net. If you’re auditing your security governance structure or setting one up from scratch, that’s the right starting point.