SaaS Business

Failure Mode and Effects Analysis: Strategic Risk Mitigation and the Anatomy of Operational Failure

Failure Mode and Effects Analysis (FMEA) serves as a systematic, proactive methodology used by organizations to identify how a product, service, or process might fail, analyze the potential consequences, and implement preventive measures before those failures manifest in the real world. By adopting a "worst-case-scenario" mindset, businesses move beyond reactive firefighting, instead utilizing a structured, quantitative approach to minimize operational risk and protect customer interests.

The necessity of such discipline was laid bare during the 2017 British Airways IT catastrophe. On May 27, 2017, a massive power failure at a UK-based data center cascaded through the airline’s infrastructure, triggering a catastrophic disruption that affected tens of thousands of travelers over a chaotic holiday weekend. This incident remains a seminal case study in how a singular, localized failure can paralyze an interconnected global operation, resulting in significant financial loss, profound reputational damage, and an urgent industry-wide reassessment of data center resilience.

Understanding the Mechanics of FMEA

FMEA: How to Prevent the £100m British Airways Catastrophe | Process Street | Compliance Operations Platform

At its core, FMEA is a collaborative exercise that requires input from cross-functional teams, including engineering, operations, and quality management. It is not merely a documentation task but a rigorous evaluation of potential failure modes—the specific ways in which a process or system can deviate from its intended function.

The process follows a logical progression: identifying potential failure points, assessing the severity of the consequences, estimating the frequency (occurrence) of the failure, and determining the likelihood of detecting the issue before it reaches the customer. These three variables—Severity (S), Occurrence (O), and Detection (D)—are traditionally used to calculate a Risk Priority Number (RPN). While the RPN provides a numerical ranking to help teams prioritize their efforts, seasoned quality professionals warn against relying on it exclusively. A low-frequency event with high-severity consequences—such as a data center power failure—must be treated with the same, if not greater, urgency as high-frequency, low-impact issues.

The 2017 British Airways Outage: A Chronology of Failure

The British Airways incident began on the morning of May 27, 2017, when an uninterruptible power supply (UPS) system at the airline’s primary data center experienced a power loss. Reports at the time suggested that the subsequent, uncontrolled restoration of power resulted in a "power surge" that physically damaged critical IT infrastructure.

FMEA: How to Prevent the £100m British Airways Catastrophe | Process Street | Compliance Operations Platform

The immediate result was a total shutdown of the airline’s communication, baggage handling, and check-in systems. Because the airline relied on centralized IT architecture, the failure prevented staff from processing passengers at airports across the globe. Over the span of three days, more than 700 flights were canceled, and approximately 75,000 travelers were stranded.

As the situation unfolded, the lack of a secondary, seamlessly functioning failover system became apparent. The recovery process was hampered by the fact that the primary systems were not just offline, but damaged, requiring extensive manual intervention to restore data integrity. By the time the systems were fully operational, the brand had suffered a significant blow to public trust.

Quantifying the Impact

The financial fallout was substantial. While early estimates circulated in the media suggested a potential impact of £100 million, International Airlines Group (IAG), the parent company of British Airways, later reported a gross cost of approximately £80 million. This figure, while lower than the initial conjecture, represents a massive operational expense, encompassing passenger compensation, hotel accommodations, logistics for rebooking flights, and the extraordinary labor costs associated with the recovery effort.

FMEA: How to Prevent the £100m British Airways Catastrophe | Process Street | Compliance Operations Platform

Beyond the balance sheet, the "soft" costs were equally damaging. The airline faced a wave of public criticism regarding its contingency planning and communication with passengers. As Willie Walsh, the then-chief executive of IAG, acknowledged in the aftermath, the event was a severe test of the brand’s resilience. It highlighted the fragility of modern, highly integrated systems that lack robust, independent, and frequently tested redundancy.

The Role of Process Controls

From an FMEA perspective, the British Airways incident serves as a diagnostic tool to identify where control failures occurred. The critical questions for an FMEA team investigating such an event would include:

  • What were the specific protocols for power maintenance, and were they followed?
  • Were there technical interlocks in place to prevent the "uncontrolled return of power" that reportedly caused the damage?
  • Was the disaster recovery plan tested under conditions that mirrored the actual failure?
  • Was there a "human in the loop" with the authority and expertise to abort the power restoration when it became clear that the hardware was being compromised?

By applying the FMEA framework, organizations can replace vague notions of "human error" with actionable process improvements. Instead of blaming individuals, companies can implement physical or software-based interlocks, automated monitoring systems that trigger alerts before damage occurs, and rigorous, mandatory verification steps for high-risk maintenance activities.

FMEA: How to Prevent the £100m British Airways Catastrophe | Process Street | Compliance Operations Platform

The Strategic Imperative for Continuous Assessment

FMEA is most effective when used as a dynamic, living document rather than a static compliance exercise. It is particularly vital during "moments of change"—such as when a business migrates to a new IT system, alters its operational procedures, or undergoes a significant scale-up. In these periods, the probability of failure increases as the "known unknowns" are replaced by new, untested variables.

Furthermore, the "normalization of deviance"—the gradual process by which unsafe practices become accepted as standard procedure—is a common precursor to major incidents. Regular FMEA reviews act as a circuit breaker for this phenomenon, forcing teams to revisit their assumptions and verify that their preventive controls remain effective in the face of changing technology and operational demands.

Implementing an Effective FMEA Workflow

FMEA: How to Prevent the £100m British Airways Catastrophe | Process Street | Compliance Operations Platform

To successfully deploy FMEA within an organization, leadership must prioritize three pillars:

  1. Inclusive Expertise: The analysis must involve those who actually execute the work. A quality manager working in isolation cannot account for the nuance of a technician’s daily reality.
  2. Evidence-Based Documentation: Every control measure must have an associated owner and verifiable evidence. If a process requires a two-person sign-off, there must be a digital record of that verification.
  3. Integration with Quality Management Systems: FMEA should not exist in a silo. It should be deeply integrated into the company’s broader Quality Management System (QMS), ensuring that when a risk is identified, it automatically triggers a corrective and preventive action (CAPA) workflow.

Modern tools, such as digital process automation platforms, allow teams to bridge the gap between theory and execution. By embedding FMEA directly into the workflow, companies can ensure that risk assessments are not just performed, but acted upon. These platforms enable real-time tracking of corrective actions, providing an audit trail that is essential for both regulatory compliance and internal accountability.

Lessons for Future Resilience

The fundamental lesson from the British Airways incident is that in a highly digitized economy, operational resilience is a competitive advantage. No system is immune to failure, but a resilient organization is one that has planned for failure at every level of its architecture.

FMEA: How to Prevent the £100m British Airways Catastrophe | Process Street | Compliance Operations Platform

By conducting thorough Failure Mode and Effects Analysis, businesses can move from a posture of vulnerability to one of prepared, proactive oversight. The process of documenting potential failures may seem like an exercise for the pessimist, but in practice, it is the most optimistic path toward long-term success. By identifying what could go wrong, organizations gain the power to ensure that it never does, thereby protecting their assets, their employees, and, most importantly, the trust of their customers.

As industries continue to rely on increasingly complex, interconnected systems, the discipline of FMEA will remain an indispensable tool for leaders who prioritize stability and excellence in their operational design. Whether in aviation, manufacturing, or digital services, the mandate remains the same: identify the risk, quantify the impact, and build the controls that turn potential disaster into a manageable, documented, and mitigated outcome.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button
PlanMon
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.