Counting Copies, Ignoring Causes: The False Confidence of Numerical Redundancy
Photo: server redundancy data center backup failover infrastructure resilience, via www.servers.com
The Comfort of Multiplication
Redundancy is one of the oldest principles in systems engineering. Duplicate the critical component, ensure the backup can assume load when the primary fails, and the system survives individual failures without interruption. The principle is sound. The problem lies in how it is applied — specifically, in the tendency to satisfy the requirement for redundancy by multiplying instances of a component while leaving the underlying conditions that would cause those instances to fail simultaneously completely unaddressed.
This pattern is sufficiently common, and its consequences sufficiently predictable, that it deserves a name. Call it redundancy theater: the organizational and architectural practice of creating the appearance of resilience through numerical duplication while the actual failure modes remain structurally intact. The result is infrastructure that passes resilience reviews, satisfies compliance checklists, and generates confident diagrams in runbook documentation — right up until the moment a shared dependency fails and takes every redundant instance down in sequence.
Shared Databases and the Single Point That Hides in Plain Sight
Consider a common architectural pattern: a stateless application tier running three replicas distributed across availability zones, fronted by a load balancer with health checks, with automated failover configured to redistribute traffic within thirty seconds of a replica failure. On paper, this is a resilient architecture. In practice, if all three replicas connect to a single relational database instance — or even a primary-replica database pair with a shared failover mechanism — the application tier's redundancy is largely decorative.
A database failure, a connection pool exhaustion event, or a schema migration that introduces a locking condition does not care how many application replicas are running. All three replicas will fail simultaneously, and they will do so in a way that the load balancer's health checks may not immediately detect, because the replicas themselves may still be responsive to HTTP probes even as they return errors to actual application requests. The redundancy that was designed to protect against application-tier failures provides no protection against the failure mode that is actually most likely to occur.
This is not a novel observation. Database high availability is a well-understood engineering problem with well-understood solutions. What makes it a persistent source of production incidents is not ignorance of the solutions but the organizational tendency to treat the application tier's redundancy as a proxy for system-wide resilience, particularly when architectural reviews are focused on component count rather than failure mode analysis.
Centralized Identity as a Universal Blast Radius
Identity providers occupy a unique position in modern infrastructure: they are the systems that everything else must reach before it can do anything. When an identity provider experiences degraded availability — whether due to network partition, certificate expiration, configuration error, or capacity constraint — the failure propagates immediately and simultaneously to every service that depends on it for authentication or authorization.
The redundancy that most organizations have implemented for their identity infrastructure is real, in the sense that it protects against individual node failures. It is frequently insufficient against the failure modes that actually threaten availability: a misconfigured access policy that denies all requests, a token signing key rotation that is not coordinated with dependent services, or a network configuration change that makes the identity provider unreachable from a specific segment of the infrastructure.
In each of these cases, running a three-node identity provider cluster rather than a single node provides no additional protection. The failure is not a node failure — it is a configuration failure, a coordination failure, or a network topology failure that affects the cluster as a whole. Numerical redundancy, applied to a component whose failure modes are systemic rather than individual, does not produce resilience. It produces a more expensive system that fails in the same ways.
Configuration Errors and the Blast Radius of Uniformity
One of the underappreciated risks of infrastructure automation is that it enables configuration errors to be applied uniformly and instantly across every redundant instance of a system. Before infrastructure-as-code became standard practice, a misconfigured firewall rule or an incorrect environment variable might affect one server while others remained functional, because manual configuration processes introduced natural variation. Today, a single erroneous Terraform change or a flawed Ansible playbook can apply the same misconfiguration to every instance in a cluster simultaneously.
This is not an argument against infrastructure automation — the benefits of automation in terms of consistency, auditability, and operational efficiency are substantial and well-established. It is an argument for recognizing that automation changes the character of configuration-related failures. The failure mode shifts from partial degradation, where some instances are affected and others continue serving traffic, to total outage, where all instances receive the same broken configuration at the same moment.
Redundancy designed to protect against hardware failures and individual node crashes does not address this failure mode. The instances are redundant. The configuration change is not.
Designing for Diversity, Not Just Duplication
True resilience requires that redundant components be capable of failing independently — which means they must differ in the dimensions along which failures actually propagate. Geographic distribution addresses hardware and network failures that are physically localized. Architectural diversity addresses the systemic failures that affect all instances sharing a common dependency or configuration baseline.
In practice, architectural diversity is harder to achieve than geographic distribution. It requires deliberate decisions to introduce variation into systems that are otherwise optimized for uniformity. Separate database instances for separate service tiers. Independent certificate authorities for different trust domains. Deployment pipelines that apply configuration changes incrementally rather than universally, with validation gates between each stage. These patterns are more expensive to design and operate than simple instance multiplication, and they resist the simplifying narratives that infrastructure diagrams typically want to tell.
But they address the failure modes that numerical redundancy cannot. A database failure that takes down one service tier does not cascade to others. A certificate rotation error that affects one trust domain does not invalidate authentication across the entire system. A misconfigured deployment that reaches the first ten percent of instances is caught before it reaches the rest.
Resilience as an Architectural Property, Not a Count
At S4Core, we hold that resilience is an architectural property that must be designed in — not a metric that can be satisfied by counting replicas. The question that should drive resilience architecture is not how many instances are running but which failure modes affect all of them simultaneously. Answering that question honestly, and then designing to address the shared vulnerabilities it reveals, is the work that actually produces systems capable of surviving the failures that matter. Everything else is theater.