Software Architecture

Building Resilient Distributed Systems: Patterns for Fault Tolerance

Build resilient distributed systems with circuit breakers, retries, and timeouts. Production patterns for handling failures, cascading errors, and maintaining availability.

Khalid Aboubakr
19 min read
Distributed SystemsResilienceFault ToleranceCircuit BreakerChaos EngineeringRetry PatternTimeout

Introduction

In distributed systems, failure is not a possibility—it's a certainty. Networks partition, services crash, databases become unavailable, and latency spikes occur. The question isn't whether failures will happen, but how your system responds when they do.

This article shares patterns and practices for building systems that remain functional even when components fail, drawn from building mission-critical systems for healthcare and government sectors.

The Eight Fallacies of Distributed Computing

Before diving into solutions, acknowledge these realities:

  1. The network is NOT reliable
  2. Latency is NOT zero
  3. Bandwidth is NOT infinite
  4. The network is NOT secure
  5. Topology does NOT remain constant
  6. There is NOT one administrator
  7. Transport cost is NOT zero
  8. The network is NOT homogeneous

Every design decision should account for these truths.

Pattern 1: Circuit Breaker

Prevent cascade failures by stopping requests to failing services.

enum CircuitState { CLOSED, // Normal operation OPEN, // Blocking requests HALF_OPEN // Testing if service recovered } class CircuitBreaker { private state = CircuitState.CLOSED; private failureCount = 0; private lastFailureTime: Date | null = null; private successCount = 0; constructor( private readonly threshold: number = 5, private readonly timeout: number = 30000, private readonly halfOpenSuccessThreshold: number = 3 ) {} async execute<T>(operation: () => Promise<T>): Promise<T> { if (this.state === CircuitState.OPEN) { if (this.shouldAttemptReset()) { this.state = CircuitState.HALF_OPEN; this.successCount = 0; } else { throw new CircuitOpenError('Circuit is open'); } } try { const result = await operation(); this.onSuccess(); return result; } catch (error) { this.onFailure(); throw error; } } private shouldAttemptReset(): boolean { return this.lastFailureTime !== null && Date.now() - this.lastFailureTime.getTime() > this.timeout; } private onSuccess(): void { if (this.state === CircuitState.HALF_OPEN) { this.successCount++; if (this.successCount >= this.halfOpenSuccessThreshold) { this.state = CircuitState.CLOSED; this.failureCount = 0; } } else { this.failureCount = 0; } } private onFailure(): void { this.failureCount++; this.lastFailureTime = new Date(); if (this.failureCount >= this.threshold) { this.state = CircuitState.OPEN; } } } // Usage const paymentCircuit = new CircuitBreaker(5, 30000); async function processPayment(order: Order): Promise<PaymentResult> { try { return await paymentCircuit.execute(() => paymentService.charge(order.customerId, order.total) ); } catch (error) { if (error instanceof CircuitOpenError) { // Fallback: queue for retry await paymentQueue.add(order); return { status: 'pending', message: 'Payment queued for processing' }; } throw error; } }

Pattern 2: Bulkhead

Isolate failures by partitioning resources.

class BulkheadExecutor { private activeRequests = 0; constructor( private readonly maxConcurrent: number, private readonly maxQueue: number, private readonly queueTimeout: number ) {} private queue: Array<{ resolve: (value: void) => void; reject: (error: Error) => void; timeoutId: NodeJS.Timeout; }> = []; async execute<T>(operation: () => Promise<T>): Promise<T> { if (this.activeRequests < this.maxConcurrent) { return this.runOperation(operation); } if (this.queue.length >= this.maxQueue) { throw new BulkheadFullError('Bulkhead queue is full'); } // Wait in queue await this.waitInQueue(); return this.runOperation(operation); } private async runOperation<T>(operation: () => Promise<T>): Promise<T> { this.activeRequests++; try { return await operation(); } finally { this.activeRequests--; this.processQueue(); } } private waitInQueue(): Promise<void> { return new Promise((resolve, reject) => { const timeoutId = setTimeout(() => { const index = this.queue.findIndex(item => item.resolve === resolve); if (index !== -1) { this.queue.splice(index, 1); reject(new BulkheadTimeoutError('Queue wait timeout')); } }, this.queueTimeout); this.queue.push({ resolve, reject, timeoutId }); }); } private processQueue(): void { if (this.queue.length > 0 && this.activeRequests < this.maxConcurrent) { const next = this.queue.shift()!; clearTimeout(next.timeoutId); next.resolve(); } } } // Usage: Separate bulkheads for different services const criticalServiceBulkhead = new BulkheadExecutor(20, 50, 5000); const analyticsServiceBulkhead = new BulkheadExecutor(5, 10, 1000); // Critical payments get more resources await criticalServiceBulkhead.execute(() => paymentService.process(order)); // Analytics can fail without affecting core functionality await analyticsServiceBulkhead.execute(() => analyticsService.track(event)) .catch(err => logger.warn('Analytics failed', err));

Pattern 3: Retry with Exponential Backoff

Handle transient failures gracefully.

interface RetryConfig { maxAttempts: number; baseDelay: number; maxDelay: number; jitterFactor: number; retryableErrors: (error: Error) => boolean; } class RetryPolicy { constructor(private config: RetryConfig) {} async execute<T>(operation: () => Promise<T>): Promise<T> { let lastError: Error; for (let attempt = 1; attempt <= this.config.maxAttempts; attempt++) { try { return await operation(); } catch (error) { lastError = error as Error; if (!this.config.retryableErrors(lastError)) { throw lastError; } if (attempt < this.config.maxAttempts) { const delay = this.calculateDelay(attempt); await this.sleep(delay); } } } throw lastError!; } private calculateDelay(attempt: number): number { // Exponential backoff: baseDelay * 2^(attempt-1) const exponentialDelay = this.config.baseDelay * Math.pow(2, attempt - 1); // Cap at maxDelay const cappedDelay = Math.min(exponentialDelay, this.config.maxDelay); // Add jitter to prevent thundering herd const jitter = cappedDelay * this.config.jitterFactor * Math.random(); return cappedDelay + jitter; } private sleep(ms: number): Promise<void> { return new Promise(resolve => setTimeout(resolve, ms)); } } // Usage const retryPolicy = new RetryPolicy({ maxAttempts: 3, baseDelay: 1000, maxDelay: 10000, jitterFactor: 0.2, retryableErrors: (error) => { // Retry network errors and 5xx responses return error instanceof NetworkError || (error instanceof HttpError && error.status >= 500); } }); const result = await retryPolicy.execute(() => externalApi.fetch(url));

Pattern 4: Timeout

Never wait indefinitely.

class TimeoutPolicy { constructor(private timeoutMs: number) {} async execute<T>(operation: () => Promise<T>): Promise<T> { const timeoutPromise = new Promise<never>((_, reject) => { setTimeout(() => { reject(new TimeoutError(`Operation timed out after ${this.timeoutMs}ms`)); }, this.timeoutMs); }); return Promise.race([operation(), timeoutPromise]); } } // Combine policies for comprehensive resilience class ResilientClient { private circuitBreaker = new CircuitBreaker(5, 30000); private bulkhead = new BulkheadExecutor(10, 20, 5000); private retryPolicy = new RetryPolicy({ maxAttempts: 3, baseDelay: 1000, maxDelay: 5000, jitterFactor: 0.2, retryableErrors: isRetryable }); private timeout = new TimeoutPolicy(10000); async call<T>(operation: () => Promise<T>): Promise<T> { return this.circuitBreaker.execute(() => this.bulkhead.execute(() => this.retryPolicy.execute(() => this.timeout.execute(operation) ) ) ); } }

Pattern 5: Graceful Degradation

Provide reduced functionality when full service is unavailable.

class ProductService { constructor( private primaryCatalog: CatalogService, private searchService: SearchService, private recommendationService: RecommendationService, private cache: Cache ) {} async getProductPage(productId: string): Promise<ProductPageData> { // Core data - required const product = await this.getProductWithFallback(productId); // Enhanced data - optional, with fallbacks const [relatedProducts, recommendations, reviews] = await Promise.allSettled([ this.getRelatedProducts(productId), this.getRecommendations(productId), this.getReviews(productId) ]); return { product, relatedProducts: this.extractOrDefault(relatedProducts, []), recommendations: this.extractOrDefault(recommendations, []), reviews: this.extractOrDefault(reviews, { items: [], summary: null }), degraded: this.isDegraded(relatedProducts, recommendations, reviews) }; } private async getProductWithFallback(productId: string): Promise<Product> { try { const product = await this.primaryCatalog.getProduct(productId); await this.cache.set(`product:${productId}`, product, 3600); return product; } catch (error) { // Try cache as fallback const cached = await this.cache.get<Product>(`product:${productId}`); if (cached) { return { ...cached, fromCache: true }; } throw error; // No fallback available } } private extractOrDefault<T>(result: PromiseSettledResult<T>, defaultValue: T): T { return result.status === 'fulfilled' ? result.value : defaultValue; } private isDegraded(...results: PromiseSettledResult<unknown>[]): boolean { return results.some(r => r.status === 'rejected'); } }

Observability for Resilience

You can't fix what you can't see.

class ResilienceMetrics { private readonly circuitBreakerState = new Gauge({ name: 'circuit_breaker_state', help: 'Current state of circuit breaker (0=closed, 1=half-open, 2=open)', labelNames: ['service'] }); private readonly requestDuration = new Histogram({ name: 'request_duration_seconds', help: 'Request duration in seconds', labelNames: ['service', 'outcome'], buckets: [0.1, 0.5, 1, 2, 5, 10] }); private readonly fallbackUsage = new Counter({ name: 'fallback_usage_total', help: 'Number of times fallback was used', labelNames: ['service', 'fallback_type'] }); recordCircuitState(service: string, state: CircuitState): void { this.circuitBreakerState.set({ service }, state); } recordRequest(service: string, duration: number, outcome: 'success' | 'failure' | 'timeout'): void { this.requestDuration.observe({ service, outcome }, duration); } recordFallback(service: string, type: string): void { this.fallbackUsage.inc({ service, fallback_type: type }); } }

Chaos Engineering

Test resilience in production.

class ChaosMonkey { private enabled = process.env.CHAOS_ENABLED === 'true'; private failureRate = parseFloat(process.env.CHAOS_FAILURE_RATE || '0.01'); async maybeInjectFailure(service: string): Promise<void> { if (!this.enabled) return; if (Math.random() < this.failureRate) { const failureType = this.selectFailureType(); switch (failureType) { case 'latency': await this.injectLatency(1000 + Math.random() * 4000); break; case 'error': throw new Error(`[Chaos] Injected failure for ${service}`); case 'timeout': await this.injectLatency(30000); // Cause timeout break; } } } private selectFailureType(): 'latency' | 'error' | 'timeout' { const rand = Math.random(); if (rand < 0.6) return 'latency'; if (rand < 0.9) return 'error'; return 'timeout'; } private injectLatency(ms: number): Promise<void> { return new Promise(resolve => setTimeout(resolve, ms)); } }

Conclusion

Building resilient distributed systems requires:

  1. Accepting that failures will happen - Design for failure from the start
  2. Implementing multiple layers of defense - Circuit breakers, bulkheads, retries, timeouts
  3. Providing graceful degradation - Some functionality is better than none
  4. Maintaining observability - You can't fix what you can't see
  5. Testing resilience actively - Chaos engineering in production

The goal isn't to prevent all failures—it's to ensure that when failures occur, the impact is minimized and recovery is automatic.

Related Articles

Backend Design17 min read

Queue-Based Architecture for Reliable Processing

Build reliable message queue systems with Redis, RabbitMQ, and AWS SQS. Covers dead letter queues, idempotency, and real-world processing patterns.

Security Engineering18 min read

API Security Hardening: A Practitioner's Guide

Secure your APIs with rate limiting, input validation, and CORS configuration. Production-tested checklist covering authentication, encryption, and error handling.

Backend Design20 min read

Database Design Patterns for Scale

Scale databases with sharding, replication, and partitioning. Covers PostgreSQL, MySQL, and MongoDB scaling patterns with real performance numbers from production systems.