Building Resilient Distributed Systems: Patterns for Fault Tolerance
Build resilient distributed systems with circuit breakers, retries, and timeouts. Production patterns for handling failures, cascading errors, and maintaining availability.
Introduction
In distributed systems, failure is not a possibility—it's a certainty. Networks partition, services crash, databases become unavailable, and latency spikes occur. The question isn't whether failures will happen, but how your system responds when they do.
This article shares patterns and practices for building systems that remain functional even when components fail, drawn from building mission-critical systems for healthcare and government sectors.
The Eight Fallacies of Distributed Computing
Before diving into solutions, acknowledge these realities:
- The network is NOT reliable
- Latency is NOT zero
- Bandwidth is NOT infinite
- The network is NOT secure
- Topology does NOT remain constant
- There is NOT one administrator
- Transport cost is NOT zero
- The network is NOT homogeneous
Every design decision should account for these truths.
Pattern 1: Circuit Breaker
Prevent cascade failures by stopping requests to failing services.
enum CircuitState { CLOSED, // Normal operation OPEN, // Blocking requests HALF_OPEN // Testing if service recovered } class CircuitBreaker { private state = CircuitState.CLOSED; private failureCount = 0; private lastFailureTime: Date | null = null; private successCount = 0; constructor( private readonly threshold: number = 5, private readonly timeout: number = 30000, private readonly halfOpenSuccessThreshold: number = 3 ) {} async execute<T>(operation: () => Promise<T>): Promise<T> { if (this.state === CircuitState.OPEN) { if (this.shouldAttemptReset()) { this.state = CircuitState.HALF_OPEN; this.successCount = 0; } else { throw new CircuitOpenError('Circuit is open'); } } try { const result = await operation(); this.onSuccess(); return result; } catch (error) { this.onFailure(); throw error; } } private shouldAttemptReset(): boolean { return this.lastFailureTime !== null && Date.now() - this.lastFailureTime.getTime() > this.timeout; } private onSuccess(): void { if (this.state === CircuitState.HALF_OPEN) { this.successCount++; if (this.successCount >= this.halfOpenSuccessThreshold) { this.state = CircuitState.CLOSED; this.failureCount = 0; } } else { this.failureCount = 0; } } private onFailure(): void { this.failureCount++; this.lastFailureTime = new Date(); if (this.failureCount >= this.threshold) { this.state = CircuitState.OPEN; } } } // Usage const paymentCircuit = new CircuitBreaker(5, 30000); async function processPayment(order: Order): Promise<PaymentResult> { try { return await paymentCircuit.execute(() => paymentService.charge(order.customerId, order.total) ); } catch (error) { if (error instanceof CircuitOpenError) { // Fallback: queue for retry await paymentQueue.add(order); return { status: 'pending', message: 'Payment queued for processing' }; } throw error; } }
Pattern 2: Bulkhead
Isolate failures by partitioning resources.
class BulkheadExecutor { private activeRequests = 0; constructor( private readonly maxConcurrent: number, private readonly maxQueue: number, private readonly queueTimeout: number ) {} private queue: Array<{ resolve: (value: void) => void; reject: (error: Error) => void; timeoutId: NodeJS.Timeout; }> = []; async execute<T>(operation: () => Promise<T>): Promise<T> { if (this.activeRequests < this.maxConcurrent) { return this.runOperation(operation); } if (this.queue.length >= this.maxQueue) { throw new BulkheadFullError('Bulkhead queue is full'); } // Wait in queue await this.waitInQueue(); return this.runOperation(operation); } private async runOperation<T>(operation: () => Promise<T>): Promise<T> { this.activeRequests++; try { return await operation(); } finally { this.activeRequests--; this.processQueue(); } } private waitInQueue(): Promise<void> { return new Promise((resolve, reject) => { const timeoutId = setTimeout(() => { const index = this.queue.findIndex(item => item.resolve === resolve); if (index !== -1) { this.queue.splice(index, 1); reject(new BulkheadTimeoutError('Queue wait timeout')); } }, this.queueTimeout); this.queue.push({ resolve, reject, timeoutId }); }); } private processQueue(): void { if (this.queue.length > 0 && this.activeRequests < this.maxConcurrent) { const next = this.queue.shift()!; clearTimeout(next.timeoutId); next.resolve(); } } } // Usage: Separate bulkheads for different services const criticalServiceBulkhead = new BulkheadExecutor(20, 50, 5000); const analyticsServiceBulkhead = new BulkheadExecutor(5, 10, 1000); // Critical payments get more resources await criticalServiceBulkhead.execute(() => paymentService.process(order)); // Analytics can fail without affecting core functionality await analyticsServiceBulkhead.execute(() => analyticsService.track(event)) .catch(err => logger.warn('Analytics failed', err));
Pattern 3: Retry with Exponential Backoff
Handle transient failures gracefully.
interface RetryConfig { maxAttempts: number; baseDelay: number; maxDelay: number; jitterFactor: number; retryableErrors: (error: Error) => boolean; } class RetryPolicy { constructor(private config: RetryConfig) {} async execute<T>(operation: () => Promise<T>): Promise<T> { let lastError: Error; for (let attempt = 1; attempt <= this.config.maxAttempts; attempt++) { try { return await operation(); } catch (error) { lastError = error as Error; if (!this.config.retryableErrors(lastError)) { throw lastError; } if (attempt < this.config.maxAttempts) { const delay = this.calculateDelay(attempt); await this.sleep(delay); } } } throw lastError!; } private calculateDelay(attempt: number): number { // Exponential backoff: baseDelay * 2^(attempt-1) const exponentialDelay = this.config.baseDelay * Math.pow(2, attempt - 1); // Cap at maxDelay const cappedDelay = Math.min(exponentialDelay, this.config.maxDelay); // Add jitter to prevent thundering herd const jitter = cappedDelay * this.config.jitterFactor * Math.random(); return cappedDelay + jitter; } private sleep(ms: number): Promise<void> { return new Promise(resolve => setTimeout(resolve, ms)); } } // Usage const retryPolicy = new RetryPolicy({ maxAttempts: 3, baseDelay: 1000, maxDelay: 10000, jitterFactor: 0.2, retryableErrors: (error) => { // Retry network errors and 5xx responses return error instanceof NetworkError || (error instanceof HttpError && error.status >= 500); } }); const result = await retryPolicy.execute(() => externalApi.fetch(url));
Pattern 4: Timeout
Never wait indefinitely.
class TimeoutPolicy { constructor(private timeoutMs: number) {} async execute<T>(operation: () => Promise<T>): Promise<T> { const timeoutPromise = new Promise<never>((_, reject) => { setTimeout(() => { reject(new TimeoutError(`Operation timed out after ${this.timeoutMs}ms`)); }, this.timeoutMs); }); return Promise.race([operation(), timeoutPromise]); } } // Combine policies for comprehensive resilience class ResilientClient { private circuitBreaker = new CircuitBreaker(5, 30000); private bulkhead = new BulkheadExecutor(10, 20, 5000); private retryPolicy = new RetryPolicy({ maxAttempts: 3, baseDelay: 1000, maxDelay: 5000, jitterFactor: 0.2, retryableErrors: isRetryable }); private timeout = new TimeoutPolicy(10000); async call<T>(operation: () => Promise<T>): Promise<T> { return this.circuitBreaker.execute(() => this.bulkhead.execute(() => this.retryPolicy.execute(() => this.timeout.execute(operation) ) ) ); } }
Pattern 5: Graceful Degradation
Provide reduced functionality when full service is unavailable.
class ProductService { constructor( private primaryCatalog: CatalogService, private searchService: SearchService, private recommendationService: RecommendationService, private cache: Cache ) {} async getProductPage(productId: string): Promise<ProductPageData> { // Core data - required const product = await this.getProductWithFallback(productId); // Enhanced data - optional, with fallbacks const [relatedProducts, recommendations, reviews] = await Promise.allSettled([ this.getRelatedProducts(productId), this.getRecommendations(productId), this.getReviews(productId) ]); return { product, relatedProducts: this.extractOrDefault(relatedProducts, []), recommendations: this.extractOrDefault(recommendations, []), reviews: this.extractOrDefault(reviews, { items: [], summary: null }), degraded: this.isDegraded(relatedProducts, recommendations, reviews) }; } private async getProductWithFallback(productId: string): Promise<Product> { try { const product = await this.primaryCatalog.getProduct(productId); await this.cache.set(`product:${productId}`, product, 3600); return product; } catch (error) { // Try cache as fallback const cached = await this.cache.get<Product>(`product:${productId}`); if (cached) { return { ...cached, fromCache: true }; } throw error; // No fallback available } } private extractOrDefault<T>(result: PromiseSettledResult<T>, defaultValue: T): T { return result.status === 'fulfilled' ? result.value : defaultValue; } private isDegraded(...results: PromiseSettledResult<unknown>[]): boolean { return results.some(r => r.status === 'rejected'); } }
Observability for Resilience
You can't fix what you can't see.
class ResilienceMetrics { private readonly circuitBreakerState = new Gauge({ name: 'circuit_breaker_state', help: 'Current state of circuit breaker (0=closed, 1=half-open, 2=open)', labelNames: ['service'] }); private readonly requestDuration = new Histogram({ name: 'request_duration_seconds', help: 'Request duration in seconds', labelNames: ['service', 'outcome'], buckets: [0.1, 0.5, 1, 2, 5, 10] }); private readonly fallbackUsage = new Counter({ name: 'fallback_usage_total', help: 'Number of times fallback was used', labelNames: ['service', 'fallback_type'] }); recordCircuitState(service: string, state: CircuitState): void { this.circuitBreakerState.set({ service }, state); } recordRequest(service: string, duration: number, outcome: 'success' | 'failure' | 'timeout'): void { this.requestDuration.observe({ service, outcome }, duration); } recordFallback(service: string, type: string): void { this.fallbackUsage.inc({ service, fallback_type: type }); } }
Chaos Engineering
Test resilience in production.
class ChaosMonkey { private enabled = process.env.CHAOS_ENABLED === 'true'; private failureRate = parseFloat(process.env.CHAOS_FAILURE_RATE || '0.01'); async maybeInjectFailure(service: string): Promise<void> { if (!this.enabled) return; if (Math.random() < this.failureRate) { const failureType = this.selectFailureType(); switch (failureType) { case 'latency': await this.injectLatency(1000 + Math.random() * 4000); break; case 'error': throw new Error(`[Chaos] Injected failure for ${service}`); case 'timeout': await this.injectLatency(30000); // Cause timeout break; } } } private selectFailureType(): 'latency' | 'error' | 'timeout' { const rand = Math.random(); if (rand < 0.6) return 'latency'; if (rand < 0.9) return 'error'; return 'timeout'; } private injectLatency(ms: number): Promise<void> { return new Promise(resolve => setTimeout(resolve, ms)); } }
Conclusion
Building resilient distributed systems requires:
- Accepting that failures will happen - Design for failure from the start
- Implementing multiple layers of defense - Circuit breakers, bulkheads, retries, timeouts
- Providing graceful degradation - Some functionality is better than none
- Maintaining observability - You can't fix what you can't see
- Testing resilience actively - Chaos engineering in production
The goal isn't to prevent all failures—it's to ensure that when failures occur, the impact is minimized and recovery is automatic.
Related Articles
Software Architecture18 min read
Event-Driven Architecture in Enterprise Systems: Patterns and Trade-offs
A practitioner's guide to implementing event-driven architecture at scale. Covers message broker selection, event schema design, eventual consistency patterns, and lessons from production systems.
Software Architecture16 min read
Microservices vs Monolith: A Practitioner's Decision Framework
Microservices vs monolith comparison with real decision criteria. Learn when microservices add value vs unnecessary complexity, with examples from production systems.
Backend Design17 min read
Queue-Based Architecture for Reliable Processing
Build reliable message queue systems with Redis, RabbitMQ, and AWS SQS. Covers dead letter queues, idempotency, and real-world processing patterns.
Security Engineering18 min read
API Security Hardening: A Practitioner's Guide
Secure your APIs with rate limiting, input validation, and CORS configuration. Production-tested checklist covering authentication, encryption, and error handling.
Backend Design20 min read
Database Design Patterns for Scale
Scale databases with sharding, replication, and partitioning. Covers PostgreSQL, MySQL, and MongoDB scaling patterns with real performance numbers from production systems.