Tekko

Language

Get in Touch →

Usually respond within 24 hours

Back to BlogArchitecture

Serverless WebSockets: Scaling Collaboration with Cloudflare Durable Objects

7 min read
CloudflareWebSocketsServerlessDistributed SystemsTypeScript
Serverless WebSockets: Scaling Collaboration with Cloudflare Durable Objects

Building real-time collaborative applications has historically been one of the most operationally expensive tasks in web development. Whether you are building a shared document editor like Google Docs, a multiplayer game, or a live financial dashboard, the requirements are always the same: low latency, high availability, and, most importantly, a consistent shared state.

In a traditional architecture, this usually involves a fleet of Node.js or Go servers maintaining persistent WebSocket connections, backed by a managed Redis instance for Pub/Sub and state synchronization. While this works, it introduces significant operational overhead. You have to manage connection draining, load balancing with sticky sessions, and the horizontal scaling of your cache layer.

Cloudflare Durable Objects (DO) change this paradigm by bringing the Actor Model to the serverless edge. In this article, we will explore how to implement state-synchronized WebSockets without a single line of Redis code.

The Stateless Constraint of Serverless

Standard serverless functions (like AWS Lambda or standard Cloudflare Workers) are ephemeral. They spin up, handle a request, and shut down. This statelessness is a feature for REST APIs—it allows for infinite horizontal scaling—but it is a bug for WebSockets.

WebSockets require a persistent connection to a specific server instance. If User A and User B are collaborating on the same document but are routed to two different serverless instances, they have no native way to talk to each other. To bridge this gap, developers typically reach for an external state provider (Redis) to act as a mailbox. This adds a network hop, increases latency, and introduces a single point of failure that you must manage and pay for.

Rethinking the Stack: Durable Objects as an Actor Model

Durable Objects solve the state problem by providing a globally unique instance of a class that combines compute and storage. Think of a Durable Object as a tiny, long-lived server that exists for a specific ID (e.g., a room ID or a document ID).

When a request comes in for room-123, Cloudflare ensures that it is routed to the exact same instance of the Durable Object, regardless of where in the world the user is. If the object isn't running, it is instantiated. If it is already running, the existing state is available in memory.

This is essentially the Actor Model. Each Durable Object is an actor that:

  1. Has its own private state.
  2. Communicates with the outside world via messages (HTTP/WebSockets).
  3. Processes messages sequentially, eliminating most race conditions.

The Architecture: Moving Beyond Managed Redis

In a traditional setup, Redis acts as the "glue." In the Durable Object model, the DO is the glue.

Traditional WebSocket Architecture

  1. Client connects to a Load Balancer.
  2. Load Balancer routes to an available Server Instance.
  3. Server Instance subscribes to a Redis channel for that specific room.
  4. When a message arrives, the server updates Redis and publishes to the channel.
  5. All other server instances receive the Redis message and push to their local connected clients.

Durable Object Architecture

  1. Client connects to a Cloudflare Worker.
  2. Worker identifies the room ID and fetches the corresponding Durable Object.
  3. Worker upgrades the connection to a WebSocket and hands it off to the Durable Object.
  4. The Durable Object maintains a list of active WebSockets in its memory and persists state to its own local disk.

This removes the need for an external Pub/Sub system entirely. Because all users for a specific "room" are connected to the same Durable Object instance, the DO can simply iterate over its internal list of connections to broadcast updates.

Practical Implementation: A Collaborative State Machine

Let's look at how we implement this. We will use TypeScript and the Cloudflare Workers wrangler environment.

1. The Worker Entry Point

The Worker's job is to act as a router. It extracts the room ID from the URL and passes the request to the Durable Object namespace.

export default { async fetch(request, env) { const url = new URL(request.url); const roomId = url.searchParams.get("room"); if (!roomId) return new Response("Missing room ID", { status: 400 }); // Get the Durable Object ID for this room name const id = env.CHAT_ROOM.idFromName(roomId); // Get the stub (the handle to the DO) const roomObject = env.CHAT_ROOM.get(id); // Forward the request to the Durable Object return roomObject.fetch(request); } };

2. The Durable Object Class

This is where the magic happens. The DO handles the WebSocket upgrade and manages the state.

export class ChatRoom { state: DurableObjectState; sessions: Set<WebSocket>; constructor(state: DurableObjectState) { this.state = state; // Keep track of active connections in memory this.sessions = new Set(); } async fetch(request: Request) { const upgradeHeader = request.headers.get("Upgrade"); if (!upgradeHeader || upgradeHeader !== "websocket") { return new Response("Expected Upgrade: websocket", { status: 426 }); } // Create a WebSocket pair (client and server) const [client, server] = Object.values(new WebSocketPair()); // Accept the connection await this.handleSession(server); return new Response(null, { status: 101, webSocket: client }); } async handleSession(ws: WebSocket) { ws.accept(); this.sessions.add(ws); ws.addEventListener("message", async (msg) => { try { // Broadcast the message to all other connected clients this.broadcast(msg.data); // Optional: Persist state to DO's transactional storage await this.state.storage.put("lastMessage", msg.data); } catch (err) { ws.close(1011, "Registration failed"); } }); ws.addEventListener("close", () => { this.sessions.delete(ws); }); } broadcast(message: string | ArrayBuffer) { for (const session of this.sessions) { try { session.send(message); } catch (err) { this.sessions.delete(session); } } } }

Optimizing for Scale with the Hibernation API

The example above keeps the Durable Object "awake" as long as a WebSocket is connected. While this is simple, it can be inefficient if you have thousands of rooms with idle connections.

Cloudflare introduced the WebSocket Hibernation API to solve this. With Hibernation, Cloudflare can serialize the state of the Durable Object and shut down the compute while keeping the WebSocket connection alive on the network edge. When a message actually arrives, the DO is "woken up," the message is processed, and the DO can go back to sleep.

To use this, you use this.state.acceptWebSocket(ws) instead of the standard ws.accept(). This allows the DO to be evicted from memory while idle, significantly reducing costs for applications with many low-activity connections.

Operational Trade-offs and Considerations

While Durable Objects simplify the stack, they are not a silver bullet. As a senior engineer, you must consider the following trade-offs:

1. Global Locality

By default, a Durable Object is created in the region closest to the first user who requests it. If you have a global team, users far from that region may experience higher latency. Cloudflare offers "Jurisdictional Restrictions" and is working on better migration tools, but for now, you should be aware that a DO lives in one place at a time.

2. Single-Threaded Nature

A single Durable Object instance is single-threaded. This is a massive advantage for data consistency (no locks needed!), but it means a single room cannot scale infinitely. If you expect 50,000 users in a single chat room, a single DO will become a bottleneck. In that scenario, you would need to implement a tree-based broadcast structure (multiple DOs acting as leaf nodes).

3. Vendor Lock-in

Using Durable Objects ties you closely to the Cloudflare ecosystem. Unlike a Dockerized Node.js app that can run on AWS, GCP, or Azure, DO code is specific to the Workers runtime. For many, the reduction in operational complexity outweighs this risk, but it must be an intentional decision.

Conclusion

The shift from managed Redis and persistent server clusters to Cloudflare Durable Objects represents a significant leap in how we build collaborative software. By leveraging the Actor Model at the edge, we can achieve:

  • Zero-Infrastructure State: No more managing Redis clusters or Pub/Sub logic.
  • Simplified Concurrency: Single-threaded execution per object removes complex locking mechanisms.
  • Lower Latency: Compute and state live closer to the user, with fewer network hops.

Actionable Next Steps:

  1. Audit your current real-time stack: Identify the cost and latency overhead of your Redis/Pub/Sub layer.
  2. Prototype a 'Room' logic: Use the wrangler CLI to deploy a basic Durable Object and test the WebSocket connection stability.
  3. Explore Hibernation: If your app has many idle users, implement the Hibernation API early to optimize for long-term cost efficiency.