My “Borrowed Architectures” series exists because good architectural solutions keep showing up as the same solution in different clothes. Getting good at noticing that, spotting patterns and carrying them over to your thinking process, is one of the best ways I know to grow as an engineer. So far, we've looked through some other disciplines: pace layers from urban planning, about drawing software boundaries by how fast each part of a system changes, and building codes, which treat a cascading failure as a design input rather than a postmortem. This time the other discipline is the other half of your own stack. Frontend and backend engineers sit next to each other and run frameworks that reinvented each other's ideas, often without noticing. A frontend engineer writes a Redux middleware to log every action. A backend engineer writes a FastAPI middleware to time every request. Both treat it as a feature of their own framework. Both have written the same function: one that takes whatever is passing through, does its job, and calls next to pass it on. When the same shape shows up across languages, frameworks, and both sides of the stack, the pattern is what's worth learning. Each framework only differs in the details. By the end of this article, you should be able to explain why middleware is so popular, and spot where it belongs in your own code. The article covers four things: The same function on both sides of the stack. Koa, FastAPI and Redux side by side, and the twelve lines of code they all share. The problem that made every framework invent it, and why copying code around or using a base class doesn't solve it. Where else it shows up in a codebase: the data layer, UI routing, test setup and LLM tool calls, plus the signs that tell you to skip it. How to build one: a booking example where the client and the server share one pipeline, and three rules for writing your own. 🧅 What middleware actually is Koa's own front page opens with a server-side middleware that times each request. Trimmed, it looks like this: // Koa (Node.js), on the server app.use(async (ctx, next) => { const start = Date.now(); await next(); ctx.set('X-Response-Time', `${Date.now() - start}ms`); }); FastAPI's middleware docs use the same job to introduce the idea in Python: # FastAPI (Python), on the server @app.middleware("http") async def add_process_time_header(request: Request, call_next): start_time = time.perf_counter() response = await call_next(request) process_time = time.perf_counter() - start_time response.headers["X-Process-Time"] = str(process_time) return response In the browser, Redux documents the same shape: // Redux, in the browser const timing = store => next => action => { const start = performance.now(); const result = next(action); console.log(action.type, `${(performance.now() - start).toFixed(1)}ms`); return result; }; All three take a thing and a next. They do some work, hand control down the chain, wait for it to come back, and then do something with what came back. Without the framework around it, a middleware is a function of two arguments: a context and a continuation. type Middleware = (ctx: C, next: (ctx: C) => R) => R; A pipeline is a few of these wrapped around one handler, each one inside the previous: mw1(ctx, () => mw2(ctx, () => handler(ctx))). When mw1 calls next, it runs mw2. When mw2 calls next, it runs the handler. The engine that builds this nesting from a list is twelve lines. export function compose(stack: Middleware[]) { return (ctx: C, terminal: (ctx: C) => R): R => { let last = -1; const dispatch = (i: number, c: C): R => { if (i dispatch(i + 1, nextCtx)); }; return dispatch(0, ctx); }; } Those twelve lines are a data structure plus an algorithm. The data structure is a plain list of layers. Express's router even keeps its middleware in an array called stack. The algorithm is a recursive walk over that list: dispatch(i) runs layer i, and when the layer calls next(), it recurses into dispatch(i + 1). Whatever a layer does before next() runs on the way down, in list order. Whatever it does after next() runs on the way back up, in reverse, as the recursion unwinds. If you've written a tree traversal, that's pre-order and post-order work on a list, and it's why Django's docs say the response passes through every layer "in reverse order on the way back out." Real frameworks do exactly this. Redux's applyMiddleware hands its middleware to a compose that comes down to one line: funcs.reduce((a, b) => (...args) => a(b(...args))). That line wraps each function around the next one, which is the same nesting you just saw. koa-compose gets there with an index and a promise, much like the dispatch above. 🧬 Where you've already used it Frontend. The browser runs every click down through the element tree and back up again. A listener on any element along the way may act on it. It has been in the DOM Level 2 Events spec since 2000. Angular's HTTP interceptors take the request and a next: HttpHandlerFn. Redux uses store => next => action. Backend. Express says it plainly: an Express application "is essentially a series of middleware function calls executed during the request-response cycle," per its docs. Koa is the same team's rebuild, made so that await next() hands control back to you. A Django middleware wraps get_response. FastAPI hands each function a call_next. In TypeScript, tRPC's middleware calls opts.next() and can pass a typed context down with it, so the procedure at the end knows ctx.user exists. ASP.NET Core does it in C#, with app.Use(async (context, next) => ...). When that many teams land on the same design independently, something in the problem keeps pushing them there. 🧩 The problem it solves All of these frameworks ran into the same problem. Every request goes through one place: the HTTP handler, Redux's dispatch, the gRPC call. And every request needs the same jobs done: auth, logging, rate limits, retries, tracing, error mapping. The auth check is a good example, because no single endpoint should own it and every endpoint needs it to happen. There are two obvious places to put that code, and both go wrong. Copy it into every handler. You paste the auth check into each endpoint. With four endpoints that's fine. With forty, someone eventually forgets to update one of them. Put it all in one shared place. That's usually a BaseHandler class that every handler extends, or a set of features built into the framework. It works well until you need something the author didn't think of. With a base class, you get one fixed order of steps, and changing it means editing a class that every handler depends on. With built-in features, you wait for the framework's maintainers to add yours. Middleware fixes this by letting anyone add a job. Each job is a small function. The pipeline is a plain list, so you can read the order in one place. Adding a job means adding one item to that list without touching anything else. That's also why each of these frameworks ended up with a big library of ready-made middleware. Their docs say so directly. Django calls middleware "a light, low-level 'plugin' system for globally altering Django's input or output." Redux calls it "a third-party extension point between dispatching an action, and the moment it reaches the reducer." A plugin system works when plugins are easy to write, and a function that takes (ctx, next) is about as easy as it gets. So other people wrote the logging, the auth, the compression and the CORS handling, and teams started choosing frameworks partly for the middleware that already existed. 🔍 Where else to look for it, and when to skip it The frameworks solved this for HTTP requests, but the same thing happens anywhere many calls go through one function. Four places are worth checking first: The same shape in four places that aren't request handlers. Each pillar is a job that belongs to every call. The data layer. Every query goes through one database client. Look for a base repository class that scopes queries to the current tenant, skips soft-deleted rows, logs the slow ones and sends reads to a replica, with one more if for each new rule. To see the same jobs written as a pipeline, read Prisma's middleware docs, where it "runs your code before and after every query, so one policy can cover your whole app." Route changes in the UI. Every navigation goes through one router. The sign here is a root layout or a ProtectedRoute wrapper that sends guests to the login page, checks roles and feature flags, and records a page view. React Router's middleware is the closest match to the code above: a function that takes the request and a next, and runs code before and after await next(). Vue Router's navigation guards teach something about API design too. Guards used to take a next argument, until it proved "a common source of mistakes," and now a guard stops navigation by returning false or a new route. Test setup. Look for a long beforeEach that every test runs through. It opens a transaction, freezes the clock, seeds a tenant and stubs the network. A matching afterEach then has to undo all of it in the right order by hand. A pytest yield fixture shows the cleaner version, and it's middleware with next() spelled yield: code before it runs on the way in, code after it runs on the way out, and pytest tears the fixtures down "in the reverse order." That's the same walk down the list and back up again. The wrapper around your LLM calls. You'll usually find one dispatcher function that every model or tool call passes through. It's full of ifs that log the call, redact secrets from the arguments, enforce a timeout and a token budget, and refuse dangerous tools in production. Most of the ones I've seen were written by people who haven't met the pattern under this name. For a version with the layers pulled apart, look at LangChain's wrap_tool_call, a hook that gets the request and a handler to call. Knowing when to skip it matters just as much. The pattern fits jobs that belong to every call, so these are the signs that you need something else: If step three decides whether step four or step five runs, you have a workflow or a state machine. A pipeline has one order, fixed when you build the list. If one step needs the result of the step before it, the two are really one handler split in half. Passing that result through the shared context hides the link between them, and it breaks the day someone reorders the list. If a job only applies to one request type, it belongs in that handler. If the list is just "log it and time it," write a wrapper function. A pipeline pays off when someone else needs to add the third thing. The mistake I see most often happens inside a middleware framework itself: a booking endpoint written as app.post('/bookings', validateSlot, holdSlot, chargeDeposit, sendConfirmation). It looks like a pipeline, but each step needs the one before it. A failed charge also has to release the hold. That's a workflow squeezed into a middleware chain. Once it spans minutes or days, like waiting for a payment webhook or a manager's approval, it belongs in a workflow engine such as Temporal, which retries failed steps and lets you run compensations when a step fails for good. Temporal still wraps every one of those steps in interceptors, each one an execute(input, next). A pipeline wraps one call. A workflow is several calls that depend on each other, spread over time. 🧪 Example: one booking feature, client and server Let's try this on a real feature and build both the server and the client with the same twelve-line compose. The feature is one step of that booking flow, the slot hold: the client asks to hold 10:00, and the server holds it or reports a conflict. The whole flow of hold, charge and confirm is still a workflow. The hold on its own is a single call, so a pipeline fits it. Server side. Every call to POST /holds goes through one function, and four jobs keep landing on it that aren't the hold's job: a request id, auth, idempotent replay and a per-tenant rate limit. type Ctx = { req: Request; state: { requestId?: string; userId?: string; tenantId?: string }; // the shared state, typed }; type Handler = Middleware>; // the response is a return value const requestId: Handler = async (ctx, next) => { ctx.state.requestId = ctx.req.headers.get('x-request-id') ?? crypto.randomUUID(); const res = await next(ctx); res.headers.set('x-request-id', ctx.state.requestId); // needs the way back return res; }; const auth: Handler = async (ctx, next) => { const user = await verify(ctx.req.headers.get('authorization')); if (!user) return new Response('unauthorized', { status: 401 }); // the stop is a value ctx.state.userId = user.id; ctx.state.tenantId = user.tenantId; return next(ctx); }; const idempotency: Handler = async (ctx, next) => { const key = ctx.req.headers.get('idempotency-key'); if (!key) return next(ctx); const scoped = `${ctx.state.userId}:${key}`; const hit = cache.get(scoped); if (hit) return hit.clone(); // stop on replay const res = await next(ctx); // needs the way back, to cache the response if (res.status { if (!(await bucket.take(ctx.state.tenantId!))) return new Response('slow down', { status: 429 }); return next(ctx); }; const holds = compose>([requestId, auth, idempotency, rateLimit]); export const POST = (req: Request) => holds({ req, state: {} }, createHold); Client side. Every call from the app goes through one fetch wrapper, and the jobs piling onto it aren't any screen's job either: a trace id, the auth header, an idempotency key for writes, and turning a 409 into a typed error. That's a pipeline too, with the same Request → Response shape as the server. withHeader copies the request with one more header, so each layer hands a new value down instead of mutating a shared one. type Call = { req: Request }; type Fetcher = Middleware>; const traceId: Fetcher = (call, next) => next({ req: withHeader(call.req, 'x-request-id', crypto.randomUUID()) }); const bearer: Fetcher = (call, next) => next({ req: withHeader(call.req, 'authorization', `Bearer ${session.token}`) }); const idempotencyKey: Fetcher = (call, next) => call.req.method === 'POST' ? next({ req: withHeader(call.req, 'idempotency-key', crypto.randomUUID()) }) : next(call); const conflicts: Fetcher = async (call, next) => { const res = await next(call); // needs the way back if (res.status === 409) throw new ConflictError(await res.text()); return res; }; const client = compose>([traceId, bearer, idempotencyKey, conflicts]); export const apiFetch = (url: string, init?: RequestInit) => client({ req: new Request(url, init) }, retrying(({ req }) => fetch(req))); retrying wraps the network call itself: it re-sends a clone of the request, up to three times, on a 502, 503 or 504. Here's where the two halves meet. Say the server holds 10:00, but its 201 gets lost on the way back and the client sees a 502 instead. The retry re-sends the same request with the same idempotency key. On the server, the idempotency layer replays the first 201 instead of holding the slot twice. Take the idempotencyKey layer out and the same retry comes back as a ConflictError: the user is told the slot is taken, by themselves. Two things stayed out of both pipelines. The actual "is this slot free?" check lives in createHold, the handler. The optimistic "pending" state lives with the mutation that asked for it, in TanStack Query's onMutate or whatever your app uses. Both belong to this one request, so neither goes in a pipeline. 🛠️ Three rules for building better middleware Some frameworks got three design choices wrong and are still paying for it in bugs. Our example gets all three right, and each rule below starts from the lines of code where it does. 1) Let next hand the result back. In the example, the idempotency layer caches the response in two lines: // idempotency, on the server const res = await next(ctx); if (res.status fetch(req)), around the network call and outside the list, since a retry has to call downstream more than once. Vue Router reached the same conclusion when it dropped next from its guards and switched to returned values. The DOM shows what the other kind of stop costs you. stopPropagation() stops an event by setting a flag on it. Every other listener depends on that same event. So one call inside a widget library can quietly disable your delegated click handler and the analytics listener on document. Nothing in your code tells you why. 3) Give the shared state a type. Layers need a place to leave things for each other, and the server example declares that place up front: state: { requestId?: string; userId?: string; tenantId?: string }; auth writes userId and tenantId, idempotency reads the first and rateLimit reads the second. That type lists every field a layer can touch. The client goes one step further and shares nothing: each layer passes a new Request down through withHeader and leaves the old one untouched. Shared state is how layers end up depending on each other, so keep it small enough to read at a glance. 🧭 One idea, many names One path, the jobs layered around it, and an outsider adding one more. If you've read the Gang of Four book, you've already met this idea as “Chain of Responsibility”. The book's own example is context-sensitive help. You ask for help on a Print button, and the button gets the request first. If it has nothing to say, it passes the request to its dialog, which can pass it on to the application. Each object either handles the request or forwards it to the next one, the same way a click bubbles from a button up to document in the DOM. Where A layer is called It passes control on with Koa, Express middleware next() FastAPI middleware call_next(request) Django middleware get_response(request) Redux middleware next(action) Angular interceptor next(req) gRPC (Go) interceptor handler(ctx, req) pytest fixture yield LangChain agents wrap_tool_call hook handler(request) Gang of Four handler a call to its successor So the next time a framework's docs say middleware, interceptor, filter, guard or hook, you know what to check: what a layer receives, what it calls to pass control on, and whether the result comes back. Try it on your own code. I'd start by handing this to your coding agent. Share the article with it and ask it to find the functions in your codebase that keep collecting jobs that belong to every call.
Middleware: An Architecture Pattern Across Stacks and Frameworks
Full Article
Original Source
Read the full article at Hackernoon →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.