Blog

Engineering at Canals

How to Implement Overload Protection for Software Engineering Teams

More traffic, more problems: Here's your guide to setting up a robust overload protection system

September 2, 2026
8
Mins read
How to Implement Overload Protection for Software Engineering Teams

Table of Contents

When an app reaches a certain scale, you will start running into incidents where traffic from a small sample of actors overloads the system. It might be a single user or job making many requests, or a heavy API that reaches production and is called by multiple users simultaneously. 

At Canals, we started running into such incidents about a year after launch: for example, one time a user submitted a huge document, became frustrated that it was taking a long time to process, and submitted it multiple times. Combined with SQS retrying the stalled jobs, it became a situation that brought us down. 

These sorts of situations are inevitable as companies grow, and if you want a resilient app, then you must have a system to handle them. That system is called overload protection.

In this post we will be covering how we implemented overload protection, how it works, and common pitfalls to watch out for.

‍

What Is Overload Protection?

Overload protection identifies offending sources of traffic and blocks them, protecting the rest of the app from performance degradation or downtime. It might make the experience worse for those users that must be blocked, but it prevents the issue from affecting all users. The goal should be to identify the smallest subset of traffic that needs to be blocked.

Our overload protection system compares a request against a dynamic set of rules. When an incoming request or job matches against any of these dynamically set rules, we drop it.

In addition, we have a service that automatically detects bad actors. It continuously tracks every single incoming API request or offline job, monitoring server and database CPU, memory usage, and latency. If any of these things spike, our system uses statistical analysis to identify the offending actor. It fires an alert, and from there a human can review and apply the appropriate overload protection rules.

‍

How to Implement Overload Protection

Define the rules

const httpOverloadProtectionBlocks = zod.object({
  url: zod.string(),
  path: zod.string(),
  queryParams: zod.string(),
  method: zod.string(),
  remoteAddress: zod.string(),
  headers: zod.record(zod.string(), zod.union([zod.string(), zod.array(zod.string()), zod.undefined()])),
});

With the exception of method, all other rules accept glob patterns. A combination of path, queryParams and method is capable of blocking any specific interaction, while remoteAddress and headers (including User-Agent and sessionId) are sufficient to block any user. The above is for http requests. Websockets are matched against messageType, as well as remoteAddress and userId. Background tasks are matched on handler and args.

Match requests

export function matchesPattern(pattern: string | undefined, value: string | undefined): boolean {
  if (pattern === undefined || pattern === '*') return true;
  if (value === undefined) return false;
  if (!pattern.includes('*')) return pattern === value;

  return new RegExp(`^${escapeRegExp(pattern).replaceAll('\\*', '.*')}$`).test(value);
}

async function isRequestBlocked(request: RequestOverloadProtectionData): Promise<boolean> {
  const {blocks} = await overloadProtectionSettings.get();

  return (blocks ?? []).some(
    (block) =>
      matchesPattern(block.url, request.url) &&
      matchesPattern(block.method, request.method) &&
      matchesPattern(block.remoteAddress, request.remoteAddress) &&
      comparePath(block.path, request.path) &&
      compareQueryParams(block.queryParams, request.queryParams) &&
      compareHeaders(block.headers, request.headers)
  );
}

All fields must match for a rule to activate. We use '' to match no params and {} to match any headers. We found this set of rules enough to cover all required cases, although due to the lack of regex support, multiple rules might need to be defined for a given scenario.

Put the matcher behind a master switch

async function isBlocked(request: RequestOverloadProtectionData): Promise<boolean> {
  const mode = await overloadProtectionDefinitions.mode.get();
  if (mode === OverloadProtectionMode.DISABLED) return false;

  if (!(await isRequestBlocked(request))) return false;

  if (mode === OverloadProtectionMode.DRY_RUN) {
    getLogger().info(request, '[OVERLOAD_PROTECTION] request would be blocked but DRY_RUN mode is set');
    return false;
  }

  return true;
}

Overload protection lives on the hottest path in the app, so you need to be able to turn it off without deleting the rules, and to try a rule out before it starts dropping real traffic.

Define a request hook

app = await fastify(...);

app.addHook('onRequest', async (request) => {
  logIncomingRequest(request);

  if (await overloadProtection.isBlocked(request))
    throw OverloadProtectionError(request);

  // ...
}

A request that reaches the database or the cache has already cost you the resource you were trying to protect, so the gate should be placed as close to the socket as the framework allows.

Detect overload automatically

During an incident, nobody can eyeball which of a thousand concurrent requests is the problem. There are three questions the detector has to answer: 

  • Is anything overloaded?
  • Who is causing it?
  • Is it worth looping in a human?

The below code snippets pertain to detecting database CPU overload. Other metrics are implemented similarly.

Attribute load to requests

export function annotateQuery(query: string): string {
  const annotations = omitUndefined({requestId: requestProperties.get(), taskId: asyncContext.get('taskId')});
  if (_.isEmpty(annotations)) return query;

  const annotationsString = Object.entries(annotations).map(([key, value]) => `${key}=${value}`).join(', ');
  return `${query} /* ${annotationsString} */`;
}

Postgres will tell you which queries are running, but not which request issued them, so we put the request id into the query text as a comment.

async function trackRequestStart(requestInfo: ActiveRequestInfo): Promise<void> {
  const requestKey = `active_requests:${requestInfo.requestId}`;
  await cache.hset(SCOPE, requestKey, requestInfo, {ttl: CACHE_TTL});
  await cache.zadd(SCOPE, ACTIVE_REQUESTS_INDEX_KEY, Date.now(), requestKey);
}

That id is all the database has, so the attributes we block on have to be recorded elsewhere. At the same entry point where we log an incoming request, we also write a record of it to Redis.

Make load attributable: Identify the offender

const OVERLOAD_SOURCES = {
  api: {attributes: ['userId', 'remoteAddress', 'cleanPath'], getActivities: getAllActiveRequests, matchAny: false},
  consumer: {attributes: ['target'], getActivities: getAllActiveJobs, matchAny: false},
  websocket: {attributes: ['userId', 'remoteAddress', 'messageType'], getActivities: () => [], matchAny: true},
};

Each source of traffic declares which of its attributes are relevant.

With that, we can take the request ids of the queries holding CPU and see if any attribute (userId, path, etc) is over-represented.

function calculateUsageRatioForRequests(requests: ActivityInfo[], attributeName: ActivityAttribute) {
  const withAttribute = requests.filter((request) => Object.hasOwn(request, attributeName));
  if (withAttribute.length === 0) return {};

  const counts = countBy(withAttribute, (request) => request[attributeName]);
  return mapValues(counts, (count) => count / withAttribute.length);
}

function findOutliers(ratios: Record<string, number>): string[] {
  const entries = Object.entries(ratios);

  // A single unique value is 100% of the load, and therefore the offender.
  if (entries.length === 1) return [entries[0][0]];

  const mean = _.mean(entries.map(([, ratio]) => ratio));
  const aboveMean = entries.filter(([, ratio]) => ratio > mean);

  // More than a handful above the mean means load is spread out: there is nobody to blame.
  return aboveMean.length <= 4 ? aboveMean.map(([value]) => value) : [];
}

The two functions doing the actual blaming are small. One turns a set of activity records into each value's share of that set. The other decides whether any share is out of line.

Example: 50 requests are holding database CPU, 40 of them belong to one user, and the rest are spread over 10 other users. That user's ratio is 0.8 against a mean of 0.09, everyone else is below the mean, and the detector emits {source: 'api', attribute: 'remoteAddress', values: ['192.163.34.20']}.

Decide whether to alert

The monitor is a periodic job. It fires only once, and is resilient to single spikes. Every threshold, the required number of consecutive readings, and the monitors themselves are runtime settings, so the detector can be retuned or switched off mid-incident without a deploy. The alert lands in Slack with the offending candidates grouped by source, and a human engineer applies the relevant rule to block them. We stopped short of blocking automatically on purpose: a false positive that locks a legitimate customer out of the product is worse than alerting a human to intervene.

‍

7 Common Pitfalls in Setting Up Overload Protection

1. Leaving some logic unprotected. Any logic that runs before the overload protection gate isn’t protected. As mentioned above, you want to place the gate as close to the socket as the framework allows.

2. Gating the enqueuing of consumer requests instead of the execution. Gating the execution (the handler) allows dropping already-enqueued tasks.

3. Sending blocked jobs to the dead letter queue (DLQ). Maybe that’s what you want, but in some situations you won’t want to flood the DLQ. Our rules allow configuring what to do with a dropped task.

4. Fetching the rules on every request. In an overload scenario, the rule store is under pressure, too. We cache the rules in memory for a minute.

5. Matching on an attribute the gate can't get for free. A userId field in the HTTP schema would be the obvious desired outcome. However, resolving a user means resolving a session against the database, making the gate depend on the resource it exists to protect. Headers (the session cookie) and remote address cover the same ground at no cost.

6. Placing the control plane living inside the app that's failing. You need to be able to configure overload protection rules after your app is already overloaded. The app itself may be unresponsive, so the control plane should be independent. That’s why our overload protection rules live in our internal DevOps app.

7. Leaving no way to disable overload protection. If your overload protection system has a bug or is itself degrading the system, you need to be able to disable it. In our overloadProtection.isBlocked function, the mode check is literally the first line, so everything below it can be isolated instantly.

‍

Implement Overload Protection for Reliability

As you scale, you will sometimes have bad actors overload the system, whether they’re external actors or internal pieces of code, or whether they’re acting intentionally or unintentionally. Overload protection is a mechanism for detecting those threats and cutting access to the bad actor to avoid impacting the entire system.

Since implementing overload protection, we’ve stopped major incidents, enhanced our  debugging process, and protected our app users.

‍

Looking for your next engineering role? Visit our Careers page to learn more about working at Canals.

Written by:

Renan S Silva

Share
Tags

Engineering at Canals

Get started with Canals
See how AI can supercharge your business
Get a Demo

Related Articles

No items found.
CONTACT

Get a Demo

See how AI can transform your business — fast.