Tutorial

Multi-Tenant Architecture: Serving Many Customers from One System Without Mixing Their Data

Silo, pool or bridge? How to isolate customers' data, enforce it with Postgres row-level security, and keep one noisy tenant from slowing everyone else.

Shobit Singh
•
8 min read
Multi-Tenant Architecture: Serving Many Customers from One System Without Mixing Their Data

A customer opens a support ticket saying they can see another company's invoices. The cause is usually mundane: one query in one endpoint is missing a WHERE tenant_id = ?. Preventing that class of bug is most of what multi-tenant architecture is about.

Multi-tenancy means serving many customers (tenants) from shared infrastructure while keeping their data, configuration and performance separate. The design question is how much to share. Share too little and you pay for it in operations. Share too much and you risk the leak above.

Where something is opinion rather than AWS or Azure guidance, it's labelled as such.


1. The isolation spectrum: silo, pool, bridge

To keep this concrete, imagine a fictional invoicing product called Ledgerly. It has 2,000 small businesses on its standard plan and a handful of large enterprise customers, some of whom have contractual or regulatory reasons to keep their data apart. We'll come back to Ledgerly throughout.

The AWS SaaS vocabulary has three models. Silo gives each tenant a dedicated application stack, with a database instance per tenant. Bridge shares the application tier but gives each tenant its own schema or database. Pool has tenants share the whole stack, including the same tables, and separates them logically.

Row-level security (RLS) is the strongest way to enforce that separation in the database, but it's an enforcement mechanism, not what defines the model. Each step trades isolation against cost: silo isolates most and costs most, while pool costs least and isolates least.

Bridge is also used more broadly for any mix of silo and pool. As systems break into smaller services, you often end up mixing them by layer. For example, a shared web tier can sit in front of per-tenant application tiers and storage.

Azure uses different names for similar ideas: automated single-tenant, fully multitenant, vertically partitioned and horizontally partitioned. A vertically partitioned setup keeps most tenants on shared infrastructure but gives dedicated infrastructure to those who need more performance or data isolation. A horizontally partitioned one separates specific components per tenant, which helps if most of your load comes from a few components.

Model

What's shared

Isolation

Cost and ops

Typical fit

Pool

Compute, database, tables (tenant_id column; RLS recommended)

Logical

Lowest cost, simplest fleet

Many small and mid-size tenants

Bridge

Compute; separate schema or database per tenant

Data-level

Moderate; migrations multiply per tenant

Compliance-sensitive data, shared app tier

Silo

Nothing (dedicated stack)

Strongest

Highest cost, most automation required

Regulated or very large enterprise tenants

Here is how Ledgerly might use all three at once:

flowchart LR
    STD["Standard plan<br/>2,000 small businesses"]
    CX["Enterprise customer X<br/>data must be kept separate"]
    CY["Enterprise customer Y<br/>own region, own contract terms"]

    subgraph shared["Shared app tier"]
        APP["Ledgerly API + workers"]
    end

    POOLDB[("POOL<br/>shared database<br/>tenant_id + RLS")]
    XDB[("BRIDGE<br/>dedicated database for X")]

    subgraph silo["SILO: fully dedicated stack"]
        YSTACK["App tier + database for Y"]
    end

    STD --> APP
    CX --> APP
    CY --> YSTACK
    APP --> POOLDB
    APP --> XDB

2. How to choose

Ask these questions in order:

  1. Do contracts or regulations require physical separation? Data residency rules, security reviews and large enterprise contracts can push some customers toward bridge or silo, and Azure's guidance treats commercial requirements for physically isolated data as a legitimate driver.

  2. How uneven are your tenants? If a few customers generate most of the load, pooling them with everyone else invites trouble (see section 6).

  3. How many tenants will you have? Per-tenant schemas and databases multiply every migration, backup and monitoring job, so the cost grows with tenant count.

  4. Can your team operate what you're designing? Azure warns against an architecture that gets more complex as you scale, or one you lack the people and skills to maintain.

Recommendation (opinion, not AWS or Azure guidance): default to pooled for the standard tier, and design so a tenant can be promoted to dedicated infrastructure later. For Ledgerly, that means the 2,000 small businesses stay pooled, and customers X and Y get promoted when their contracts require it.

AWS describes a similar hybrid, with siloed infrastructure for high-traffic customers and bridge or pool for quieter ones, but "pool first, promote the exceptions" as a default is a judgment call. Promotion is only cheap if tenant identity is built in from day one. Retrofitting it is expensive.

3. Two planes: control and application

Mature SaaS systems split into two planes. The application plane serves tenants' requests. The control plane handles everything about tenants themselves: onboarding, identity, tenant registry, billing and metering. In Ledgerly, creating a new customer account and provisioning their data is control plane work. Generating one of their invoices is application plane work.

AWS's reference architecture shows how the pieces fit. A workflow service such as Step Functions orchestrates tenant onboarding and provisioning as a durable process. STS issues short-lived credentials scoped to a tenant by mapping a tenant tag into the session. KMS can encrypt each tenant's data under a per-tenant key. Other clouds have equivalents, so the pattern carries over: automated onboarding, credentials scoped per tenant, and, where practical, separate encryption keys per tenant.

4. Carrying tenant context through every layer

A multi-tenant system fails when some code path forgets which tenant it is serving, so tenant context has to be explicit at every layer. Resolve it once at the edge from verified identity, such as a signed token claim or an authenticated subdomain, and never from a client-supplied header or body field.

Then carry it through logs, traces, metrics, cache keys, queue messages, object-storage prefixes, background jobs and webhooks. A withTenant(...) wrapper (shown below) helps because developers don't have to remember a WHERE clause on every query.

5. Working example: pooled Postgres with row-level security

RLS moves the tenant filter into the database, so every query is checked instead of relying on each one to include it.

Schema and policy

-- Migrations run as app_owner; the application connects as app_user.
CREATE ROLE app_user LOGIN NOSUPERUSER NOBYPASSRLS;

CREATE TABLE invoices (
  id           uuid PRIMARY KEY DEFAULT gen_random_uuid(),
  tenant_id    uuid NOT NULL,
  amount_cents bigint NOT NULL,
  created_at   timestamptz NOT NULL DEFAULT now()
);

-- Lead every index with tenant_id
CREATE INDEX invoices_tenant_created ON invoices (tenant_id, created_at DESC);

ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;
ALTER TABLE invoices FORCE ROW LEVEL SECURITY;  -- owner is NOT exempt

CREATE POLICY tenant_isolation ON invoices
  USING      (tenant_id = NULLIF(current_setting('app.current_tenant', true), '')::uuid)
  WITH CHECK (tenant_id = NULLIF(current_setting('app.current_tenant', true), '')::uuid);

GRANT SELECT, INSERT, UPDATE, DELETE ON invoices TO app_user;

The NULLIF(..., '') makes the policy fail closed. If the setting is missing or empty, the comparison yields NULL and the query returns no rows, instead of raising a cast error or matching something unintended.

Application helper (Node/TypeScript, pg)

import { Pool, PoolClient } from "pg";
const pool = new Pool({ connectionString: process.env.DATABASE_URL });

export async function withTenant<T>(
  tenantId: string,                       // from the verified token, never from user input
  fn: (db: PoolClient) => Promise<T>
): Promise<T> {
  const client = await pool.connect();
  try {
    await client.query("BEGIN");
    // third argument `true` = local to this transaction (same as SET LOCAL)
    await client.query("SELECT set_config('app.current_tenant', $1, true)", [tenantId]);
    const result = await fn(client);
    await client.query("COMMIT");
    return result;
  } catch (err) {
    await client.query("ROLLBACK");
    throw err;
  } finally {
    client.release();
  }
}

Pitfalls

  • Connection pooling can leak tenants. Session-level state is attached to the physical connection. When the pool hands that connection to a different tenant's request, the new request inherits the old tenant's context. The fix is to scope tenant context to the transaction rather than the connection. That is also what makes it work with transaction-mode pooling.

  • Some roles skip RLS. By default the table owner bypasses RLS, and FORCE ROW LEVEL SECURITY closes that gap. Superusers and roles with BYPASSRLS are always exempt, and FORCE doesn't change that.

  • SECURITY DEFINER functions use their owner's privileges. They run with their owner's privileges rather than the caller's. If the owner is a superuser or has BYPASSRLS, the function runs without RLS no matter who calls it, so a function that returns data can expose every tenant's rows to any caller. If the owner is an ordinary role and the table has FORCE ROW LEVEL SECURITY, the policies still apply inside the function. Run the app as a restricted role, give these functions a non-exempt owner, and audit them.

  • Constraints can reveal rows. A unique index isn't subject to the policy, so a colliding insert across tenants can confirm that another tenant's row exists. Either scope uniqueness with tenant_id or translate the error into a generic failure.

  • RLS governs rows, not columns. A user can pass every row policy and still change a column they shouldn't, such as upgrading their own plan. Put column-level restrictions and entitlement checks in separate layers.

6. Noisy neighbors

In shared infrastructure, one tenant's spike degrades everyone else. One engineering team describes it as inevitable in a traditional multi-tenant setup, and data and storage services are especially prone to it. Seasonality makes it worse, because a subset of tenants can see load jump suddenly.

Mitigations, roughly from cheapest to most expensive:

  1. Tag every metric with tenant_id. You can't fix a noisy tenant you can't identify.

  2. Per-tenant rate limits and quotas at the API gateway.

  3. Fair queuing: per-tenant concurrency caps on workers, so one tenant's bulk import can't starve the queue.

  4. Query guardrails: statement timeouts and row limits, with read-heavy reporting on replicas.

  5. Promoting the outliers to dedicated resources, sometimes called the "supertenant" approach.

7. Scaling out with stamps (cells)

Past a certain size, one big shared environment becomes a single blast radius. The Deployment Stamps pattern repeats a complete copy of the stack, and each copy serves a group of tenants.

A stamp dedicated to one tenant gives the highest isolation and avoids noisy neighbors, but it is expensive. Stamps shared by many tenants cost less, though they still need their own patterns for managing sharing inside the stamp, and noisy neighbors can still occur there. Either way, an outage or a bad release stays within the customers on that stamp, and you can keep adding stamps as you grow.

Stamps depend on automation. Azure's guidance calls for automated deployment, typically declarative infrastructure as code such as Bicep or Terraform. Stamps can also be grouped into rings, so changes reach early-adopter cohorts first.

8. Plan for the tenant lifecycle

Azure's lifecycle guidance calls out trials, tenants who upgrade to a tier that requires a dedicated deployment, and rebalancing a tenant to a different deployment to even out load. If your design can't relocate a tenant, the last two become custom migrations. A tenant-to-deployment routing table in the control plane, plus a tenant export/import tool you've rehearsed before you need it, make them routine.

9. Test for leaks continuously

The main test surfaces are cross-tenant data leaks, tenant-ID propagation through every code path, noisy-neighbor behavior and per-tenant quotas. A minimal leak test belongs in CI:

test("tenant B cannot see tenant A's invoices", async () => {
  await withTenant(A, db =>
    db.query("INSERT INTO invoices (tenant_id, amount_cents) VALUES ($1, 100)", [A]));
  const res = await withTenant(B, db => db.query("SELECT * FROM invoices"));
  expect(res.rowCount).toBe(0);
});

Add a startup or CI check that every tenant-scoped table has RLS enabled and forced. Also call your SECURITY DEFINER functions as the app role with a tenant set, and assert they return only that tenant's data.


Checklist

  • Tenant resolved from verified identity, once, at the edge

  • Tenant context in every log line, metric, cache key, queue message and job

  • RLS enabled and forced on every tenant table, with explicit USING and WITH CHECK

  • App runs as a non-owner, non-superuser role without BYPASSRLS; SECURITY DEFINER functions have a non-exempt owner and are audited

  • Transaction-scoped tenant context (safe with pgBouncer transaction mode)

  • tenant_id leads indexes and uniqueness constraints

  • Per-tenant rate limits, quotas and metering

  • Automated onboarding and a tested path to move a tenant to dedicated infrastructure

  • Cross-tenant leak tests in CI

  • Everything deployable by IaC, ready for stamps when you need them

Conclusion

Most of the work comes down to three decisions: how much to share, how to carry tenant context through every layer, and how to move a tenant when its needs change. Starting pooled, automating the control plane early and testing for leaks in CI keeps those options open.


Sources

Comments (0)

Join the discussion by logging into your account.

No comments yet. Be the first to comment!

Shobit Singh

Passionate developer sharing knowledge about modern web technologies and best practices.

Subscribe to Shobit Singh's Newsletter

Direct email dispatches when new stories are published. Zero algorithms.