CodeSapient
← All articles

Multi-Tenant SaaS Architecture: Tenant Isolation Beyond the Database

A practical guide to multi-tenant SaaS architecture: choosing shared tables, schema per tenant, or database per tenant, enforcing isolation with PostgreSQL row-level security, and closing leaks in caches, jobs, files, and search.

By CodeSapient Labs ·
Risograph-style illustration of an apartment building with a locked door on every unit standing on a shared platform foundation, representing multi-tenant SaaS architecture

The cross-tenant data leak in a SaaS product is rarely in the query you wrote carefully. It is in the one you forgot: a report endpoint added on a Friday without a tenant_id filter, a cache key built from a product ID alone, a background job that ran with whichever tenant the worker handled last. One customer sees another customer's invoices, and the conversation stops being about features.

Multi-tenant SaaS architecture is the set of decisions that let many customers share one product, one codebase, and usually one infrastructure while each of them experiences it as their own private system. This guide covers the choice you make first (how to store tenant data), the control that does the most work (letting PostgreSQL enforce isolation), and the places teams most often forget once the database is handled: caches, queues, files, search, logs, and noisy neighbours.

Risograph-style illustration of an apartment building with locked doors on every unit standing on a foundation labelled shared platform, representing multi-tenant SaaS
Conceptual illustration: one shared platform, a locked door for every tenant.

What multi-tenancy actually means

A tenant is the customer boundary in your product, usually a company, workspace, store, or account that owns data and users. Multi-tenancy means those tenants are served by the same running application rather than by a separate copy of the software per customer. It is what keeps onboarding, upgrades, monitoring, and billing as a single operation instead of one per customer.

AWS's SaaS guidance gives the vocabulary most architecture discussions now use: silo, pool, and bridge models. In a silo, a tenant gets dedicated resources. In a pool, tenants share resources and isolation is enforced logically. A bridge mixes the two, pooling some layers and siloing others. The useful point in that framing is that the choice is made per layer, not once for the whole system. Your web tier can be pooled while one regulated customer's database is siloed.

Choosing a data model: shared tables, schemas, or databases

The database is where tenancy decisions are hardest to reverse, so it deserves the most thought up front.

Diagram comparing three ways to store tenant data: pool with shared tables and a tenant_id column, bridge with a schema per tenant in one database, and silo with a database per tenant, with isolation and cost increasing from left to right
Conceptual diagram: isolation and operating cost both rise as you move from pool to silo.

Shared tables with a tenant_id column (pool)

Every tenant-owned table carries a tenant_id, and every read and write is scoped to it. There is one schema, one migration run per release, and one connection pool. Cross-tenant reporting for your own team is a normal SQL query. The weakness is that isolation is logical: one missing filter can expose data, and one heavy tenant shares buffers, locks, and I/O with everyone else. Both weaknesses can be managed, and the rest of this article is largely about how.

Schema per tenant (bridge)

Each tenant gets its own PostgreSQL schema with identical tables, and the application switches the search_path per request. In Django, django-tenants implements this pattern, with shared apps in the public schema and tenant-specific apps in each tenant schema. Per-tenant backup and restore become simpler, and an unscoped query cannot read across schemas by accident. The cost is operational: each migration runs once per schema, a failed migration can leave tenants on different versions, and thousands of schemas mean a very large system catalog for tools and backups to work through. It still shares one database server, so it does not solve noisy neighbours.

Database per tenant (silo)

Each tenant has its own database or instance. This gives the strongest separation, the simplest per-tenant restore, a clean story for data residency, and the option to size hardware per customer. It also multiplies everything you operate: migrations, connection pools, monitoring, upgrades, and cost for small tenants who barely use the product.

How to choose

  • Start pooled for most B2B SaaS products with many small and mid-sized customers, and add database-enforced isolation from day one.
  • Silo selectively when a contract, regulator, or data-residency requirement demands physical separation, or when one tenant is large enough to justify dedicated capacity.
  • Use schema per tenant deliberately, for example when tenant counts stay modest and you need per-tenant restores or per-tenant schema variations, not because it feels like a safe middle ground.
  • Plan for a hybrid. Keep a tenant directory that records where each tenant's data lives, so moving a tenant from the pool to its own database later is a routing change plus a data move, not a rewrite.

Make the database enforce isolation with row-level security

In a pooled design, relying on every developer to remember WHERE tenant_id = ... on every query is how leaks happen. PostgreSQL's row security policies move that rule into the database: once row-level security (RLS) is enabled on a table, normal access must be allowed by a policy, and if no policy exists the default is deny.

-- Run as the table owner / migration role
ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;
ALTER TABLE invoices FORCE ROW LEVEL SECURITY;

CREATE POLICY tenant_isolation ON invoices
  USING (tenant_id = NULLIF(current_setting('app.tenant_id', true), '')::uuid)
  WITH CHECK (tenant_id = NULLIF(current_setting('app.tenant_id', true), '')::uuid);

-- The application connects as a role that does not own the table
GRANT SELECT, INSERT, UPDATE, DELETE ON invoices TO app_user;

-- Per request, inside the transaction:
BEGIN;
SELECT set_config('app.tenant_id', '3f2b8c1e-5a7d-4e9b-9c40-1d2e3f4a5b6c', true);
SELECT * FROM invoices WHERE status = 'open';  -- only this tenant's rows
COMMIT;

A few details in that snippet matter more than they look:

  • Who bypasses RLS. Superusers and roles with BYPASSRLS always bypass policies, and table owners do too unless you use FORCE ROW LEVEL SECURITY. Run the application as a separate, non-owner role and keep the migration role for migrations.
  • Missing context fails closed. current_setting(name, true) returns NULL instead of an error when the setting does not exist, and NULLIF handles an empty value. A request that never set a tenant sees zero rows rather than every row.
  • Transaction-local settings. Passing true as the last argument to set_config scopes the value to the current transaction, like SET LOCAL. That is what keeps tenant context from leaking between requests on a shared connection, and it is why this pattern works with PgBouncer's transaction mode, as covered in our guide to PostgreSQL connection pooling for SaaS APIs.

Schema details that RLS does not cover

  • Uniqueness is per tenant. An invoice number or a slug is unique within a tenant, so the constraint should be UNIQUE (tenant_id, number), not UNIQUE (number). A global constraint also reveals, through its error, that another tenant already uses that value.
  • Foreign keys can cross tenants. The PostgreSQL documentation notes that referential integrity checks bypass row security. If invoice.customer_id references customers(id) alone, nothing stops an invoice in tenant A pointing at a customer in tenant B. Composite keys such as FOREIGN KEY (tenant_id, customer_id) REFERENCES customers (tenant_id, id) close that gap.
  • Indexes lead with tenant_id. Almost every query filters by tenant, so composite indexes like (tenant_id, created_at) usually serve real access patterns better than single-column ones.
  • Keep policies simple. The docs describe policies that only compare values in the current row as the simplest and best-performing case. Policies that look up other tables on every row are slower and, as the docs explain, can introduce race conditions.

RLS is a backstop, not a replacement for application scoping. Keep tenant filters in your ORM layer as well, so queries are efficient and intent is visible in code, and let the database catch what the application misses.

Resolve the tenant once, then carry it everywhere

Every request needs exactly one answer to "which tenant is this?" before any business logic runs. Common sources are the subdomain or custom domain, a claim in a signed session or token, or an explicit workspace switch in the UI. Whatever the source, verify that the authenticated user is actually a member of that tenant. A tenant ID taken from a header or request body and trusted without that check turns your isolation model into a URL parameter.

In Django, this is typically one middleware that resolves the tenant, checks membership, stores it in request-scoped context, and opens each database transaction with the tenant setting applied. The harder part is everything that runs outside that request.

Hub diagram showing tenant context resolved per request and carried into six layers: database RLS policy per transaction, cache key, job queue task arguments, file storage path prefix, search index filter, and logs and metrics
Conceptual diagram: tenant context has to reach every layer, not only the database.

Where tenant data leaks outside the database

Caches

A cache key like product:42 works in a single-tenant app and becomes a leak in a multi-tenant one if IDs are not globally unique or if the cached value is a rendered page with tenant-specific prices. Put the tenant in every key. Django's cache framework lets you set a KEY_FUNCTION so the prefix is applied centrally instead of by convention, as described in the Django cache documentation. Check HTTP caching too: a CDN or reverse proxy that caches authenticated, tenant-specific responses by path alone can serve one tenant's page to another.

Background jobs

Workers have no request, so they have no tenant unless you give them one. Pass tenant_id explicitly in every task's arguments, set the database context at the start of the task, and clear it when the task ends so the next task on that worker starts clean. Scheduled jobs that run "for everyone" should loop over tenants and process each one in its own scoped transaction, rather than running a single unscoped query. If you are building on Celery, our guide to Django Celery and Redis background jobs covers the task patterns this builds on.

File and object storage

Store uploads under a tenant prefix such as tenants/{tenant_id}/..., derive that prefix on the server from the resolved tenant, and serve files through short-lived signed URLs generated after an authorization check. Never build a storage path from a filename or ID the client sent without confirming it belongs to the current tenant.

Search indexes and AI retrieval

Search engines and vector stores usually sit outside PostgreSQL, so RLS does not protect them. Store the tenant on every document and apply the tenant filter in the query layer that every caller must go through, rather than trusting each feature to remember it. This matters more as products add AI assistants: a retrieval step that searches across all tenants' documents can surface another customer's content in an answer.

Webhooks and integrations

Outbound webhooks should use per-tenant endpoints and per-tenant signing secrets, and inbound webhooks from third parties must be mapped to a tenant before processing. A shared endpoint with ambiguous tenant mapping is one of the failure modes covered in our article on SaaS webhook reliability.

Logs, metrics, and error reports

Add tenant_id to every structured log line, trace, and error report. It makes incidents debuggable and lets you see per-tenant usage. Be careful in the other direction as well: error trackers and logs often capture request payloads, so avoid logging customer data that your support tooling would then expose to staff who should not see it.

Noisy neighbours and fairness

In a pooled system, one tenant's bulk import can slow everyone else. Isolation is about performance as well as data:

  • Apply rate limits per tenant, not only per IP or per user, so one integration cannot exhaust shared API capacity.
  • Route heavy or bulk work to separate queues so large tenants do not starve small ones of worker time.
  • Use database safeguards such as statement_timeout for interactive traffic so a runaway query fails instead of holding resources.
  • Measure usage per tenant. When one tenant consistently dominates load, that is the signal to move it to a bridge or silo tier, which is straightforward when tenant_id is already on every row.

Turning a single-tenant app into a multi-tenant one

Many products start with one customer and need multi-tenancy later. A safe order of work looks like this:

  1. Create a tenants table and a tenant membership model for users.
  2. Add a nullable tenant_id to every tenant-owned table, backfill it with the original customer's tenant, then make it NOT NULL.
  3. Replace global unique constraints and foreign keys with tenant-scoped composite versions.
  4. Add tenant resolution middleware and scope ORM queries through a shared manager or repository layer.
  5. Enable RLS table by table, starting with the most sensitive data, and run the application as a non-owner role.
  6. Work through caches, jobs, storage, search, and integrations from the previous section.
  7. Onboard a second internal test tenant before the second real customer, and use it to hunt for leaks.

Testing tenant isolation

Isolation needs tests that try to break it, not only tests that check happy paths. Create two tenants with overlapping data in your fixtures, then assert that tenant A cannot read, update, or delete tenant B's records through each API endpoint. Run integration tests against PostgreSQL as the same non-owner role production uses, otherwise RLS is silently bypassed and the tests prove nothing. A simple CI check can flag tables that have a tenant_id column but no row security, using the relrowsecurity and relforcerowsecurity flags in pg_class.

Common mistakes

  • Choosing schema or database per tenant by default and discovering the migration and operations cost only after onboarding hundreds of customers.
  • Running the application as the table owner or a superuser, which quietly disables RLS.
  • Setting tenant context at session level on pooled connections instead of per transaction.
  • Trusting a tenant ID sent by the client without checking membership.
  • Protecting the database and forgetting caches, workers, file storage, and search.
  • Admin and support tools that query across tenants with no audit trail.

FAQ

Is row-level security enough on its own?

No. It is a strong backstop for the database, but it does not protect caches, object storage, search indexes, or third-party services. Combine it with application-level scoping and the per-layer controls above.

Does row-level security slow PostgreSQL down?

Simple policies that compare a column with a session setting add little work, especially when indexes lead with tenant_id. Policies that query other tables per row are where performance problems appear. Measure with EXPLAIN ANALYZE on real query patterns rather than assuming either way.

When should we move a tenant to its own database?

When a contract, regulation, or residency requirement calls for physical separation, or when that tenant's load meaningfully affects others. Having tenant_id everywhere and a tenant directory in place makes the move a planned data migration rather than a redesign.

Can Django do multi-tenancy without a third-party package?

Yes. A pooled design needs a tenant model, a middleware, scoped managers, and RLS policies applied in migrations. Packages like django-tenants are useful when you specifically want schema-per-tenant.

Getting the foundation right

Multi-tenancy is cheapest to get right before the second customer and most expensive to retrofit after the fiftieth. Start pooled unless you have a clear reason not to, let PostgreSQL enforce the tenant boundary, carry tenant context through every layer outside the database, and keep a path open to silo the few tenants that will eventually need it.

If you are planning a new SaaS product or need to convert an existing app to serve many customers safely, CodeSapient builds and reviews this kind of architecture as part of our custom software and SaaS development services. Get in touch to talk through your tenancy model, data isolation, and scaling plan.

← Back to Blog