Shopify Case Study: Infrastructure for Millions of Merchants
A modular Rails monolith on top of tenant-isolated database pods, with storefront rendering split out and deployed close to buyers
This is an independent teardown based entirely on Shopify's own public engineering writing. KarmaKoders did not build or consult on this system.
Who: Shopify, the commerce platform behind millions of independent online stores. Scale: during Black Friday–Cyber Monday 2025, merchants sold $14.6 billion, the platform processed 90 petabytes of data and peaked at 489 million requests per minute; the year before, Shopify reported 10.5 trillion database queries and 80 million requests per minute on app servers alone — and says that level of traffic is now a normal day.
The constraint: every store is a tenant on shared infrastructure, and traffic is wildly spiky. A single creator's product drop or a flash sale can multiply one shop's load in seconds, and BFCM multiplies everyone's at once. The platform has to absorb one tenant's spike without slowing the rest, keep checkout up when a database fails, and do all of it on a Ruby on Rails codebase that started in the mid-2000s and is worked on by thousands of engineers. Rewriting into microservices was the obvious move. Shopify mostly did not make it.
Rendering diagram…
Before
After
Decisions
| Option A | Option B | What Shopify chose | Why it fits | |
|---|---|---|---|---|
| Codebase shape | Break into microservices | Keep one monolith, enforce internal boundaries | Modular monolith, with selective extraction | Keeps one deploy, one test suite and shared tooling while giving teams ownership; avoids distributed-systems overhead for code that changes together |
| Scaling the database | Bigger primary / read replicas | Shard by tenant into isolated pods | Pods keyed on shop_id | Commerce data is naturally partitioned by merchant — one store's orders never join another's — so isolation is almost free and caps blast radius |
| What to extract | Extract many domains | Extract only paths with a different performance profile | Storefront Renderer split out | Read-heavy, latency-critical buyer traffic benefits from its own runtime, caching and regional deployment; admin and checkout stay in the core |
| Language performance | Rewrite hot paths in another language | Make the runtime faster | Invest in YJIT for Ruby | Speeds up the whole fleet without a rewrite; Shopify uses Rust selectively where a systems language genuinely helps |
| Peak readiness | Load test before the event | Year-round platform-wide scale tests plus chaos drills | Nine-month BFCM readiness program | Failures that only appear when everything runs at capacity together cannot be found by testing components alone |
The pattern across every row is the same: isolate where the data is naturally separate, unify where the code changes together. Tenants are separated at the data layer because they never interact. Teams are separated by component boundaries, not network boundaries, because their code still ships together. The one real extraction — the storefront — happened because buyer-facing rendering has a different latency budget and deployment geography from the admin, not because microservices were fashionable.
Shopify has been candid that the modular monolith is a work in progress: in its own 2020 write-up, the main monolith had 37 components and Packwerk enforced boundaries on only about a third of them. The lesson is not that the architecture was ever finished — it is that incremental boundary enforcement let a very old codebase keep scaling while the business grew.
# config/database.yml defines shards: pod_1, pod_2, ... each with its own primary.
class ApplicationRecord < ActiveRecord::Base
self.abstract_class = true
connects_to shards: {
pod_1: { writing: :pod_1, reading: :pod_1_replica },
pod_2: { writing: :pod_2, reading: :pod_2_replica }
}
end
# A tiny, globally replicated lookup: which pod holds this tenant?
class PodDirectory
def self.pod_for(shop_id)
Rails.cache.fetch("pod:#{shop_id}", expires_in: 10.minutes) do
ShopPlacement.find_by!(shop_id: shop_id).pod_name.to_sym
end
end
end
# Every request and every background job runs inside its tenant's pod.
class TenantScope
def self.with_shop(shop_id, &block)
ActiveRecord::Base.connected_to(shard: PodDirectory.pod_for(shop_id), &block)
end
end
# Usage in a job — note shop_id is part of the job payload, never looked up globally.
class FulfillOrderJob < ApplicationJob
def perform(shop_id:, order_id:)
TenantScope.with_shop(shop_id) do
Order.find(order_id).fulfill!
end
end
endDesigning a multi-tenant platform that has to survive its own Black Friday?
We help founders and CTOs decide what to isolate, what to keep in one codebase, and when sharding is actually worth it. Book a 30-minute architecture review.
Next step
Book a 15-min Architecture Review
Walk through your system with a lead architect and leave with a scoped recommendation.
Schedule on Cal.com