
Architecture debates often get framed as a choice between two camps. One side wants things simple. The other wants things independent. Both are responding to real problems, but the argument gets easier to reason about once you separate design boundaries from deployment boundaries.
The first question is not “Should this be a monolith or microservices?” It is:
What problem are we trying to solve, and does that problem require a process boundary?
If the problem is poor structure, adding a network does not fix it. It turns a local coupling problem into a distributed one, with latency, partial failure, versioning and operational work attached.
This post compares modular monoliths and microservices, looks closely at what Segment, Prime Video and Shopify actually learned from their systems, and ends with a decision rule you can use without picking a camp first.
The short version
A monolith is primarily a deployment shape. A Big Ball of Mud is a design failure. They are not the same thing.
Good boundaries matter in either architecture. A modular monolith lets you find and change those boundaries while the cost of doing so is still low.
Microservices become compelling when independent deployment, independent scaling, security isolation, team autonomy or technology choice is valuable enough to justify distributed-system costs.
The practical strategy is often: start modular, enforce the boundaries, and extract only where a real forcing function appears.
Monolith is a deployment shape. A mess is not.
The word monolith is overloaded. In the useful sense, it describes a system that is deployed as one unit. That says something about the deployment boundary, not about whether the code inside that boundary is well designed.
A Big Ball of Mud is a different problem. Foote and Yoder described it as a casually and haphazardly structured system whose organization is driven more by expediency than design. A large monolith can be carefully modular; a smaller monolith can be a mess. Foote & Yoder — Big Ball of Mud
Microservices is also less precise than the name suggests. Fowler and Lewis describe the style as building one application from small services that run in separate processes, communicate over lightweight mechanisms, are organized around business capabilities, and are independently deployable through automated deployment. They also explicitly note that there is no single precise definition of the term. Fowler & Lewis — Microservices
A modular monolith keeps one deployment unit but puts strong internal boundaries around business capabilities. The code remains in one process, yet modules expose narrow interfaces and keep their implementation details private. The important part is not the folder structure; it is whether the boundary is real enough that callers do not need to reach inside.
So the useful comparison is not “simple architecture versus advanced architecture.” It is:
Where are the boundaries, and what enforces them?
In a modular monolith, the boundary is enforced by code organization, visibility rules, dependency rules, architecture tests and tooling. In microservices, the boundary is also a process and network boundary. That extra wall is powerful, but it comes with a price.
What microservices buy you
Microservices are not valuable because having many deployables is inherently better. They are valuable because separation can solve problems that are awkward to solve inside one deployment unit.
Independent deployment
A team can change one service and deploy it without rebuilding or releasing every other service. For many organizations, this is the strongest reason to adopt the style: the architecture can reduce coordination required for changes when service boundaries and team ownership line up.
But “many services” does not automatically mean “independent delivery.” If five services must be released together, share a database, or require constant cross-team coordination, you have acquired much of the complexity of distribution without getting the independence you wanted.
Independent scaling
If one capability has a radically different resource profile, it can make sense to scale it separately instead of scaling the entire application.
This is especially useful when a component is unusually CPU-heavy, memory-heavy, latency-sensitive, or high-throughput compared with the rest of the system.
Stronger isolation
A separate process can create a harder boundary around resources, failures, security and ownership. That can be useful, but it is not magic: a service that depends synchronously on three other services is still coupled to them at runtime.
Technology choice
A service can use a different runtime, language, or persistence technology when that difference has a concrete benefit. That freedom is real, but it should be treated as a consequence of independent ownership rather than as a goal by itself.
What the network costs
The network changes the failure model, not just the syntax of the call.
Remote calls are slower and less predictable
An in-process call is usually a local operation. A remote call involves serialization, transport, queuing and another process. Latency becomes variable, and a timeout is now a normal outcome that callers have to handle.
Fowler describes this as one of the central trade-offs of microservices: distributed calls introduce latency and additional complexity, and the system has to be designed for failure. Fowler — Microservice Trade-Offs
Every call can fail
A local function call generally succeeds or throws because of something in the current process. A remote call can fail because the other process is overloaded, the network is unavailable, the request timed out, a proxy failed, or the response was lost after the other side processed the request.
That creates a new class of design work: timeouts, retries, idempotency, circuit breaking, backpressure, observability and recovery.
Data boundaries become real system boundaries
Once data is owned by separate services, you lose the convenience of freely joining everything together. Cross-service consistency usually needs explicit protocols, such as events, sagas or compensating actions, rather than assuming one local transaction can cover the entire workflow.
A modular monolith can use one physical database while still enforcing logical ownership. A stricter design can give each module its own schema or even its own database. The important rule is that other modules should not treat another module's tables as their own API.
Operations become part of the architecture
A single deployable is one thing to build, deploy and monitor. A service fleet requires more automation, more health checks, more dashboards, more deployment coordination and more places where a failure can hide.
This is the microservice premium Fowler describes: the architecture introduces costs of deployment automation, monitoring, failure handling and distributed consistency. Those costs can be worthwhile, but they need a problem large enough to repay them. Fowler — Microservice Premium
Three real systems, three different lessons
Case studies are useful only when you keep their context. None of the examples below proves that one architecture is universally better.
Segment: microservices solved a real problem, then became the bigger problem
Segment's server-side destinations pipeline originally suffered from head-of-line blocking: if one destination was slow or unhealthy, it could hold up work for the other destinations. The team responded by giving each destination its own service and queue, which isolated failures between destinations. That was a sensible reason to introduce stronger boundaries. Twilio Segment — Goodbye Microservices
The design then accumulated a different kind of cost. They added more than 50 destinations, which meant more repositories and more shared-library coordination. Eventually there were more than 140 destination services, and three full-time engineers were spending most of their time keeping the system alive.
Segment consolidated those destinations into one service and one repository. The team reported that one engineer could deploy the service in minutes, and that shared-library improvements rose from 32 during the microservice period to 46 in the following year. They were also explicit about the trade-offs they accepted, including weaker fault isolation and less effective in-memory caching. Twilio Segment — Goodbye Microservices
The useful lesson is not “Segment proved monoliths are better.” It is that the original boundary solved a real problem, and the operational cost eventually outweighed that benefit for this particular part of the product.
Prime Video: distribution itself was the bottleneck
Prime Video's example needs even more context. The system was a video quality analysis tool, not the Prime Video streaming platform itself.
The team initially built the tool from distributed serverless components, using orchestration and intermediate storage between processing steps. At higher scale, the architecture supported only about 5% of the expected load. The team identified large amounts of S3 read/write traffic and a high number of Step Functions state transitions as important cost and scaling problems. InfoQ — Prime Video Switched from Serverless to EC2 and ECS to Save Costs
They consolidated the business logic into a single application process, moved the workload onto ECS running on EC2, and passed intermediate frames in memory instead of through S3. The revised architecture reduced operational costs by about 90%. Detector instances were still distributed across ECS tasks, so the final design was not simply “one giant process everywhere.” InfoQ — Prime Video Switched from Serverless to EC2 and ECS to Save Costs
The lesson is more precise than “monoliths are cheaper”: moving large amounts of data across boundaries has a cost, and sometimes the boundary itself is the wrong place to separate the work.
Shopify: a huge monolith can be made more modular without splitting everything
Shopify's 2020 account is a useful counterexample to the assumption that a very large codebase must immediately become a fleet of services. Its core monolith had more than 2.8 million lines of Ruby and 500,000 commits, with hundreds of developers working on it. Shopify had been breaking the application into components designed around business responsibilities. Shopify — Under Deconstruction: The State of Shopify's Monolith
The interesting part is that a public interface alone was not enough. When Shopify analyzed the dependency graph, every component depended on more than half of the others, and there were many circular dependencies. The interfaces existed, but the dependency graph still made the system hard to reason about. Shopify — Under Deconstruction: The State of Shopify's Monolith
Shopify built Packwerk for static dependency analysis and integrated it into the pull-request workflow. At the time of the article, Packwerk covered about a third of the components, and Shopify's long-term goal was an acyclic dependency graph. The company also said it was deliberate about extracting services only when there was a good reason. Examples included high-throughput, read-only storefront rendering and credit-card vaulting, where sensitive data isolation mattered. Shopify — Under Deconstruction: The State of Shopify's Monolith
That is a much stronger argument for modularity than for any particular deployment style: make the boundaries useful first, then decide which boundaries deserve their own process.
What a real modular monolith looks like
A modular monolith is more than a directory tree with names like Orders and Payments. Four properties make the idea useful.
Property | What it means | Why it matters |
|---|---|---|
Business-oriented modules | Group code by business capability and keep high-cohesion changes together. | Features stay closer to the code they change. |
Explicit public interfaces | Other modules depend on a small, intentional API rather than internal classes. | Internals can change without creating a chain reaction. |
Owned data | A module owns the data behind its behavior. Other modules use its API or events instead of querying its tables directly. | Prevents hidden coupling through the database. |
Enforced boundaries | Use compiler/project rules, architecture tests, static analysis or framework tooling. | A rule that lives only in documentation will eventually be bypassed. |
The data rule deserves a nuance: there is no single database layout that defines a modular monolith. You can keep one database and isolate modules by schema, or use separate databases. What matters is ownership and access, not whether every module has a different database server. Kamil Grzybek describes both a shared physical database with isolated schemas and stronger module-level data ownership as viable approaches. Kamil Grzybek — Modular Monolith: Integration Styles
Spring Modulith shows the enforcement side clearly: it can verify that module dependencies have no cycles and that external code accesses a module through its API packages rather than its internals. Spring Modulith — Verifying Application Module Structure
Shopify's experience adds an important refinement: encapsulation and dependency direction need to improve together. A clean-looking public API does not help much if everything depends on everything else. Their 2020 post explicitly emphasizes both a simpler dependency graph and high cohesion, including “change locality” — code that tends to change together should live together. Shopify — Under Deconstruction: The State of Shopify's Monolith
A small example in C#
Suppose Orders needs to charge a payment. The important part is not the full payment implementation; it is the shape of the boundary.
MyShop/
├── src/
│ ├── Orders/
│ │ ├── MyShop.Orders/
│ │ └── MyShop.Orders.Contracts/
│ ├── Payments/
│ │ ├── MyShop.Payments/
│ │ └── MyShop.Payments.Contracts/
│ └── MyShop.Host/ one deployable
└── tests/
└── MyShop.ArchitectureTests/Orders may reference Payments.Contracts, but it should not reference MyShop.Payments or reach into the Payments database.
The contract should also be coarse enough that it could survive a future process boundary:
public interface IPayments
{
Task<PaymentResult> ChargeAsync(
ChargeRequest request,
CancellationToken ct);
}
public record ChargeRequest(
Guid OrderId,
decimal Amount,
string Currency,
string IdempotencyKey);
public record PaymentResult(
Guid PaymentId,
PaymentStatus Status);The IdempotencyKey is intentional for a payment operation, but the example is not a complete payment system. A production payment workflow also needs a recovery strategy for the gap between persisting local state and receiving the gateway result.
For example, the module can persist a Pending payment before calling the gateway, use the same idempotency key with a provider that explicitly supports idempotent requests, and reconcile unresolved Pending records asynchronously.
That distinction matters because an idempotency key does not make a distributed transaction atomic. It gives the system a way to safely recognize repeated attempts when the downstream provider guarantees the semantics.
The caller depends only on IPayments. It does not know how Payments stores cards, talks to a provider, or models its internal entities:
// Orders depends on the contract, not the Payments implementation.
public sealed class PlaceOrderHandler(IPayments payments)
{
// ... call payments.ChargeAsync(...)
}A small architecture test can make the dependency rule visible:
[Fact]
public void Orders_does_not_reference_Payments_internals()
{
var references = typeof(OrdersModule)
.Assembly
.GetReferencedAssemblies()
.Select(x => x.Name);
Assert.DoesNotContain("MyShop.Payments", references);
}That test is intentionally small, not a complete architecture-checking framework. For richer rules — cycles, namespace access, allowed dependencies and package-level constraints — use tools such as ArchUnit, NetArchTest or framework-specific verification such as Spring Modulith.
Extraction day
Now suppose Payments develops a real forcing function. Maybe it needs to scale independently, needs stronger security isolation, or needs a different runtime.
A modular boundary makes the migration smaller because the caller already depends on an abstraction. But one important detail changes when crossing a process boundary: do not assume the exact same source-level interface or assembly should be shared by both applications. The in-process abstraction is an implementation detail; the HTTP or messaging contract needs versioning and compatibility rules of its own.
The caller can keep its local abstraction:
public sealed class HttpPayments(HttpClient http) : IPayments
{
public async Task<PaymentResult> ChargeAsync(
ChargeRequest request,
CancellationToken ct)
{
using var response = await http.PostAsJsonAsync(
"/v1/payments/charges",
request,
ct);
response.EnsureSuccessStatusCode();
return (await response.Content
.ReadFromJsonAsync<PaymentResult>(ct))!;
}
}And the new service can host the Payments implementation independently:
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddPaymentsModule(
builder.Configuration.GetConnectionString("Payments")!);
var app = builder.Build();
app.MapPost(
"/v1/payments/charges",
async (
ChargeRequest request,
IPayments api,
CancellationToken ct) =>
Results.Ok(await api.ChargeAsync(request, ct)));
app.Run();There is a subtle but important retry issue here. Current .NET guidance says the standard resilience handler retries transient failures for all HTTP methods by default, and Microsoft explicitly warns that retrying methods such as POST can be harmful because the operation may be performed more than once. It provides DisableForUnsafeHttpMethods() for cases where such retries are not appropriate. Microsoft Learn — Build resilient HTTP apps
For a payment endpoint, do not blindly add a generic retry policy and call it safe because the request contains an idempotency key. Either disable automatic POST retries or enable them only when the payment provider documents the idempotency semantics you are relying on.
For example:
builder.Services
.AddHttpClient<IPayments, HttpPayments>(client =>
{
client.BaseAddress = new Uri(paymentsUrl);
})
.AddStandardResilienceHandler(options =>
{
options.Retry.DisableForUnsafeHttpMethods();
});Extraction still is not free. Once the network appears, the compiler cannot protect you from timeouts, partial failure, duplicate delivery, incompatible deployments, or a service that is temporarily unavailable. The Payments data also has to move behind its new ownership boundary, and the service needs its own deployment, monitoring and operational lifecycle.
The point of the modular monolith is not that extraction will be easy. It is that the expensive distributed-system work starts when there is a reason to pay for it.
Side by side
Modular monolith | Microservices | |
|---|---|---|
Deployment | One deployment unit | Many independently deployable units |
Calls between modules | Usually in-process | Network calls with latency and partial failure |
Scaling | Usually scale the application as a whole | Scale services independently |
Data ownership | Can share one physical database, but modules should own their data | Each service typically owns its data and exposes access through APIs/events |
Changing a boundary | Refactoring within one codebase | May require API/versioning, data migration and deployment coordination |
Boundary enforcement | Tooling and discipline | Process and network boundaries add stronger isolation, but do not remove all coupling |
Team autonomy | Less deployment independence | Can support independent team deployment when boundaries are stable |
Operational load | Lower by default | Higher by default as service count grows |
Technology choice | Usually easier to keep one stack | Independent technology choices are possible per service |
Failure model | Failures are often local to a process, except for external dependencies | Partial failure is a normal architectural concern |
Neither column is a guarantee. A badly designed modular monolith can be tightly coupled, and a microservice architecture can become a distributed monolith if services still have to change and deploy together.
The decision rule
The two camps actually agree on more than the debate suggests.
Fowler argues that the microservice premium becomes worthwhile when system complexity is high enough that a monolith is becoming difficult to manage. He also points out that successful microservice stories he had seen often began with a monolith that grew too large, and that service boundaries are difficult to get right early. Fowler — Monolith First
Sam Newman similarly argues that microservices should not be the default and that the goal of decomposition is independent deployability rather than decomposition for its own sake. InfoQ — Decomposing a Monolith Does Not Require Microservices
Stefan Tilkov's counterargument is worth keeping: the network boundary can protect teams from coupling in a way that an internal convention cannot. That is a real benefit, especially for organizations where internal boundaries repeatedly erode. The disagreement is less about whether boundaries matter and more about whether you need a process boundary to make yours stick. Stefan Tilkov — Don't start with a monolith
A practical decision rule follows.
1. Do you have a forcing function?
Good reasons to split one capability out include:
Forcing function | Typical reason |
|---|---|
Independent deployment | Teams need to release without coordinating every change. |
Different scaling profile | One capability has very different traffic or resource needs. |
Security or compliance isolation | A capability must be isolated because of sensitive data or stricter controls. |
Different technology requirement | Another runtime or storage technology provides a meaningful advantage. |
These are reasons to consider a boundary, not automatic proof that a service is the answer.
2. Can you enforce the boundary before you split it?
If you stay in one deployable, can you make the build fail when one module reaches into another module's internals? Can you prevent direct database access? Can you make ownership visible? Can you keep the dependency graph from becoming a cycle?
Shopify's experience is a useful warning here: public interfaces without a healthy dependency graph can simply add another layer of indirection to an already tangled system. Shopify — Under Deconstruction: The State of Shopify's Monolith
3. If one capability needs separation, extract that capability — not the whole application
A forcing function for Payments does not automatically imply separate services for Orders, Catalog, Search, Notifications and Users. Split the part that has a reason to be independent and leave the rest together until they develop their own reasons.
This is also where incremental patterns such as branching by abstraction and the strangler-fig approach become useful: the migration can happen one boundary at a time instead of becoming a rewrite. Newman describes both techniques in his decomposition guidance. InfoQ — Decomposing a Monolith Does Not Require Microservices
The decision in one sentence
No forcing function + strong boundary discipline → modular monolith.
Clear forcing function → extract the affected capability.
No forcing function + no boundary discipline → fix the design first.
A deployment style cannot rescue a system whose boundaries are already broken.
Common mistakes
Mistake | What goes wrong | Better approach |
|---|---|---|
Splitting by technical layer | A “data service” and a “UI service” make every feature cross a network boundary. | Split around business capabilities and change locality. |
Building a distributed monolith | Services still share databases or must deploy together. | Aim for real independent ownership and deployment before multiplying services. |
Reaching across module data | One module queries another module's tables directly. | Treat data ownership as part of the boundary. |
Boundaries on paper | Interfaces exist, but dependencies still form cycles. | Enforce dependency direction in the build or CI. |
Too many modules too soon | Developers have a harder time understanding the system than they did before. | Start with coarse, high-cohesion boundaries and split when change patterns justify it. |
Chatty APIs | Many tiny calls are awkward in-process and painful over the network. | Prefer coarse operations, plain data and contracts designed for remote use. |
Blind retries | A retry turns an uncertain outcome into a duplicate side effect. | Design idempotency explicitly and retry only when the downstream contract makes it safe. |
A quick sanity check
Before choosing microservices, ask:
Can I name the owner, responsibility and public API of each module?
Can I change one module's internals without changing its callers?
Does each module treat other modules' data as private?
Would CI fail if someone crossed an internal boundary?
Do most changes stay within one or two modules?
If one module had to become a service next year, could I describe the migration as an adapter, an endpoint, a data move and a deployment — rather than a rewrite?
Mostly yes means you have boundaries worth preserving, regardless of whether they remain in-process.
The final question is the useful one. It turns “we might need microservices someday” into an engineering test: how expensive would it actually be to separate this capability?
The point is not the monolith. It is the boundary.
A modular monolith is not a promise that you will never use microservices. Microservices are not a badge that a system has matured.
The architectural decision that matters first is whether a capability has a boundary that is coherent, owned and enforceable. A process boundary can strengthen that separation when the business or operational problem justifies it. But a network should be the consequence of a reason, not the substitute for one.
Draw the boundaries first.
Then decide whether the walls should be made of namespaces and build rules or of processes and network calls.
Sources
Primary and first-party sources
Martin Fowler and James Lewis, Microservices — March 2014.
Martin Fowler, Monolith First — June 2015.
Martin Fowler, Microservice Premium — May 2015.
Martin Fowler, Microservice Trade-Offs — 2015.
Stefan Tilkov, Don't start with a monolith — 2015.
Alexandra Noonan, Goodbye Microservices: From 100s of problem children to 1 superstar — Twilio Segment, July 2018.
Philip Müller, Under Deconstruction: The State of Shopify's Monolith — Shopify, September 2020.
Prime Video: Scaling up the audio/video monitoring service and reducing costs by 90% — Prime Video Tech, 2023. The original page may redirect; see the InfoQ summary below.
Spring Modulith — Verifying Application Module Structure — Spring, current documentation.
Kamil Grzybek, Modular Monolith: Integration Styles.
Kamil Grzybek, Modular Monolith: Domain-Centric Design.
Brian Foote and Joseph Yoder, Big Ball of Mud — PLoP '97.
Supporting sources
Thomas Betts, Decomposing a Monolith Does Not Require Microservices — Sam Newman at QCon London — InfoQ, May 2020.
Rafał Gancarz, Prime Video Switched from Serverless to EC2 and ECS to Save Costs — InfoQ, May 2023.
Microsoft, Build resilient HTTP apps: Key development patterns — current .NET guidance for
AddStandardResilienceHandlerand HTTP retries.
Source note
The Segment, Shopify and Prime Video examples are historical engineering reports from 2018–2023. They are useful evidence about the trade-offs those teams encountered, not universal benchmarks or proof that one architecture is always cheaper, faster or more reliable.
Comments (0)
Join the discussion by logging into your account.
No comments yet. Be the first to comment!