There’s a quiet, persistent tension in the way we build software. It sits between the machine’s appetite for throughput and the developer’s need to keep a coherent mental model of what the code actually does. We call one “engineering for scale” and the other “engineering for understanding.” They pull in opposite directions more often than we like to admit. A system can hum along at ten thousand requests per second, perfectly load-balanced and elegantly sharded, yet still be a black box to the people who maintain it. When the internal model drifts too far from the business domain, you don’t have a platform—you have a liability that ships changes at a crawl.

Abstract visualization of interconnected nodes representing system architecture
Complexity grows silently until it becomes a barrier to understanding.

Two Axes of System Quality

Scale engineering is obsessed with numbers: requests per second, p99 latency, throughput under peak load. It’s a discipline with a rich toolset—load balancers, sharding strategies, backpressure protocols like those described in RFC 9419. These are measurable, optimizable, and deeply satisfying to tune. Understanding, on the other hand, is squishier. It lives in the distance between a user story and the code that fulfills it, in the clarity of module boundaries, in how long it takes a new hire to trace a single request path without wanting to quit. When a business concept like “customer credit limit” gets smeared across three microservices, a Redis cache, and a stored procedure, the system’s cognitive load spikes. You’ve built a distributed puzzle that only a handful of people can solve—and even they forget the solution after a few months away.

When Scale Tactics Scatter the Domain

Take a straightforward payment flow. In a monolithic codebase, it might be a single function: PaymentResult process(PaymentIntent intent). You can read it, test it, and reason about its edge cases in an afternoon. But to hit five-nines availability and handle ten thousand payments a second, you decompose it. Now you have a PaymentRequested event on a Kafka topic, a fraud check worker, a payment execution worker, a reconciliation job that cross-references gateway logs, and a dead-letter queue for the stragglers. Each piece scales independently. Each piece is a small miracle of resilience engineering. But the original concept—“process a payment”—has evaporated. It’s an emergent property of six components, none of which tell the whole story. A developer asking “what happens when a payment fails?” has to spelunk through multiple repositories, internalize at-least-once delivery semantics, and hope the dead-letter handler is actually wired up. The system scales. It also resists comprehension.

Developer sketching domain boundaries on a whiteboard
Whiteboard sessions remain one of the best tools for aligning on domain boundaries.

Understanding as a First-Class Requirement

We don’t hesitate to set SLOs for latency or uptime. Why not for understandability? One practical proxy is time-to-first-meaningful-contribution: how many days before a new team member can ship a small, safe change to a core business rule? If the answer is measured in weeks, your system has a comprehension deficit. Another signal is the number of modules touched per user story. A tax calculation update that ripples through three services and a shared library is a red flag. The business rule is scattered, and every future change will be a game of whack-a-mole.

Domain-Driven Design’s bounded contexts offer a way out. When a service boundary aligns with a domain boundary—a “Payment Context” rather than a “Database Write Service”—the mapping from user need to code location stays intuitive. The system’s structure tells a story that matches the business narrative. That’s not just aesthetics; it’s a maintenance survival strategy.

The Modular Monolith: A Sensible Default

Before you reach for Kubernetes and a message broker, consider the modular monolith. It’s a single deployable with well-enforced internal boundaries. Modules communicate through interfaces, not direct database queries. You get the encapsulation benefits of services without the network debugging nightmares. Shopify has written candidly about this: they use Packwerk to enforce module boundaries and only extract a service when the operational need is undeniable. Extraction is a last resort, not a rite of passage. The monolith gives you fast feedback, simple transactions, and a codebase you can actually navigate. When the domain is still shifting under your feet, that’s worth more than hypothetical scalability.

When Distribution Is Earned

Some problems demand distribution. A global CDN, a real-time bidding platform, a telemetry ingestion pipeline—these aren’t vanity projects. The scale is real, and the architecture must match. The danger zone is aspirational distribution: teams adopting microservices because they hope for Netflix traffic, not because they have it. The result is often a distributed monolith—services so tightly coupled that independent deployment is a fantasy, but you still pay the debugging tax of network hops and eventual consistency.

When distribution is earned, the challenge becomes preserving understanding despite the sprawl. A few tactics that help:

  • Observability that tells a business story. Distributed traces named after user journeys—“PlaceOrder,” not “handleRequest”—let developers follow a narrative through the system.
  • Schema-first design. A versioned schema registry becomes the source of truth when the code is too fragmented to read. Your events and APIs are the contract; the implementation is secondary.
  • Domain-aligned service boundaries. Services should mirror bounded contexts, not technical layers. A “Payment Service” makes sense. A “Database Write Service” obscures more than it reveals.
Close-up of tangled network cables in a server rack
Distributed systems can tangle domain logic as thoroughly as physical cables.

Heuristics for Keeping Systems Comprehensible

Here are five rules of thumb I lean on during design reviews. They’re not academic—they’ve been forged in the mess of real codebases.

  1. The New Developer Test. Can someone new to the team explain a core business flow by reading the code in one afternoon? If they can’t, the structure is hiding the domain.
  2. The Single Change Point Rule. A business rule change should require a change in exactly one place. If updating a tax calculation touches three services, the rule is scattered.
  3. Event Storming Before Event Sourcing. Model events on a whiteboard with domain experts before you encode them in Kafka. The events should reflect what the business cares about, not just what the database emits.
  4. Scale In, Not Out. Before splitting a service, ask: can we scale vertically or optimize data access? Premature distribution is the root of much accidental complexity.
  5. Documentation as Code. Architecture decision records and README files that explain the why behind a design choice prevent future developers from cargo-culting a pattern that no longer applies.

FAQ

What is the main difference between engineering for scale and engineering for understanding?

Engineering for scale focuses on quantitative system properties like throughput, latency, and fault tolerance. Engineering for understanding focuses on qualitative properties like code clarity, domain alignment, and the cognitive load required to maintain and extend the system. The two often conflict because patterns that improve scalability—such as event sourcing and microservices—can scatter business logic across many components.

How can I tell if my system has an understanding problem?

Signs include: new team members taking weeks to make simple changes, frequent misunderstandings about system behavior during incidents, and business rules duplicated across multiple services. A practical test is to ask a developer to trace a single user request end-to-end; if they need to consult five different codebases and three message queues, understanding has been sacrificed.

When should I choose a modular monolith over microservices?

Choose a modular monolith when your scaling bottlenecks are not yet at the level that requires independent deployment of components, when your team is small enough to manage a single codebase, and when the domain is still evolving rapidly. The modular monolith preserves the ability to extract services later while keeping the cognitive overhead low during early development.

Can you measure understandability like you measure latency?

There is no single metric as precise as p99 latency, but you can use proxies: time to onboard a new developer, number of modules touched per user story, and the ratio of business-logic changes to total changes. Teams can also run regular “architecture katas” where they trace a hypothetical change through the system and measure how many components are affected.