There’s a quiet fracture running through most software teams, and it’s not about tabs versus spaces. It’s the difference between engineering for scale and engineering for understanding. The first mode optimizes for throughput, uptime, and resource efficiency under load. The second optimizes for the speed at which a new developer—or your future self—can form an accurate mental model of the system, locate a bug, or extend a feature without breaking adjacent logic. Both are legitimate engineering disciplines. The damage happens when a team applies the tools, rituals, and abstractions of scale-engineering to a problem that is fundamentally a crisis of understanding, or vice versa. This article maps the boundary between these two modes, shows how their design tradeoffs diverge, and offers concrete heuristics for choosing which one a given piece of work actually needs.
I’ve seen a startup spend three months building a Kafka-based event sourcing pipeline for an internal admin panel used by four people. I’ve also seen a payments platform handle a Black Friday peak by relying on a single-threaded Python service whose primary design virtue was that a junior engineer could read the entire codebase in an afternoon. Both teams were smart. Both had misdiagnosed their problem. The first team was engineering for scale when they needed understanding. The second was engineering for understanding when they needed scale. The scars from those misdiagnoses are what this article is about.

The Two Modes, Defined Through Their Constraints
Engineering for scale is a response to non-functional requirements: throughput, latency, availability, fault tolerance, and cost per operation. Its primary constraint is the physics of distributed systems. The unit of progress is a service-level objective (SLO), and the primary design tension is between consistency and availability under partition. The literature is vast—Google’s SRE book, the Dynamo paper, the Tail at Scale—and the patterns are well-catalogued: sharding, backpressure, circuit breakers, leader election, CRDTs.
Engineering for understanding is a response to a different constraint: the finite working memory of the human brain. The unit of progress is the time it takes a developer to answer “what happens when I change this line?” with confidence. The primary design tension is between local reasoning and global efficiency. The literature is thinner but no less important: The Programmer’s Brain by Felienne Hermans, the cognitive dimensions framework, and decades of empirical research on code comprehension. The patterns include information hiding, the principle of least astonishment, and the rule of explicit invariants.
The confusion arises because both modes use the same vocabulary—modularity, decoupling, abstraction—but mean radically different things by them. In scale-engineering, decoupling means independent failure domains and asynchronous boundaries. In understanding-engineering, decoupling means you can reason about a module without loading the entire program into your head. These two definitions sometimes align. Often they do not.
When Scale Abstractions Become Comprehension Liabilities
Consider a team that adopts a microservice architecture. The stated goal is scalability: independent deployment, independent scaling, fault isolation. The team carves the monolith into a dozen services, each with its own repository, CI/CD pipeline, and data store. The system now scales beautifully. It also takes a new engineer three weeks to set up a local development environment, and debugging a single request requires correlating logs across six services with different logging formats and clock skews. The team has gained operational scale at the cost of cognitive scale—the ability of a human mind to hold the system in working memory.
This tradeoff is not inherently wrong. If the system genuinely serves millions of requests per second, the cognitive cost may be worth paying. But I’ve seen this pattern applied to internal tools with a dozen users, where the real bottleneck was never throughput but the fact that nobody fully understood the data model. The team was optimizing a constraint that did not bind, while ignoring the one that did.
A concrete example: a content management system I worked on used an event-sourced architecture with a custom projection engine. The write path was elegantly scalable—append-only events, async projections, CQRS. But when an editor asked “why is this article showing the wrong author?”, tracing the bug meant reconstructing the event stream, understanding the projection logic, and accounting for eventual consistency. The answer took two days. The system served 200 requests per day. The architecture was a textbook example of scale-engineering applied to a problem whose binding constraint was debuggability.
Abstractions That Serve the Reader, Not Just the Machine
Engineering for understanding does not mean avoiding abstraction. It means choosing abstractions that reduce the working-set size of the developer’s brain. A well-designed module for understanding has a small interface surface, predictable behavior, and a clear mapping between the domain concept and the code. When something goes wrong, the abstraction fails in a way that points toward the fault, not away from it.
Take the example of a simple repository pattern in a web application. From a scale perspective, wrapping database queries in a repository interface might seem like unnecessary indirection—a thin veneer over SQL that adds latency. But from an understanding perspective, that interface is a contract. It tells the reader: “Here is where persistence happens. Everything you need to know about data access is behind this boundary.” When a query is slow, you know exactly which file to open. When a new developer joins, they can learn the system one repository at a time. The abstraction is not for the machine; it is for the human.
This is why I’m skeptical of frameworks that collapse these boundaries in the name of “simplicity.” An ORM that auto-generates migrations from model classes saves keystrokes but obscures the database schema. A framework that auto-wires dependencies saves configuration but makes it impossible to trace the object graph without running the application. These tools optimize for the first hour of development at the expense of the thousandth hour of maintenance. They are scale-engineering tools masquerading as understanding-engineering tools.

The RFC as a Diagnostic Tool
One of the most underused instruments for distinguishing between these two modes is the internal RFC (Request for Comments) process. When done well, an RFC is not a design document; it is a decision record that makes tradeoffs explicit. A good RFC for a new service should answer two questions separately:
- What are the scale constraints this design addresses? (Throughput, latency, data volume, availability targets.)
- What are the understanding constraints this design addresses? (Team size, onboarding time, expected reader-to-writer ratio, debugging surface area.)
If the second question gets a hand-wavy answer or is treated as a subset of the first, the design is almost certainly over-indexed on scale. I’ve started asking teams to include a “Day 2 Debugging Scenario” in their RFCs: a concrete walkthrough of how an on-call engineer would diagnose a specific failure mode in the proposed system. This exercise often reveals that a beautifully scalable architecture is a nightmare to troubleshoot, and it forces the team to either invest in observability tooling or simplify the design.
The RFC process itself is a tool for understanding. A well-structured RFC, reviewed by peers who ask “how would I debug this?” rather than just “does this handle 10x load?”, acts as a forcing function for cognitive empathy. It is the difference between designing a system that works and designing a system that can be understood to work.
Code Structure as a Signal of Intent
You can often diagnose which mode a team is operating in by looking at their code structure—not the architecture diagram, but the actual directory tree and import graph. Scale-oriented codebases tend to organize around infrastructure concerns: handlers, producers, consumers, repositories, connectors. Understanding-oriented codebases tend to organize around domain concepts: orders, inventory, pricing, shipping.
Here is a heuristic I use when reviewing a new service. Open the top-level directory. If you see folders named controllers/, services/, models/, and utils/, the team is likely thinking in terms of layers—a scale pattern. If you see folders named orders/, inventory/, and pricing/, the team is likely thinking in terms of domain boundaries—an understanding pattern. Neither is universally correct, but the choice should be deliberate. A layered structure makes it easy to apply cross-cutting concerns (logging, auth, rate limiting) but scatters domain logic across the codebase. A domain structure co-locates related logic but can make cross-cutting changes tedious. The question is: which type of change does this system experience most often?
I once inherited a codebase for a billing system that was organized by layer. The “services” folder contained a single 4,000-line file called BillingService.java. The team had chosen a layered structure out of habit, but the system’s primary change vector was domain-specific: new pricing rules, new invoice formats, new tax calculations. Every change touched that monolithic service file. Reorganizing around domain concepts—TaxEngine, InvoiceGenerator, PricingRules—reduced the average change size by 60% and cut onboarding time from weeks to days. The system did not get faster. It got more legible.
Testing Strategies Diverge
The two modes also demand different testing strategies. Scale-engineering emphasizes resilience testing: chaos experiments, load tests, soak tests, and fault injection. The goal is to verify that the system degrades gracefully under stress. Understanding-engineering emphasizes specification testing: characterization tests, property-based tests, and tests that serve as executable documentation. The goal is to verify that the system behaves as the developer expects, and that the tests themselves communicate intent.
A property-based test for a JSON parser might assert that parse(serialize(x)) == x for any valid input. This test is not about throughput; it is about invariant preservation. It tells a future maintainer: “This is a fundamental truth about the parser. If your change breaks this, you have violated the contract.” A load test tells you nothing about the contract. It tells you about the system’s behavior under stress. Both are valuable, but they answer different questions. Confusing them leads to test suites that are slow, flaky, and uninformative—the worst of both worlds.

When the Two Modes Collide: The Case of API Design
API design is the arena where the tension between scale and understanding becomes most visible. A REST API designed for scale might use sparse fields, cursor-based pagination, and conditional requests to minimize payload size and server load. An API designed for understanding might use rich, self-describing responses, simple offset pagination, and consistent resource shapes that mirror the domain model. The scale-optimized API is a joy for a high-throughput mobile client on a flaky network. The understanding-optimized API is a joy for a developer exploring the system with cURL or Postman.
The mistake is assuming you must choose one and apply it uniformly. A more careful approach is to version the API surface by consumer: an internal “debug” endpoint that returns verbose, self-describing payloads, and a production endpoint that strips optional fields and uses compact representations. GraphQL is an interesting case here: it shifts the burden of understanding from the server to the client by letting the client specify exactly what data it needs. This is a scale win for the server but can be a comprehension loss for developers who now need to understand the entire graph to write a single query. The tool is neutral; the context determines whether it helps or hurts.
Heuristics for Choosing Your Mode
Here are the questions I ask when a team is debating whether to invest in scale-engineering or understanding-engineering for a given system or feature:
- What is the binding constraint? If the system is already struggling under load, scale-engineering is non-negotiable. If the system is stable but nobody on the team can explain how it works, understanding-engineering is the priority.
- What is the reader-to-writer ratio? Code that is read 100 times more often than it is written should be optimized for reading. This includes most internal libraries, shared modules, and core business logic.
- What is the cost of a misunderstanding? In a payments system, a misunderstood invariant can cost millions. In a social media feed, it might cost a few annoyed users. The higher the cost, the more you should bias toward understanding.
- Can you afford the cognitive overhead? Distributed systems introduce inherent complexity. If your team is small and your domain is deep, adding infrastructure complexity may exceed the team’s cognitive budget.
- Is the abstraction paying rent? Every abstraction should either improve scale or improve understanding. If it does neither, delete it. If it improves one at the expense of the other, make that tradeoff explicit.
FAQ
Can a system be engineered for both scale and understanding simultaneously?
Yes, but it requires deliberate effort and often a layered approach. The core domain logic can be written for understanding—clear, well-documented, with strong invariants—while the infrastructure layer around it handles scale concerns like replication, sharding, and backpressure. The key is to keep the boundary between these layers explicit and to avoid letting scale abstractions leak into the domain code. When scale concerns infect the domain logic, you get code that is neither scalable nor understandable.
How do I convince my team to invest in understanding when we are under pressure to ship features?
Frame it as a velocity argument, not a quality argument. Show concrete examples where poor understanding caused delays: a bug that took days to diagnose, an onboarding that took weeks, a refactor that broke unrelated features. Measure the time spent on these activities and compare it to the time that would have been spent on clearer code, better tests, or better documentation. Understanding-engineering is not about slowing down to be “clean”; it is about removing friction from the development process. When the team sees that friction as a tax on their own productivity, the investment becomes easier to justify.
What is the single highest-impact practice for improving understanding in a codebase?
In my experience, it is explicit invariant documentation. For every module, class, or function that contains non-trivial logic, document the invariants it maintains and the preconditions it expects. This can be as simple as a comment at the top of a file: “All Order objects in this module have a non-null customerId and at least one line item.” When these invariants are explicit, a developer reading the code can verify them locally rather than tracing through the entire system. When they are violated, the failure is close to the cause. This practice alone has saved me more debugging hours than any other single technique.
Is microservices an anti-pattern for understanding?
Not inherently, but the default microservices playbook often ignores understanding concerns. A microservice that is truly independent—with its own data store, its own team, and a stable, well-documented API—can be easier to understand than a monolith, because the cognitive surface area is smaller. The problem arises when services are split along technical boundaries rather than domain boundaries, or when the inter-service communication patterns are so complex that understanding any single service requires understanding the entire mesh. The heuristic is: if you cannot draw the system’s data flow on a whiteboard from memory, it is too complex to understand, regardless of how well it scales.