Engineering for Scale vs. Engineering for Understanding: A False Dichotomy
We have a bad habit in this industry of treating “scale” as the only engineering virtue that matters. Systems that handle millions of requests per second, databases that replicate across continents, architectures that survive entire data-center outages—these are the war stories we tell at conferences. But there’s another kind of scale that gets far less attention: the scale of understanding. It’s the ability of a codebase, an API, or a system design to be grasped by the next developer, the new team member, or even your future self at 3 a.m. when a pager screams. We often frame engineering for scale and engineering for understanding as opposing forces, a trade-off you have to make. I think that framing isn’t just wrong—it’s actively harmful. The real craft is designing systems that are both operationally scalable and cognitively accessible. When you sacrifice understandability on the altar of throughput, you’re not building a scalable system. You’re building a fragile monument to cleverness that crumbles the moment the original authors leave the room.

The Sociotechnical Gap: Where Systems Meet Minds
Mark Ackerman coined “sociotechnical gap” to describe the divide between the social needs of users and the technical capabilities of systems. I’m going to repurpose it here for a related problem: the gap between the mental models of the engineers who build a system and the actual runtime behavior of that system. When we engineer for scale, we often introduce abstractions—caches, queues, eventual consistency, sharding strategies—that widen this gap. The system in production no longer behaves like the code on your laptop. The sociotechnical gap becomes a cognitive gap, and every incident turns into an archaeology expedition into layers of undocumented assumptions.
Consider a typical microservices architecture. The promise is independent deployability and fault isolation. The reality, often, is a distributed monolith where understanding a single user request requires tracing a call graph across seventeen services, three message brokers, and a Redis cluster that someone added to fix a performance problem six months ago. The system scales operationally—you can spin up more instances—but it has failed to scale cognitively. The number of engineers who truly understand the end-to-end flow can be counted on one hand. That’s a bus-factor problem, a hiring problem, and an incident-response problem all rolled into one.
Concrete Symptoms of a Cognitive Scaling Failure
Before we can fix the problem, we need to recognize it. Here are the signs I look for, drawn from years of untangling systems that grew faster than their documentation:
- The “Who Owns This?” Loop: A production issue arises, and the incident channel cycles through five teams because no one is sure which service is the root cause. Each team points to the next, citing API contracts that are technically correct but semantically ambiguous.
- Configuration as a Secret Language: The system’s behavior is governed not by its source code, but by a labyrinth of YAML files, feature flags, and environment variables. Changing a timeout value in one place causes a cascading failure in a seemingly unrelated component. The configuration is the real source code, and it has no tests.
- Onboarding as a Multi-Month Ordeal: New engineers are told it takes six months to become productive. That’s not a sign of a deep domain; it’s a sign of a system that has externalized its complexity onto the humans who maintain it. The domain might be complex, but the system should encapsulate that complexity, not radiate it.
- RFCs as Fiction: The original design documents describe a clean, layered architecture. The current system has accreted so many tactical fixes that the RFCs are now historical artifacts, not living documents. The gap between the documented intent and the running system is a breeding ground for misunderstandings.

API Design as a Cognitive Interface
If there’s one place where the tension between scale and understanding is most visible, it’s in API design. An API isn’t just a contract between services; it’s a user interface for developers. A well-designed API reduces the cognitive load on the caller. A poorly designed one forces the caller to internalize the implementation details of the service behind it.
Take pagination as an example. Offset-based pagination (?page=2&limit=50) is simple to implement and understand. It scales poorly for large datasets because the database has to scan and discard rows. Cursor-based pagination (?cursor=eyJsYXN0X2lkIjoxMDAwfQ==) scales beautifully but forces the client to manage opaque tokens. The “engineering for scale” choice is cursor-based. The “engineering for understanding” choice is offset-based. The right choice is to provide a cursor-based API with a clear, documented rationale, and to include a migration path or a compatibility layer that lets simple clients use offset-based access for small result sets. You don’t have to choose; you have to design.
Another example: error responses. A service under load might start returning 503 Service Unavailable with a Retry-After header. That’s a scale-oriented design. But if the error body is an empty JSON object, you’ve failed the developer who needs to debug a client issue. A cognitively scaled API includes a machine-readable error code, a human-readable message, and a link to the relevant documentation. The extra bytes are negligible; the reduction in support tickets and debugging time is not.
RFC 7807 and the Art of Telling Developers What Went Wrong
RFC 7807 defines a standard format for problem details in HTTP APIs. It suggests fields like type, title, detail, and instance. Adopting this is a trivial engineering effort. The payoff is that any developer who has seen one RFC 7807 response can immediately understand an error from any service that uses it. That’s cognitive scale: a pattern that makes the whole ecosystem more understandable, not just one service faster. I’ve seen teams resist this because “our errors are unique” or “we need a custom format for our tooling.” That’s the siren song of local optimization. The global optimum—the one that scales understanding across teams and over time—is consistency with well-known standards.
Architecture Diagrams as Rhetorical Devices
I’m a firm believer that an architecture diagram is not a technical artifact; it’s a rhetorical device. Its purpose is to tell a story about the system, to guide the viewer’s attention to the relationships that matter. A diagram that shows every microservice, every database, every queue, and every bidirectional arrow is a map of the territory at 1:1 scale—useless. A diagram that abstracts away the details to show the flow of a single critical request, or the trust boundaries between domains, is a tool for understanding.
When I review system designs, I ask for two diagrams. The first is the “scale diagram”: the physical topology, the instance counts, the network zones. The second is the “understanding diagram”: the logical flow of a key business operation, annotated with the mental model a developer should hold. If the second diagram is too complex to draw, the system is too complex to operate. This isn’t a soft skill; it’s a hard constraint. A system you can’t diagram is a system you can’t debug under pressure.

Heuristics for Closing the Gap
I don’t believe in universal laws of software engineering. Context is everything. But I do believe in heuristics—rules of thumb that, when applied thoughtfully, tilt the odds in your favor. Here are the ones I use when I want to ensure a system scales not just in throughput, but in comprehension.
1. The New Hire Test
Can a new engineer—not a genius, just a competent one—understand the core request flow within their first week? If not, the system’s cognitive load is too high. This doesn’t mean the system must be simple; it means the interface to the system’s complexity must be simple. The internals can be as complex as necessary, but the entry points should be obvious, the error messages clear, and the documentation a map, not a maze.
2. The “One Hour to Yes” Rule for APIs
When you design an API, the goal should be that a developer can go from zero to a successful, meaningful API call in under an hour. This includes reading the docs, getting credentials, and making the request. If it takes longer, your API isn’t self-service; it’s a gatekeeper. Stripe’s API documentation is the gold standard here. Their quickstart gets you charging a credit card in minutes, not because payments are simple, but because they invested heavily in the developer experience layer.
3. Prefer Standards with Network Effects
Every time you invent a custom authentication scheme, a bespoke serialization format, or a unique retry strategy, you’re taxing the understanding of every future developer who touches your system. Use OAuth 2.0, even if it feels heavy. Use JSON, even if Protobuf is faster. Use exponential backoff with jitter, as described in the AWS Architecture Blog. The cognitive cost of a custom solution is almost always higher than the performance cost of a standard one, unless you’re operating at a scale that very few organizations actually reach.
4. Make the Implicit Explicit
Every system has implicit assumptions: “this service is never called more than 100 times per second,” “this cache key is always populated before the request arrives,” “this database column is never null.” When these assumptions are violated—and they will be—the system fails in ways that are baffling to anyone who didn’t write the original code. My rule: if an assumption is load-bearing, it must be explicit. That means assertions in code, checks in CI/CD pipelines, and documentation that’s generated from the code, not written separately and left to rot.
FAQ: Scaling Understanding in Practice
Isn’t this just “write better documentation”?
No. Documentation is a symptom, not a cure. If your system requires a novel’s worth of documentation to be understood, the system itself is the problem. Good documentation explains the why and the what, but the how should be evident from the code and the API design. The best documentation is the one you don’t need to read because the system’s behavior is predictable and consistent. Aim for that, and then write documentation for the edge cases and the rationale.
How do you convince a performance-obsessed team to care about understandability?
Speak their language: measure it. Track metrics like “time to first meaningful commit” for new engineers, “mean time to resolve” for incidents, and the number of services touched per incident. When you can show that a tangled architecture is directly increasing downtime and slowing down feature delivery, you’ve made the cost of incomprehension visible. Performance engineers respect data. Give them data about cognitive performance.
Doesn’t engineering for understanding slow down initial development?
Yes, sometimes. Writing clear error messages, designing consistent APIs, and keeping diagrams up to date takes time. But the trade-off isn’t between speed now and speed later; it’s between speed now and sustained speed over the lifetime of the system. A system that’s hard to understand accumulates technical debt at a faster rate. Every hack added because someone didn’t understand the existing code makes the next change even harder. The initial investment in clarity pays compound interest. The initial rush to ship pays a debt that grows exponentially.
What is the single biggest mistake teams make when trying to scale?
They optimize for the machine’s constraints instead of the human’s constraints. They worry about milliseconds of latency and megabytes of memory while creating systems that take weeks to debug. The machine’s constraints are well-understood and easy to measure. The human’s constraints—working memory, attention, the need for mental models—are just as real, but they’re invisible in your monitoring dashboards. The biggest mistake is pretending they don’t exist.
Closing the Loop
Engineering for scale and engineering for understanding aren’t opposing forces. They’re two dimensions of the same problem: building systems that work reliably over time. A system that’s fast but incomprehensible will eventually become slow, because no one will dare to optimize it. A system that’s understandable but can’t handle load will frustrate users. The craft is in finding the designs that satisfy both dimensions—and in recognizing that, in most cases, the bottleneck isn’t the CPU. It’s the developer staring at the screen, trying to figure out what the hell is going on.
Next in this series: “The API Contract as a Social Contract”—why your OpenAPI spec is a promise, not just a document.