I’ve spent the last decade bouncing between two extremes: building systems that had to survive millions of requests per second, and untangling codebases that had been prematurely optimized for a scale they never reached. The scars from both are real. But the deeper wound—the one that keeps me up at night—is watching teams apply the patterns of hyperscale engineering to problems that desperately need something else: engineering for understanding.
This isn’t a manifesto against performance. It’s a careful dissection of a false dichotomy. When we talk about “scale,” we usually mean throughput, latency, and resource utilization. When I talk about “understanding,” I mean the cognitive load required for a developer to hold the system’s structure, state, and data flow in their head accurately enough to make a safe change. The gap between these two goals is where most production incidents are born.

The Two Architectures: A Side-by-Side Dissection
Let’s get concrete. Engineering for scale typically optimizes for resource efficiency and fault tolerance. You reach for patterns like event sourcing, CQRS, microservices, and asynchronous message passing. You introduce Kafka, service meshes, and distributed sagas. The code becomes a choreography of eventually consistent state machines. This is the architecture of Netflix’s playback control system or LinkedIn’s newsfeed—and it’s exactly right for those problems.
Engineering for understanding optimizes for a different set of constraints: local reasoning, explicit state transitions, and minimal dependencies. The ideal is that a developer can open a single module, read it top-to-bottom, and form a correct mental model of what happens when a specific endpoint is called. This often means monolithic code organization, synchronous request-reply flows, and a relational database that serves as the single source of truth. It’s the architecture of a well-factored Rails application or a carefully layered Spring Boot service.
The tension arises when we apply the first set of patterns to problems that don’t have the scale constraints to justify them. I’ve seen a 12-person startup build a microservice mesh with 40+ services, each with its own CI/CD pipeline, because “that’s how the big players do it.” The result wasn’t infinite scalability. It was infinite debugging sessions trying to trace a single user request across 17 network hops, each with its own retry logic and partial failure modes.
The Cognitive Load Budget
Every architectural decision imposes a tax on the team’s cognitive load. A simple synchronous call between two modules has a near-zero tax: you can follow the control flow in a debugger. Replace that with an asynchronous message over a queue, and you’ve just added several new concepts the developer must hold in their head: message serialization, delivery guarantees, potential reordering, and the state of the consumer at the time of processing.
Here’s a heuristic I use: the number of places you need to look to understand a single request’s behavior is the primary measure of a system’s understandability. In a well-layered monolith, that number might be 3-5 files. In a microservice mesh with event sourcing, it can easily be 15-20, spread across multiple repositories, each with its own configuration for retries, timeouts, and circuit breakers.
This isn’t just a junior-developer problem. I’ve watched principal engineers stare at whiteboards for hours trying to reason about a saga’s compensation logic after a partial failure. The system scaled beautifully to 50,000 requests per second. It also scaled beautifully to 50,000 ways to fail silently.
The Hidden Assumptions in API Design
APIs are where the sociotechnical gap becomes visible. A RESTful API that returns a clean JSON document is a contract that says, “Here is the state of the world right now.” A GraphQL API says, “Ask for what you need, and I’ll assemble it.” An event-driven API says, “Something happened; you figure out what it means.”
Each of these carries assumptions about the consumer’s ability to reason about the system. The RESTful approach assumes the consumer can make a single request and get a consistent snapshot. The event-driven approach assumes the consumer can reconstruct state from a stream of deltas—a task that is notoriously difficult to get right, especially when events arrive out of order or are duplicated.
Consider this snippet from a real codebase I worked on. It’s a handler for a UserRegistered event in a system that used Kafka to propagate state changes:
func handleUserRegistered(event UserRegisteredEvent) error {
user, err := userService.GetUser(event.UserID)
if err != nil {
// The user might not exist yet because the projection
// that creates the user record hasn't processed the event.
// We'll retry with exponential backoff.
return err
}
// Now we can send the welcome email.
return emailService.SendWelcomeEmail(user.Email)
}
This code is a trap. It looks simple, but it hides a distributed systems problem inside what appears to be a straightforward function. The comment is a confession: we’ve built a system where the order of operations is non-deterministic from the perspective of any single service. The developer who wrote this knew it was fragile. The developer who inherits it will learn the hard way.
Compare this to a synchronous approach:
func registerUser(req RegisterUserRequest) error {
user, err := userService.CreateUser(req)
if err != nil {
return err
}
return emailService.SendWelcomeEmail(user.Email)
}
This version is boring. It doesn’t scale to 100,000 registrations per minute. But for a system handling 100 registrations per minute, it’s correct, debuggable, and the failure modes are obvious. The tradeoff is explicit: you’re choosing understandability over throughput. That’s a choice you should make consciously, not accidentally by cargo-culting Netflix’s architecture.

When the RFCs Lead You Astray
The IETF and other standards bodies have given us beautifully specified protocols. HTTP’s semantics are a masterclass in designing for interoperability. But the way we implement those specs often undermines understandability.
Take content negotiation. RFC 7231 defines Accept and Content-Type headers with precision. A server can support multiple representations of a resource, and a client can express preferences. In practice, I’ve seen teams implement content negotiation that branches on Accept: application/json vs. Accept: application/xml in ways that are scattered across middleware, serializers, and controller logic. The result is that the actual response format for a given request becomes an emergent property of the system, not something you can determine by reading the endpoint’s code.
This is a classic case of specification compliance at the expense of local reasoning. The RFC is correct. The implementation is correct. But the developer experience is broken because the system’s behavior is no longer locally predictable. A better approach for most applications is to version the API explicitly in the URL path (/v1/users.json) and reject requests that don’t match. It’s less elegant. It’s more understandable.
The Team Factor: Conway’s Law as a Design Tool
Conway’s Law states that organizations design systems that mirror their communication structures. The corollary is that if you want a system that’s easy to understand, you need a team structure that supports understanding. A team of five people cannot effectively own 20 microservices. The cognitive load of context-switching between services, each with its own deployment pipeline and monitoring dashboard, will overwhelm them.
I’ve seen this play out in a team that adopted the “you build it, you run it” DevOps model without adjusting their service granularity. Each developer was on call for 8-10 services. When an alert fired at 3 a.m., the on-call engineer had to re-learn the service’s architecture, its dependencies, and its failure modes—all while half-asleep and under pressure. The result was a mean time to recovery (MTTR) measured in hours, not minutes.
The fix wasn’t better monitoring or more runbooks. It was merging services back into larger, coherent modules that a single person could fully understand. The team’s velocity increased, and their MTTR dropped by an order of magnitude. They traded theoretical scalability for practical operability.
The Hidden Cost of “Best Practices”
Many “best practices” in software engineering are context-free prescriptions that optimize for scale at the expense of understanding. Consider the advice to “always use a message queue for inter-service communication.” This is excellent advice if you need to decouple services for independent scaling or if you need to handle traffic spikes gracefully. It’s terrible advice if you have two services that are always deployed together and have tightly coupled lifecycles.
I’ve seen teams introduce Kafka between a user service and a notification service that were always deployed as a single unit. The result was that every notification was delayed by the polling interval, and debugging a missed notification required checking the producer, the Kafka cluster, and the consumer. A direct HTTP call would have been simpler, faster to develop, and easier to debug. The only thing the queue added was latency and complexity.
This is not an argument against message queues. It’s an argument against adopting patterns without understanding the tradeoffs they impose on your team’s ability to reason about the system. Every piece of infrastructure you add is a piece of infrastructure you have to operate, monitor, and debug.

Practical Heuristics for Choosing Understandability
After years of making these mistakes myself and cleaning up after others, I’ve settled on a set of heuristics that help me decide when to prioritize understandability over scale. They’re not rules. They’re questions that force you to be explicit about your assumptions.
1. The “Single Developer” Test
Can a single developer—not your most senior, not your most junior—understand the entire request path for a critical user flow without consulting external documentation? If the answer is no, you’ve probably over-abstracted. The system might scale, but your team’s ability to change it safely is compromised.
2. The “3 a.m.” Test
When an alert fires at 3 a.m., can the on-call engineer diagnose the problem using only the information available in your observability tools and their own understanding of the system? If they need to wake up another engineer who “owns” a different service, you’ve built a distributed system that requires distributed knowledge—and that’s a fragile state.
3. The “New Hire” Test
How long does it take a new engineer to make their first production change? In a system optimized for understanding, it should be days, not weeks. If your onboarding process requires a multi-week tour of service ownership and architecture diagrams, your system is too complex for its actual requirements.
4. The “Delete a Service” Test
If you can’t explain what would break if you deleted a particular microservice, you have too many services. This is a sign that the service boundaries don’t align with the business domain boundaries—they’re artifacts of an architectural pattern applied without sufficient context.
When Scale Actually Matters
There are times when engineering for scale is non-negotiable. If you’re building a real-time bidding system that processes millions of requests per second with sub-millisecond latency requirements, you need asynchronous communication, careful resource management, and likely a custom protocol. If you’re building a distributed database, the consistency and partition tolerance constraints will force you into complex consensus algorithms like Raft or Paxos.
But these are specialized systems built by teams with deep expertise in distributed computing. The vast majority of business applications—even those serving millions of users—can be built with a well-structured monolith backed by a relational database and a read-replica for scaling queries. Shopify still runs a modular monolith, and they handle Black Friday traffic just fine.
The key is to understand the actual bottlenecks in your system before you introduce complexity to solve them. Profile your database queries before you add a cache. Measure your request latency before you introduce a message queue. Optimize for the constraints you actually have, not the ones you might have someday.
FAQ
Isn’t a monolith just technical debt waiting to happen?
Only if it’s poorly structured. A well-modularized monolith with clear domain boundaries and enforced dependency rules can be easier to refactor than a tangled mesh of microservices. The key is discipline: packages that don’t depend on each other unnecessarily, interfaces that hide implementation details, and a build system that enforces these constraints. The monolith isn’t the problem; the lack of modularity is.
How do I convince my team to prioritize understandability when everyone wants to use the latest distributed systems patterns?
Start by measuring the cost of the complexity you already have. Track the time spent debugging cross-service issues, the onboarding time for new engineers, and the number of incidents caused by misunderstandings of system behavior. Present these numbers alongside your actual scaling requirements. Often, the data will show that you’re paying a complexity tax for scale you don’t need. Frame the conversation around tradeoffs, not dogma.
What’s the right time to break a monolith into services?
When you have a clear scaling bottleneck that can’t be solved by vertical scaling or read replicas, and when you have a team structure that can support independent service ownership. The trigger should be a concrete, measured problem—not a fear of future problems. A good rule of thumb: if you can’t point to a specific database query or CPU-bound computation that’s causing user-facing latency, you’re not ready to break things apart.
Doesn’t event-driven architecture make systems more decoupled and therefore easier to understand?
It makes them more decoupled at runtime, but it often makes them harder to understand at development time. Decoupling is a double-edged sword: it reduces the blast radius of failures, but it also obscures the causal relationships between components. A developer reading an event handler has no easy way to see what produced the event or what other handlers will react to it. The system’s behavior becomes an emergent property of the event graph, which is notoriously difficult to reason about.
Closing Thoughts
The sociotechnical gap between software construction and user needs isn’t just about understanding the end user. It’s about understanding the developers who will maintain and extend the system. Every abstraction we add, every layer of indirection, every asynchronous boundary—these are costs we impose on future maintainers. Sometimes those costs are justified by the scale we need to achieve. Often, they’re not.
The next time you reach for a message queue, a microservice, or an event-sourced architecture, ask yourself: am I solving a real scaling problem, or am I just making the system harder to understand? The answer will tell you everything you need to know about whether your architecture is serving your team or your ego.