There are two kinds of systems work that look identical on a sprint board and diverge almost immediately in production. The first is engineering for scale: designing for throughput, replication, failover, and cost per request. The second is engineering for understanding: designing so that a new engineer, a tired on-call developer, or a future version of yourself can reconstruct why the system behaves the way it does. Scale optimizes for load. Understanding optimizes for legibility. Most teams claim to do both. Most teams actually do the first and then wonder why every incident review ends with the phrase “we need better documentation.”
This is not a soft-skills essay. It is about concrete practices: API design, documentation engineering, developer experience research, and the politics of technical decision-making. The distance between what developers build and what users actually need is often the same distance between a system that can handle a million requests and a system that can be understood by the person who has to fix it at 3 a.m.
I have spent enough years on both sides of that distance to have scars. Some of them are from writing services that scaled beautifully and were operationally opaque. Some are from writing services that were a joy to debug and a disaster under load. The point is not to choose one. The point is to know which one you are doing, and when.

Scale Is a Property of the System; Understanding Is a Property of the Team
When someone says “this needs to scale,” they usually mean one of three things: more requests per second, more data per node, or more engineers touching the same code. The first two are measurable. The third is political. A system that scales technically but cannot be understood by the people who maintain it will eventually be rewritten, abandoned, or wrapped in a layer of “temporary” tooling that becomes permanent.
Understanding is not the same as documentation. Documentation is a lagging artifact. Understanding is a property of the system itself: the shape of the API, the naming of the modules, the error messages, the logs, the way state is represented. A system can have excellent documentation and still be impossible to understand because the code and the docs describe different realities. A system can have no formal docs and still be legible because the design itself carries the explanation.
Consider two versions of the same endpoint. Version A:
POST /v1/process
{
"data": "...",
"mode": 2
}
Version B:
POST /v1/orders/{order_id}/fulfillment
{
"strategy": "partial",
"items": [...]
}
Version A scales fine. It is a generic pipe. But the word process and the integer mode carry no meaning. Every consumer has to maintain a private mapping of what mode=2 means. Every new engineer has to ask. Version B is slightly more verbose, slightly more constrained, and dramatically more legible. The URL names a resource. The body names a strategy. The domain language is in the interface, not in a wiki page that nobody updates.
This is the core tradeoff. Scale often pushes toward generality: fewer endpoints, fewer concepts, fewer words. Understanding pushes toward specificity: more names, more constraints, more explicit intent. The best systems find a way to do both, but only if someone is explicitly responsible for the second half.
The RFC Is a Political Document
Engineering for understanding has a formal home: the RFC process. Not the IETF kind, though the lineage matters. I mean the internal request-for-comments document that many teams use before building something non-trivial. The RFC is where scale and understanding negotiate.
An RFC that only discusses throughput, latency, and storage is a scale document. An RFC that only discusses naming, module boundaries, and error semantics is a design document. The useful ones do both, and they do it in a specific order: first the problem, then the constraints, then the proposed shape, then the alternatives, then the operational story.
I have seen RFCs that were technically brilliant and politically naive. They proposed a clean architecture that would require three teams to change their code in ways that would break their quarterly OKRs. The RFC was approved in a meeting and ignored in practice. The system that emerged was a hybrid: the new service existed, but the old interfaces remained, and the “temporary” adapter layer became the de facto API. The scale was fine. The understanding was worse than before, because now there were two ways to do everything and no single source of truth.
The politics of technical decision-making is not a distraction from engineering. It is the medium in which engineering happens. If you want a system to be understandable, you have to make understanding a political priority, not just a technical one. That means saying no to features that would make the system faster but harder to reason about. It means defending a naming convention in a review even when it feels pedantic. It means writing the RFC so that a person who was not in the room can reconstruct the decision six months later.
Documentation Engineering Is Not Writing; It Is Designing
Most documentation is written after the fact, by someone who is tired, against a deadline, with no clear idea of who will read it. That is not documentation engineering. That is regret, formatted as Markdown.
Documentation engineering treats docs as part of the system. The same way you would design an API for consistency, you design the docs for consistency. The same way you would test a service for edge cases, you test the docs against real user questions. The same way you would version an API, you version the docs. The same way you would monitor a service for errors, you monitor the docs for confusion.
Concretely, this means:
- Docs are built from the same source of truth as the code. If the API spec is OpenAPI, the docs are generated from it, not written alongside it. If the error codes are defined in a schema, the docs reference that schema, not a copy-pasted table.
- Every example is runnable. A code snippet that has never been executed is a lie waiting to be discovered. The best docs are tested in CI. The second-best docs are tested by a human who actually runs the command.
- Error messages are documentation. A user who gets
Error 500: something went wrongwill not read your docs. They will open a support ticket. A user who getsError 422: order_id must be a UUID, received 'abc'has just received a micro-lesson in your API contract. - The docs have a single entry point for each persona. A new user, an integrating partner, and an on-call engineer need different paths. If they all start at the same page, the docs are not designed; they are accumulated.
None of this is about writing better sentences. It is about designing a system where the explanation is a first-class artifact, not a byproduct.

Developer Experience Research Is Not a Survey
Developer experience research is the practice of watching real developers try to use your API, your SDK, your CLI, or your docs, and then actually changing the thing based on what you saw. It is not a quarterly survey. It is not a Net Promoter Score. It is not a focus group where everyone nods politely.
The most useful DX research I have done was embarrassingly simple: sit next to a developer who has never seen the system, give them a task, and shut up. Do not help. Do not explain. Do not say “oh, that’s a known issue.” Just watch. Take notes. Count the number of times they look at the docs, the number of times they guess wrong, the number of times they say “wait, what?”
Every one of those moments is a bug in the system’s legibility. The developer is not the problem. The system is. If a competent engineer cannot figure out how to authenticate in under five minutes, the authentication docs are broken. If a competent engineer cannot tell whether an operation is idempotent, the API design is broken. If a competent engineer cannot find the error code they just received, the error handling is broken.
Scale thinking says: “We’ll add a FAQ.” Understanding thinking says: “We’ll change the API so the question never arises.” The second is more expensive in the short term and dramatically cheaper in the long term, because every question a developer has to ask is a support ticket, a Slack interruption, or a silent abandonment.
The Architecture Diagram as a Rhetorical Device
Architecture diagrams are not neutral. They are arguments. A diagram that shows five boxes and three arrows is making a claim about what matters and what can be ignored. A diagram that shows fifty boxes and a hundred arrows is making a different claim: that the system is too complex to be understood, and you should just trust the people who drew it.
I have seen diagrams that were technically accurate and completely useless. They showed every service, every queue, every database, every cache, every load balancer, every firewall, every VPN tunnel. They were beautiful. They were also unreadable. The person who drew them understood the system. The person who needed to understand the system did not.
The best architecture diagrams are situated. They answer a specific question: “What happens when a user submits an order?” or “What happens when the payment service goes down?” They show the happy path and the failure path. They label the arrows with the actual data that flows, not just “HTTP” or “gRPC.” They include the human actors: the on-call engineer, the support agent, the customer. They are drawn for a reader, not for a poster.
This is the same principle as API design. A diagram that tries to show everything shows nothing. A diagram that shows one thing clearly is a tool for understanding. A diagram that shows everything is a tool for status.
Scale Hides; Understanding Reveals
There is a reason these two goals conflict. Scale is about hiding complexity. A load balancer hides the fact that there are twelve instances. A cache hides the fact that the database is slow. A queue hides the fact that the downstream service is down. All of these are good things, until something goes wrong. Then the hiding becomes the problem.
Understanding is about revealing complexity. A good error message reveals the exact point of failure. A good log line reveals the exact state that led to the failure. A good API reveals the exact contract that was violated. A good architecture diagram reveals the exact path that the request took. The tension is not between good and bad engineering. It is between two different kinds of good engineering, and the best systems are the ones that know when to hide and when to reveal.
Concretely: hide the number of instances behind a load balancer, but reveal the instance ID in the response headers. Hide the cache internals, but reveal the cache hit/miss status in the logs. Hide the queue depth from the user, but reveal it on the dashboard. Hide the retry logic, but reveal the retry count in the error message. Every one of these is a small decision that costs almost nothing and pays off enormously when something breaks.
What This Looks Like in Practice
Let me give you a concrete example from a system I worked on. We had a service that processed incoming webhooks. It scaled fine: it could handle thousands of events per second, it had retries, it had a dead-letter queue, it had all the operational bells and whistles. But when a webhook failed, the error message was processing failed. The logs had a stack trace, but the stack trace was from a library deep inside the service, not from our code. The on-call engineer had to grep through the codebase to figure out what processing failed actually meant.
We fixed it by changing the error contract. Every webhook failure now returned a structured error with three fields: stage (which part of the pipeline failed), reason (why it failed, in domain language), and event_id (which event failed). The scale did not change. The throughput did not change. The understanding changed completely. The on-call engineer could now see, in the alert, exactly what happened and where to look. The support team could now tell a customer, in plain English, why their webhook was rejected. The docs could now reference the error contract instead of a vague paragraph about “common issues.”
That is the difference. Scale is about making the system fast enough. Understanding is about making the system legible enough. Both are engineering. Both are hard. But only one of them is usually forgotten.

Heuristics for Choosing
Here are the heuristics I use when I am not sure which mode I am in:
- If the change makes the system faster but harder to explain, it is a scale change. That is fine. Just say so. Do not pretend it is also a clarity improvement.
- If the change makes the system easier to explain but slightly slower, it is an understanding change. That is also fine. Just measure the slowdown and make sure it is acceptable.
- If a new engineer cannot explain the system after a week, the system is not understandable. No amount of documentation will fix that. The design itself has to change.
- If an on-call engineer cannot diagnose a failure from the alert alone, the system is not observable enough. Observability is a subset of understanding. Logs, metrics, and traces are the interface between the system and the person who has to fix it.
- If the API docs and the API behavior disagree, the docs are wrong. Always. The code is the source of truth. The docs are a derived artifact. If you cannot keep them in sync, generate the docs from the code or delete the docs.
- If a decision was made in a meeting and not written down, it was not made. It was a conversation. Conversations do not scale. RFCs scale. Decision records scale. Write it down or accept that you will re-litigate it in six months.
None of these heuristics are original. They are the accumulated scar tissue of watching systems scale beautifully and fail legibly, or scale poorly and explain themselves perfectly. The goal is not to avoid the scars. The goal is to learn from them before they happen to you.
FAQ
Is engineering for understanding just another term for documentation?
No. Documentation is one output of engineering for understanding, but the practice is broader. It includes API design, error message design, logging strategy, architecture diagramming, RFC writing, and developer experience research. A system can have excellent documentation and still be hard to understand if the code, the API, and the docs tell different stories. Understanding is a property of the system, not just the prose around it.
How do you measure whether a system is understandable?
You measure it the same way you measure any other quality: with a proxy. Time-to-first-successful-call for a new developer is a good proxy. Time-to-diagnosis for an on-call engineer is another. Number of support tickets that could have been answered by better error messages is a third. None of these are perfect, but they are all better than “I think the docs are pretty good.” The most reliable signal is watching a competent developer try to use the system without help and counting the moments of confusion.
Can a system be both scalable and understandable?
Yes, but not by accident. The two goals pull in different directions: scale pushes toward generality and hiding, understanding pushes toward specificity and revealing. The systems that achieve both do so because someone explicitly negotiated the tradeoff at every decision point. That usually means an RFC process that treats legibility as a first-class requirement, a documentation pipeline that is tested in CI, and a team culture that treats “I can’t explain this” as a bug, not a personality trait.
What is the first thing a team should do if their system is hard to understand?
Stop adding features for a week and do a legibility audit. Pick the three most common failure modes, the three most common integration tasks, and the three most confusing parts of the API. For each one, ask: Can a new engineer explain this? Can an on-call engineer diagnose this from the alert alone? Can a user figure this out without asking a human? The answers will tell you where to start. Usually the first fix is not more docs. It is better error messages, better naming, or a simpler API surface.
This article is part of an ongoing series on the gap between what developers build and what users actually need. The next piece will look at how API versioning decisions are really made, and why the technical arguments are usually the least important part of the conversation.