Scale engineering is the practice of designing systems that stay coherent under massive quantitative stress—millions of requests, terabytes of data, thousands of concurrent edits. Sense-making engineering is the practice of designing systems that stay coherent under qualitative stress—ambiguous user intent, evolving domain language, conflicting mental models. The two get lumped together because both involve “architecture,” but the materials are different. Scale engineering works with throughput, latency, and fault tolerance. Sense-making engineering works with affordances, conceptual models, and information scent. When a team optimizes for the former without investing in the latter, the result is a fast, reliable, and utterly baffling product. This article is for the developers, tech leads, and API designers who have felt that gap—the sociotechnical chasm between what the system can do and what the user can understand.
The Two Architectures: Runtime Topology vs. Conceptual Topology
Every software system has at least two architectures. The first is the runtime topology: services, databases, caches, queues, and the wires between them. This is the architecture we draw in C4 diagrams, monitor with dashboards, and defend with SLOs. The second is the conceptual topology: the nouns, verbs, and relationships the system exposes to a developer or end user. This architecture lives in API documentation, SDK method signatures, CLI flags, and the mental models people build when they read a quickstart guide.
Scale engineering primarily concerns itself with the runtime topology. It asks: can this service handle a 10x traffic spike? Will this database partition gracefully? Sense-making engineering concerns itself with the conceptual topology. It asks: does a new user understand what this endpoint does within 30 seconds of reading its description? Does the error message point toward a corrective action or just a stack trace?
The tension arises because these architectures are coupled but not aligned. A beautifully sharded, eventually consistent data store can present a conceptual model so riddled with caveats—stale reads, write conflicts, session affinity—that the developer experience becomes hostile. The system scales technically; it fails to scale cognitively.

Where the Gap Hurts Most: API Design as a Cognitive Interface
APIs are the most literal manifestation of the sociotechnical gap. An API is simultaneously a technical contract (bytes in, bytes out, status codes) and a cognitive contract (“I expect this resource to behave like the others”). When an API is designed purely for implementation convenience, the cognitive contract breaks.
Consider a pagination scheme. From a scale perspective, cursor-based pagination is superior: it’s stateless on the server, avoids offset drift, and handles large datasets efficiently. But from a sense-making perspective, cursor-based pagination is opaque. A developer integrating the API cannot easily jump to page 5 or estimate how many pages remain. The system scales; the developer’s understanding does not. The fix is not to abandon cursor-based pagination but to supplement it with metadata—total counts, next/prev links, human-readable hints—that bridge the conceptual gap. The runtime topology remains efficient; the conceptual topology becomes navigable.
This pattern repeats across domains. A GraphQL schema that mirrors the database exactly is easy to build but forces clients to understand internal table relationships. A REST endpoint that returns raw error codes from a legacy monolith preserves the backend’s semantics at the expense of the consumer’s sanity. In each case, the engineering team has optimized for the system they own, not the system the user experiences.
The RFC as a Sense-Making Artifact
I’ve started treating RFCs (Requests for Comments) not just as design documents but as sense-making probes. When I write an RFC for a new endpoint or a breaking change, I include a section called “Developer Experience Impact.” It’s a structured narrative: what will a developer think this does on first reading? What will they try first? What will confuse them? This section is not about runtime performance; it’s about cognitive performance.
For example, in a recent RFC for a batch operation endpoint, the initial design accepted an array of resource IDs and returned a map of results. The runtime topology was clean: a single POST, a single database query with an IN clause, a single response. But the conceptual topology was a mess. Partial failures were represented as null values in the map, forcing the client to iterate and check. The revised design returned an array of result objects, each with a status field. The runtime cost was identical. The cognitive cost dropped sharply. The RFC’s “Developer Experience Impact” section made that tradeoff explicit and won the argument.

Error Messages Are the Front Line of Sense-Making
If you want to diagnose whether your team prioritizes scale over understanding, look at your error messages. A scale-optimized error message tells the developer what went wrong in the system: “Connection refused,” “NullPointerException,” “409 Conflict.” A sense-making error message tells the developer what went wrong in their intent and how to fix it: “The resource was modified by another request. Fetch the latest version and retry.”
I once worked on a system where a misconfigured environment variable produced the error: “Invalid configuration. See logs.” The logs were in a different tool, behind a different authentication layer, and rotated every hour. The error message was technically accurate—the configuration was indeed invalid—but it was a sense-making failure. The developer had to context-switch, authenticate, search, and correlate timestamps just to discover they’d misspelled a key name. A better error message would have been: “Unknown configuration key ‘DATABASE_URL’. Did you mean ‘DATABASE_URL’?” That message costs nothing in runtime resources. It saves enormous cognitive resources.
This is not about being “nice” to developers. It’s about reducing the total cost of ownership of the software. Every minute a developer spends deciphering a cryptic error is a minute not spent building features, fixing bugs, or improving reliability. The sociotechnical gap has a measurable tax.
Tooling That Bridges the Gap
Some tools explicitly address this gap. OpenAPI (formerly Swagger) lets you document API semantics alongside HTTP semantics. JSON Schema can annotate fields with descriptions, examples, and deprecation notices. But these are only as good as the annotations. A generated OpenAPI spec from code that lacks comments is just a machine-readable version of a bad cognitive model. The tooling is necessary but insufficient; the real work is the editorial discipline of explaining why a field exists and when it should be used.
I’ve also seen teams use decision records (ADRs) to capture the rationale behind API choices. An ADR that says “We chose cursor-based pagination because our dataset is append-only and offset pagination would miss new records” is a sense-making artifact. It helps future maintainers—and future users—understand the constraints that shaped the interface. Without it, someone will file an issue asking for offset pagination, and the team will have forgotten why they didn’t do it in the first place.

When Scale Engineering Undermines Sense-Making
There are specific architectural patterns that, while excellent for runtime scale, actively degrade the user’s ability to reason about the system. Eventual consistency is the classic example. A system that accepts a write and returns success before the data is visible to all readers is a marvel of distributed systems engineering. But to a developer who just called the write endpoint and then the read endpoint, it’s a bug. The system says “OK” and then lies by omission. The fix is not to abandon eventual consistency but to surface the inconsistency: expose a read-your-writes token, provide a wait_for_consistency parameter, or document the propagation delay explicitly.
Another pattern is microservice decomposition by technical layer rather than by bounded context. A team might split a monolith into a “frontend service,” a “business logic service,” and a “data service.” This scales development teams but fractures the conceptual model. A single user action now requires coordinating three services, each with its own error modes, rate limits, and authentication. The user’s mental model—“I’m updating my profile”—collides with the system’s model—“You’re making seven RPC calls across three trust boundaries.” The sociotechnical gap widens.
The alternative is to decompose by bounded context, aligning service boundaries with domain boundaries. A “Profile Service” owns the entire profile concept. It may internally call other services, but the API consumer sees a coherent interface. This is the core insight of Domain-Driven Design: the conceptual topology should drive the runtime topology, not the other way around.
Practical Heuristics for Closing the Gap
After years of oscillating between these two mindsets, I’ve settled on a set of heuristics that help me—and the teams I work with—keep both architectures in view.
1. The “New Developer” Test
For any API endpoint, UI component, or CLI command, ask: Can a developer who has never seen this system before use it correctly without reading the implementation? If the answer is no, the conceptual model is leaking runtime details. Add documentation, rename parameters, or restructure the response until the answer is yes. This is not about dumbing things down; it’s about making the interface self-describing.
2. Error Budgets for Cognitive Load
We use error budgets for reliability—how much downtime is acceptable before we halt feature work. I propose a parallel concept: a cognitive error budget. Every time a developer encounters an unexpected error, a confusing parameter name, or missing documentation, it consumes the budget. When the budget is exhausted, the team stops building new features and invests in sense-making improvements. This makes the invisible cost visible to product managers and stakeholders.
3. The “Why” Annotation Rule
Every field in an API response, every configuration option, every CLI flag must have a non-obvious “why” documented. Not just what it is, but why a developer would use it. “Timeout: the maximum time in milliseconds the client will wait for a response” is a what. “Timeout: set this higher for batch operations that may take several seconds; lower for interactive requests where the user is waiting” is a why. The why bridges the system’s behavior and the developer’s intent.
4. Conceptual Compression
Good sense-making engineering compresses concepts. A REST API that exposes 50 endpoints for 50 variations of the same resource is a scale engineering success—each endpoint is optimized. But it’s a sense-making failure because the developer must learn 50 things. A better design might expose 5 endpoints with query parameters that compose. The runtime cost is slightly higher; the cognitive cost is dramatically lower. Conceptual compression is the art of reducing the number of distinct concepts a user must hold in their head to use the system effectively.
5. The “What Happens If” Protocol
Before shipping any feature, run a “What happens if…” session. What happens if the user passes an empty array? What happens if they call this endpoint before that one? What happens if the network drops mid-request? Document the answers. If the answers are “undefined behavior” or “it depends on internal state,” you’ve found a sense-making gap. Fill it before you ship.
FAQ
What is the sociotechnical gap in software engineering?
The sociotechnical gap refers to the disconnect between a system’s technical implementation and the way users—including developers consuming an API—understand and interact with it. A system can be technically sound but cognitively inaccessible, forcing users to build flawed mental models that lead to errors, frustration, and increased support costs.
How do I know if my API is optimized for understanding rather than just scale?
Look at your support tickets and developer forums. Are users repeatedly confused by the same concepts? Do they misuse endpoints in predictable ways? Also, examine your error messages: do they explain what the user should do next, or do they just describe what went wrong internally? If your error messages read like stack traces, you’re optimizing for the wrong audience.
Doesn’t focusing on developer experience slow down feature delivery?
In the short term, yes—it requires additional design work, documentation, and iteration. But in the long term, it reduces the total cost of ownership by lowering support burden, decreasing integration time for new consumers, and preventing breaking changes that arise from misunderstood interfaces. The investment pays for itself in reduced cognitive debt, much like addressing technical debt pays off in reduced maintenance overhead.
What’s the difference between good documentation and good sense-making design?
Good documentation explains a confusing system. Good sense-making design makes the system less confusing in the first place. Documentation is a patch; sense-making design is a preventative measure. You need both, but the goal should be to reduce reliance on documentation by making the system’s behavior more predictable and its concepts more intuitive.
How do I convince my team to invest in sense-making engineering?
Frame it in terms they already care about. If your team values reliability, show how confusing interfaces lead to operator error and incidents. If they value velocity, show how unclear APIs slow down internal consumers and increase integration time. If they value hiring, point out that a system with a steep learning curve makes onboarding new engineers slower and more expensive. Sense-making engineering is not a separate concern; it’s a multiplier on every other engineering investment.