
I’ve watched smart engineers burn days chasing elegant, wrong theories. They spot a stack trace, a memory spike, or a race condition, and their brain instantly serves up a neat, plausible story. “The connection pool is exhausted.” “The cache invalidation is lagging.” “It’s a GC pause.” They latch onto that story, open a profiler, and start hunting for evidence to confirm it. That’s not debugging. That’s storytelling. And it’s a painfully slow way to find a root cause.
The best debuggers I know do something counterintuitive. They don’t start with a hypothesis. They start with a question. A genuine, open-ended question that forces them to observe the system’s actual behavior before their brain fills in the gaps with assumptions. The difference sounds subtle, but in practice, it’s the difference between a 30-minute fix and a three-day yak shave.
The Hypothesis Trap
Forming a hypothesis feels productive. It gives you direction. You pop open your tools, you look for the thing you expect to see, and often you find something that looks like it. The trouble is, complex systems are full of red herrings. A thread dump showing 200 threads waiting on a database connection pool looks like a smoking gun. Your hypothesis—”the database is slow”—seems confirmed. You spend the next two hours tuning queries, only to realize the real issue was a deadlocked configuration update that prevented those threads from ever releasing their connections. The threads were waiting, sure, but not for the reason you assumed.
When you start with a hypothesis, you’re essentially asking a yes/no question: “Is the database slow?” The system will almost always give you a “yes” to something, because a production system under load is always slow somewhere. You’ll find a slow query, a saturated index, a disk I/O spike. You’ll fix it, feel good, and the bug will still be there. You’ve optimized a symptom, not the cause.
Questions That Expose the System’s Actual State
A good debugging question is open-ended and focuses on what the system is actually doing, not what you think it should be doing. Instead of “Is the database slow?”, ask “What are the threads doing right now?” The difference is critical. The first question sends you looking for a specific condition. The second forces you to dump thread stacks, sort them by state, and read them without a preconceived narrative.
Here’s a real example from a memory pressure incident I worked on. The service was restarting every few hours with an OutOfMemoryError. The team’s immediate hypothesis was a memory leak in the application code. They spent a day profiling heap dumps, looking for objects that weren’t being garbage collected. They found some, patched them, deployed, and the service still crashed.
I joined and asked a different question: “What is consuming the heap right before the crash?” Not “where is the leak?” but “what’s in the heap?” We took a heap dump 30 seconds before the OOM kill, loaded it into Eclipse MAT, and ran a simple histogram. The top consumer wasn’t a leaked domain object. It was a 1.8GB byte array allocated by a single thread. That thread was deserializing a payload from an upstream service that had silently changed its contract, sending a 2GB blob instead of a 2MB one. No leak. Just a massive, legitimate allocation the JVM couldn’t handle. The fix was a payload size check, not a code refactor. The question shaped the outcome.

Questions as a Forcing Function for Observability
Starting with a question also exposes gaps in your observability. If you can’t answer the question, you don’t have the right telemetry. That’s valuable information in itself. When an engineer forms a hypothesis first, they often try to answer it with the telemetry they already have, even if it’s the wrong telemetry. They’ll squint at CPU graphs to diagnose a thread contention issue, or grep application logs to understand a kernel-level packet drop.
Good questions force you to instrument the right thing. “What is the distribution of response times for this endpoint?” requires percentiles, not averages. “Which threads are in a BLOCKED state right now?” requires a thread dump, not a heap dump. “What system calls is this process making?” requires strace or eBPF, not application logs. The question comes first, then the tool. Not the other way around.
Example: The Case of the Missing Milliseconds
An API endpoint had a p99 latency of 2 seconds, but the p50 was 50ms. The team’s hypothesis was a slow downstream service. They had dashboards showing downstream latency, and sure enough, the downstream’s p99 was 1.8 seconds. Case closed? Not quite. The question I asked was: “What is the exact difference between the time our service receives the request and the time it sends the downstream call?”
They didn’t have that metric. They had client-side latency for the downstream, but not the internal gap. We added a single timing span. The result: the downstream call was fast, but the service was spending 1.5 seconds deserializing a large request body before making the call. The deserialization was CPU-bound and blocked the event loop. The downstream was a victim, not the culprit. The question revealed a blind spot that the hypothesis had papered over.
How to Formulate a Debugging Question
This isn’t a soft skill. It’s a technical discipline. A well-formed debugging question has three properties:
- It’s specific to a boundary. “What is happening?” is too vague. “What is the state of the connection pool when the error rate spikes?” is a question about a specific component at a specific time.
- It demands a quantitative answer. Not “is it slow?” but “what is the 99th percentile latency, and what is its breakdown?” Numbers force precision and prevent hand-waving.
- It’s answerable with the system’s current instrumentation, or it reveals what instrumentation is missing. If you can’t answer it, you’ve just identified a critical observability gap. That’s a win.
Let’s apply this to a common scenario: a Kubernetes pod that occasionally restarts. The hypothesis-driven approach: “It’s probably the liveness probe failing because the app is slow during garbage collection.” You tweak the probe timeout, increase the heap, and wait. The question-driven approach: “What is the exact exit code and reason for the last 10 container restarts?” You run kubectl describe pod and see exit code 137 (SIGKILL). That’s not a failed probe; that’s the OOM killer. The question immediately narrows the problem space to memory, not health checks. You just saved a day of tuning the wrong knob.

Questions Prevent the Blame Game
There’s a social benefit here too. Hypotheses often carry an implicit accusation. “The database is slow” points a finger at the DBA team. “The network is dropping packets” blames the infrastructure folks. These statements trigger defensive responses and waste time in war rooms. A question is neutral. “What is the TCP retransmit rate between service A and service B?” is a fact-finding mission. It invites collaboration. The network engineer can pull that data without feeling attacked, and you might both learn that the retransmit rate is zero—redirecting the investigation elsewhere without bruised egos.
When a Hypothesis Is Actually Useful
I’m not saying hypotheses are useless. They’re essential for designing experiments once you’ve narrowed the problem space. After you’ve asked questions, gathered data, and identified a suspicious component, then you form a hypothesis: “If I increase the connection pool size, the BLOCKED thread count will drop.” That’s a testable, falsifiable statement. You make the change, measure the result, and confirm or reject. That’s the scientific method. But the scientific method starts with observation and a question, not with the hypothesis. The hypothesis is step three, not step one.
In practice, the sequence looks like this:
- Observe the symptom. “Users report timeouts on the checkout page.”
- Ask a question to scope the observation. “What is the error rate and latency distribution for the checkout service in the last 15 minutes?”
- Gather data to answer the question. Pull metrics, traces, logs.
- Form a hypothesis based on the data. “The latency spike correlates with a deployment of the inventory service. The checkout service is waiting on inventory calls.”
- Test the hypothesis with a controlled change or further targeted observation.
Most engineers jump from step 1 to step 4, skipping the critical narrowing that steps 2 and 3 provide. That’s the habit to break.
Building a Question-First Culture
If you lead a team, you can model this. When someone reports a bug, don’t ask “What do you think is causing it?” Ask “What’s the most surprising thing you’ve observed about the system’s behavior?” This reframes the discussion around evidence, not speculation. In postmortems, instead of “What went wrong?” ask “What question would have led us to the root cause in 5 minutes, and why didn’t we have the data to answer it?” This turns every incident into an observability investment, not just a fix.
Debugging is fundamentally a learning process. You’re trying to understand a system that is behaving in a way you didn’t expect. The fastest way to learn is to admit you don’t know and ask a precise question. The slowest way is to pretend you do know and chase a phantom. Next time you’re staring at a broken system, resist the urge to explain it. Just ask it what it’s doing. The answer will surprise you.
Frequently Asked Questions
Why do engineers naturally jump to hypotheses?
Our brains are pattern-matching machines. We’ve seen similar symptoms before, and we’re rewarded for quick answers. In many engineering cultures, saying “I don’t know” feels like a weakness, while proposing a theory—even a wrong one—feels like progress. It takes deliberate practice to override that instinct and sit with the discomfort of an unanswered question.
How do I know if my question is good enough?
A good question should point you to a specific piece of telemetry. If you can’t name the metric, log, or trace that would answer it, the question is still too broad. Refine it until it maps to a concrete query you can run against your observability stack. If that query doesn’t exist, you’ve just found a gap worth filling.
Does this approach work for intermittent, hard-to-reproduce bugs?
It’s especially powerful for those. With intermittent bugs, hypotheses are almost always guesses because you can’t observe the system in the failing state on demand. A question-first approach forces you to instrument the system to capture its state when the bug occurs, so you’re ready next time. The question becomes: “What should we log or trace so that the next occurrence gives us a complete picture?”