I still remember the first time I had to debug a VPC peering issue without the console. No pretty diagrams, no drag-and-drop route tables. Just a terminal, tcpdump, and a sinking realization that I’d spent three years deploying cloud infrastructure without really understanding what lived beneath the aws_vpc resource. The cloud promised to abstract away complexity. Instead, it abstracted away our competence.

We’ve turned into wizards of YAML and JSON, conjuring entire network topologies with a few dozen lines of Terraform or CloudFormation. But hand us a misbehaving BGP session or a VLAN mismatch on a bare-metal switch, and the wizard robe slips. We’re not engineers anymore; we’re API callers. And that distinction bites hard when things break.

The Golden Age of Abstraction

Let’s be fair: cloud abstractions are a genuine triumph of software engineering. AWS, Azure, and GCP took the arcane rituals of racking servers, crimping cables, and configuring spanning tree and turned them into a few clicks or a terraform apply. That democratization is real. A startup with two developers can deploy a globally distributed application in an afternoon. That’s honestly remarkable.

The problem isn’t the abstraction itself. It’s what happened to the people using it. When you never have to calculate a subnet mask by hand, you lose the intuition for why 10.0.0.0/8 and 10.0.0.0/16 are different beasts. When security groups magically handle stateful filtering, you forget that a stateless ACL requires explicit return rules. The knowledge doesn’t just atrophy—it never forms in the first place.

Abstract network visualization with glowing nodes

The Subnet Mask Litmus Test

I’ve started asking a simple question in technical interviews: “Given an IP of 192.168.1.50/27, what’s the network address, broadcast address, and usable host range?” The blank stares I get from candidates with “Senior Cloud Engineer” on their résumés are terrifying. These aren’t trick questions. This is the absolute minimum required to understand why two instances in the same VPC can’t talk to each other when someone fat-fingers a route table entry.

Let’s break it down, because apparently we need to re-teach this. A /27 mask means 27 bits are reserved for the network portion, leaving 5 bits for hosts. That gives you 32 addresses per subnet (2^5). Subtract the network address and broadcast address, and you’ve got 30 usable IPs. For 192.168.1.50, the network address is 192.168.1.32, broadcast is 192.168.1.63, and the usable range is .33 through .62. This isn’t arcane knowledge; it’s the foundation of packet forwarding. Yet cloud engineers stare at it like it’s hieroglyphics.

The cloud consoles hide this beautifully. You type a CIDR block into a field, and the platform auto-calculates everything. But when you’re troubleshooting a VPN tunnel that won’t establish Phase 2 because of a mismatched proxy ID, that console is useless. You need to understand that the proxy ID is essentially a subnet pair, and if your on-premises firewall expects 10.0.0.0/24 but your cloud VPN is configured for 10.0.0.0/16, the tunnel will flap until you fix it. No amount of clicking around the AWS VPN dashboard will surface that mismatch clearly.

How We Got Here

The shift started innocently enough. Around 2015, as cloud adoption hit the mainstream, the industry narrative changed. “NoOps” became a buzzword. The idea was that developers could own infrastructure because the cloud made it simple. What actually happened was that developers learned just enough to be dangerous, and dedicated network engineers were sidelined as “legacy” thinkers.

I watched a team spend two days debugging a “network outage” that was actually a DNS resolution failure because their custom DHCP option set in a VPC wasn’t propagating correctly. They’d never touched a DHCP configuration file in their lives. They didn’t know that DHCP options are applied at instance boot and cached, so changing them mid-flight requires a lease renewal. They just kept restarting instances and praying. Two days. For a dhclient -r command.

Server rack with glowing lights and cables

The Terraform Trap

Infrastructure as Code (IaC) is a double-edged sword. On one hand, it enforces reproducibility and version control. On the other, it lets you deploy complex network architectures without understanding a single component. You can write:

resource "aws_vpc" "main" {
  cidr_block = "10.0.0.0/16"
}

resource "aws_subnet" "public" {
  vpc_id     = aws_vpc.main.id
  cidr_block = "10.0.1.0/24"
}

resource "aws_route_table" "public" {
  vpc_id = aws_vpc.main.id

  route {
    cidr_block = "0.0.0.0/0"
    gateway_id = aws_internet_gateway.igw.id
  }
}

This works. terraform apply returns green. You’ve just built a routable public subnet. But do you know why the route table needs an entry for 0.0.0.0/0 pointing to the IGW? Do you understand that without it, the subnet’s implicit router (the VPC’s virtual routing layer) has no default route, so return traffic from the internet never finds its way back to your instance? Most people I ask just say, “That’s how you make it public.” That’s not understanding. That’s memorizing a recipe.

I’ve seen Terraform modules that deploy transit gateways with complex route propagation settings, and the engineers who wrote them couldn’t explain the difference between static routes and propagated routes. They just copied a module from a registry and tweaked variables until it stopped throwing errors. When a production issue arose—a spoke VPC suddenly unable to reach an on-premises network—they had no mental model to diagnose it. The transit gateway was propagating a 10.0.0.0/8 route from one VPC that overlapped with a more specific 10.0.0.0/16 static route, and the routing priority bit them. Basic longest-prefix-match rule. But they’d never heard of it.

The OSI Model Isn’t Just a Poster

Somewhere along the way, the OSI model became a joke. “Please Do Not Throw Sausage Pizza Away” is funny, but the actual layers matter. When a cloud engineer can’t distinguish between a Layer 3 routing problem and a Layer 4 firewall issue, troubleshooting becomes a game of random clicking. I’ve seen security groups (Layer 4 stateful firewalls) blamed for what was actually a missing route in a route table (Layer 3). I’ve seen people try to fix a TLS handshake failure (Layer 6/7) by adjusting network ACLs (Layer 4). The cloud abstracts these layers into separate console sections, but the underlying reality doesn’t change. Packets still flow the same way they did in 1981.

Let’s get concrete. Suppose you have an EC2 instance that can’t reach an S3 endpoint. The cloud engineer’s first instinct is to check the security group. Outbound rules allow all traffic. Then they check the VPC endpoint policy. It allows s3:GetObject. Then they’re stuck. What they’re missing is that the instance’s route table doesn’t have a route to the VPC endpoint’s prefix list. Without that route, traffic destined for S3 goes to the IGW, gets NATed, and leaves AWS entirely before trying to re-enter via the public S3 endpoint—which might be blocked by firewall rules or simply add latency. The fix is a single route entry pointing pl-xxxxx to the VPC endpoint ID. But if you don’t understand that VPC endpoints inject prefix lists into route tables, you’ll never find it.

Fiber optic cables with light signals

BGP: The Protocol Everyone Uses and Nobody Learns

Border Gateway Protocol is the backbone of the internet and the default dynamic routing protocol for most cloud-to-on-premises connections. Yet I’d estimate that fewer than 10% of cloud engineers can configure a BGP session manually on a Cisco or Juniper device. They rely on the cloud provider’s VPN wizard to set up the tunnel and BGP peering, and when it works, they move on. When it doesn’t, they open a support ticket.

Here’s a real scenario: a Direct Connect link with BGP keeps flapping. The cloud engineer sees “BGP session down” alerts but has no idea how to interpret the BGP state machine. They don’t know that Idle means the router hasn’t even attempted a TCP connection to the peer, while Active means it’s trying but failing. They don’t know to check for a firewall blocking TCP port 179. They don’t know how to read a BGP update message to see if the AS path is malformed. They just escalate to the network team, who then has to explain that the on-premises router is advertising a route with a private AS number that the cloud provider filters by default. A single local-as override command would fix it. But the cloud engineer never learned that command exists.

Security Groups: Stateful Magic, Stateless Confusion

Cloud security groups are stateful. That’s a gift. You allow outbound traffic, and return traffic is automatically permitted. This is so convenient that engineers forget stateful firewalling isn’t universal. When they encounter network ACLs (stateless) or on-premises firewalls (often stateless for certain rules), they create rules that allow outbound but forget the corresponding inbound rule for return traffic. The result: one-way communication failures that are maddeningly intermittent because some protocols (like ICMP) might slip through while others (like TCP) get blocked.

I once saw a hybrid cloud setup where the cloud-side security group allowed all outbound TCP to an on-premises server, but the on-premises firewall had a stateless rule that only allowed inbound TCP from the cloud subnet. Return traffic from the on-premises server was blocked because the firewall didn’t have an explicit outbound rule for the cloud subnet. The cloud engineers couldn’t understand why the TCP three-way handshake completed (SYN, SYN-ACK) but then data transfer failed. They’d never seen a stateless firewall before. They thought all firewalls worked like security groups.

Reclaiming Competence

I’m not advocating we all go back to manually configuring switches. That’s absurd. The cloud is here to stay, and its abstractions are valuable. But we need to stop treating those abstractions as a substitute for knowledge. They’re a convenience, not a crutch. If you can’t explain what happens to a packet at each layer as it travels from an EC2 instance to an on-premises server via a VPN tunnel, you’re not a network engineer. You’re a cloud console operator.

Here’s my prescription: spend time in the weeds. Set up a lab with a couple of old routers or virtual machines running FRRouting. Configure BGP by hand. Break it. Fix it. Use tcpdump to watch the TCP handshake and the BGP OPEN messages. Calculate subnet masks until you can do it in your head. Read the actual RFCs—RFC 1918 for private addressing, RFC 4271 for BGP, RFC 793 for TCP. They’re dense, but they’re the source code of the internet. The cloud didn’t rewrite them; it just hid them behind a pretty UI.

When you understand the fundamentals, the cloud becomes a tool rather than a mystery. You’ll know why a VPN tunnel won’t establish Phase 2 when the proxy IDs don’t match. You’ll know why a route isn’t propagating even though the BGP session is up. You’ll know why a packet is being dropped even though the security group allows it. You’ll stop being a YAML jockey and start being an engineer.

FAQ

Why should I learn subnetting if the cloud console calculates it for me?

Because the console won’t help you when you’re troubleshooting a routing issue at 2 AM and need to understand why two subnets that “look fine” in the UI can’t communicate. Subnet math gives you the mental model to spot overlaps, misconfigured route tables, and VPN proxy ID mismatches instantly. Without it, you’re guessing.

Isn’t BGP overkill for most cloud engineers?

Not if you’re working with hybrid cloud or multi-cloud architectures. Direct Connect, ExpressRoute, and site-to-site VPNs all rely on BGP for dynamic routing. If you can’t read a BGP table or understand AS path prepending, you’re dependent on someone else to fix your infrastructure. That’s a career-limiting move.

How can I practice networking fundamentals without buying physical hardware?

Use virtual labs. GNS3, EVE-NG, and even Docker containers with FRRouting let you build complex topologies on your laptop. You can simulate BGP peering, OSPF areas, VLAN trunking, and firewall rules. Break things intentionally and fix them. The cloud is a production environment; your laptop is a playground. Use it.

What’s the most common networking mistake you see in cloud deployments?

Assuming that security groups are the only traffic control mechanism. I constantly see engineers forget about route tables, network ACLs, and the fact that traffic leaving the VPC to the internet gets NATed unless you’ve set up an egress-only internet gateway or a NAT instance correctly. They focus on the shiny security group rules and ignore the plumbing underneath.