The Host Owns the Wires: a GCP Shared-VPC landing zone that proves its own guardrails
Every enterprise cloud diagram makes the same promise. The network team owns the network. Service teams own their workloads and nothing else. Traffic enters through one controlled front door and leaves through one controlled back door. The boxes are labeled, the arrows are drawn, and everyone in the meeting nods.
Almost nobody tests it. The promise lives in IAM bindings scattered across projects, folders, subnets, and load balancer policies, and the usual evidence that it holds is that nobody has complained yet. “Service team A cannot touch the shared network” is a falsifiable claim — you could impersonate service team A and try — but in most organizations that experiment has never been run.
This lab runs it. It stands up a miniature enterprise GCP landing zone
from an empty organization — Google’s enterprise foundations
blueprint
condensed to the smallest shape that still has all the moving parts:
two folders, one Shared-VPC host project, two service projects, one
least-privilege service account per project, centralized ingress
through a global external Application Load Balancer with Cloud Armor,
and centralized egress through a Secure Web Proxy. Then a scripted
verifier logs in as each project’s service account and attempts five
things a well-run landing zone must refuse or allow in a very specific
way. Every proof is a real API call against real GCP resources; the run
exits 0 only if all five guarantees hold.
Everything below comes from the runnable lab linked above, and the run described is a real one against a real GCP organization: six Terraform stages applied in order (about 15 minutes), the full proof suite green, a single-pass reverse-order destroy (about 11 minutes), and a bill in single-digit dollars.
Here is the whole topology, from the lab’s
labs/lab-gcp-shared-vpc-foundation/README.md:
organization
│
├── folder: networking
│ └── project: svpc-host-<suffix> (Shared VPC host)
│ ├── VPC: svpc
│ │ ├── snet-host 10.10.0.0/24 (host resources, incl. SWP IP)
│ │ ├── snet-service-a 10.10.1.0/24 (networkUser: sa-service-a, not sa-service-b)
│ │ ├── snet-service-b 10.10.2.0/24 (networkUser: sa-service-b, not sa-service-a)
│ │ └── snet-proxy-only 10.10.4.0/23 (REGIONAL_MANAGED_PROXY, required for the SWP)
│ ├── Global external ALB + Cloud Armor (edge-allowlist policy)
│ │ └── hybrid NEGs → VM IPs in the service projects
│ └── Secure Web Proxy (swp, TCP 443 on the host subnet)
│
└── folder: workloads
├── project: svpc-service-a-<suffix> (service project A)
│ └── vm-service-a (10.10.1.10, attached to snet-service-a)
└── project: svpc-service-b-<suffix> (service project B)
└── vm-service-b (10.10.2.10, attached to snet-service-b)
data paths:
client --80--> [ALB + Cloud Armor, host project] --hybrid NEG--> vm-service-a / vm-service-b
vm-service-a/b --443--> [Secure Web Proxy, host project] --allowlisted domains only--> internet
The org chart is the security model
Enterprise GCP does not start with VMs. It starts with folders and
projects, because in GCP the resource hierarchy is the security
boundary: IAM roles granted on a project stop at that project’s edge,
and the whole landing-zone pattern is about exploiting that edge
deliberately. The lab’s first Terraform stage builds exactly that
bedrock — a networking folder for the host project, a workloads
folder for the two service projects — and nothing else.
labs/lab-gcp-shared-vpc-foundation/terraform/01_foundation/05_projects.tf:
resource "random_id" "suffix" {
byte_length = 3
}
resource "google_project" "host" {
project_id = "svpc-host-${random_id.suffix.hex}"
name = "svpc-host"
folder_id = google_folder.networking.name
billing_account = var.billing_account
auto_create_network = false
deletion_policy = "DELETE"
}
resource "google_project" "service_a" {
project_id = "svpc-service-a-${random_id.suffix.hex}"
name = "svpc-service-a"
folder_id = google_folder.workloads.name
billing_account = var.billing_account
auto_create_network = false
deletion_policy = "DELETE"
}
# ...
Two small flags carry weight here. auto_create_network = false
suppresses the default VPC GCP would otherwise helpfully create in
every project — in a Shared-VPC world, service projects owning any
network of their own is exactly the failure mode being designed out.
And the random_id suffix on every project ID exists because project
IDs are globally unique and soft-deleted projects squat on their names
for 30 days; without it, running the lab twice in a month would
collide with the ghost of the previous run.
The lab splits this and everything after it into six staged Terraform
roots — 01_foundation through 06_workloads — each independently
applyable, each reading the previous stage’s outputs through a
terraform_remote_state data source pointed at a local backend. A
single root would apply faster, but the layering is the lesson:
folders and projects, then identities, then the shared network, then
ingress, then egress, then finally the workloads that exercise it all.
The host owns the wires
The “Shared” in Shared VPC is one resource on the host side and one attachment per service project, and the entire pattern hangs on them.
labs/lab-gcp-shared-vpc-foundation/terraform/03_network/04_shared_vpc.tf:
resource "google_compute_shared_vpc_host_project" "host" {
project = local.host_project_id
}
resource "google_compute_shared_vpc_service_project" "service_a" {
host_project = google_compute_shared_vpc_host_project.host.project
service_project = local.service_a_project_id
}
resource "google_compute_shared_vpc_service_project" "service_b" {
host_project = google_compute_shared_vpc_host_project.host.project
service_project = local.service_b_project_id
}
After this, the VPC, every subnet, and every firewall rule live in the host project — the service projects contain compute and nothing else. Each service project gets exactly one subnet, sized and numbered by the host. And before any workload exists, the host closes the network in both directions:
labs/lab-gcp-shared-vpc-foundation/terraform/03_network/08_firewall_baseline.tf:
resource "google_compute_firewall" "deny_all_ingress" {
project = local.host_project_id
name = "deny-all-ingress"
network = google_compute_network.svpc.id
direction = "INGRESS"
priority = 65534
deny {
protocol = "all"
}
source_ranges = ["0.0.0.0/0"]
}
resource "google_compute_firewall" "deny_all_egress" {
project = local.host_project_id
name = "deny-all-egress"
network = google_compute_network.svpc.id
direction = "EGRESS"
priority = 65534
deny {
protocol = "all"
}
destination_ranges = ["0.0.0.0/0"]
}
This default-deny baseline is the posture everything later has to justify itself against: the ingress and egress stages don’t “configure access”, they punch specific, host-owned holes through an otherwise closed wall. There is no Cloud NAT and no external IP anywhere in this lab — a VM that wants to reach the internet has precisely one option, and it belongs to the host.
The IAM that makes the diagram true
None of the above means anything without the role bindings, because Shared VPC does not by itself stop a service project’s identity from doing anything — IAM does. The lab creates one service account per project and grants each the narrowest role set that still lets it do its job.
labs/lab-gcp-shared-vpc-foundation/terraform/02_iam/02_locals.tf:
locals {
# ...
host_roles = ["roles/compute.networkAdmin", "roles/compute.securityAdmin"]
service_roles = ["roles/compute.instanceAdmin.v1", "roles/iam.serviceAccountUser"]
}
The host’s service account can administer networks and firewall rules — on the host project only. Each service account can administer VM instances — in its own project only. No service account holds any role on the host project or on the other service project. That asymmetry is the entire landing-zone contract in two lines: the host can wire, the services can compute, and neither can do the other’s job.
The subtlest binding is one level below the project: subnet-level IAM.
Attaching a VM to a Shared-VPC subnet requires
roles/compute.networkUser on that subnet, and the host grants it
per subnet, per tenant:
labs/lab-gcp-shared-vpc-foundation/terraform/03_network/07_subnet_iam.tf:
resource "google_compute_subnetwork_iam_member" "service_a_sa" {
project = local.host_project_id
region = var.region
subnetwork = google_compute_subnetwork.service_a.name
role = "roles/compute.networkUser"
member = "serviceAccount:${local.service_a_sa_email}"
}
# ...
Service-A’s account is a networkUser on snet-service-a and — this
is the load-bearing part — not on snet-service-b. “Each service
project rents exactly one subnet” is not a naming convention or a
diagram label; it is this binding’s absence on every subnet but one.
Proof 3 will try to violate it directly.
One front door: Cloud Armor at the edge
Both workload VMs serve a small HTTP page, but neither is reachable directly: no external IPs, and the VPC firewall admits only Google’s load-balancer and health-check ranges. The only way in is a global external Application Load Balancer in the host project, and bolted to its backend services is a Cloud Armor policy that is closed by default.
labs/lab-gcp-shared-vpc-foundation/terraform/04_ingress/06_cloud_armor.tf:
resource "google_compute_security_policy" "edge" {
project = local.host_project_id
name = "edge-allowlist"
rule {
action = "deny(403)"
priority = 2147483647
description = "Default rule: deny everything the host has not allowed."
match {
versioned_expr = "SRC_IPS_V1"
config {
src_ip_ranges = ["*"]
}
}
}
dynamic "rule" {
for_each = { for i, cidr in var.allowlist_cidrs : i => cidr }
content {
action = "allow"
priority = 1000 + tonumber(rule.key)
description = "Host-approved caller."
match {
versioned_expr = "SRC_IPS_V1"
config {
src_ip_ranges = [rule.value]
}
}
}
}
}
var.allowlist_cidrs defaults to an empty list, so the freshly
deployed front door answers 403 to everyone on the planet. Opening
it is a host-project Terraform apply — which is exactly how the verify
flow will do it later.
There is one genuinely awkward constraint hiding in this stage. A GCP backend service and its backends must live in the same project, and Cloud Armor attaches to the backend service — so if the backends were ordinary instance groups in the service projects, the backend services (and with them the allowlist) would have to live there too, outside host control. The lab’s way out is hybrid network endpoint groups: host-owned NEGs that address the service-project VMs as bare IP:port endpoints.
labs/lab-gcp-shared-vpc-foundation/terraform/04_ingress/05_negs.tf:
resource "google_compute_network_endpoint_group" "vm" {
for_each = local.services
project = local.host_project_id
name = "neg-${each.key}"
zone = var.zone
network = data.terraform_remote_state.network.outputs.network_self_link
network_endpoint_type = "NON_GCP_PRIVATE_IP_PORT"
default_port = 80
}
resource "google_compute_network_endpoint" "vm" {
for_each = local.services
project = local.host_project_id
zone = var.zone
network_endpoint_group = google_compute_network_endpoint_group.vm[each.key].name
ip_address = each.value.vm_ip
port = 80
}
Those ip_address values are computed with cidrhost() from the
subnet CIDRs back in stage 03 — 10.10.1.10 and 10.10.2.10 — so the
endpoints can be registered before the VMs exist, and stage 06 later
assigns the same addresses to the VMs’ NICs. The trade-off is real and
worth stating plainly: to the load balancer these are opaque IP:port
pairs, not instances, so there’s no autohealing or instance-identity
integration. That’s the price of keeping the entire ingress plane —
forwarding rule, URL map, backend services, and above all the Cloud
Armor policy — in the host project.
One back door: the egress proxy
Egress gets the same treatment as ingress: one host-owned choke point, closed by default, opened by explicit policy. The choke point is a Secure Web Proxy — an explicit proxy the VMs must deliberately send their traffic to — with a gateway security policy that allows exactly the domains the host names.
labs/lab-gcp-shared-vpc-foundation/terraform/05_egress/05_policy.tf:
resource "google_network_security_gateway_security_policy" "swp" {
project = local.host_project_id
name = "swp-policy"
location = var.region
}
resource "google_network_security_gateway_security_policy_rule" "allow" {
for_each = { for i, domain in var.allowed_domains : domain => i }
project = local.host_project_id
location = var.region
gateway_security_policy = google_network_security_gateway_security_policy.swp.name
name = "allow-${replace(each.key, ".", "-")}"
priority = 100 + each.value
enabled = true
session_matcher = "host() == '${each.key}'"
basic_profile = "ALLOW"
}
The default allowlist is one domain (github.com), matched by
SNI/host — no TLS inspection, no CA pool; the proxy never sees inside
the connection, it only decides which destinations exist at all. Any
domain without a matching rule has no path out.
But a proxy alone enforces nothing — a VM could simply not use it. The enforcement is the pairing with stage 03’s deny-all egress rule, plus a single allow:
labs/lab-gcp-shared-vpc-foundation/terraform/05_egress/07_firewall_egress.tf:
resource "google_compute_firewall" "allow_egress_to_swp" {
project = local.host_project_id
name = "allow-egress-to-swp"
network = data.terraform_remote_state.network.outputs.network_id
direction = "EGRESS"
priority = 1000
allow {
protocol = "tcp"
ports = ["443"]
}
destination_ranges = ["${local.swp_ip}/32"]
target_tags = ["web"]
}
From a workload VM, the reachable internet is one IP and one port: the proxy’s. Everything else dies at the VPC firewall before it ever leaves the network. The proxy then applies the domain policy. Two layers, both host-owned, and the verifier will test both.
Five proofs, each a real API call
Here is where the lab stops describing its security model and starts
interrogating it. A TypeScript verifier (on Bun) impersonates the
lab’s service accounts through IAM Credentials — the runner’s own
identity holds serviceAccountTokenCreator on each SA, granted in
stage 02 for exactly this purpose:
labs/lab-gcp-shared-vpc-foundation/verifier/src/gcp/impersonate.ts:
export async function impersonatedAuth(
targetPrincipal: string,
): Promise<GoogleAuth> {
const baseAuth = new GoogleAuth({ scopes: [CLOUD_PLATFORM_SCOPE] });
const sourceClient = await baseAuth.getClient();
const client = new Impersonated({
sourceClient,
targetPrincipal,
targetScopes: [CLOUD_PLATFORM_SCOPE],
lifetime: 3600,
});
return new GoogleAuth({ authClient: client });
}
So when the verifier “acts as service team A”, that’s not a
simulation — it holds a real short-lived token for sa-service-a and
calls the real Compute API with it. The five proofs:
- The host owns the network. The host SA creates and deletes a probe firewall rule on the shared VPC (must succeed); service-A’s SA attempts the same insert (must be denied).
- No cross-project workloads. Service-A’s SA attempts to create a VM in service project B (must be denied).
- No subnet borrowing. Service-A’s SA attempts to create a VM in its own project, attached to service B’s subnet (must be denied — by the subnet-level IAM from stage 03).
- Ingress only when the host allows. The load balancer answers
403from Cloud Armor while the allowlist is empty, and200from both VMs once the host allowlists the runner’s IP. - Egress only through the host’s proxy. From inside the VMs: direct internet access fails, the allowlisted domain succeeds through the proxy, a non-allowlisted domain is denied by the proxy.
The negative proofs have a subtlety that’s easy to get wrong: IAM
propagation delay. A permission granted (or revoked) seconds ago can
flicker, so “it was denied once” doesn’t prove policy — it might just
be a binding that hasn’t propagated. The verifier’s retry policy is
therefore asymmetric: allowed-checks retry with backoff until they
succeed, but denied-checks must be denied consistently, across
spaced attempts, and anything other than the expected
PERMISSION_DENIED fails the proof rather than counting toward it.
labs/lab-gcp-shared-vpc-foundation/verifier/src/lib/retry.ts:
export async function expectConsistentDenial(
op: () => Promise<unknown>,
opts: {
isExpectedDenial: (err: unknown) => boolean;
attempts?: number;
delayMs?: number;
sleep?: (ms: number) => Promise<void>;
},
): Promise<{ deniedConsistently: boolean; detail: string }> {
const {
isExpectedDenial,
attempts = 3,
delayMs = 20000,
sleep = defaultSleep,
} = opts;
for (let attempt = 1; attempt <= attempts; attempt += 1) {
try {
await op();
return {
deniedConsistently: false,
detail: `attempt ${attempt}/${attempts} unexpectedly succeeded`,
};
} catch (err) {
if (!isExpectedDenial(err)) {
return {
deniedConsistently: false,
detail: `attempt ${attempt}/${attempts} failed with a non-denial error: ${String(err)}`,
};
}
}
if (attempt < attempts) await sleep(delayMs);
}
return {
deniedConsistently: true,
detail: `denied on all ${attempts} attempts`,
};
}
A denial caused by a typo’d project ID or an expired token would be a
NOT_FOUND or an auth error — treated as a failed proof, not a passed
one. A test that can’t distinguish “denied by policy” from “broken
differently” isn’t testing the policy.
Proof 4: the host flips the switch
The ingress proof is the one that plays out in two acts, because its
whole point is a before/after. The verify flow first confirms the
closed door — the load balancer must answer the Cloud Armor 403,
stably, across multiple samples. Then the host opens it, and the
mechanism matters: not a console click, not a gcloud one-liner, but
a Terraform apply of the ingress stage with one new tfvars file.
labs/lab-gcp-shared-vpc-foundation/scripts/verify.sh:
echo "==> Allowlisting this runner's IP at the edge (host-project apply)"
runner_ip="$(curl -fsS https://api.ipify.org)"
cat > "$tf/04_ingress/allowlist.auto.tfvars" <<EOF
allowlist_cidrs = ["$runner_ip/32"]
EOF
"$root/scripts/deploy_cloud.sh" --stage 04_ingress
The policy change goes through the same pipeline as the
infrastructure itself, so there is no drift between what Terraform
believes and what the edge enforces — and make destroy deletes the
tfvars file again, so the next deploy starts closed. After the apply,
phase 2 requires 200 from both /a and /b, each serving a page
that names the VM behind it. In the verified run, the closed door held
403 on every sample, and both paths answered 200 with the right
page after the flip.
Proof 5: inside the walls
The egress proof has to run inside the VMs — the claim is about what a workload can reach, not what the runner can. But the VMs have no external IPs, so the verifier SSHes in through IAP TCP tunnels — and even that path is a small demonstration of the thesis, because the IAP range is admitted by a host-owned firewall rule; the host controls the maintenance door too. Over SSH, three commands per VM:
labs/lab-gcp-shared-vpc-foundation/verifier/src/proofs/egress.ts:
function buildChecks(
config: VerifierConfig,
proxy: string,
vmLabel: string,
): EgressCheck[] {
return [
{
expectation: `direct internet access from ${vmLabel} fails (default-deny egress)`,
command: `curl -sS -m 8 -o /dev/null https://${config.allowedDomain}`,
timeoutMs: 30000,
passWhen: (code) => code !== 0,
},
{
expectation: `the allowlisted domain succeeds through the host's proxy from ${vmLabel}`,
command: `curl -sS -m 20 -o /dev/null -w '%{http_code}' -x ${proxy} https://${config.allowedDomain}`,
timeoutMs: 40000,
passWhen: (code) => code === 0,
},
{
expectation: `a non-allowlisted domain is denied by the host's proxy from ${vmLabel}`,
command: `curl -sS -m 20 -o /dev/null -x ${proxy} https://${config.deniedDomain}`,
timeoutMs: 40000,
passWhen: (code) => code !== 0,
},
];
}
Note that the first and second checks request the same domain —
github.com, which the proxy allows. The only variable between them
is the path: direct versus through the proxy. That isolates the thing
being tested. If the direct request fails and the proxied one
succeeds, the difference cannot be the destination’s fault; it is the
host’s egress policy, working. In the verified run the exit codes told
the story precisely: the direct request died with curl’s timeout exit
code 28 (the packets left the VM and were silently eaten by the
deny-all egress rule), the proxied github.com request returned
200, and the proxied example.com request was refused by the proxy
at the CONNECT stage — denied not by the network but by the policy,
exactly one layer further in.
Failing in the expected way is the standard each check applies. The
verifier’s summary line reports pass/fail per expectation, and make verify exits non-zero if any of them — allowed or denied — comes out
wrong.
Run it, then try to break it yourself
From a fresh clone, per
labs/lab-gcp-shared-vpc-foundation/README.md:
make deploy # ~15-25 min; stages apply in order 01 -> 06
make verify # phase 1 (closed door), edge allowlist apply, phase 2 (all five proofs)
make destroy # reverse-order teardown 06 -> 01
The honest caveat: this lab needs a real GCP organization, because
folders and Shared VPC have no emulator and no sandbox — and a
personal Gmail account doesn’t have an organization. The README
documents a free path via Cloud Identity (a domain you own, a TXT
record, one sign-in; 30–60 minutes, once) and the four org-level roles
you need: Folder Admin, Project Creator, Billing Account User, and
Shared VPC Admin. One finding from the verified run worth passing on:
Organization Admin is not a superset here — it doesn’t include
folders.create, so Folder Admin and the Shared-VPC admin role had to
be granted explicitly, exactly as the prerequisites list says.
The stages can also be applied one at a time (make foundation,
make iam, …, each previewable with --dry-run), which is the better
way to internalize the layering: watch the org chart appear, then the
identities, then the closed network, then the two doors, and only then
the workloads. Teardown runs the stages in reverse and deletes all
three projects; the random suffix means the next run starts clean.
What transfers out of the lab is less the Terraform than the habit. Every guarantee in this landing zone — host-only network authority, per-subnet tenancy, the closed-by-default edge, proxy-only egress — exists as a binding or a policy that some future refactor could silently drop, and none of the five proofs would be hard to port to a real landing zone’s CI. “Service teams can’t touch the network” is either a sentence in your architecture doc or a test that ran this morning. The gap between those two is the entire difference, and closing it costs about twenty minutes and a few dollars a run.