~/tech-with-ugur

Where is your outbound traffic actually going? Build an Envoy egress gateway that routes and meters every request

2026-09-02 cybersecurityobservability

Run the companion lab

Ask a team about inbound traffic and you get a carefully argued answer: the load balancer, the WAF, the security groups reviewed line by line. Ask the same team where their services are allowed to connect out to, and the answer is usually a shrug and a default route to the internet.

That asymmetry is not academic. It is the path a data exfiltration takes on the way out, and it is the path a surprise five-figure bandwidth bill takes on the way to the invoice. Nobody notices either one, because nothing is counting.

AWS made the same argument at length this summer in Prevent data exfiltration: AWS egress controls for cloud workloads — outbound is the under-attended direction, it is how a compromised workload reaches its command-and-control, and the fix is layered: network inspection with domain filtering, DNS-level blocking, data perimeters at the API layer. The coverage that followed extended the concern to agentic systems that can be talked into sending data outward. Every one of those write-ups recommends the same architecture, and it is one nobody demonstrates end to end on a laptop: default-deny egress through a controlled chokepoint, with per-destination allow-listing.

So this lab builds it, small enough to run in an evening. Two workloads sit on a private Docker network with internal: true — no gateway, no route to the host, no route anywhere. One dual-homed Envoy container joins that network and the network the destinations live on, so it is the only door. It allow-lists by FQDN, answers everything else 403 itself, and meters every byte per destination into Prometheus and Grafana. Then it tries to break its own control twice rather than asserting it works.

Everything runs offline: the “internet” is three mock HTTP servers in containers, and every destination is a name under the reserved example.com documentation domain, so nothing here can resolve to or reach anything real. Every code block below is copied verbatim from the lab.

Three networks, one door

The whole architecture rests on one Compose keyword.

labs/lab-envoy-egress-gateway/compose.yaml:

networks:
  # The private workload subnet. `internal: true` means Docker creates no
  # gateway for it: containers attached to it have no route to the host, to
  # other networks, or to the internet. This is what makes the whole lab honest
  # - the workloads are not merely configured to use the proxy, they have no
  # alternative.
  workload_net:
    internal: true
  # The far side of the gateway, standing in for the internet.
  egress_net:
  # Operations: the gateway's admin interface, Prometheus and Grafana. A fixed
  # subnet lets egress-proxy take a stable address here so Envoy's admin
  # listener can bind to it specifically instead of every interface.
  ops_net:
    ipam:
      config:
        - subnet: 10.30.0.0/24

internal: true tells Docker not to create a gateway interface for that bridge. A container attached to it has no default route at all — not to the internet, not to the host, not to another Docker network. That is an honest stand-in for a VPC private subnet with no NAT gateway, which is exactly how you would build this on a cloud provider, rather than a hand-wave.

The workloads are on that network and nothing else. No ports:, no second interface, no network_mode:

  client-checkout:
    build: ./client
    environment:
      CLIENT_NAME: client-checkout
      HTTP_PROXY: http://egress-proxy:3128
      RUN_SECONDS: ${RUN_SECONDS:-120}
      REQUEST_PLAN: http://api.payments.example.com/orders@1000,http://assets.cdn.example.com/bundle.js@20000
      DENIED_URL: http://exfil.shadow-analytics.example.com/collect
      BYPASS_HOST: payments-direct.example.com
      BYPASS_PORT: "8080"
    networks:
      - workload_net

egress-proxy is the only container on more than one traffic network, which is what makes it a gateway rather than a suggestion:

  egress-proxy:
    image: envoyproxy/envoy:v1.39.1
    volumes:
      - ./proxy/envoy.yaml:/etc/envoy/envoy.yaml:ro
    networks:
      workload_net: {}
      egress_net: {}
      ops_net:
        ipv4_address: 10.30.0.10

The three destinations live on egress_net behind network aliases equal to the FQDNs they answer as — which is why, further down, the gateway’s route config reads exactly as it would in production. They are one image deployed three times, differing only in how much they send back:

  upstream-payments:
    build: ./upstream
    environment:
      DESTINATION_NAME: api.payments.example.com
      PAYLOAD_BYTES: "2048"
    networks:
      egress_net:
        aliases:
          - api.payments.example.com
          - payments-direct.example.com

2 KB for the API archetype, 64 KB for assets.cdn.example.com, and 512 KB for telemetry.metrics.example.com — the endpoint somebody turned on once and nobody remembers. Hold on to that second alias on the payments container; it is the whole basis of the bypass test later.

The allow-list is route config

There is no allow-list file in this lab and no plugin. The allow-list is the gateway’s HTTP route table: one virtual host per permitted destination, matched on the request’s :authority.

labs/lab-envoy-egress-gateway/proxy/envoy.yaml:

                route_config:
                  name: egress_routes
                  # The allow-list. One virtual host per destination we permit,
                  # matched on the request's authority. Add a destination by
                  # adding a virtual host and a cluster - nothing else.
                  virtual_hosts:
                    - name: payments
                      domains: ["api.payments.example.com", "api.payments.example.com:*"]
                      routes:
                        - name: payments_route
                          match: { prefix: "/" }
                          route: { cluster: payments }

Each virtual host names a cluster, and the cluster says where that destination actually is:

    - name: payments
      type: STRICT_DNS
      connect_timeout: 2s
      track_cluster_stats: { request_response_sizes: true }
      load_assignment:
        cluster_name: payments
        endpoints:
          - lb_endpoints:
              - endpoint:
                  address:
                    socket_address: { address: api.payments.example.com, port_value: 8080 }

STRICT_DNS means Envoy resolves that name itself, on egress_net, and keeps the resolution fresh. track_cluster_stats: { request_response_sizes: true } is the switch that makes Envoy emit per-destination body-size histograms — remember it, because half the metering story depends on it.

Everything that matches none of the three destination virtual hosts falls through to the last one, whose domains is ["*"]:

                    # Default deny. Anything that did not match a virtual host
                    # above is answered here by the gateway itself and never
                    # reaches any upstream.
                    - name: denied
                      domains: ["*"]
                      routes:
                        - name: denied_route
                          match: { prefix: "/" }
                          direct_response:
                            status: 403
                            body:
                              inline_string: "egress denied: destination not in allow-list\n"

direct_response means the gateway writes that answer out itself. There is no cluster behind denied — so an off-list request has nowhere it could be forwarded, even if the route were misconfigured. The failure mode of a typo is “denied”, not “silently allowed”, and that property is worth more than it looks.

Adding a destination is those two blocks and nothing else: a virtual host and a cluster. That is the entire change-management story for this control, and it is a large part of why building it on route config beats a bespoke filter.

A detour: making the workloads speak forward-proxy

HTTP_PROXY in the environment is a convention, not a mechanism, and the Node runtime is a good example of how thin that convention is. Node’s fetch() ignores the variable entirely, and undici’s ProxyAgent establishes a CONNECT tunnel — which would hide the authority and the body sizes from the gateway and quietly destroy the metrics this lab is about. So the workloads issue the forward-proxy request themselves.

labs/lab-envoy-egress-gateway/client/src/egress/proxyClient.ts:

// A forward proxy expects the *absolute-form* request target: the request line
// carries the whole URL, so the gateway sees the :authority pseudo-header and
// can route on the destination FQDN. Node's fetch() does not honour HTTP_PROXY
// and undici's ProxyAgent tunnels with CONNECT, which would hide the authority
// and the body sizes from the gateway - so the request is made explicitly here.
export function requestViaProxy(args: {
  proxyHost: string;
  proxyPort: number;
  url: string;
  timeoutMs: number;
}): Promise<ProxyResponse> {
  const target = new URL(args.url);
  return new Promise((resolve) => {
    const req = http.request(
      {
        host: args.proxyHost,
        port: args.proxyPort,
        method: "GET",
        path: args.url,
        headers: { host: target.host },
        timeout: args.timeoutMs,
      },
      # ...

path is the full URL rather than /orders. That is absolute-form, and it is the entire difference between a request the gateway can route on and an opaque tunnel it can only count bytes through.

What a denial looks like from both sides

Both workloads spend most of a run on their allow-listed destinations. Five seconds in, each one also asks for exfil.shadow-analytics.example.com, which is not on the list. From the workload’s side, in the structured summary it prints before it exits:

labs/lab-envoy-egress-gateway/README.md, from a default 120-second run:

  "deniedCount": 1,
  "failureCount": 0,
  "denied": {
    "url": "http://exfil.shadow-analytics.example.com/collect",
    "status": 403,
    "bodyPreview": "egress denied: destination not in allow-list\n"
  },

From the gateway’s side, the same two attempts in its access log:

egress-proxy-1  | {"authority":"exfil.shadow-analytics.example.com","bytes_received":0,"bytes_sent":45,"downstream":"172.18.0.3:56942","duration_ms":0,"response_code":403,"route_name":"denied_route","upstream_cluster":null}
egress-proxy-1  | {"authority":"exfil.shadow-analytics.example.com","bytes_received":0,"bytes_sent":45,"downstream":"172.18.0.4:57080","duration_ms":0,"response_code":403,"route_name":"denied_route","upstream_cluster":null}

route_name: denied_route with upstream_cluster: null is the gateway saying “I answered this myself; it went nowhere.” Two lines with two different downstream addresses: both workloads tried.

That null cluster has a consequence for the metrics. A denial never reaches a cluster, so no cluster statistic can count it — the denial count lives on the listener instead, keyed by response-code class. It is a small thing to get wrong when you build a dashboard for this and then wonder why the denial panel is permanently empty.

Proving the control

Configuring a workload to use a proxy proves nothing. An attacker who owns the process simply does not use it. So both workloads spend part of every run trying to get out around the gateway, and the lab asserts that they cannot — twice, at two different layers.

By name. Each workload resolves payments-direct.example.com and opens a plain TCP connection to port 8080, with no proxy in the path.

labs/lab-envoy-egress-gateway/client/src/egress/bypass.ts:

// The same attempt, starting from a name. On the workload network the name of a
// destination that lives on the far side of the gateway does not resolve at all,
// so the bypass usually dies one step earlier than the connection attempt.
export async function attemptDirectByName(args: {
  host: string;
  port: number;
  timeoutMs: number;
  resolve4?: Resolve4;
}): Promise<BypassOutcome> {
  const resolve4 =
    args.resolve4 ??
    ((host: string) =>
      new Resolver({ timeout: args.timeoutMs, tries: 1 }).resolve4(host));
  let addresses: string[];
  try {
    addresses = await resolve4(args.host);
  } catch (err) {
    return { blocked: true, stage: "dns", code: errorCode(err) };
  }
  # ...

It dies at the first step, and the workload records why:

labs/lab-envoy-egress-gateway/README.md:

"bypass": { "host": "payments-direct.example.com", "blocked": true, "stage": "dns", "code": "ESERVFAIL" }

The destination genuinely exists — payments-direct.example.com is that second alias on the upstream-payments container. But it is an alias on egress_net, and the workload is on workload_net, where Docker’s embedded DNS will not answer for it and the internal network gives the resolver nowhere to forward the question.

The alias is also deliberately absent from the gateway’s route config. It is not a virtual host and not a cluster, so if that name ever showed up in the access log it would mean a workload had found a route the lab did not intend. Its absence is therefore evidence, which is precisely what the end-to-end gate checks:

labs/lab-envoy-egress-gateway/e2e.sh:

log_hits=$(docker compose logs --no-log-prefix egress-proxy | grep -c 'payments-direct.example.com' || true)
check "$([ "$log_hits" = "0" ] && echo true || echo false)" \
  "the bypass attempt left no trace in the gateway's access log" "lines=$log_hits"

By raw IP. DNS is the easy layer to blame, so the gate goes a step lower. It reads the destination container’s actual address off egress_net with docker inspect, then runs a probe from a workload that connects to that literal address with no name resolution involved at all:

# Same claim one layer lower: hand a workload the destination's raw IP address,
# with no name resolution involved, and it still has no route to it. The address
# lookup is deliberately not allowed to abort the script: if it fails, that is a
# failed assertion like any other, not a silent exit under `set -e`.
direct_ip=$(docker inspect -f '{{ (index .NetworkSettings.Networks (printf "%s_egress_net" (index .Config.Labels "com.docker.compose.project"))).IPAddress }}' \
  "$(docker compose ps -q upstream-payments)" 2>/dev/null) || direct_ip=""

The kernel answers ENETUNREACH — network is unreachable. Not a missing name, not a filtered packet: no route, and nothing the process could have done differently. That is the difference between a workload that has been asked to use the gateway and a workload that has no alternative.

The interface that nearly undid it

There is a third probe in that gate, and it exists because of a mistake worth repeating.

Envoy’s admin interface is how this lab reads its own meter: Prometheus scrapes /stats/prometheus off it directly, no exporter and no sidecar anywhere. The obvious way to make that reachable is to bind admin to 0.0.0.0:9901. That was the original design, and it is wrong — because 0.0.0.0 on a container attached to three networks means all three, including the workload network. The unauthenticated admin interface, which can do considerably more than serve statistics, would have been one HTTP request away from the very workloads the lab claims are boxed in.

labs/lab-envoy-egress-gateway/proxy/envoy.yaml:

admin:
  # Bound to the gateway's ops_net address, not 0.0.0.0: only Prometheus, which
  # is attached to ops_net, can reach this unauthenticated interface. The
  # workload network has no route here at all. In production you go one step
  # further and bind admin to localhost, shipping statistics out through a
  # stats sink instead.
  address:
    socket_address: { address: 10.30.0.10, port_value: 9901 }

Binding to a specific address means the socket does not exist on the workload interface at all. And since a claim you have not tried to break is not a claim, the gate fires the same raw-IP probe at it:

labs/lab-envoy-egress-gateway/e2e.sh:

# The admin interface is the meter, and it can do a great deal more than serve
# statistics. It is bound to the ops network, which the workloads are not on, so
# the same probe must fail against it too.
if docker compose run --rm --no-deps -T client-checkout npm run bypass-probe -- 10.30.0.10 9901 > /tmp/egress-admin-probe.log 2>&1; then
  pass "a workload cannot reach the gateway's admin interface on the ops network"
else
  fail "a workload cannot reach the gateway's admin interface on the ops network"
  cat /tmp/egress-admin-probe.log
fi

That fixed 10.30.0.10 is why ops_net has a hardcoded subnet: Envoy has to know the address to bind it, before Docker would otherwise have assigned one.

Reading the meter

Now the half nobody builds. The gateway is the chokepoint, so it is also the only place in the system that can count. Four statistics carry the whole story, and Prometheus scrapes them straight off the admin interface every five seconds.

How many requests went where. A cluster is a destination, so envoy_cluster_upstream_rq_total is requests per destination:

labs/lab-envoy-egress-gateway/README.md:

curl -sG --data-urlencode 'query=sum by (envoy_cluster_name) (envoy_cluster_upstream_rq_total)' \
  http://localhost:9090/api/v1/query | jq -r '.data.result[] | "\(.metric.envoy_cluster_name) \(.value[1])"'
cdn 12
payments 119
telemetry 11

How many bytes. envoy_cluster_upstream_cx_rx_bytes_total is what came back from each destination — the number an egress bill is computed from:

curl -sG --data-urlencode 'query=sum by (envoy_cluster_name) (envoy_cluster_upstream_cx_rx_bytes_total)' \
  http://localhost:9090/api/v1/query | jq -r '.data.result[] | "\(.metric.envoy_cluster_name) \(.value[1])"'
cdn 788424
payments 263347
telemetry 5769005

Put those two results side by side and the point of the whole exercise falls out. Payments is by far the chattiest destination — 119 requests against telemetry’s 11 — and by far the cheapest. Telemetry, on eleven requests across two minutes, moved more than five and a half megabytes: twenty-two times the payments traffic, from an endpoint that would never appear in a request-count dashboard as anything but noise. Request counts and byte counts answer different questions, and only one of them is on the invoice.

How big a typical response is. envoy_cluster_upstream_rs_body_size_bucket is a histogram, so quantiles come out of it:

curl -sG --data-urlencode 'query=histogram_quantile(0.5, sum by (envoy_cluster_name, le) (rate(envoy_cluster_upstream_rs_body_size_bucket[5m])))' \
  http://localhost:9090/api/v1/query | jq -r '.data.result[] | "\(.metric.envoy_cluster_name) \(.value[1])"'
cdn 49152
payments 1536
telemetry 393216

Those are interpolations within the bucket holding each exact payload (2048, 65536 and 524288 bytes), which is the best a histogram can do — and it is only that good because the buckets were replaced. Envoy’s defaults are far too coarse to tell a 2 KB API response from a 512 KB upload; with them, all three destinations land in one bucket and the panel says nothing at all.

labs/lab-envoy-egress-gateway/proxy/envoy.yaml:

stats_config:
  # Envoy's default histogram buckets are far too coarse to tell a 2 KB API
  # response apart from a 512 KB telemetry upload. These buckets bracket each of
  # the three payload archetypes exactly.
  histogram_bucket_settings:
    - match:
        prefix: "cluster."
      buckets: [512, 1024, 2048, 4096, 8192, 16384, 32768, 65536, 131072, 262144, 524288, 1048576]

Powers of two bracketing each archetype. This was the part of the build flagged as most likely to need a fallback, and in the end the buckets were enough: the end-to-end gate asserts each destination’s median lands in the bucket (PAYLOAD_BYTES/2, PAYLOAD_BYTES] and all three hold.

How often the allow-list held. Denials never reach a cluster, so — as the null upstream in that access-log line already told us — they are counted on the listener, by response-code class:

labs/lab-envoy-egress-gateway/README.md:

curl -sG --data-urlencode 'query=sum(envoy_http_downstream_rq_xx{envoy_http_conn_manager_prefix="egress",envoy_response_code_class="4"})' \
  http://localhost:9090/api/v1/query | jq -r '.data.result[0].value[1]'
2

Two — one per workload. The egress in that label is the HTTP connection manager’s stat_prefix, a statistics namespace that has nothing to do with the listener’s own name, which is a genuinely easy hour to lose.

A provisioned nine-panel Grafana dashboard ships in the repo and draws all four: stat tiles for requests, bytes, distinct destinations and denials; the same split per destination over time; a share-of-volume bar that makes the telemetry endpoint’s dominance visually undeniable; and the body-size percentiles. It comes up with anonymous viewer access, so docker compose up and localhost:3000 is the entire reader experience.

The per-request ledger

Aggregate statistics are keyed by destination. They cannot tell you which workload made a request, and during an incident that is the first thing anyone asks. So the gateway writes one JSON line per request as well.

labs/lab-envoy-egress-gateway/proxy/envoy.yaml:

                # The per-request ledger. Aggregate statistics are keyed by
                # destination and cannot tell you *which workload* made a
                # request; this can, via the downstream address.
                access_log:
                  - name: envoy.access_loggers.stdout
                    typed_config:
                      "@type": type.googleapis.com/envoy.extensions.access_loggers.stream.v3.StdoutAccessLog
                      log_format:
                        json_format:
                          authority: "%REQ(:AUTHORITY)%"
                          route_name: "%ROUTE_NAME%"
                          upstream_cluster: "%UPSTREAM_CLUSTER%"
                          response_code: "%RESPONSE_CODE%"
                          bytes_sent: "%BYTES_SENT%"
                          bytes_received: "%BYTES_RECEIVED%"
                          duration_ms: "%DURATION%"
                          downstream: "%DOWNSTREAM_REMOTE_ADDRESS%"

authority is the destination as the workload wrote it. route_name is which allow-list entry matched. upstream_cluster is where it went, or null when the gateway answered on its own. And downstream — the workload’s own address and port — is the only per-client identity anywhere in the record.

That is also the reconciliation point. The workloads keep their own books, and the gate compares them against the gateway’s independent count:

labs/lab-envoy-egress-gateway/e2e.sh:

gateway_total=$(promq 'sum(envoy_cluster_upstream_rq_total)')
client_total=$(( $(echo "$checkout_summary" | jq -r .totalSuccesses) + $(echo "$batch_summary" | jq -r .totalSuccesses) ))
check "$(awk -v a="$gateway_total" -v b="$client_total" 'BEGIN { print (a == b) ? "true" : "false" }')" \
  "the gateway's request count matches what the workloads report" "gateway=$gateway_total workloads=$client_total"

Exact equality, not a tolerance. A control you cannot reconcile against the thing it controls is not a control.

Where the visibility ends

Four limits, stated plainly, because a post that only lists what a control catches is advertising.

The statistics cannot tell you which workload. Every Envoy cluster statistic is keyed by destination. envoy_cluster_upstream_rq_total{envoy_cluster_name="cdn"} is the sum across both workloads and no label splits it. Per-client attribution here lives in the access log’s downstream field, one line at a time — not on the dashboard. Getting it into the metrics means something heavier: a listener per workload with its own stat_prefix, a header-derived label, or an identity-aware sidecar. That is where a production build of this spends its next week.

This is plaintext HTTP forward proxying. The workloads speak absolute-form HTTP to the gateway, which is why it can see the path, route on the authority and measure body sizes. Real egress is mostly HTTPS, where a forward proxy gets a CONNECT tunnel and nothing more. The destination and the byte counts survive that change — CONNECT names its target, and bytes are bytes — but the paths and the per-response body-size histograms do not. Keeping those under TLS means termination and a trusted CA in every workload, which is a different and much heavier lab.

internal: true is a private subnet, not a data perimeter. It models one real thing well — a subnet with no NAT gateway and no route out. It models none of the rest. No DNS-level blocking, no IAM condition on which principals may call which service, no VPC endpoint policy, no inspection of what an allowed destination does with the data once it has it. The AWS advisory that motivated this describes all of those layers together; this lab is the network one.

The admin interface is exposed on purpose. It is bound to 10.30.0.10:9901 so Prometheus can scrape it. Unreachable from the workloads, unreachable from your host — but reachable by anything on the ops network. In production you bind admin to localhost and ship statistics out through a stats sink, so nothing on any network can reach it at all.

Run it

docker compose up -d --build
docker compose logs -f client-checkout client-batch

Both workloads run for 120 seconds and exit; the gateway, the destinations, Prometheus and Grafana stay up. Then http://localhost:3000 for the dashboard and http://localhost:9090 for the raw series. Or ./e2e.sh for the whole thing unattended — it resets to a clean stack, waits out the run, makes 22 assertions about every claim above that can be asserted, and tears down after itself. About two and a half minutes once the images are built.

Two things can trip you up, both documented in the lab rather than papered over. The 10.30.0.0/24 on ops_net is hardcoded, so a corporate VPN handing out 10.x addresses will collide with it — the README names the error and every one of the four places the subnet is written down. And RUN_SECONDS=600 docker compose up recreates the workloads but not the gateway, so its counters carry the previous run forward; tear the stack down first if you want the numbers to describe a single run.

The lab is here: labs/lab-envoy-egress-gateway. The statistics reference for everything above is Envoy’s own cluster statistics documentation.

Two hours of work gets you an answer to a question most teams cannot answer at all: how many requests left, for where, carrying how many bytes — and proof that nothing left any other way.