Where is your outbound traffic actually going? Build an Envoy egress gateway that routes and meters every request
Ask a team about inbound traffic and you get a carefully argued answer: the load balancer, the WAF, the security groups reviewed line by line. Ask the same team where their services are allowed to connect out to, and the answer is usually a shrug and a default route to the internet.
That asymmetry is not academic. It is the path a data exfiltration takes on the way out, and it is the path a surprise five-figure bandwidth bill takes on the way to the invoice. Nobody notices either one, because nothing is counting.
AWS made the same argument at length this summer in Prevent data exfiltration: AWS egress controls for cloud workloads — outbound is the under-attended direction, it is how a compromised workload reaches its command-and-control, and the fix is layered: network inspection with domain filtering, DNS-level blocking, data perimeters at the API layer. The coverage that followed extended the concern to agentic systems that can be talked into sending data outward. Every one of those write-ups recommends the same architecture, and it is one nobody demonstrates end to end on a laptop: default-deny egress through a controlled chokepoint, with per-destination allow-listing.
So this lab builds it, small enough to run in an evening. Two workloads sit on a
private Docker network with internal: true — no gateway, no route to the host,
no route anywhere. One dual-homed Envoy container joins that network and the
network the destinations live on, so it is the only door. It allow-lists by
FQDN, answers everything else 403 itself, and meters every byte per
destination into Prometheus and Grafana. Then it tries to break its own control
twice rather than asserting it works.
Everything runs offline: the “internet” is three mock HTTP servers in
containers, and every destination is a name under the reserved example.com
documentation domain, so nothing here can resolve to or reach anything real.
Every code block below is copied verbatim from the lab.
Three networks, one door
The whole architecture rests on one Compose keyword.
labs/lab-envoy-egress-gateway/compose.yaml:
networks:
# The private workload subnet. `internal: true` means Docker creates no
# gateway for it: containers attached to it have no route to the host, to
# other networks, or to the internet. This is what makes the whole lab honest
# - the workloads are not merely configured to use the proxy, they have no
# alternative.
workload_net:
internal: true
# The far side of the gateway, standing in for the internet.
egress_net:
# Operations: the gateway's admin interface, Prometheus and Grafana. A fixed
# subnet lets egress-proxy take a stable address here so Envoy's admin
# listener can bind to it specifically instead of every interface.
ops_net:
ipam:
config:
- subnet: 10.30.0.0/24
internal: true tells Docker not to create a gateway interface for that bridge.
A container attached to it has no default route at all — not to the internet,
not to the host, not to another Docker network. That is an honest stand-in for a
VPC private subnet with no NAT gateway, which is exactly how you would build
this on a cloud provider, rather than a hand-wave.
The workloads are on that network and nothing else. No ports:, no second
interface, no network_mode:
client-checkout:
build: ./client
environment:
CLIENT_NAME: client-checkout
HTTP_PROXY: http://egress-proxy:3128
RUN_SECONDS: ${RUN_SECONDS:-120}
REQUEST_PLAN: http://api.payments.example.com/orders@1000,http://assets.cdn.example.com/bundle.js@20000
DENIED_URL: http://exfil.shadow-analytics.example.com/collect
BYPASS_HOST: payments-direct.example.com
BYPASS_PORT: "8080"
networks:
- workload_net
egress-proxy is the only container on more than one traffic network, which is
what makes it a gateway rather than a suggestion:
egress-proxy:
image: envoyproxy/envoy:v1.39.1
volumes:
- ./proxy/envoy.yaml:/etc/envoy/envoy.yaml:ro
networks:
workload_net: {}
egress_net: {}
ops_net:
ipv4_address: 10.30.0.10
The three destinations live on egress_net behind network aliases equal to the
FQDNs they answer as — which is why, further down, the gateway’s route config
reads exactly as it would in production. They are one image deployed three
times, differing only in how much they send back:
upstream-payments:
build: ./upstream
environment:
DESTINATION_NAME: api.payments.example.com
PAYLOAD_BYTES: "2048"
networks:
egress_net:
aliases:
- api.payments.example.com
- payments-direct.example.com
2 KB for the API archetype, 64 KB for assets.cdn.example.com, and 512 KB for
telemetry.metrics.example.com — the endpoint somebody turned on once and
nobody remembers. Hold on to that second alias on the payments container;
it is the whole basis of the bypass test later.
The allow-list is route config
There is no allow-list file in this lab and no plugin. The allow-list is the
gateway’s HTTP route table: one virtual host per permitted destination, matched
on the request’s :authority.
labs/lab-envoy-egress-gateway/proxy/envoy.yaml:
route_config:
name: egress_routes
# The allow-list. One virtual host per destination we permit,
# matched on the request's authority. Add a destination by
# adding a virtual host and a cluster - nothing else.
virtual_hosts:
- name: payments
domains: ["api.payments.example.com", "api.payments.example.com:*"]
routes:
- name: payments_route
match: { prefix: "/" }
route: { cluster: payments }
Each virtual host names a cluster, and the cluster says where that destination actually is:
- name: payments
type: STRICT_DNS
connect_timeout: 2s
track_cluster_stats: { request_response_sizes: true }
load_assignment:
cluster_name: payments
endpoints:
- lb_endpoints:
- endpoint:
address:
socket_address: { address: api.payments.example.com, port_value: 8080 }
STRICT_DNS means Envoy resolves that name itself, on egress_net, and keeps
the resolution fresh. track_cluster_stats: { request_response_sizes: true } is
the switch that makes Envoy emit per-destination body-size histograms — remember
it, because half the metering story depends on it.
Everything that matches none of the three destination virtual hosts falls
through to the last one, whose domains is ["*"]:
# Default deny. Anything that did not match a virtual host
# above is answered here by the gateway itself and never
# reaches any upstream.
- name: denied
domains: ["*"]
routes:
- name: denied_route
match: { prefix: "/" }
direct_response:
status: 403
body:
inline_string: "egress denied: destination not in allow-list\n"
direct_response means the gateway writes that answer out itself. There is no
cluster behind denied — so an off-list request has nowhere it could be
forwarded, even if the route were misconfigured. The failure mode of a typo is
“denied”, not “silently allowed”, and that property is worth more than it looks.
Adding a destination is those two blocks and nothing else: a virtual host and a cluster. That is the entire change-management story for this control, and it is a large part of why building it on route config beats a bespoke filter.
A detour: making the workloads speak forward-proxy
HTTP_PROXY in the environment is a convention, not a mechanism, and the Node
runtime is a good example of how thin that convention is. Node’s fetch()
ignores the variable entirely, and undici’s ProxyAgent establishes a
CONNECT tunnel — which would hide the authority and the body sizes from the
gateway and quietly destroy the metrics this lab is about. So the workloads
issue the forward-proxy request themselves.
labs/lab-envoy-egress-gateway/client/src/egress/proxyClient.ts:
// A forward proxy expects the *absolute-form* request target: the request line
// carries the whole URL, so the gateway sees the :authority pseudo-header and
// can route on the destination FQDN. Node's fetch() does not honour HTTP_PROXY
// and undici's ProxyAgent tunnels with CONNECT, which would hide the authority
// and the body sizes from the gateway - so the request is made explicitly here.
export function requestViaProxy(args: {
proxyHost: string;
proxyPort: number;
url: string;
timeoutMs: number;
}): Promise<ProxyResponse> {
const target = new URL(args.url);
return new Promise((resolve) => {
const req = http.request(
{
host: args.proxyHost,
port: args.proxyPort,
method: "GET",
path: args.url,
headers: { host: target.host },
timeout: args.timeoutMs,
},
# ...
path is the full URL rather than /orders. That is absolute-form, and it is
the entire difference between a request the gateway can route on and an opaque
tunnel it can only count bytes through.
What a denial looks like from both sides
Both workloads spend most of a run on their allow-listed destinations. Five
seconds in, each one also asks for exfil.shadow-analytics.example.com, which is
not on the list. From the workload’s side, in the structured summary it prints
before it exits:
labs/lab-envoy-egress-gateway/README.md, from a default 120-second run:
"deniedCount": 1,
"failureCount": 0,
"denied": {
"url": "http://exfil.shadow-analytics.example.com/collect",
"status": 403,
"bodyPreview": "egress denied: destination not in allow-list\n"
},
From the gateway’s side, the same two attempts in its access log:
egress-proxy-1 | {"authority":"exfil.shadow-analytics.example.com","bytes_received":0,"bytes_sent":45,"downstream":"172.18.0.3:56942","duration_ms":0,"response_code":403,"route_name":"denied_route","upstream_cluster":null}
egress-proxy-1 | {"authority":"exfil.shadow-analytics.example.com","bytes_received":0,"bytes_sent":45,"downstream":"172.18.0.4:57080","duration_ms":0,"response_code":403,"route_name":"denied_route","upstream_cluster":null}
route_name: denied_route with upstream_cluster: null is the gateway saying
“I answered this myself; it went nowhere.” Two lines with two different
downstream addresses: both workloads tried.
That null cluster has a consequence for the metrics. A denial never reaches a
cluster, so no cluster statistic can count it — the denial count lives on the
listener instead, keyed by response-code class. It is a small thing to get wrong
when you build a dashboard for this and then wonder why the denial panel is
permanently empty.
Proving the control
Configuring a workload to use a proxy proves nothing. An attacker who owns the process simply does not use it. So both workloads spend part of every run trying to get out around the gateway, and the lab asserts that they cannot — twice, at two different layers.
By name. Each workload resolves payments-direct.example.com and opens a
plain TCP connection to port 8080, with no proxy in the path.
labs/lab-envoy-egress-gateway/client/src/egress/bypass.ts:
// The same attempt, starting from a name. On the workload network the name of a
// destination that lives on the far side of the gateway does not resolve at all,
// so the bypass usually dies one step earlier than the connection attempt.
export async function attemptDirectByName(args: {
host: string;
port: number;
timeoutMs: number;
resolve4?: Resolve4;
}): Promise<BypassOutcome> {
const resolve4 =
args.resolve4 ??
((host: string) =>
new Resolver({ timeout: args.timeoutMs, tries: 1 }).resolve4(host));
let addresses: string[];
try {
addresses = await resolve4(args.host);
} catch (err) {
return { blocked: true, stage: "dns", code: errorCode(err) };
}
# ...
It dies at the first step, and the workload records why:
labs/lab-envoy-egress-gateway/README.md:
"bypass": { "host": "payments-direct.example.com", "blocked": true, "stage": "dns", "code": "ESERVFAIL" }
The destination genuinely exists — payments-direct.example.com is that second
alias on the upstream-payments container. But it is an alias on egress_net,
and the workload is on workload_net, where Docker’s embedded DNS will not
answer for it and the internal network gives the resolver nowhere to forward the
question.
The alias is also deliberately absent from the gateway’s route config. It is not a virtual host and not a cluster, so if that name ever showed up in the access log it would mean a workload had found a route the lab did not intend. Its absence is therefore evidence, which is precisely what the end-to-end gate checks:
labs/lab-envoy-egress-gateway/e2e.sh:
log_hits=$(docker compose logs --no-log-prefix egress-proxy | grep -c 'payments-direct.example.com' || true)
check "$([ "$log_hits" = "0" ] && echo true || echo false)" \
"the bypass attempt left no trace in the gateway's access log" "lines=$log_hits"
By raw IP. DNS is the easy layer to blame, so the gate goes a step lower.
It reads the destination container’s actual address off egress_net with
docker inspect, then runs a probe from a workload that connects to that
literal address with no name resolution involved at all:
# Same claim one layer lower: hand a workload the destination's raw IP address,
# with no name resolution involved, and it still has no route to it. The address
# lookup is deliberately not allowed to abort the script: if it fails, that is a
# failed assertion like any other, not a silent exit under `set -e`.
direct_ip=$(docker inspect -f '{{ (index .NetworkSettings.Networks (printf "%s_egress_net" (index .Config.Labels "com.docker.compose.project"))).IPAddress }}' \
"$(docker compose ps -q upstream-payments)" 2>/dev/null) || direct_ip=""
The kernel answers ENETUNREACH — network is unreachable. Not a missing name,
not a filtered packet: no route, and nothing the process could have done
differently. That is the difference between a workload that has been asked to
use the gateway and a workload that has no alternative.
The interface that nearly undid it
There is a third probe in that gate, and it exists because of a mistake worth repeating.
Envoy’s admin interface is how this lab reads its own meter: Prometheus scrapes
/stats/prometheus off it directly, no exporter and no sidecar anywhere. The
obvious way to make that reachable is to bind admin to 0.0.0.0:9901. That was
the original design, and it is wrong — because 0.0.0.0 on a container attached
to three networks means all three, including the workload network. The
unauthenticated admin interface, which can do considerably more than serve
statistics, would have been one HTTP request away from the very workloads the
lab claims are boxed in.
labs/lab-envoy-egress-gateway/proxy/envoy.yaml:
admin:
# Bound to the gateway's ops_net address, not 0.0.0.0: only Prometheus, which
# is attached to ops_net, can reach this unauthenticated interface. The
# workload network has no route here at all. In production you go one step
# further and bind admin to localhost, shipping statistics out through a
# stats sink instead.
address:
socket_address: { address: 10.30.0.10, port_value: 9901 }
Binding to a specific address means the socket does not exist on the workload interface at all. And since a claim you have not tried to break is not a claim, the gate fires the same raw-IP probe at it:
labs/lab-envoy-egress-gateway/e2e.sh:
# The admin interface is the meter, and it can do a great deal more than serve
# statistics. It is bound to the ops network, which the workloads are not on, so
# the same probe must fail against it too.
if docker compose run --rm --no-deps -T client-checkout npm run bypass-probe -- 10.30.0.10 9901 > /tmp/egress-admin-probe.log 2>&1; then
pass "a workload cannot reach the gateway's admin interface on the ops network"
else
fail "a workload cannot reach the gateway's admin interface on the ops network"
cat /tmp/egress-admin-probe.log
fi
That fixed 10.30.0.10 is why ops_net has a hardcoded subnet: Envoy has to
know the address to bind it, before Docker would otherwise have assigned one.
Reading the meter
Now the half nobody builds. The gateway is the chokepoint, so it is also the only place in the system that can count. Four statistics carry the whole story, and Prometheus scrapes them straight off the admin interface every five seconds.
How many requests went where. A cluster is a destination, so
envoy_cluster_upstream_rq_total is requests per destination:
labs/lab-envoy-egress-gateway/README.md:
curl -sG --data-urlencode 'query=sum by (envoy_cluster_name) (envoy_cluster_upstream_rq_total)' \
http://localhost:9090/api/v1/query | jq -r '.data.result[] | "\(.metric.envoy_cluster_name) \(.value[1])"'
cdn 12
payments 119
telemetry 11
How many bytes. envoy_cluster_upstream_cx_rx_bytes_total is what came back
from each destination — the number an egress bill is computed from:
curl -sG --data-urlencode 'query=sum by (envoy_cluster_name) (envoy_cluster_upstream_cx_rx_bytes_total)' \
http://localhost:9090/api/v1/query | jq -r '.data.result[] | "\(.metric.envoy_cluster_name) \(.value[1])"'
cdn 788424
payments 263347
telemetry 5769005
Put those two results side by side and the point of the whole exercise falls out. Payments is by far the chattiest destination — 119 requests against telemetry’s 11 — and by far the cheapest. Telemetry, on eleven requests across two minutes, moved more than five and a half megabytes: twenty-two times the payments traffic, from an endpoint that would never appear in a request-count dashboard as anything but noise. Request counts and byte counts answer different questions, and only one of them is on the invoice.
How big a typical response is. envoy_cluster_upstream_rs_body_size_bucket
is a histogram, so quantiles come out of it:
curl -sG --data-urlencode 'query=histogram_quantile(0.5, sum by (envoy_cluster_name, le) (rate(envoy_cluster_upstream_rs_body_size_bucket[5m])))' \
http://localhost:9090/api/v1/query | jq -r '.data.result[] | "\(.metric.envoy_cluster_name) \(.value[1])"'
cdn 49152
payments 1536
telemetry 393216
Those are interpolations within the bucket holding each exact payload (2048, 65536 and 524288 bytes), which is the best a histogram can do — and it is only that good because the buckets were replaced. Envoy’s defaults are far too coarse to tell a 2 KB API response from a 512 KB upload; with them, all three destinations land in one bucket and the panel says nothing at all.
labs/lab-envoy-egress-gateway/proxy/envoy.yaml:
stats_config:
# Envoy's default histogram buckets are far too coarse to tell a 2 KB API
# response apart from a 512 KB telemetry upload. These buckets bracket each of
# the three payload archetypes exactly.
histogram_bucket_settings:
- match:
prefix: "cluster."
buckets: [512, 1024, 2048, 4096, 8192, 16384, 32768, 65536, 131072, 262144, 524288, 1048576]
Powers of two bracketing each archetype. This was the part of the build flagged
as most likely to need a fallback, and in the end the buckets were enough: the
end-to-end gate asserts each destination’s median lands in the bucket
(PAYLOAD_BYTES/2, PAYLOAD_BYTES] and all three hold.
How often the allow-list held. Denials never reach a cluster, so — as the
null upstream in that access-log line already told us — they are counted on the
listener, by response-code class:
labs/lab-envoy-egress-gateway/README.md:
curl -sG --data-urlencode 'query=sum(envoy_http_downstream_rq_xx{envoy_http_conn_manager_prefix="egress",envoy_response_code_class="4"})' \
http://localhost:9090/api/v1/query | jq -r '.data.result[0].value[1]'
2
Two — one per workload. The egress in that label is the HTTP connection
manager’s stat_prefix, a statistics namespace that has nothing to do with the
listener’s own name, which is a genuinely easy hour to lose.
A provisioned nine-panel Grafana dashboard ships in the repo and draws all four:
stat tiles for requests, bytes, distinct destinations and denials; the same
split per destination over time; a share-of-volume bar that makes the telemetry
endpoint’s dominance visually undeniable; and the body-size percentiles. It
comes up with anonymous viewer access, so docker compose up and
localhost:3000 is the entire reader experience.
The per-request ledger
Aggregate statistics are keyed by destination. They cannot tell you which workload made a request, and during an incident that is the first thing anyone asks. So the gateway writes one JSON line per request as well.
labs/lab-envoy-egress-gateway/proxy/envoy.yaml:
# The per-request ledger. Aggregate statistics are keyed by
# destination and cannot tell you *which workload* made a
# request; this can, via the downstream address.
access_log:
- name: envoy.access_loggers.stdout
typed_config:
"@type": type.googleapis.com/envoy.extensions.access_loggers.stream.v3.StdoutAccessLog
log_format:
json_format:
authority: "%REQ(:AUTHORITY)%"
route_name: "%ROUTE_NAME%"
upstream_cluster: "%UPSTREAM_CLUSTER%"
response_code: "%RESPONSE_CODE%"
bytes_sent: "%BYTES_SENT%"
bytes_received: "%BYTES_RECEIVED%"
duration_ms: "%DURATION%"
downstream: "%DOWNSTREAM_REMOTE_ADDRESS%"
authority is the destination as the workload wrote it. route_name is which
allow-list entry matched. upstream_cluster is where it went, or null when the
gateway answered on its own. And downstream — the workload’s own address and
port — is the only per-client identity anywhere in the record.
That is also the reconciliation point. The workloads keep their own books, and the gate compares them against the gateway’s independent count:
labs/lab-envoy-egress-gateway/e2e.sh:
gateway_total=$(promq 'sum(envoy_cluster_upstream_rq_total)')
client_total=$(( $(echo "$checkout_summary" | jq -r .totalSuccesses) + $(echo "$batch_summary" | jq -r .totalSuccesses) ))
check "$(awk -v a="$gateway_total" -v b="$client_total" 'BEGIN { print (a == b) ? "true" : "false" }')" \
"the gateway's request count matches what the workloads report" "gateway=$gateway_total workloads=$client_total"
Exact equality, not a tolerance. A control you cannot reconcile against the thing it controls is not a control.
Where the visibility ends
Four limits, stated plainly, because a post that only lists what a control catches is advertising.
The statistics cannot tell you which workload. Every Envoy cluster statistic
is keyed by destination. envoy_cluster_upstream_rq_total{envoy_cluster_name="cdn"}
is the sum across both workloads and no label splits it. Per-client attribution
here lives in the access log’s downstream field, one line at a time — not on
the dashboard. Getting it into the metrics means something heavier: a listener
per workload with its own stat_prefix, a header-derived label, or an
identity-aware sidecar. That is where a production build of this spends its next
week.
This is plaintext HTTP forward proxying. The workloads speak absolute-form
HTTP to the gateway, which is why it can see the path, route on the authority
and measure body sizes. Real egress is mostly HTTPS, where a forward proxy gets
a CONNECT tunnel and nothing more. The destination and the byte counts survive
that change — CONNECT names its target, and bytes are bytes — but the paths and
the per-response body-size histograms do not. Keeping those under TLS means
termination and a trusted CA in every workload, which is a different and much
heavier lab.
internal: true is a private subnet, not a data perimeter. It models one
real thing well — a subnet with no NAT gateway and no route out. It models none
of the rest. No DNS-level blocking, no IAM condition on which principals may
call which service, no VPC endpoint policy, no inspection of what an allowed
destination does with the data once it has it. The AWS advisory that motivated
this describes all of those layers together; this lab is the network one.
The admin interface is exposed on purpose. It is bound to 10.30.0.10:9901
so Prometheus can scrape it. Unreachable from the workloads, unreachable from
your host — but reachable by anything on the ops network. In production you bind
admin to localhost and ship statistics out through a stats sink, so nothing on
any network can reach it at all.
Run it
docker compose up -d --build
docker compose logs -f client-checkout client-batch
Both workloads run for 120 seconds and exit; the gateway, the destinations,
Prometheus and Grafana stay up. Then http://localhost:3000 for the dashboard
and http://localhost:9090 for the raw series. Or ./e2e.sh for the whole
thing unattended — it resets to a clean stack, waits out the run, makes 22
assertions about every claim above that can be asserted, and tears down after
itself. About two and a half minutes once the images are built.
Two things can trip you up, both documented in the lab rather than papered over.
The 10.30.0.0/24 on ops_net is hardcoded, so a corporate VPN handing out
10.x addresses will collide with it — the README names the error and every one
of the four places the subnet is written down. And RUN_SECONDS=600 docker compose up recreates the workloads but not the gateway, so its counters carry
the previous run forward; tear the stack down first if you want the numbers to
describe a single run.
The lab is here:
labs/lab-envoy-egress-gateway.
The statistics reference for everything above is Envoy’s own cluster
statistics
documentation.
Two hours of work gets you an answer to a question most teams cannot answer at all: how many requests left, for where, carrying how many bytes — and proof that nothing left any other way.