Isolating unsafe shell execution with Kubernetes Jobs
A shell runner can work perfectly and still expose everything mounted beside its output directory. If it also has network access, an attacker can download another program before returning a harmless-looking result.
This lab deliberately accepts shell commands. Two Hono servers create Kubernetes Jobs and collect their output through PVCs. The insecure and secure releases use the same server and worker image bytes. Their Job configurations determine what the command can see and reach.
The comparison asks two concrete questions: can a worker retrieve a seeded private canary, and can it download a harmless script from a local HTTP fixture? Both releases must first prove they can execute an ordinary command. A failed request alone says little about isolation.
Follow one request
The lab runs on a single-node kind cluster with Cilium enforcing NetworkPolicy. The insecure release listens through a localhost port forward on port 3000; the secure release uses port 3001. Each release has its own server and PVC. A third namespace hosts the download fixture.
The tested environment is macOS arm64 with Node 22.23.2, npm 10.9.8, Docker Desktop Engine 28.3.2 with 8 GiB available, kind 0.32.0, Helm 4.2.4, and kubectl 1.36.3. Bootstrap checks the prerequisites, creates the named cluster with default CNI disabled, installs Cilium 1.20.1, builds and loads the images, deploys both releases, and starts the owned port forwards. Other platforms have not been verified.
From the lab directory, the reader workflow is:
From labs/lab-kubernetes-job-isolation/README.md:
npm ci
npm run bootstrap
npm run e2e
Setup downloads dependencies and images. Runtime E2E checks use the local cluster and fixture; they need no cloud credentials or real malware. Keep the unauthenticated endpoints on localhost.
From labs/lab-kubernetes-job-isolation/README.md:
curl -sS http://127.0.0.1:3000/execute \
-H 'content-type: application/json' \
-d '{"message":"printf '\''hello\\n'\''"}'
Send the same request to port 3001 for the secure comparison. Each response contains a different server-generated UUID, exit code 0, and the captured hello line.
The request passes through five steps:
- Hono validates the JSON and body size, then waits for an execution slot.
- The server generates a UUID, prepares storage, and creates a fixed Job. The request contributes command text; it cannot choose the image, mount, service account, or security context.
- Kubernetes starts an Ubuntu-based worker. Its Node launcher invokes
/bin/sh -cwith the supplied text and captures combined stdout and stderr. - The launcher flushes the UUID-named Markdown output file, then publishes a separate exit-metadata file.
- The server waits for Job completion, validates and reads both files, and returns the captured text and command exit code.
Passing the message as JSON configuration keeps it out of the launcher’s construction. The launcher still intentionally executes that message as shell syntax. Infrastructure must contain the resulting process.
Collection success and command success are separate. A command that writes stderr and exits 7 returns HTTP 200 with exit code 7. The launcher exits successfully after saving that result. Missing or invalid files are infrastructure failures and return generic HTTP 500. Deadlines return generic HTTP 504 after awaited deletion. Jobs have no retries, so a failed collection does not repeat command side effects.
Extract private data in the insecure release
The server seeds a clearly synthetic record at /data/private/pii.json. Its private directory is root-owned with mode 0700; the file has mode 0600 and a fresh FAKE-PII- canary. These permissions are useful only if the worker lacks root access and cannot see unrelated storage.
The insecure worker runs as root and mounts the whole release PVC at /home/runner/data. The E2E submits a command that prints an execution marker, searches the mounted data for pii.json, and returns its contents. The check compares the returned bytes with the actual seeded record obtained independently through an operator inspection.
The secure worker must return the execution marker without the canary. The operator checks that its separate release also contains a correctly permissioned synthetic record. This rules out an empty or missing target as the explanation for denial.
Private records are only one target. The suite first creates a result containing a unique marker, confirms that result remains on the PVC, then gives a subsequent execution its known path. The insecure worker can read the earlier result. The secure worker cannot retrieve that marker. Additional parent-path probes check traversal behavior, and operator inspection confirms that the prior result remains intact afterward.
Restrict the worker’s view
The secure worker receives one server-selected PVC subdirectory:
From labs/lab-kubernetes-job-isolation/src/kubernetes/template.ts:
volumeMounts: config.secure
? [
{
name: "data",
mountPath: "/home/runner/data",
subPath: `runs/${id}`,
},
{ name: "tmp", mountPath: "/tmp" },
]
: [{ name: "data", mountPath: "/home/runner/data" }],
Inside the secure container, /home/runner/data is that execution’s directory. Its parent is a container filesystem directory; stepping upward does not reveal the parent of the PVC subPath. The server retains its whole-volume mount to seed private data and collect results.
The identity configuration complements the mount:
From labs/lab-kubernetes-job-isolation/src/kubernetes/template.ts:
securityContext: config.secure
? {
runAsUser: 10001,
runAsGroup: 10001,
runAsNonRoot: true,
seccompProfile: { type: "RuntimeDefault" },
}
: { runAsUser: 0, runAsGroup: 0 },
The root server assigns the secure run directory to UID/GID 10001. It does not add an fsGroup that would make the private fixture group-readable. The worker also drops every capability, disables privilege escalation, and uses a read-only root filesystem with a separate 64 MiB writable /tmp emptyDir. Its service-account token is not mounted.
| Control | Insecure worker | Secure worker |
|---|---|---|
| PVC visibility | Whole release volume | Current execution subPath |
| UID/GID | 0 | 10001 |
| Capabilities | Container defaults | All dropped |
| Privilege escalation | Allowed | Disabled |
| Root filesystem | Writable | Read-only |
| Service-account token | Mounted | Automount disabled |
| Network policy | No worker isolation policy | No ingress or egress allowances |
| Seccomp | Runtime defaults | Explicit RuntimeDefault |
The suite combines runtime commands with live Pod specifications. It checks the secure UID, zero capability sets, NoNewPrivs, token absence, and a root-filesystem write returning EROFS. Neither release uses privileged workers, host namespaces, direct hostPath mounts, or container-engine sockets.
Prove outbound blocking without depending on DNS
The fixture serves a fixed harmless /marker.sh. Tests download and inspect its exact bytes; they never execute it. A request log supplies an observation independent of the worker’s output.
The secure release installs this worker-selected policy:
From labs/lab-kubernetes-job-isolation/deploy/chart/templates/policy.yaml:
{{- if .Values.secure }}
{"apiVersion":"networking.k8s.io/v1","kind":"NetworkPolicy","metadata":{"name":"deny-worker-network","namespace":{{ .Release.Namespace | toJson }}},"spec":{"podSelector":{"matchLabels":{"app.kubernetes.io/component":"worker"}},"policyTypes":["Ingress","Egress"],"ingress":[],"egress":[]}}
{{- end }}
The Job builder places the worker component label on both Jobs and Pods. The server remains outside that selector and retains Kubernetes API access. Secure workers need no network callbacks because results travel through storage.
A hostname download failure could mean only that DNS is blocked. The suite therefore tests three routes separately: the fixture hostname, its actual Service IP, and its actual Pod IP. For each route, an insecure download succeeds before the secure attempt and again afterward. Each successful control must return the marker, leave its exact bytes on disk, and add one fixture request event.
For a secure direct-IP denial, the evidence predicate is:
From labs/lab-kubernetes-job-isolation/src/e2e/network-download.ts:
export function blockedIp(evidence: {
exitCode: number | undefined;
before: number;
after: number;
positiveBefore: boolean;
positiveAfter: boolean;
markerExists: boolean;
}): boolean {
return (
evidence.exitCode === 28 &&
evidence.before === evidence.after &&
evidence.positiveBefore &&
evidence.positiveAfter &&
!evidence.markerExists
);
}
That combines curl’s timeout exit code, no downloaded file, unchanged fixture request count, and successful controls on both sides of the denial. The target was reachable during the comparison, and DNS was unnecessary for the attempted route. A separate bounded resolution check measures DNS itself.
The policy declares both ingress and egress denial, but these probes measure outbound fixture access. They do not establish ingress isolation, every node-address path, or general internet blocking. Those claims would require additional evidence.
Treat result files as attacker-controlled
A private mount is not sufficient if the controller follows a worker-created symlink when collecting output. The submitted command can tamper with its own output and metadata paths, so collection is another trust boundary.
The server’s reader opens with no-follow and nonblocking flags, then checks the opened inode:
From labs/lab-kubernetes-job-isolation/src/storage/safe-files.ts:
const handle = await open(
path,
constants.O_RDONLY | constants.O_NOFOLLOW | constants.O_NONBLOCK,
);
try {
// Require a single-link regular file within the expected byte limit.
const before = await handle.stat();
if (!before.isFile() || before.nlink !== 1 || before.size > maximum)
throw new ExecutionError("infrastructure");
The rest of the reader uses a bounded allocation, reads one extra byte to detect growth beyond the limit, and rechecks the descriptor against the named path. It always closes the file handle. The launcher also validates result paths before publishing metadata.
Deployed probes replace output or metadata with symlinks, FIFOs, directories, invalid metadata, or missing files. They require generic infrastructure errors without a leaked canary, followed by a normal successful execution. Filesystem unit tests separately exercise the server reader: a deployed rejection by the launcher alone cannot prove the reader is safe.
Compare evidence and respect the bounds
The machine-readable artifacts/e2e-report.json records each operation, expected outcome, observed response and output, and supporting live Job/Pod evidence. npm run e2e exits nonzero if an assertion fails.
| Observation | Insecure | Secure |
|---|---|---|
| Ordinary command and known nonzero exit | Preserved | Preserved |
| Seeded synthetic private record | Exact canary exposed | Command runs without canary |
| Known retained result | Prior marker exposed | Prior marker unavailable |
| Fixture download by Service/Pod IP | Exact bytes, fixture hit | Timeout, no file, no hit |
| Token availability | Nonempty token file, bytes not returned | Read fails |
| Root-filesystem write | Succeeds | EROFS |
| Invalid collection paths | Generic error and recovery | Generic error and recovery |
Both modes bound the request body to 8 KiB and combined capture to 64 KiB. Capture overflow sets an explicit truncation indication and stops the process group. Each Job has a 30-second active deadline. The 60-second request budget reserves five seconds for confirmed cleanup, leaving 55 seconds for queueing, scheduling, and execution. Two slots per release limit concurrent executions.
The suite inspects timeout Job and Pod deletion and known output-path cleanup, then verifies recovery. It also executes 130 retention filler Jobs per release to measure the 128-completed-result cap and checks that old sentinels are pruned. The one-hour age limit has unit-test and live TTL coverage; the suite does not wait an hour.
These bounds cover captured results and known resources. Commands can create extra files in writable storage. Retention bookkeeping is in memory, so server restarts can leave older PVC output outside that accounting until teardown. A process escaping the original process group relies on container termination for final cleanup. None of this establishes comprehensive denial-of-service protection.
Logging has its own exposure boundary. Structured operation logs include submitted command text in arguments. A secret embedded literally in a command can therefore appear in logs. Captured output and token bytes are excluded from operation logs; the local E2E report intentionally retains synthetic test evidence. Backend diagnostics keep safe error categories and codes while omitting raw exception bodies that could echo commands or results.
The attacker controls command text. The server, fixed Job template, administrator, runtime, enforcement components, host, and kernel remain trusted. A compromised controller creating unrestricted Jobs and kernel exploits are outside this demonstration.
Use npm run teardown when finished. It stops the recorded port forwards and removes only the positively identified lab cluster and its PVC data. The companion lab contains the implementation, prerequisites, and full evidence checklist.