~/tech-with-ugur

A code-execution sandbox for a chat agent on kind

2026-09-18 aiclouddevops

Run the companion lab

Give a language model three café receipts and ask what a coffee costs, and it will answer with three confident numbers. Sometimes they are even correct. The model is predicting a plausible continuation, not solving a linear system, and nothing in its output separates the two cases.

A code-execution tool changes the question the model is being asked. Instead of producing the answer, it produces a program that computes the answer, and something else runs that program and hands back real stdout. The model still has to model the problem and pick a solver; the arithmetic belongs to CPython and numpy.

This lab builds that loop end to end on a local kind cluster: an assistant-ui chat front end, a Hono server hosting a LangGraph agent, a Python sandbox service with numpy, pandas, scipy, sympy and scikit-learn baked into its image, and Postgres holding the conversation checkpoints. It runs with no API key by default. A deterministic stand-in model plays the same café problem every time, and builds its final answer from whatever the sandbox actually returned, so the whole path — tool call, execution, structured result, rendered answer — is exercised without a model provider anywhere in the picture.

Four services on one cluster

The cluster runs four workloads in one namespace, applied as plain manifests with kustomize:

Every step is a Makefile target, and the reader workflow is two commands:

From labs/lab-ai-code-sandbox-kind/README.md:

make up          # cluster, images, secrets, deploy, port-forward
open http://localhost:3000

make up and make down never change the current kubectl context: every command targets kind-ai-code-sandbox explicitly, and the context is saved and restored around cluster creation and deletion even when a step fails partway through. The tested environment is macOS on Apple silicon with Docker Desktop engine 28.3.2 (at least 8 GiB allotted), kind v0.32.0 with its node image pinned by digest, kubectl v1.36.3, and Node 24 for the test harness. Other platforms have not been verified.

A tool description the model can trust

A tool description is a contract the model reads once and then acts on for the rest of the conversation. If it says scipy is available and the image does not have it, every failure looks like a model mistake. Hand-written descriptions drift from the image they describe, quietly, on the first dependency bump.

So this one is not hand-written. The sandbox reports what it actually is, and the server turns that report into the description when it starts:

From labs/lab-ai-code-sandbox-kind/server/src/sandbox/tool-description.ts:

// Generated from the sandbox's own /capabilities so every claim the model
// reads is true of the image that will run its code.
export function describeCodeExecutor(capabilities: Capabilities): string {
  const { limits } = capabilities;
  return [
    "Runs a complete Python program in an isolated sandbox and returns its status, exit code, stdout, stderr and an optional structured result.",
    "Use for: every calculation — arithmetic, solving equations or linear systems, optimisation, fitting, statistics, symbolic maths and checking a proposed answer.",
    "Do NOT use for: fetching data from the internet, installing packages, keeping state between calls, or work that needs longer than the timeout.",
    `Environment: Python ${capabilities.pythonVersion}, standard library plus the modules below.`,
    installedModulesLine(capabilities),
    `Limits: wall-clock timeout ${limits.executionTimeoutSeconds} s; stdout truncated after ${limits.maxStdoutBytes} bytes and stderr after ${limits.maxStderrBytes} bytes; code at most ${limits.maxCodeBytes} bytes; result.json at most ${limits.maxResultBytes} bytes; at most ${limits.maxConcurrentExecutions} executions run at once.`,
    capabilities.network,
    capabilities.persistence,
    capabilities.structuredResult,
    "Response: JSON with status (succeeded | failed | timed_out), exitCode, stdout, stderr, result (parsed result.json or null), resultError, durationMs and truncated. If the sandbox is busy or unreachable you get {error, message}: retry once, then tell the user.",
  ].join("\n");
}

The three sentences pulled straight from capabilities — network, persistence and structuredResult — are written once, in the sandbox, next to the code that enforces them. Everything else is interpolated from the same document the sandbox serves. The result the model sees, read live from the running server, is:

From labs/lab-ai-code-sandbox-kind/README.md:

Runs a complete Python program in an isolated sandbox and returns its status, exit code, stdout, stderr and an optional structured result.
Use for: every calculation — arithmetic, solving equations or linear systems, optimisation, fitting, statistics, symbolic maths and checking a proposed answer.
Do NOT use for: fetching data from the internet, installing packages, keeping state between calls, or work that needs longer than the timeout.
Environment: Python 3.12.14, standard library plus the modules below.
Installed modules: numpy (numpy 2.5.3), pandas (pandas 3.0.5), scipy (scipy 1.18.1), sympy (sympy 1.14.0), sklearn (scikit-learn 1.9.1).
Limits: wall-clock timeout 30 s; stdout truncated after 65536 bytes and stderr after 65536 bytes; code at most 65536 bytes; result.json at most 65536 bytes; at most 20 executions run at once.
No network use is expected: do not download data, install packages or call external services.
Every execution starts a fresh interpreter in a fresh, empty working directory; no files, variables or imports survive between executions.
To return structured data, write one JSON document to result.json in the working directory (at most 65536 bytes); it comes back parsed as `result`, or `resultError` explains why it was rejected.
Response: JSON with status (succeeded | failed | timed_out), exitCode, stdout, stderr, result (parsed result.json or null), resultError, durationMs and truncated. If the sandbox is busy or unreachable you get {error, message}: retry once, then tell the user.

That is the whole contract: which modules exist and at which versions, every limit as a number, how to hand back structured data, and what the response looks like when the sandbox is busy or unreachable. The code_executor tool itself is a DynamicStructuredTool whose single code field is a zod string bounded by the sandbox’s own maxCodeBytes, described as “a complete, self-contained Python 3 program”.

Generating the description is only half of it; the other half is proving it did not drift. One end-to-end test reads the description off the live server and the capabilities document off the live sandbox and compares them:

From labs/lab-ai-code-sandbox-kind/e2e/src/scripted/tool-description.e2e.test.ts:

    const modulesLine =
      description
        .split("\n")
        .find((l) => l.startsWith("Installed modules: ")) ?? "";
    const described = modulesLine
      .replace("Installed modules: ", "")
      .replace(/\.$/, "")
      .split(", ");
    expect(described).toEqual(
      capabilities.modules.map(
        (m) => `${m.importName} (${m.distribution} ${m.version})`,
      ),
    );

It goes on to assert every one of the six limits, the Python version, both usage lines and all three statements. Bump numpy in the image without redeploying the server, or add a limit the description does not mention, and the suite fails.

One synchronous endpoint

The obvious design for an execution service is a job queue: POST /execute returns an id, the caller polls, results live in a store with a TTL. That machinery exists to solve a problem this lab does not have. A tool call is already one caller waiting on one result, and the model cannot do anything useful until the program finishes.

So POST /execute simply awaits the run and answers on the same connection:

From labs/lab-ai-code-sandbox-kind/sandbox/src/app/api/app_factory.py:

        if not slots.try_acquire():
            log.warning(
                "Refusing execute request: all slots busy.", in_use=slots.in_use
            )
            return JSONResponse(
                {
                    "error": "sandbox_busy",
                    "message": "all execution slots are in use; retry shortly",
                },
                status_code=429,
                headers={"Retry-After": "1"},
            )
        try:
            outcome = await run_code(
                code, settings=settings, runs_root=runs_root, log=log
            )
        finally:
            slots.release()
        return _outcome_response(outcome, log=log)

Because the result travels back only on the request that submitted it, there is no execution store, no ids, no polling and no TTL to manage. The trade-off is that the caller — the server, and behind it the browser — has to stay connected for as long as the code runs, up to the timeout.

Concurrency is a counter rather than a queue. ExecutionSlots.try_acquire takes a slot if one is free and returns False otherwise, so a request that arrives when all twenty slots are busy gets a 429 with Retry-After immediately instead of occupying a held connection while it waits.

The run itself is one awaited subprocess:

From labs/lab-ai-code-sandbox-kind/sandbox/src/app/execution/runner.py:

    process = await asyncio.create_subprocess_exec(
        sys.executable,
        "-I",
        "-u",
        "main.py",
        cwd=workdir,
        env=_minimal_env(workdir),
        stdin=asyncio.subprocess.DEVNULL,
        stdout=asyncio.subprocess.PIPE,
        stderr=asyncio.subprocess.PIPE,
        start_new_session=True,
    )
    stdout_pipe = _require_pipe(process.stdout, "stdout")
    stderr_pipe = _require_pipe(process.stderr, "stderr")
    stdout = BoundedCapture(limit=settings.max_output_bytes)
    stderr = BoundedCapture(limit=settings.max_output_bytes)
    drains = asyncio.gather(stdout.drain(stdout_pipe), stderr.drain(stderr_pipe))
    timed_out = False
    try:
        await asyncio.wait_for(
            process.wait(), timeout=settings.execution_timeout_seconds
        )
    except TimeoutError:
        timed_out = True
    finally:
        _kill_group(process.pid)
        await process.wait()

python -I -u isolates the interpreter from user site directories and environment-based configuration and turns off output buffering. The environment is built from scratch — PATH, HOME, TMPDIR, LANG and the BLAS thread limits, nothing inherited. start_new_session=True puts the program in its own process group so the timeout path can os.killpg the whole group rather than just the Python process, which is what makes a program that spawned children actually stop. Output is captured through bounded pipes, so a runaway print truncates instead of exhausting the pod’s memory, and the finally kills the group on the normal path too. A single uvicorn worker serves all of it: the event loop is free while the subprocesses run.

One detail that only matters when things go wrong: the server’s HTTP timeout to the sandbox has to be strictly longer than the sandbox’s own execution timeout, or a slow run turns into a client-side abort and the model is told the sandbox is unreachable when in fact it was about to report timed_out with the program’s partial output. The defaults are 30 s and 45 s, and the server asserts the ordering at startup and refuses to start if it is wrong.

Proving a run only sees itself

“Each execution is isolated” is easy to write and easy to get wrong, so the suite fires twenty executions at once, each printing and writing its own UUID canary after a randomised sleep, and checks what came back:

From labs/lab-ai-code-sandbox-kind/e2e/src/scripted/isolation.e2e.test.ts:

    // Serial execution would take at least 20 × 1.5 s = 30 s.
    expect(wallMs).toBeLessThan(12_000);
    const cwds = new Set<string>();
    runs.forEach((run, i) => {
      const own = canaries[i] as string;
      expect(run.status).toBe("succeeded");
      expect(run.stderr).toBe("");
      expect(run.result).toEqual({ canary: own });
      const [printed, report] = run.stdout.trim().split("\n");
      expect(printed).toBe(own);
      const details = JSON.parse(report as string) as {
        cwd: string;
        files: string[];
        siblings: unknown;
      };
      expect(details.files).toEqual([
        `canary-${own}`,
        "main.py",
        "result.json",
      ]);
      expect(details.siblings).toBe("denied");
      cwds.add(details.cwd);
      for (const other of canaries) {
        if (other !== own) expect(JSON.stringify(run)).not.toContain(other);
      }
    });
    expect(cwds.size).toBe(CONCURRENT);

The wall-clock assertion is the part that keeps the test honest: twenty serialised 1.5-second runs would take at least 30 seconds, so passing under 12 proves the single worker really did overlap them, which is the only condition under which the isolation assertions mean anything. Each response must carry its own canary in stdout and in result, exactly three files in its own working directory, a distinct cwd, and none of the other nineteen canaries anywhere in the response. Two further tests show that an attribute one execution sets on builtins is gone in the next, and that a finished execution’s directory no longer exists by the time another one looks for it.

siblings is "denied" because the shared runs directory has mode 0o300 — writable and enterable, not listable. The comment above that constant is careful about what it buys:

From labs/lab-ai-code-sandbox-kind/sandbox/src/app/execution/workspace.py:

_RUNS_DIR = "runs"
# Owner may create and enter entries but not list them. That keeps a casual
# os.listdir("..") from listing other runs' directories, but it is not a
# boundary against code that tries: submitted code runs as the UID that owns
# this directory, so it can chmod it back and list it, and it can find other
# runs through /proc/<pid>/cwd.
_RUNS_MODE = 0o300

That is the honest version. It stops a casual os.listdir(".."); it is not a boundary against code that means to get past it.

Walking the café problem

The fixed demo is a three-unknown linear system dressed up as three receipts:

From labs/lab-ai-code-sandbox-kind/README.md:

A café sells coffee, tea and sandwiches. Anna pays €13 for 2 coffees, 1 tea and 1 sandwich. Ben pays €19 for 1 coffee, 3 teas and 2 sandwiches. Cleo pays €18 for 3 coffees, 2 teas and 1 sandwich. What does each item cost?

The system prompt is what turns a chat model into something that reaches for the tool instead of guessing:

From labs/lab-ai-code-sandbox-kind/server/src/agent/system-prompt.ts:

// The date is pinned so the model never mistakes its training cutoff for today.
export function buildSystemPrompt(today: Date): string {
  const date = today.toISOString().slice(0, 10);
  return [
    `Today's date is ${date}.`,
    "You solve maths word problems for the user.",
    "You MUST call the code_executor tool for every calculation. Never do arithmetic in your head.",
    "Work in this order:",
    "1. Model the problem explicitly: name the unknowns and write the equations, constraints or objective.",
    "2. Choose a solver that fits: for example numpy.linalg.solve for linear systems, scipy.optimize for optimisation or curve fitting, sympy for exact symbolic answers.",
    "3. Run it with code_executor and write the solution to result.json.",
    "4. Verify it in code: substitute the solution back, compute residuals or check the constraints.",
    "Answer with exactly these four markdown sections, in this order: ### Model, ### Solver, ### Solution (a markdown table), ### Verification.",
    "Take every number in your answer from the tool results. If the tool fails, say so instead of guessing.",
  ].join("\n");
}

It is rebuilt before every model call rather than baked in once at startup, through a dynamic system prompt middleware, so the date it carries is today’s rather than whatever day the pod happened to start — or, worse, the model’s training cutoff.

In scripted mode the stand-in model answers the café problem with a fixed program:

From labs/lab-ai-code-sandbox-kind/server/src/agent/cafe-script.ts:

export const CAFE_SOLVER_CODE = [
  "import json",
  "import numpy as np",
  "",
  "# Rows: Anna, Ben, Cleo. Columns: coffee, tea, sandwich.",
  "A = np.array([[2, 1, 1], [1, 3, 2], [3, 2, 1]], dtype=float)",
  "b = np.array([13, 19, 18], dtype=float)",
  "x = np.linalg.solve(A, b)",
  "max_residual = float(np.max(np.abs(A @ x - b)))",
  'solution = {name: round(float(v), 6) for name, v in zip(["coffee", "tea", "sandwich"], x)}',
  'print("solution:", solution)',
  'print("max residual:", max_residual)',
  'with open("result.json", "w") as f:',
  '    json.dump({"solution": solution, "maxResidual": max_residual}, f)',
].join("\n");

Three receipts become a 3×3 coefficient matrix, numpy.linalg.solve does the work, the residual check is computed in the same program rather than asserted in prose, and both travel back as result.json. The stand-in model then fills its four-section answer template — Model, Solver, Solution, Verification — only from that tool result. If the tool path breaks, the answer says the sandbox produced no solution rather than quietly printing the right prices from memory, which is exactly what makes the keyless end-to-end test meaningful.

Getting that to the browser is a short bridge. The LangGraph stream becomes an AI SDK UI message stream and is returned as the HTTP response:

From labs/lab-ai-code-sandbox-kind/server/src/chat/chat-routes.ts:

      const stream = await agent.stream(
        { messages: [new HumanMessage(text)] },
        {
          configurable: { thread_id: threadId },
          streamMode: [...STREAM_MODE],
          signal: c.req.raw.signal,
        },
      );
      return createUIMessageStreamResponse({
        stream: toUIMessageStream(stream, {
          onFinish: () => log.info("Streaming chat turn succeeded."),
          onError: (err) => log.error({ err }, "Streaming chat turn failed."),
          onAbort: () => log.warn("Streaming chat turn aborted by the client."),
        }),
      });

On the client, useChatRuntime and AssistantChatTransport point at that endpoint and carry the thread id on every request. A custom tool UI renders each code_executor call as a card with the generated Python, its stdout and the parsed result, so the answer and the thing that produced it sit next to each other.

History is not stored as UI messages. Reloading the page asks the server for the thread, which reads the LangGraph checkpoint through agent.graph.getState and converts it with the same mapper the stream uses — one test runs the scripted model against an in-memory checkpointer and shows the rebuilt assistant message matches the streamed one on its text and tool parts.

Implementation turned up one consequence of streaming a turn over a single request that is worth knowing about before you build this yourself. A turn lives only as long as its HTTP request: reload the page or hit “New chat” while code_executor is running and the turn stops after the model’s tool call has been checkpointed but before its tool result has. The checkpoint is then left holding a tool call with no answer — which the Anthropic API rejects on the next request, permanently breaking that thread in live mode. The fix repairs the request rather than the history:

From labs/lab-ai-code-sandbox-kind/server/src/agent/interrupted-tool-calls-middleware.ts:

// A turn that is cut off while a tool runs (a reload, "New chat", a closed
// tab) aborts the graph after the model's tool call was checkpointed but
// before its tool message was. Providers such as Anthropic reject a request
// that contains an unanswered tool call, so every later turn on that thread
// would fail. This repairs only the request sent to the model; the stored
// history is left exactly as it was.
export function createInterruptedToolCallsMiddleware({
  logger,
}: {
  logger: Logger;
}) {
  return createMiddleware({
    name: "InterruptedToolCallsMiddleware",
    wrapModelCall: (request, handler) => {
      const messages = answerInterruptedToolCalls(request.messages);
      if (messages === request.messages) return handler(request);
      logger.warn(
        { answered: messages.length - request.messages.length },
        "Answering tool calls left unanswered by an interrupted turn.",
      );
      return handler({ ...request, messages: [...messages] });
    },
  });
}

The synthetic message uses the same shape as the tool’s other structured errors, so the model reads it like any failed sandbox call. Nothing is written back to the checkpoint; the interrupted card stays visibly unfinished in the history, which is the truthful rendering of what happened.

What this sandbox does not protect against

This is a teaching sandbox, not a hostile-code isolation system, and the distinction is worth being precise about rather than gesturing at.

What it does have is ordinary pod hardening and no egress:

From labs/lab-ai-code-sandbox-kind/deploy/sandbox.yaml:

      automountServiceAccountToken: false
      # Without an init process the sandbox app is PID 1, and a child that a
      # timed-out program spawned is re-parented to it and never reaped: it
      # stays behind as a zombie after every such run. Sharing the pod's
      # process namespace makes the pause container PID 1, and it reaps them.
      shareProcessNamespace: true
      securityContext:
        runAsNonRoot: true
        runAsUser: 10001
        runAsGroup: 10001
        fsGroup: 10001
        seccompProfile: { type: RuntimeDefault }

From labs/lab-ai-code-sandbox-kind/deploy/sandbox-networkpolicy.yaml:

# Code in the sandbox has no reason to open connections. Deny all egress
# from sandbox pods (DNS included); incoming requests from the server and
# the kubelet's probes are unaffected.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: sandbox-deny-egress
spec:
  podSelector:
    matchLabels: { app: sandbox }
  policyTypes: [Egress]
  egress: []

The container also drops every capability, disables privilege escalation, runs with a read-only root filesystem over a 512 Mi emptyDir at /work, and is capped at 2 cores and 1 GiB. The egress policy is not decorative: kind has shipped kube-network-policies since v0.24, so it is actually enforced, and one of the end-to-end tests proves it by having sandbox code fail to reach Postgres both by IP and by name while the server still can.

What it does not have:

What the suite checks

make e2e runs seven scripted test files against the deployed cluster with no API key: the sandbox HTTP contract (exact module versions, an exact stdout and result for a linear system, a traceback on failure, a timeout that leaves no process behind, output truncation, an unparsable result.json, malformed requests, and the 429), per-request isolation, the generated tool description, the agent loop through the real chat API, persistence across a server-pod restart with no leakage between threads, the network policy, and a headless Chromium run through the actual UI that sends the problem, reads the code card and the answer table, reloads to confirm both come back, and starts a new chat. All 21 tests pass in about 62 seconds on Apple silicon, and the run writes e2e/reports/e2e-report.json plus a full-page screenshot of the finished answer.

make e2e-live runs the same shape of checks against Claude instead of the stand-in — the café problem, a least-squares fit chosen to check the model picks a fitting solver rather than reciting a memorised answer, and the browser flow, all asserting on behaviour rather than wording. It type-checks and lints but has not yet been run against the real model; the keyless suite is the verified gate.

The whole thing tears down with make down, which stops the port-forward and deletes only this lab’s kind cluster.