A self-hosted coding agent: Qwen Code with Superpowers on one AWS GPU VM
Coding agents usually come with a hosted model behind them. You install the CLI, sign in, and every token goes to someone else’s API. I wanted to see what is left when you take that part away: an open-weight coding model on hardware you rent by the hour, an open-source agent CLI, and an open-source skills framework on top.
The lab puts all of it on one AWS GPU VM. vLLM serves Qwen/Qwen3-Coder-Next-FP8. The Qwen Code CLI runs in a container next to it, with the Superpowers skills extension installed. You hand the agent a task file and watch it work. Then a separate verifier, with tests the agent never saw, decides whether the result is any good.
Two questions drove the lab. Does a skills framework that was written for Claude Code and frontier models still steer a much smaller open-weight model, running with nobody around to answer its questions? And once an agent says it is done, how do you know?
The reader workflow is six Makefile targets.
From labs/lab-qwen-code-superpowers-aws/README.md:
make doctor
make cloud-plan
make cloud-up
make run TASK=log-summary
make verify
make destroy
The laptop needs Terraform, the AWS CLI v2, ssh, rsync and Make. There is no local mode: the lab costs about USD 5.72 per hour while the VM exists, and there is no automatic shutdown. make destroy is part of the workflow, not an afterthought.
Picking the VM
The account I used has a quota of 8 vCPUs for “Running On-Demand G and VT instances” in eu-central-1. That means exactly one 2xlarge at a time, so the question was not “which GPU is best” but “which GPU fits in 8 vCPUs”. These were the candidates offered in the region, with on-demand prices from the AWS price list at the time:
| Instance | GPU | GPU memory | Host RAM | USD/hour |
|---|---|---|---|---|
g6.2xlarge | L4 | 22 GiB | 32 GiB | 1.22 |
g5.2xlarge | A10G | 22 GiB | 32 GiB | 1.52 |
g6e.xlarge | L40S | 44 GiB | 32 GiB | 2.33 |
g6e.2xlarge | L40S | 44 GiB | 64 GiB | 2.80 |
g7e.2xlarge | RTX PRO 6000 Blackwell | 96 GiB | 64 GiB | 5.72 |
The sizing comes from the model, not from a benchmark table. Qwen3-Coder-Next has 80B parameters in total, 3B of them active per token. In FP8 the weights are 80.4 GB on disk, about 75 GiB. That rules out everything with 22 or 44 GiB before context is even considered.
Context is the second half of the arithmetic, and here the model is unusual. Only 12 of its 48 layers use full attention; the others are linear-attention layers with a fixed-size state. The KV cache therefore grows only in those 12 layers: 12 layers × 2 KV heads × 256 dimensions × 2 (key and value) × 2 bytes = 24 KiB per token. At the full context of 262,144 tokens that is about 6 GiB. Weights plus full context fit on one 96 GiB GPU, which leaves g7e.2xlarge as the only 8-vCPU option. g6e.2xlarge with the smaller Qwen/Qwen3.5-35B-A3B-FP8 is kept as a documented fallback, but only the default pair was run end to end.
The live measurement agreed with the arithmetic. With vLLM told to use 90% of the GPU, it reserved 87,443 of 97,887 MiB.
No open port and no stored key
The VM runs a coding agent that auto-approves every tool call, so I wanted the VM itself to be as uninteresting to an attacker as possible. It has no IAM role, no key pair, and no security group rule that admits traffic from the internet. It still has a public IP, but only for outbound downloads.
SSH goes through an EC2 Instance Connect Endpoint instead. The two security groups only point at each other.
From labs/lab-qwen-code-superpowers-aws/terraform/04_network.tf:
resource "aws_vpc_security_group_egress_rule" "endpoint_to_vm_ssh" {
security_group_id = aws_security_group.endpoint.id
referenced_security_group_id = aws_security_group.vm.id
ip_protocol = "tcp"
from_port = 22
to_port = 22
}
resource "aws_vpc_security_group_ingress_rule" "vm_ssh_from_endpoint" {
security_group_id = aws_security_group.vm.id
referenced_security_group_id = aws_security_group.endpoint.id
ip_protocol = "tcp"
from_port = 22
to_port = 22
}
The endpoint can open port 22 on the VM, and the VM accepts port 22 only from the endpoint. Who may use the endpoint is decided by IAM, not by an IP allowlist. There is no /32 to keep up to date when your laptop changes networks.
The VM definition carries the rest of the hardening.
From labs/lab-qwen-code-superpowers-aws/terraform/06_vm.tf:
# No instance profile and no key pair: the VM holds no AWS credentials and
# accepts only short-lived keys pushed through EC2 Instance Connect.
resource "aws_instance" "vm" {
ami = data.aws_ami.dlami.id
instance_type = var.instance_type
subnet_id = aws_subnet.public.id
vpc_security_group_ids = [aws_security_group.vm.id]
associate_public_ip_address = true
user_data = file("${path.module}/../vm/host/boot.sh")
user_data_replace_on_change = true
metadata_options {
http_endpoint = "enabled"
http_tokens = "required"
http_put_response_hop_limit = 1
}
# ...
The hop limit of 1 matters because of Docker. A request from inside a container passes one extra network hop, so the IMDSv2 token answer never reaches it. Containers on this VM cannot talk to instance metadata. Since there is no instance role, there would be no credentials to steal anyway, but the check costs nothing.
On the laptop, one script knows how to connect. Every call makes a throwaway key, pushes it for 60 seconds, and tunnels SSH through the endpoint.
From labs/lab-qwen-code-superpowers-aws/scripts/vm_ssh.sh:
work_dir="$(mktemp -d)"
trap 'rm -rf "$work_dir"' EXIT
ssh-keygen -q -t ed25519 -N "" -C "lab-ephemeral" -f "$work_dir/key"
cat >"$work_dir/config" <<EOF
Host lab
HostName $instance_id
User ubuntu
IdentityFile "$work_dir/key"
IdentitiesOnly yes
UserKnownHostsFile "$ssh_dir/known_hosts"
StrictHostKeyChecking $host_key_checking
HostKeyAlias $instance_id
ServerAliveInterval 30
ConnectTimeout 30
LogLevel ERROR
ProxyCommand aws ec2-instance-connect open-tunnel --region $region --instance-id %h
EOF
aws ec2-instance-connect send-ssh-public-key \
--region "$region" \
--instance-id "$instance_id" \
--instance-os-user ubuntu \
--ssh-public-key "file://$work_dir/key.pub" \
--output text >/dev/null
Nothing is written to ~/.ssh. Host key checking is strict: bootstrap reads the VM’s host key from its EC2 console output and pins it in a lab-local known_hosts before the first connection. On this AMI the console output always had the key, so the accept-new fallback was never needed.
A tunnel through the endpoint lives for at most one hour. That would be a problem if an agent run lived inside the SSH session, so it does not. make run starts the agent as a detached process on the VM and only the live view runs over SSH. I killed the local SSH connection 30 seconds into a run to check: make status still showed it running, and make watch reattached and rendered the run from turn 1.
Serving the model
vLLM runs from the official container image, pinned by tag and digest, in the same Compose project as everything else.
From labs/lab-qwen-code-superpowers-aws/vm/compose.yaml:
vllm:
image: vllm/vllm-openai:v0.30.0@sha256:8a69ffad015f138d7170c4ddc429e230a3bc1c1719f67e14324749df200a4b90
entrypoint: ["sh", "-c"]
command:
- >-
exec vllm serve "$$MODEL_ID"
--revision "$$MODEL_REVISION"
--served-model-name "$$SERVED_MODEL_NAME"
--max-model-len "$$MAX_MODEL_LEN"
--gpu-memory-utilization "$$GPU_MEMORY_UTILIZATION"
--max-num-seqs "$$MAX_NUM_SEQS"
--enable-auto-tool-choice --tool-call-parser qwen3_coder
--enable-prefix-caching
--host 0.0.0.0 --port 8000
$$SERVE_EXTRA_ARGS
# ...
# Loopback only: model-check on the VM can reach it, nothing else outside Docker.
ports:
- "127.0.0.1:8000:8000"
# ...
networks: [model, hf]
--max-num-seqs was not in the first version, and its absence was the first real failure of the lab. The weights downloaded in about three minutes, loaded in about 75 seconds, and then vLLM stopped while it set up its cache. The error said that max_num_seqs (1024) exceeded the available Mamba cache blocks (514), and that each decode sequence requires one block.
This is the other side of the architecture that makes long context cheap. Each of the linear-attention layers keeps one fixed-size state per sequence that is being decoded. With 90% of the GPU and a 262,144-token context, only 514 of those state slots fit, and vLLM’s default is to allow 1024 sequences at once. For one agent a limit of 16 is plenty.
From labs/lab-qwen-code-superpowers-aws/terraform/01_variables.tf:
variable "max_num_seqs" {
type = number
default = 16
description = "Max sequences vLLM decodes at once. Qwen3-Coder-Next is a hybrid model: its linear-attention layers keep one fixed-size Mamba cache state per in-flight sequence, so this is capped by the state slots vLLM preallocates, not by GPU memory alone. One agent plus a handful of subagents needs far fewer than vLLM's default of 1024."
}
The lesson is a small one, but not obvious: on a hybrid model, the limit on concurrent requests comes from the model’s state slots, not only from GPU memory.
The networks do as much for isolation as the port binding. model is an internal Docker network with no route out. vLLM reaches Hugging Face through its own hf network, which nothing else joins.
From labs/lab-qwen-code-superpowers-aws/vm/compose.yaml:
networks:
# The agent reaches the model here. Internal: no route out.
model:
internal: true
# vLLM's own way out, to download weights. Nothing else joins it.
hf: {}
# Package installs for the agent and the verifier. The model is not on it.
egress: {}
There was one more fix before the lab was stable, and it was about disk, not GPU. The Deep Learning AMI mounts its 1.9 TB local NVMe at /opt/dlami/nvme, and the boot script moved Docker’s data root there. After the first deploy, the 100 GB root volume was still 80% full. Docker 29 on this AMI stores images through containerd, so they landed in /var/lib/containerd on the root volume, whatever Docker’s own data root said. The boot script now moves containerd’s root to the NVMe too, before any image exists, which brought root usage from 80 down to 50 GB.
make cloud-up blocks until the model answers a real tool-call request. This is the check it runs, and its output from the fresh-clone run.
From labs/lab-qwen-code-superpowers-aws/README.md:
ok model qwen3-coder-next is served
ok tool call parsed: get_weather {'city': 'Berlin'}
info 512 tokens in 3.2 s = 162 tokens/s (single request)
info GPU NVIDIA RTX PRO 6000 Blackwell Server Edition, 87443 MiB, 97887 MiB
From make cloud-up to that output took about 13 minutes, on both full deploys.
The agent container
Qwen Code talks to any OpenAI-compatible server. Its settings point it at vLLM over the internal network.
From labs/lab-qwen-code-superpowers-aws/vm/coder/settings.json:
"modelProviders": {
"openai": [
{
"id": "qwen3-coder-next",
"name": "Qwen3-Coder-Next on vLLM",
"envKey": "VLLM_API_KEY",
"baseUrl": "http://vllm:8000/v1",
"generationConfig": {
"timeout": 300000,
"maxRetries": 2,
"contextWindowSize": 262144,
"samplingParams": { "temperature": 1.0, "top_p": 0.95, "max_tokens": 8192 }
}
}
]
},
The API key is a placeholder; vLLM does not check one, and the port is not reachable from outside anyway.
The agent runs in --approval-mode yolo: every file write and every shell command goes through without asking. That makes the container the only boundary, so it is built to hold nothing and reach little.
From labs/lab-qwen-code-superpowers-aws/vm/compose.yaml:
coder:
build: ./coder
image: qwen-lab/coder
profiles: [tools]
# The container is the sandbox: every tool call is auto-approved.
read_only: true
tmpfs:
- /run/qwen:exec,mode=1777,size=1g
- /tmp:exec,mode=1777,size=8g
cap_drop: [ALL]
security_opt: ["no-new-privileges:true"]
mem_limit: 16g
pids_limit: 1024
# ...
networks: [model, egress]
A read-only root filesystem, no capabilities, a non-root user, and memory and process limits so a runaway agent cannot starve vLLM. The only writable bind mount is the run’s own workspace. There is no Docker socket, no AWS credential and no SSH material in it. make containment-check tests this from both sides, and it passed 9 of 9 checks in the fresh-clone run.
The trade-off should be said plainly. The container keeps outbound internet on egress, because an agent that builds a TypeScript project has to run npm install. A badly steered agent could therefore still download code or send data out, just as npm install on your laptop could. What the containment buys is narrower: no path to the AWS account, the Docker host or a stored credential.
Installing Superpowers, and a fix I did not need
Superpowers is installed at image build time from a pinned tag. It is cloned first and installed from the local path, because Qwen Code’s own --ref install downloads a tarball and records no commit.
From labs/lab-qwen-code-superpowers-aws/vm/coder/Dockerfile:
USER node
# Superpowers from the pinned tag. Installing from a local clone records
# what was installed; installing with --ref would download a tarball with
# no commit to check.
ENV QWEN_HOME=/opt/qwen-seed
RUN git clone --quiet --depth 1 --branch v6.4.2 https://github.com/obra/superpowers.git /opt/superpowers \
&& qwen extensions install /opt/superpowers:superpowers --consent
COPY --chown=node:node settings.json /opt/qwen-seed/settings.json
The more interesting part was how Superpowers’ bootstrap reaches the model. In Claude Code, a SessionStart hook injects it. Before any cloud spend, reading Qwen Code’s source showed that this hook does nothing under Qwen Code by default: Qwen substitutes ${CLAUDE_PLUGIN_ROOT} into the hook command, but does not export the variable, so the hook prints its JSON in a shape Qwen ignores. Exporting the variable yourself changes the output shape, and Qwen applies it. So the first image exported it.
On the VM, it turned out that this “fix” was not needed. Installed from a local clone, the extension names GEMINI.md as its context file, and Qwen Code loads every extension’s context file at the start of every session. GEMINI.md imports the using-superpowers skill and a tool-name table. With the export, the model got the same bootstrap twice. So the export was dropped, and the build now checks that the path the bootstrap really uses is still there.
From labs/lab-qwen-code-superpowers-aws/vm/coder/check-extension.sh:
expected_commit="8ca22dba9a94f28898bbce59f2537ff4d87c747d"
actual_commit="$(git -C /opt/superpowers rev-parse HEAD)"
[ "$actual_commit" = "$expected_commit" ] || { echo "Superpowers is $actual_commit, expected $expected_commit" >&2; exit 1; }
# ...
grep -qF "skills/using-superpowers/SKILL.md" "$extension_dir/GEMINI.md" \
|| { echo "GEMINI.md does not import skills/using-superpowers/SKILL.md" >&2; exit 1; }
The hook mismatch is real, and worth knowing about if you port a Claude Code plugin to Qwen Code. It just was not the path that mattered here.
Handing over a task, and the gate that stopped everything
The sample task is a TypeScript CLI that reads a web-server access log and prints a JSON summary: counts per status class, the top paths, the p95 latency and the number of malformed lines. The task file fixes everything a verifier needs to agree on: the Node version, the exact dev dependencies, the command that runs the program, the output schema and the percentile method. And its first lines say that nobody is there.
From labs/lab-qwen-code-superpowers-aws/tasks/log-summary/task.md:
# Task: access-log summary CLI in TypeScript
No human is available to answer questions during this task. Where these
requirements leave something open, choose the most reasonable option,
write it down in `ASSUMPTIONS.md`, and continue.
That was not enough. The first live run took 3 turns and 7 seconds and built nothing. The agent classified the task, asked clarifying questions, proposed three architectures, and ended with “Please confirm or suggest an alternative before I proceed with the implementation plan.” In headless mode that is the end of the session. Superpowers is built around a human approving each step, and it did exactly what it was built to do.
So the lab got two additions. The first is a short contract, appended to the system prompt of every session next to today’s date.
From labs/lab-qwen-code-superpowers-aws/vm/coder/contract.md:
You are running unattended. No human will read or answer your messages until the task is finished.
Follow the Superpowers workflow, and wherever a skill asks a human for input or approval, decide yourself:
1. In brainstorming, accept your own recommended option for every question and every choice of approach, and continue.
2. When brainstorming is complete, write the spec immediately and save it under docs/superpowers/specs/.
3. When the spec is written, write the implementation plan immediately and save it under docs/superpowers/plans/.
4. When the plan is written, execute it immediately with superpowers:subagent-driven-development, until every task in the plan is complete and verified.
When everything is done and verified, end your final message with the line: ALL TASKS COMPLETE
The second is a driver. It runs qwen -p with the task, and when a session ends without reporting completion, it resumes the same session with one scripted reply. The reply is chosen from what the workspace already contains, like an owner who always says “yes, go on”.
From labs/lab-qwen-code-superpowers-aws/vm/coder/drive-agent.sh:
stage_reply() { # picks the next scripted reply from the workspace state
if ! compgen -G "/workspace/docs/superpowers/specs/*.md" >/dev/null; then
echo "spec|No human is available. Accept your recommended option for every open question and approach, write the spec now under docs/superpowers/specs/, then continue with the plan."
elif ! compgen -G "/workspace/docs/superpowers/plans/*.md" >/dev/null; then
echo "plan|The spec is approved. Write the implementation plan now under docs/superpowers/plans/, then continue."
elif [ "$implement_sent" = no ]; then
echo "implement|The plan is approved. Execute it now with superpowers:subagent-driven-development and continue until every task is complete and verified."
else
echo "continue|Continue until every task in the plan is complete and verified. When everything is done, end your final message with the line: ALL TASKS COMPLETE"
fi
}
Each reply is written into the transcript as its own driver event, so the scripted steering is visible in the live view and in the verdict, never hidden. The loop is bounded: 200 turns per round, 90 minutes of wall time shared across all rounds, and at most 8 rounds.
The detail that needed the most care is how the driver decides that a round is done. Searching the transcript for ALL TASKS COMPLETE is not good enough. The string shows up in places that do not mean completion: the agent announcing that it will write it later, a todo item, a file that happens to contain it.
From labs/lab-qwen-code-superpowers-aws/vm/coder/drive-agent.sh:
round_is_complete() {
node -e '
const fs = require("fs");
const lines = fs.readFileSync(process.argv[1], "utf8").split("\n");
let last = null;
for (const line of lines) {
const trimmed = line.trim();
if (!trimmed) continue;
let obj;
try { obj = JSON.parse(trimmed); } catch { continue; }
if (obj && obj.type === "result") last = obj;
}
if (!last || last.is_error !== false || typeof last.result !== "string") {
process.exit(1);
}
process.exit(last.result.trim().endsWith("ALL TASKS COMPLETE") ? 0 : 1);
' "$1"
}
Only the round’s own last result line counts, it has to be a success, and its text has to end with the marker. A round cut off by the turn or time limit produces no result line at all, so it is never complete, whatever the transcript said before.
Watching it work
The transcript is Qwen Code’s stream-json output. A small TypeScript viewer renders it live, in the style of a coding terminal: text as it streams, each tool call as a block, file writes as diffs, skill loads highlighted. This is an excerpt from the first driven run, with the scripted reply in the middle.
From labs/lab-qwen-code-superpowers-aws/README.md:
── turn 2/200 ──
★ Skill(superpowers:brainstorming)
⎿ skill loaded
…
── turn 11/200 ──
⏺ Now let me create the parser module:
● Write(src/parser.ts)
+ import { ParsedLine } from './types.js';
+
+ const LOG_PATTERN =
+ /^(\S+)\s+(\S+)\s+(\S+)\s+\[([^\]]*)\]\s+"(\S+)\s+(\S+)\s+(\S+)"\s+(\d{3})\s+(\d+|-)\s+(\d+(?:\.\d+)?)\s*$/;
+
+ export function parseLine(line: string): ParsedLine | null {
+ const match = line.match(LOG_PATTERN);
+ if (!match) {
+ return null;
… 36 more lines
── turn 18/200 ──
● Write(src/cli.test.ts)
+ import { describe, it, expect, vi, beforeEach } from 'vitest';
+ import { readFileSync } from 'fs';
+
+ // Mock fs module
+ vi.mock('fs', async () => ({
+ readFileSync: vi.fn(),
+ }));
+
+ describe('CLI', () => {
+ beforeEach(() => {
+ vi.resetAllMocks();
+ });
+
+ it('should handle missing file', async () => {
… 1 more lines
…
● Shell(npm test)
│ > [email protected] test
│ > vitest run
│
│
│ RUN v5.0.2 /workspace
│
│ ✓ src/cli.test.ts (1 test) 3ms
│ ✓ src/parser.test.ts (11 tests) 5ms
│ ✓ src/summarizer.test.ts (5 tests) 11ms
│
│ Test Files 3 passed (3)
│ Tests 17 passed (17)
…
▶ owner (scripted, round 2, spec): No human is available. Accept your recommended option for every open question and approach, write the spec now under docs/superpowers/specs/, then continue with the plan.
── session resumed ──
── turn 43/200 ──
● Write(docs/superpowers/specs/2026-09-28-access-log-cli-design.md)
+ # Access Log Summary CLI - Design Specification
+
+ ## Project Overview
+
+ A command-line tool that reads web-server access logs and outputs a JSON summary to stdout.
+
+ ## Requirements
+
… 47 more lines
The same excerpt shows something the verifier later measures: the agent had written its parser and run its tests before it wrote a spec, and only wrote the spec when the driver asked for it.
Trust, but verify
When the agent ends with ALL TASKS COMPLETE, that is a claim. The verifier is what decides. It mounts the run’s workspace read-only, copies it into scratch space, and runs six hard checks there: the agent’s exit code, npm ci, the type check, the agent’s own tests, hidden acceptance tests, and the CLI on hidden log fixtures compared with known answers. The fixtures differ from anything in the task file, and the acceptance folder is never mounted into the agent’s container.
Next to the verdict, the verifier reports process evidence it reads from the transcript. These are measured, never graded: which skills loaded, whether a test file was written before the code it covers, whether the agent ran its tests and the program before it stopped. The test-first rule is simple and deliberately literal.
From labs/lab-qwen-code-superpowers-aws/vm/tools/src/verify/measured-signals.ts:
/** Finds the first test write, the first source write, and the last source write, by order. */
export function analyzeFiles(calls: ToolCall[]): {
firstTestOrder: number | null;
firstSourceOrder: number | null;
lastSourceOrder: number | null;
} {
let firstTestOrder: number | null = null;
let firstSourceOrder: number | null = null;
let lastSourceOrder: number | null = null;
for (const call of calls) {
const kind = classifyFileWrite(call);
if (kind === "test") {
if (firstTestOrder === null) firstTestOrder = call.order;
} else if (kind === "source") {
if (firstSourceOrder === null) firstSourceOrder = call.order;
lastSourceOrder = call.order;
}
}
return { firstTestOrder, firstSourceOrder, lastSourceOrder };
}
This is the verdict of the first driven run.
From labs/lab-qwen-code-superpowers-aws/README.md:
VERDICT: PASS
pass agent-exit
pass install
pass typecheck
pass own-tests
pass acceptance
pass fixtures
Process evidence (measured, not graded)
skills loaded: superpowers:brainstorming
test before code: no
ran tests: yes
ran program: yes
checked after last change: yes
turns: 47
tool calls: edit:2, glob:5, list_directory:2, read_file:3, run_shell_command:17, skill:1, write_file:15
rounds: 2
replies: spec:1
spec written: yes
plan written: yes
subagent calls: 0
finish reason: complete
What happened over five runs
After the driver was in place, the lab ran log-summary five times: two runs during development, and three from a fresh clone, following the README only.
| Run | Duration | Turns | Rounds | Scripted replies | Skills loaded | Test before code |
|---|---|---|---|---|---|---|
| 1 | 102 s | 47 | 2 | spec | brainstorming | no |
| 2 | 242 s | 96 | 1 | none | brainstorming, writing-plans, subagent-driven-development, executing-plans, test-driven-development | no |
| 3 | 161 s | 68 | 1 | none | brainstorming | no |
| 4 | 176 s | 73 | 1 | none | brainstorming, writing-plans, subagent-driven-development, executing-plans | yes |
| 5 | 110 s | 52 | 1 | none | none | no |
All five passed all six hard checks. The contract alone carried four of the five runs through every gate in a single session; the scripted reply was needed once. By the only measure that decides the verdict, the model did the job every time.
The process evidence tells a less tidy story.
Test-first happened once in five runs. Run 2 even loaded the test-driven-development skill, and its first test file still came after the source file it covered. Loading a skill is not the same as following it turn by turn. That is exactly why the lab measures this separately from the verdict.
One run loaded no Superpowers skill at all. Run 5 read the task, built the CLI, ran its tests and finished in 110 seconds without calling a single skill. It passed like the others. For a task this size, the bootstrap in the context was sometimes a suggestion, not a rule.
Subagents never ran. Runs 2 and 4 reached the subagent-driven-development skill, as the contract asked. In run 2 it reported “No ordinary background subagents are available in this session” and fell back to executing the plan inline. Across all five runs there were zero subagent calls. In headless Qwen Code 0.24.6, this part of the Superpowers workflow did not happen.
Five runs of one small task are not a benchmark. They do not say how the model does on a large codebase or on a task with real ambiguity. What they do say is that the pipeline works, that an open-weight model on one rented GPU can finish a well-specified task unattended and pass tests it never saw, and that “passes” and “followed the method” are two different questions with two different answers.
What it cost, and tearing it down
Building and testing the lab took three sessions on g7e.2xlarge, about 2.3 VM hours in total, for about USD 13.20. The VM was destroyed after each session. A fresh deploy, from make cloud-up to a model that answers, took about 13 minutes each time; the driven runs themselves took two to four minutes each.
make destroy runs terraform destroy and then a leftover check that searches the region for anything still tagged with the lab’s name: instances, volumes, the Instance Connect Endpoint, security groups, VPCs and elastic IPs. The fresh-clone run ended with 13 resources destroyed and the line Nothing left behind.
If I built it again, I would start with the driver. The approval gate is not a bug in Superpowers; it is the whole point of brainstorming with a human. Running it unattended means deciding, in advance and in writing, what “yes” sounds like, and making every scripted “yes” visible in the transcript.