~/tech-with-ugur

The Agent Is in the Page. The Key Is Not.

2026-09-09 ai

Run the companion lab

Your internal B2B app has forty screens and your users know four of them. The usual answers to that are better docs, better search, or a chatbot that tells people where to click. Alibaba’s page-agent is a different answer: an agent that lives inside the page and clicks for them.

It is not Playwright and it is not browser-use. There is no headless browser, no extension, no Python sidecar. You npm install it, hand it an OpenAI-compatible endpoint, and your users can say what they want instead of hunting for the right screen. It reads the page purely through the DOM — no screenshots, no vision model — and clicks, types, selects, and submits.

Getting that working is an afternoon. What took the rest of the lab is everything around it. The agent runs in the browser, which means the model credential must never be there. Its LLM traffic has to go out through your own backend on the user’s session token, and the moment you do that you are operating an authenticated /chat/completions relay on your users’ behalf, with everything that implies.

This post is that build: a maintenance logbook for a solar-panel fleet, where a service engineer types one sentence and watches the app fill in its own form.

The lab is here: lab-page-agent-solar-maintenance. It runs with docker compose up and needs no credentials at all.

What the run looks like

One sentence in, twelve steps, a filed maintenance report:

The engineer typed:

replaced the string 3 inverter fan on the Almeria roof array this morning, took two hours, panel 14 still shows a hotspot

The agent navigated to the right site, opened the report form, chose a component from a dropdown, typed four fields, ticked a checkbox, filled the field that ticking it revealed, and pressed Submit. The app landed on /reports with a new row.

One detail in that recording matters more than it looks: the very first action is wait. I will come back to why, because it is the most realistic failure mode an in-page agent has and it is invisible when you drive the app by hand.

An honest note on what drove it

The lab ships two transports. LLM_MODE=gemini forwards to Gemini’s OpenAI-compatible endpoint with a server-held key. LLM_MODE=scripted — the default — is a deterministic fake model that speaks the same protocol and needs no key at all.

Everything shown and measured in this post is the scripted transport. I did not get to a verified run against a live Gemini key, so I am not going to claim a real model drove it. What that costs the post is the “can the model actually reason about this form” question, which stays open. What it does not cost is anything structural: the agent library is real, the DOM extraction is real, the relay is real, the browser traffic is real, and the credential boundary — the thing this post is actually about — is verified end to end in a real Chromium tab.

It turns out the fake model earns its place anyway. Writing one against a real agent’s protocol is the clearest possible way to show what that protocol is.

The wire format is one tool

Everything in this section was read out of the published package ([email protected], MIT) rather than the docs site, because the wire contract is what you have to build against.

On every step the library POSTs to ${baseURL}/chat/completions with exactly two messages — a system prompt and one assembled user prompt — and exactly one tool offered, named AgentOutput, with tool_choice naming it, parallel_tool_calls: false, and no streaming. The tool’s arguments are:

{
  "evaluation_previous_goal": "...",   // all three optional
  "memory": "...",
  "next_goal": "...",
  "action": { "<tool_name>": { /* that tool's input */ } }
}

There are nine built-in action tools: click_element_by_index, input_text, select_dropdown_option, scroll, scroll_horizontally, wait, ask_user, execute_javascript (off by default), and done.

Notice what is missing. There is no navigate tool. The agent cannot go to a URL. Route changes happen only because it clicked a link, exactly as a person’s would. That is a design statement rather than a gap, and it has a direct architectural consequence: your app has to be a client-side-routed SPA, because a full page load would tear down the agent mid-task.

The other thing worth internalising is that there is no growing assistant/tool message chain. Each step rebuilds one user message containing the original request, a step counter, a Current time: line, a history block summarising prior steps, and the current browser state. The browser state header begins Current Page: [<title>](<url>) and the body is the simplified DOM as indexed lines.

That Current time: line is not decoration — it is the only reason a model can resolve “this morning” into a date.

The credential problem

Here is the whole security story in one file.

frontend/src/agent/setupAgent.ts:

import { PageAgent } from "page-agent";

/**
 * The agent talks to our own origin, not to a model provider. There is no
 * API key here — not an empty one, not a placeholder, none — so the library
 * sends no Authorization header of its own and customFetch supplies ours.
 */
export function createSolarAgent(getToken: () => string | null): PageAgent {
  return new PageAgent({
    baseURL: "/api/agent/v1",
    // The backend pins the real model; this name is only what the library
    // puts in the request body, and the relay overwrites it.
    model: "relayed",
    maxSteps: 20,
    // Extract the whole page, not just the viewport, so the agent never has to
    // scroll to find a form field.
    viewportExpansion: -1,
    customFetch: async (input, init) => {
      const headers = new Headers(init?.headers);
      const token = getToken();
      if (token) headers.set("Authorization", `Bearer ${token}`);
      return await fetch(input, { ...init, headers });
    },
  });
}

Two things make this work, and only one of them is obvious.

The obvious one is customFetch: the library lets you supply the fetch it uses, so you can attach your app’s own session token.

The non-obvious one is that apiKey is genuinely optional. In the library’s source the Authorization header is spread in conditionally — omit apiKey and it sends no Authorization header at all, leaving the slot free for customFetch to fill. If the library had defaulted to an empty string or a placeholder, customFetch would be fighting it, and the honest version of this post would be “you can’t do this cleanly”. It doesn’t, so you can.

I verified that empirically rather than by reading source: driving the real @page-agent/llms client against a stub with apiKey omitted and inspecting the outgoing request init, there is no Authorization header on it.

baseURL points at our own origin, so the request goes to /api/agent/v1/chat/completions — same-origin, session cookie rules, our backend. There is no provider hostname anywhere in the bundle.

The backend hop, and what it can honestly control

An authenticated /chat/completions passthrough is an open LLM relay if you skip any part of it. Four controls actually hold.

Authenticate. backend/src/auth/middleware.ts:

export function requireAuth(secret: string): MiddlewareHandler {
  return async (c, next) => {
    const header = c.req.header("Authorization") ?? "";
    const token = header.startsWith("Bearer ")
      ? header.slice("Bearer ".length)
      : "";
    if (token === "") return c.json({ error: "Missing session token." }, 401);
    try {
      c.set("user", await verifySession(token, secret));
    } catch {
      return c.json({ error: "Invalid session token." }, 401);
    }
    await next();
  };
}

Pin and allowlist. The relay decides the model, not the client. backend/src/agent/guard.ts:

  // Only these four fields survive. `model` and `max_tokens` are ours, and
  // anything the client sent that is not named here — api_key, stream, user,
  // a different endpoint's parameters — is gone.
  return {
    ok: true,
    request: {
      model: config.geminiModel,
      max_tokens: config.agentMaxTokens,
      messages: messages as ChatMessage[],
      ...(tools === undefined ? {} : { tools }),
      ...(tool_choice === undefined ? {} : { tool_choice }),
      ...(parallel_tool_calls === undefined ? {} : { parallel_tool_calls }),
    },
  };

There is a subtle zod trap buried in that function. The incoming message schema is z.object({ role: z.string() }).loose(). Without .loose(), zod’s default strict object would silently strip content from every message — the relay would forward perfectly well-formed requests with no actual prompt in them, which is a spectacularly annoying thing to debug. Verified against real zod 4.5.4 before it could bite.

Cap. The size check covers everything forwarded, not just the messages, further up the same backend/src/agent/guard.ts:

  // Size EVERYTHING the client controls that we would forward, not just the
  // messages. `tools` and `tool_choice` are opaque to us, so a caller could
  // otherwise slip a tiny message array past the cap alongside a gigantic tool
  // schema and have us relay it upstream on our key. The count cap runs first
  // so an obviously bad request is rejected before we serialise anything.
  const forwarded = JSON.stringify({ messages, tools, tool_choice });
  if (Buffer.byteLength(forwarded, "utf8") > config.agentMaxRequestBytes) {
    return { ok: false, status: 413, error: "Request too large." };
  }

Budget. A per-user sliding window, checked before the body is read, in backend/src/agent/routes.ts:

    // Budget FIRST, before the body is read or parsed: a caller who is over
    // budget must not be able to make us do the expensive work anyway.
    const decision = budget.tryConsume(user.id);
    if (!decision.allowed) {
      logger.warn(
        { userId: user.id },
        "Rejecting an agent step: budget spent.",
      );
      c.header("Retry-After", String(decision.retryAfterS));
      return c.json({ error: "Agent call budget exhausted." }, 429);
    }

And on the way out, the upstream headers are built from scratch — nothing from the inbound request is spread into them, so the browser’s Authorization never reaches Gemini and Gemini’s key never reaches the browser. backend/src/agent/gemini.ts:

      const response = await fetch(url, {
        method: "POST",
        headers: {
          "Content-Type": "application/json",
          Authorization: `Bearer ${config.geminiApiKey}`,
        },
        body: JSON.stringify(request),
      });

That file also scrubs upstream error bodies before they hit the logs, because a misconfigured endpoint that echoes your request back will otherwise write a live key straight into your log file.

The control I thought I had, and don’t

When I planned this lab I listed “drop any client-supplied system prompt override” as one of the relay’s controls. It cannot be one, and pretending otherwise would be the dishonest version of this post.

With an in-page agent, the entire prompt — system message included — is composed in the browser by the library. There is nothing server-side to compare it against. The relay has no authored prompt of its own, so it has no baseline from which to detect tampering. A user who opens devtools can send whatever messages they like to your relay, and it will forward them, on your key, to your model.

What actually holds is identity, model pinning, size caps, and budget. Those are real and they are worth having. But the residual risk is that you are running a model-access endpoint for authenticated users, and the correct mental model is “metered API I am reselling to my own users”, not “internal implementation detail of my UI”.

The e2e suite pins the parts that do hold. e2e/tests/hardening.spec.ts:

test("the browser sends the session token and no model credential", async ({
  page,
}) => {
  const relayHeaders: Record<string, string>[] = [];
  page.on("request", (request) => {
    if (request.url().includes("/api/agent/v1/"))
      relayHeaders.push(request.headers());
  });

  await page.goto("/login");
  await page.getByLabel("Email").fill("[email protected]");
  await page.getByLabel("Password").fill("solar");
  await page.getByRole("button", { name: "Sign in" }).click();
  // AgentMount installs window.pageAgent from a post-navigation effect, which
  // does not resolve within the click's own action wait; evaluating too early
  // races it and finds nothing to call.
  await page.waitForFunction(() => window.pageAgent !== undefined);
  await page.evaluate(async () => {
    await window.pageAgent!.execute("open the Almeria Roof Array site");
  });

  expect(relayHeaders.length).toBeGreaterThan(0);
  for (const headers of relayHeaders) {
    expect(headers.authorization).toMatch(/^Bearer eyJ/); // our JWT, not a provider key
    expect(JSON.stringify(headers)).not.toMatch(/AIza|sk-|x-goog-api-key/i);
  }
});

Your markup is now an API

This is the part I did not expect to be the most useful section of the lab.

page-agent has no vision model and takes no screenshots. The DOM is its only input. That quietly promotes your accessibility work from “the right thing to do” to “the interface contract”, and it means the agent perceives your page in ways that will surprise you.

I captured the library’s own getBrowserState() against the real report form in headless Chromium, before and after filling it. Empty form:

[0]<label for=component>Component />
[1]<select name=component>Choose a component
Inverter
Panel string />
[2]<label for=component-ref>Unit />
[3]<input type=text name=component-ref />
...
[11]<input type=checkbox checked=false name=follow-up-required />
[12]<button type=submit name=submit-report>Submit report />

After selecting a component, typing into two fields, and ticking the box:

[11]<input type=checkbox checked=true name=follow-up-required />
*[12]<label for=follow-up-note>Follow-up note />
*[13]<input type=text name=follow-up-note />
[14]<button type=submit name=submit-report>Submit report />

Five things fall out of that, all of which changed the code.

The agent cannot read back what it typed. The component reference input was filled with “String 3 inverter” and the select was set to “Inverter”. Neither line carries a value= attribute. The walker reads HTML attributes, not live DOM properties, and a React-controlled input never writes its value back to the attribute — a <textarea> has no value attribute at all, and a <select>’s selection is not one either. The agent only knows what it typed from its own step history.

The single exception is deliberate: for checkboxes and radios the walker overwrites checked with the live property, which is why you can see it go false → true above.

Indices shift under you. submit-report moved from [12] to [14] the moment ticking the checkbox revealed two new elements. Hold onto that one; the next section is built on it.

Labels get their own indices. <label> elements are interactive as far as the walker is concerned, so the element list is about twice as long as you would guess, and “match on the visible text” is not the safe shortcut it appears to be.

id vanishes when it duplicates name. The walker drops attributes whose value duplicates an earlier attribute’s, and name precedes id in its include list. So id is not a reliable matching key. name is.

Attribute values are truncated to 20 characters. Short, distinct names matter more than they would for a human reading your markup.

And one rule the library enforces on your behalf: the dropdown has to be a native <select>. select_dropdown_option throws “Element is not a select element” on anything else, so a beautifully styled div-based combobox is simply undrivable. frontend/src/routes/ReportForm.tsx:

        <label htmlFor="component">Component</label>
        <select
          id="component"
          name="component"
          required
          value={component}
          onChange={(e) => setComponent(e.target.value)}
        >
          <option value="">Choose a component</option>
          {COMPONENTS.map((option) => (
            <option key={option} value={option}>
              {option}
            </option>
          ))}
        </select>

Nothing exotic. That is rather the point — the markup an agent can operate is the markup a screen reader could already operate.

A fake model, not a recorded tape

The default transport had to make an LLM-driven end-to-end test deterministic and keyless. The obvious way to do that is to replay a canned sequence of tool calls.

That does not work here, and the DOM capture above is exactly why. Every action addresses its target by index, and indices are assigned by DOM walk order at the moment of extraction. A tape that says {"click_element_by_index":{"index":12}} clicks the wrong element the first time a row is added, a field is revealed, or a heading changes — which in this form happens on the second-to-last step, every single run.

So LLM_MODE=scripted is a deterministic fake model rather than a recording. Each request, it parses the current URL and the indexed element lines out of the user prompt, walks an ordered list of intents expressed against stable semantics instead of indices, picks the first one that applies and is not already satisfied, resolves it to a live index, and emits a well-formed AgentOutput tool call.

An intent looks like this. backend/src/agent/scripted/intents.ts:

export const INTENTS: Intent[] = [
  {
    goal: "Open the Almeria Roof Array site",
    when: onSites,
    satisfied: (view) => !onSites(view.url),
    locate: (view) => findByText(view, "Almeria Roof Array"),
    act: (element) => ({ click_element_by_index: { index: element.index } }),
  },

And the resolver is a pure function from one request to one action. backend/src/agent/scripted/resolver.ts:

/**
 * The scripted transport's whole brain: a pure function from one request to one
 * action. It re-reads the live DOM every step and resolves intents to indices
 * then and there, so it is unaffected by elements appearing, disappearing, or
 * being renumbered — which a recorded tape of tool calls would not be.
 *
 * Success is confirmed POSITIVELY — we are on the reports list and the submit
 * step actually succeeded. The tempting alternative, "no intent applies to this
 * page, so we must be finished", reports a filed report for any page it fails
 * to recognise, including a request it could not parse at all.
 */

That second paragraph is a bug I nearly shipped. “No intent applies, so we must be done” is the natural way to write the terminal condition, and it reports success for any page the resolver fails to recognise — including a request it could not parse at all. Confirming success positively costs one extra condition and turns a silent false pass into a visible failure.

The failure that taught me the most

The agent worked on the first try when I drove it by hand. Then I ran it from a script that signs in and calls execute() immediately, with no pause:

=== AGENT RUN (0.6s) ===
success: false | steps: 1
actions: done
said   : Could not find the next control to use on http://localhost:5173/sites.

One step, straight to failure. The only difference was timing. /sites renders immediately and then fetches its data in a useEffect, so an agent that observes the page within a few hundred milliseconds of arriving sees a heading, a nav link, and no sites at all. The intent applied, was unsatisfied, and could not find its element — so the resolver concluded there was nothing left to do.

This is the most realistic failure an in-page agent has, and it is invisible when you test by hand. A person clicking through is always slower than the fetch. An agent is not. The race here measured about 30 ms.

The library already anticipates it. The built-in wait tool’s own description is “Wait for x seconds. Can be used to wait until the page or data is fully loaded” — a real model reading that would plausibly call it. My scripted model had no such judgement, so it needed the instinct spelled out. The distinction that fixed it is the interesting part:

Back in backend/src/agent/scripted/resolver.ts:

  const step = nextStep(view, today);
  if (step === NOT_RENDERED) {
    // Count prior waits out of the history, so this stays a pure function.
    const waits = view.history.filter((s) => s.goal === WAIT_GOAL).length;
    if (waits < MAX_WAITS) {
      return { nextGoal: WAIT_GOAL, action: { wait: { seconds: 1 } } };
    }
    return done(
      `Waited, but never found the next control on ${view.url}.`,
      false,
    );
  }

Bounded to five attempts, counted out of the agent’s own history so the resolver stays a pure function. A control that is genuinely absent still fails — just five seconds later, with a message that says what happened.

Re-run at the same speed that broke it:

=== AGENT RUN (13.2s) ===
success: true | steps: 12 | waits: 1
actions: wait -> click -> click -> select_dropdown_option
         -> input_text x4 -> click -> input_text -> click -> done
said   : Filed the maintenance report for the Almeria Roof Array.
URL    : http://localhost:5173/reports

String 3 inverter — Inverter, 2026-09-09, 2h, by Rosa Iglesias
Replaced the string 3 inverter fan.
Follow-up: Panel 14 still shows a hotspot.

The wait is the first action, which is the whole story in one line: the agent arrived, looked, recognised nothing, waited a second, and found the list.

Security properties observed on that same run rather than argued: the only origin the browser contacted was http://localhost:5173; all twelve relay calls carried the app’s own session JWT; and no AIza, no x-goog, no provider credential of any kind left the page.

What I would not ship yet

The relay is a model-access endpoint you are operating. Not a UI detail. It needs the budget, the caps, and a line in whatever document says what your service does with user data — because your users’ free-text now reaches a model provider through your key.

Page content is model input. The agent reads the DOM, and everything in that DOM goes into the prompt. On this lab it is seeded fixture data. On a real logbook it would be other engineers’ free-text notes, which means untrusted text arrives in the model’s context by design. The library ships declarative controls for this — data-page-agent-ignore drops a subtree entirely, data-page-agent-not-interactive leaves text visible but unclickable — and I documented them in the lab README, but this lab deliberately does not build an injection scenario, so I am not going to tell you those controls are sufficient. I have not tested them under attack.

There is a direct eval in the bundle. It implements the execute_javascript tool. This lab never enables that tool, but the eval still ships, so a page with a CSP that omits unsafe-eval inherits the problem regardless of your configuration.

The model half is unproven here. Scripted mode says nothing about whether a real model reliably maps a messy sentence onto six form fields. That is the question I would want answered before putting this in front of engineers, and it is the obvious next run.

The takeaway

The in-page agent is a genuinely different shape from browser automation, and page-agent makes the easy part easy. Two things surprised me.

The first is that the security work dominates. “Add an AI assistant to our app” sounds like a frontend task and turns out to be “operate an authenticated LLM relay”, with a control you would assume you have — authorship of the prompt — that you structurally cannot have when the prompt is composed in the browser.

The second is nicer. The agent’s API surface is your accessible markup. Every fix that made the agent work — a native <select>, a real <label>, a meaningful name — is a fix a screen-reader user would also have wanted. For once, the AI feature and the accessibility backlog point in exactly the same direction.

The full lab, including the scripted transport and the e2e suite, is at lab-page-agent-solar-maintenance.