~/tech-with-ugur

One agent, four open-weight models: Amazon Bedrock with LangChain in TypeScript

2026-09-28 aicloud

Run the companion lab

Most “getting started with Bedrock” walkthroughs begin the same way: create an IAM user, create an access key, paste it into a .env file. The key never expires, it sits in plain text next to the code, and the tutorial ends before anyone deletes it.

Since AWS CLI 2.32.0 there is a shorter path. aws login reuses the sign-in you already have for the AWS console: you approve a request in the browser, and the CLI stores a short-lived session on your machine. No key is created, so there is nothing to leak and nothing to clean up.

This lab builds on that. It is a small TypeScript service that answers plain-English questions about a shop database. A LangChain agent does the work by calling three typed tools, one per table, and every model call goes through Amazon Bedrock. One field in the request picks which of four open-weight models answers: MiniMax M2.5, NVIDIA Nemotron 3 Super 120B, DeepSeek V3.2 or Moonshot AI Kimi K3. The agent code is the same for all four.

The whole reader workflow is five Makefile targets.

From labs/lab-bedrock-langchain-agent/README.md:

make login
make up
make ask MODEL=kimi-k3 Q="How many customers live in Vienna?"
make e2e
make down

The host needs Docker with Compose, Make and the AWS CLI. Node.js is not needed; the app, the tests and the end-to-end script all run in containers. There is no offline mode: the lab calls real models, and one make e2e run costs a few cents.

Signing in without an access key

make login runs aws login on the host. The app container gets no credentials of its own. Two bind mounts let it use the session the CLI stored.

From labs/lab-bedrock-langchain-agent/compose.yaml:

    volumes:
      # Read-only: the app only needs to know which login session to use.
      - type: bind
        source: ${HOME}/.aws/config
        target: /home/node/.aws/config
        read_only: true
        bind: { create_host_path: false }
      # Writable: the SDK writes refreshed 15-minute credentials back here.
      - type: bind
        source: ${HOME}/.aws/login/cache
        target: /home/node/.aws/login/cache
        bind: { create_host_path: false }

That is the whole credential setup. The app never constructs credentials; the AWS SDK’s default credential chain finds the login session by itself. Nothing is copied into the image and nothing is written to the repo.

It is worth being exact about what these two mounts expose:

So sign in with an identity that can invoke Bedrock models and little else, not with an administrator.

The cache mount has to be writable, and the reason is the refresh. The credentials in the cache last 15 minutes. The SDK in the container refreshes them and writes the new ones back to the same file. This was tested on its own before any app code existed: a container made a round of calls, waited, and made a second round 17 minutes later, with no AWS CLI use on the host in between. The cache file was rewritten in the same second as the second round of calls, and it kept mode 600 and its owner on the host.

create_host_path: false closes a small trap. Without it, Docker creates a missing host path as an empty directory, and the app would start with an empty “config file” that is really a folder. With it, Compose fails instead. make up also runs a preflight first, which checks the CLI version, the sign-in, the Region and access to each of the four models, with one line per check.

Two limits of this part. The lab was run on macOS with Docker Desktop only; the README has a Linux note about user ids, but that path is untested. And when the 12-hour session ends, make login should be enough without restarting the app, because the SDK reads the sign-in files again on the next request. That follows from the code and was not observed in a run.

Why open-weight models

All four models are open-weight: their weights are published, so you could also run them on your own hardware, if you have the hardware. Bedrock hosting them means a laptop is enough to try four of them on the same task in an afternoon.

They also make a good test for the agent code. Four models from four different vendors do not behave the same way, even behind one API. The differences that showed up are in the results section below, and two of them changed the code.

The dataset: random-looking, but always the same

The database has three tables: 20 customers, 15 products and 60 orders. Orders point to customers and products through enforced foreign keys. The data is fake and looks random, but it is generated from a fixed seed.

From labs/lab-bedrock-langchain-agent/src/db/seed-data.ts:

/**
 * Generates the whole dataset from a fixed seed. The data looks random but is
 * identical on every machine, which is what makes known-answer checks possible.
 */
export function generateSeedData(seed: number = SEED): SeedData {
  const faker = new Faker({ locale: [en] });
  faker.seed(seed);
  faker.setDefaultRefDate(REFERENCE_DATE);
  return {
    customers: makeCustomers(faker),
    products: makeProducts(faker),
    orders: makeOrders(faker),
  };
}

The reference date matters as much as the seed. Without it, the generated dates depend on the day the generator runs, and the dataset changes overnight.

Reproducible data is what makes it possible to check an agent at all. If the customer Marcella Cormier exists on every machine with the same email address, a test can ask for that address and compare. The generator also shapes the data so that questions have exactly one answer. Stock values are unique, so “lowest stock in a category” is never a tie. And one customer gets exactly one order.

From labs/lab-bedrock-langchain-agent/src/db/seed-data.ts:

function makeOrders(faker: Faker): SeedData["orders"] {
  return Array.from({ length: ORDER_COUNT }, (_, index) => ({
    id: index + 1,
    // The last customer gets exactly one order (the last one), which gives
    // the e2e a question with a single unambiguous answer.
    customerId:
      index === ORDER_COUNT - 1
        ? CUSTOMER_COUNT
        : faker.number.int({ min: 1, max: CUSTOMER_COUNT - 1 }),
    productId: faker.number.int({ min: 1, max: PRODUCT_COUNT }),
    quantity: faker.number.int({ min: 1, max: 5 }),
    status: faker.helpers.arrayElement(ORDER_STATUSES),
    orderedAt: faker.date.between({ from: YEAR_START, to: REFERENCE_DATE }),
  }));
}

The app applies the migrations and seeds the database before it starts listening. So when GET /readyz answers 200, the schema and the data are there.

Three tools, and no SQL for the model

The agent never writes SQL. It gets three tools, query_customers, query_products and query_orders, and each tool takes a set of optional filters. Here is the one for orders.

From labs/lab-bedrock-langchain-agent/src/agent/tools/query-orders.ts:

const schema = z.object({
  id: positiveIntegerField("Exact order id, a whole number such as 12."),
  customer_id: positiveIntegerField(
    "Id of the customer who placed the order, a whole number. Get it from query_customers.",
  ),
  product_id: positiveIntegerField(
    "Id of the ordered product, a whole number. Get it from query_products.",
  ),
  status: z
    .enum(ORDER_STATUSES)
    .nullish()
    .describe("Order status: pending, shipped, delivered or cancelled."),
  limit: limitField,
});

Every field has a description, because the description is all the model knows about the field. The tool itself is an explicit DynamicStructuredTool.

From labs/lab-bedrock-langchain-agent/src/agent/tools/query-orders.ts:

export function createQueryOrdersTool({ db, logger }: ToolDeps) {
  return new DynamicStructuredTool({
    name: "query_orders",
    description:
      "Looks up orders of the shop. Use for: finding what a customer ordered, how many units, the order status, or which orders contain a product. Rows contain customer_id and product_id, not names. Do NOT use for: looking up names, emails or prices; use query_customers and query_products with the ids. All filters are optional and are combined with AND.",
    schema,
    func: async (filters) => {
      const condition = toCondition(filters);
      return readRows({
        db,
        logger,
        tool: "query_orders",
        limit: filters.limit,
        query: (tx, limit) =>
          tx
            .select({
              id: orders.id,
              customer_id: orders.customerId,
              product_id: orders.productId,
              quantity: orders.quantity,
              status: orders.status,
              ordered_at: orders.orderedAt,
            })
            .from(orders)
            .where(condition)
            .orderBy(orders.id)
            .limit(limit),
        count: (tx) => tx.$count(orders, condition),
      });
    },
  });
}

Three decisions in this file are deliberate.

Filters, not SQL. The model’s input becomes the value of a parameter in a query that Drizzle builds. It never becomes part of the SQL text. A model that is confused, or a user who tries to talk the model into something, can at worst ask for rows the tool was built to return.

Ids only. query_orders returns customer_id and product_id, not names. To answer “what did Brionna Ebert order”, the agent has to chain three calls: find the customer, find the orders, find the product. That is the point of the lab’s hardest question.

One result shape. All three tools return their rows through the same helper.

From labs/lab-bedrock-langchain-agent/src/agent/tools/result.ts:

  let result: ToolResult<Row>;
  try {
    result = await db.transaction(async (tx) => {
      const rows = await query(tx, limit);
      const total = await count(tx);
      return { ok: true as const, rows, total, truncated: total > rows.length };
    }, READ_ONLY);
  } catch (err) {
    // Deliberately not re-thrown: the model gets a structured error and can
    // try again. The details stay in the log and are not sent to the model.
    logger.error({ err, tool }, "Reading rows failed.");
    result = {
      ok: false,
      error: {
        code: "DATABASE_ERROR",
        message: "The database query failed. Try again or use other filters.",
      },
    };
  }
  return JSON.stringify(result);

The query runs in a read-only transaction, so Postgres itself refuses a write, whatever the tool code does. total counts all matching rows, not only the returned ones, so a counting question is answered correctly even when the rows were cut off by the limit. And an error is a value, not an exception: one failed query should not end the whole request.

The system prompt is short and the same for every model.

From labs/lab-bedrock-langchain-agent/src/agent/prompt.ts:

export function buildSystemPrompt(today: Date): string {
  const date = today.toISOString().slice(0, 10);
  return [
    "You answer questions about a shop's customers, products and orders.",
    `Today's date is ${date}.`,
    "",
    "Rules:",
    "1. You MUST call at least one tool before you answer. Never answer from memory.",
    "2. Answer only from tool results. If the tools return no rows, say that nothing was found.",
    "3. Repeat exact values (names, email addresses, prices, quantities) exactly as the tools returned them.",
    "4. Write counts and quantities as digits, for example 4.",
    "5. Orders contain ids. Use query_customers and query_products to turn ids into names.",
    "6. Keep the answer to one or two sentences.",
    "",
    "Tools: query_customers, query_products, query_orders.",
  ].join("\n");
}

The date is stated because a model otherwise assumes that “today” is the end of its training data.

One agent, four models

The models live in one registry. It is the only place in the app where a Bedrock identifier appears.

From labs/lab-bedrock-langchain-agent/src/agent/models.ts:

export const MODELS = {
  "minimax-m2.5": {
    key: "minimax-m2.5",
    bedrockId: "minimax.minimax-m2.5",
    kind: "in-region-model-id",
    sendTemperature: true,
  },
  "nemotron-super-3": {
    key: "nemotron-super-3",
    bedrockId: "nvidia.nemotron-super-3-120b",
    kind: "in-region-model-id",
    sendTemperature: true,
  },
  "deepseek-v3.2": {
    key: "deepseek-v3.2",
    bedrockId: "deepseek.v3.2",
    kind: "in-region-model-id",
    sendTemperature: true,
  },
  "kimi-k3": {
    key: "kimi-k3",
    bedrockId: "global.moonshotai.kimi-k3",
    kind: "global-inference-profile",
    sendTemperature: false,
  },
} as const satisfies Record<string, ModelEntry>;

The kind field records a difference that is easy to miss. Three of the models are called by a model ID, which reaches the model in the Region the app is configured for. The Region therefore has to offer the model. Five Regions were verified to offer all three: us-east-1, us-east-2, us-west-2, eu-west-2 and eu-north-1. In any other Region the preflight fails and the stack does not start.

Kimi K3 is called through an inference profile. With a profile, Bedrock routes the request to a Region that has capacity. A global. profile may use any Region worldwide, so you do not control where the request is processed. If that matters, for EU data residency for example, this is the line to change. A us. profile for the same model exists and keeps requests in US Regions.

The agent is built per request from one registry entry.

From labs/lab-bedrock-langchain-agent/src/agent/agent.ts:

  const model = new ChatBedrockConverse({
    model: entry.bedrockId,
    region,
    maxTokens: MAX_OUTPUT_TOKENS,
    ...(entry.sendTemperature ? { temperature: 0 } : {}),
  });
  return createAgent({
    model,
    tools,
    systemPrompt: buildSystemPrompt(today),
    middleware: [createStepLoggingMiddleware({ logger })],
  });

ChatBedrockConverse talks to Bedrock’s Converse API, which has one request and response shape for all models, tool calls included. That is what you get for free: no vendor SDK per model, no separate tool-call format per model. There are no credentials in this code either, only a Region and an identifier.

What you do not get for free is identical behaviour. The sendTemperature flag exists because Kimi K3 rejects a request that sets a temperature. The other three run at temperature 0.

A request may use at most 10 model turns and 60 seconds. The step limit stops a model that keeps calling tools, and the timeout stops a request that hangs.

Asking a question and reading the log

This question needs all three tools.

From labs/lab-bedrock-langchain-agent/README.md:

make ask MODEL=deepseek-v3.2 Q="Which product did the customer Brionna Ebert order, and how many units?"
{
  "answer": "Brionna Ebert ordered 3 units of the Bespoke Gold Car.",
  "model": "deepseek-v3.2",
  "toolCalls": [
    { "name": "query_customers", "args": { "name": "Brionna Ebert" }, "ok": true, "total": 1 },
    { "name": "query_orders", "args": { "customer_id": 20 }, "ok": true, "total": 1 },
    { "name": "query_products", "args": { "id": 15 }, "ok": true, "total": 1 }
  ],
  "durationMs": 6820
}

The response carries the tool calls next to the answer, so you can see how the answer came about and not only what it is. The same steps appear in the app’s log as they happen. A LangChain middleware wraps every model turn and every tool call and logs its start, its success or its failure as one JSON line. These are lines of the request above.

From labs/lab-bedrock-langchain-agent/README.md:

{"level":30,"time":"2026-09-27T20:46:30.825Z","appName":"shop-agent","requestId":"a2040fe2-3408-4dca-95d8-3807f075aed4","model":"deepseek-v3.2","queryLength":71,"msg":"Answering question..."}
{"level":30,"time":"2026-09-27T20:46:32.721Z","appName":"shop-agent","requestId":"a2040fe2-3408-4dca-95d8-3807f075aed4","model":"deepseek-v3.2","turn":1,"durationMs":1893,"toolCallsRequested":["query_customers"],"usage":{"input_tokens":1478,"output_tokens":77,"total_tokens":1555},"msg":"Model turn succeeded."}
{"level":30,"time":"2026-09-27T20:46:32.724Z","appName":"shop-agent","requestId":"a2040fe2-3408-4dca-95d8-3807f075aed4","model":"deepseek-v3.2","tool":"query_customers","args":{"name":"Brionna Ebert"},"msg":"Tool call..."}
{"level":30,"time":"2026-09-27T20:46:32.728Z","appName":"shop-agent","requestId":"a2040fe2-3408-4dca-95d8-3807f075aed4","model":"deepseek-v3.2","tool":"query_customers","total":1,"durationMs":4,"msg":"Tool call succeeded."}
{"level":30,"time":"2026-09-27T20:46:37.645Z","appName":"shop-agent","requestId":"a2040fe2-3408-4dca-95d8-3807f075aed4","model":"deepseek-v3.2","toolCalls":3,"durationMs":6820,"msg":"Answering question succeeded."}

The log shows where the time goes. The tool call took 4 milliseconds and the model turn before it took 1.9 seconds. It also shows the token usage per turn. It does not contain the rows that came back, nor the question text, only its length. Every line carries the request id, so one request can be picked out of a busy stream.

This log is what explained the one real bug of the lab, which comes next.

How the answers are checked

make e2e asks five questions on each of the four models, 20 requests, one after another. The questions and their expected answers are committed together.

From labs/lab-bedrock-langchain-agent/e2e/questions.ts:

  {
    id: "orders-across-tables",
    question:
      "Which product did the customer Brionna Ebert order, and how many units? Answer with the product name and the quantity as digits.",
    expected: ["Bespoke Gold Car", "3"],
  },

A unit test derives the same questions and answers from the seed and compares them with this file. If someone changes the seed, that test fails and prints the new values.

A request passes on two conditions.

From labs/lab-bedrock-langchain-agent/e2e/run.ts:

    if ((body.toolCalls ?? []).length === 0) {
      return { ...cell, passed: false, reason: "no tool call" };
    }
    const missing = question.expected.filter(
      (value) => !answerContains(body.answer ?? "", value),
    );

The first condition means that a right answer without a tool call fails. A model cannot pass by guessing. The second one compares text, and that needs some care with numbers. A plain substring check for 3 also accepts “order 13” or “product id 35”. So a number only counts when it stands alone.

From labs/lab-bedrock-langchain-agent/e2e/normalize.ts:

  const standsAlone = new RegExp(
    `(?<![a-z0-9_])(?<!\\d[.,])${escapeRegExp(needle)}(?![a-z0-9_])(?!\\.\\d)`,
  );
  return standsAlone.test(haystack);

Each cell of the grid is asked once. There are no retries, and the run exits non-zero unless all 20 pass. That rule is strict on purpose, and it has a price, as the results show.

Results

This is the grid of a passing run.

From labs/lab-bedrock-langchain-agent/README.md:

question              minimax-m2.5      nemotron-super-3  deepseek-v3.2     kimi-k3
-----------------------------------------------------------------------------------
customer-lookup       PASS              PASS              PASS              PASS
customer-filter       PASS              PASS              PASS              PASS
product-lookup        PASS              PASS              PASS              PASS
product-filter        PASS              PASS              PASS              PASS
orders-across-tables  PASS              PASS              PASS              PASS

20 of 20 passed

All four models pass all five questions, also in a run from a fresh clone that followed the README step by step. But the first live run did not look like this, and not every later run did either. The history is more useful than the grid.

The first run ended at 19 of 20. DeepSeek V3.2 failed the question that needs all three tools. The step log showed why: the model sent its numbers as text, "customer_id": "20" instead of 20. The tool’s schema rejected that, the model got a generic error, and it repeated the same call until the turn limit ended the request.

The fix has two parts. First, the numeric fields accept a whole number sent as text.

From labs/lab-bedrock-langchain-agent/src/agent/tools/result.ts:

/**
 * Some models send numbers as text (`"20"`), so plain digits become a number.
 * A value that is too large is rejected by the bounds of the field's schema.
 */
function digitsToNumber(value: number | string): number | string {
  if (typeof value === "number") return value;
  const trimmed = value.trim();
  return DIGITS_ONLY.test(trimmed) ? Number(trimmed) : value;
}

Second, when arguments really do not fit, the model now learns which field was wrong and why. The middleware checks the rejected arguments against the tool’s own schema.

From labs/lab-bedrock-langchain-agent/src/agent/invalid-arguments.ts:

export function describeInvalidArguments(
  schema: unknown,
  args: unknown,
): string | null {
  if (!(schema instanceof z.ZodType)) return null;
  const result = schema.safeParse(args);
  if (result.success) return null;
  return result.error.issues
    .map((issue) => {
      const field = issue.path.map(String).join(".") || "arguments";
      return `${field}: ${issue.message}`;
    })
    .join("; ")
    .slice(0, MAX_MESSAGE_LENGTH);
}

The model gets this back as a tool result with the code INVALID_ARGUMENTS, and the request goes on. The lesson is wider than this one model: an error message for a model has to be something it can act on. “The call failed” led to the same call again. “customer_id: expected a whole number” does not.

After the fix, two runs in a row ended at 20 of 20, with no failed tool call on any model.

Kimi K3 rejects the temperature setting. This showed up before any agent code existed, in the first test calls, and it is why the registry has a sendTemperature flag. Kimi K3 also handled the counting question in its own way: it asked for a single row with limit: 1 and read the count from total. That is a correct use of the result shape, and it saves tokens.

MiniMax M2.5 starts its answers with line breaks. The app trims the answer before it returns it.

One run lost two cells to the service. In a later run, two requests on Nemotron failed, and neither failure was the model’s answer. One was an internal server error from Bedrock on the first model turn. The other was a timeout: model turns that normally take about 2 seconds took 17, 13 and more than 30 seconds. With one attempt per cell and no retries, that fails the run. It happened in one of six runs. A cell that fails with BEDROCK_ERROR or REQUEST_TIMEOUT is the service, and running make e2e again is the right reaction. A cell that fails with “expected …, got …” or “no tool call” is the model’s own answer.

The session ended in the middle of a run. Two more runs failed for a simpler reason: the 12-hour aws login session was over. The app answered every request with 502 AWS_LOGIN_REQUIRED and the hint to run make login. That is the designed behaviour, and it was good to see it happen for real and not only in a unit test.

Over the two clean runs, the average request took about 2.0 seconds on Nemotron, 3.1 on Kimi K3, 5.1 on DeepSeek V3.2 and 6.5 on MiniMax M2.5. Read these as a first impression, not as a measurement: ten questions per model, one Region, one evening.

What this lab is not

It is not a benchmark. Five questions with known answers show whether a model can use three tools correctly on a small task. They do not show which model is better, cheaper or faster in general.

It is also not a service to put on a network. The endpoint has no authentication, which is why Compose publishes the port on 127.0.0.1 only. The container holds a session with all permissions of the identity you signed in with. And the tools read from a database that contains nothing worth protecting.

What it does show is that the distance between “one model works” and “four models work” is small when the pieces are in the right place: one registry for the identifiers, one API for all models, tools that are strict about what they accept and clear about what they reject, and a log that shows each step. make down stops the stack and deletes the database volume. It does not touch ~/.aws.