~/tech-with-ugur

Managed RAG on Google Cloud: how do you prove Vertex AI Search actually retrieved anything?

2026-09-04 aicloud

Run the companion lab

Most RAG tutorials hand you a pile of moving parts — a chunker, an embedding model, a vector database, a reranker, a prompt template — and by the time the plumbing works, the only question that actually mattered has quietly gone unanswered: did retrieval do anything at all?

You ask the system a question. You get a fluent, confident, well-structured answer. That proves nothing. A large language model produces fluent, confident, well-structured answers about topics it has never retrieved a single byte on. The demo looks identical either way.

Google’s Vertex AI Search is the opposite bargain from the DIY stack: you hand it documents, it owns parsing, chunking, embedding, indexing, retrieval, reranking and grounded answer generation, and it hands back an answer with per-claim citations and grounding scores. Almost nothing is left for you to get wrong — which makes it the perfect place to stop worrying about the plumbing and start worrying about the proof.

This lab stands the whole thing up with Terraform, loads ten documents into it with a small TypeScript CLI, and then spends most of its energy on an adversarial verification suite whose entire job is to make the system fail if it is bluffing. Every code block below is copied verbatim from the lab; clone it and reproduce all of it.

First: which Google product is this, actually?

This is worth thirty seconds, because the naming is genuinely a mess and it will cost you an afternoon otherwise. The console calls it Vertex AI Search. The documentation now lives under an “Agent Search” heading. The API is discoveryengine.googleapis.com. The Terraform resources are google_discovery_engine_data_store and google_discovery_engine_search_engine. Those are all the same product, and search results will hand you all four names as if they were different things.

There are also two genuinely different neighbours that solve a similar problem:

Vertex AI Search is the fully managed one: point it at documents, get grounded answers with citations out of one API call. That’s what this lab uses, and the reason it’s the right starting point is precisely that it removes every variable except the one under test.

Four resources, and two settings that decide everything

The Terraform is small: enable two APIs, create a bucket, create a data store, create a search app. But two of those settings are load-bearing in a way nothing warns you about — get them wrong and the answer API simply refuses to exist for you.

labs/lab-vertex-ai-search-rag/terraform/07_search_engine.tf:

# The two settings that decide whether grounded answers exist at all:
# the Enterprise tier and the LLM add-on. Without both, answerQuery fails.
resource "google_discovery_engine_search_engine" "app" {
  project        = var.project_id
  engine_id      = local.engine_id
  collection_id  = "default_collection"
  location       = google_discovery_engine_data_store.corpus.location
  display_name   = "Vertex AI Search RAG lab app"
  data_store_ids = [google_discovery_engine_data_store.corpus.data_store_id]

  search_engine_config {
    search_tier    = "SEARCH_TIER_ENTERPRISE"
    search_add_ons = ["SEARCH_ADD_ON_LLM"]
  }
}

SEARCH_TIER_ENTERPRISE plus SEARCH_ADD_ON_LLM. Drop either one and you still get a perfectly functional search app — it just can’t generate grounded answers, which is the entire point.

The data store has its own trap, and this one only shows up on your second terraform plan:

labs/lab-vertex-ai-search-rag/terraform/06_data_store.tf:

resource "google_discovery_engine_data_store" "corpus" {
  project           = var.project_id
  location          = var.search_location
  data_store_id     = local.data_store_id
  display_name      = "Vertex AI Search RAG lab corpus"
  industry_vertical = "GENERIC"

  # CONTENT_REQUIRED is what makes this an unstructured-document store:
  # the documents themselves are the data, not rows of metadata.
  content_config = "CONTENT_REQUIRED"
  solution_types = ["SOLUTION_TYPE_SEARCH"]

  # The lab is disposable; let terraform destroy really destroy it.
  deletion_policy = "DELETE"

  # The API fills this in for you at creation time. Leave it out of the
  # configuration and the next plan reads it as drift and wants to replace the
  # data store — which the API refuses while a search app still points at it.
  document_processing_config {
    default_parsing_config {
      digital_parsing_config {}
    }
  }

  depends_on = [google_project_service.discoveryengine]
}

That document_processing_config block is pure defensive configuration. The API populates it server-side whether you write it or not; omit it and the next plan sees a field that appeared out of nowhere, decides the resource needs replacing, and then discovers the API won’t delete a data store that a search app still points at. You end up wedged for no reason at all.

Three separate things stop the very first deploy

None of these are in the quickstart, and all three fired in order on a fresh project:

  1. Application Default Credentials carry no quota project, and Discovery Engine refuses calls without one: 403 ... requires a quota project, which is not set by default. Fixed with gcloud auth application-default set-quota-project.
  2. Terraform ignores that quota project anyway. The second apply failed identically until both provider blocks were told to override:

labs/lab-vertex-ai-search-rag/terraform/00_main.tf:

# The Discovery Engine API refuses calls that carry no quota project, and
# Terraform does not pick one up from your Application Default Credentials on
# its own. These two settings tell it to bill the API call to your own project.
provider "google" {
  project               = var.project_id
  billing_project       = var.project_id
  user_project_override = true
}
  1. google_project_service cannot bootstrap itself. With the override on, the Service Usage calls now bill to the lab project — which fails with “Cloud Resource Manager API has not been used in project … before or it is disabled”. You need the API that enables APIs to be enabled by hand, once, before Terraform can enable anything:
gcloud services enable cloudresourcemanager.googleapis.com serviceusage.googleapis.com --project <your-project-id>

It is a genuinely circular prerequisite, and it is a one-liner, and nothing tells you about it until the apply fails.

Getting documents in

The import job does not run as you. It runs as the Discovery Engine service agent, and that agent needs more on your bucket than the obvious grant:

labs/lab-vertex-ai-search-rag/terraform/05_bucket.tf:

# The import runs as the Discovery Engine service agent, not as you, and it
# needs two things on this bucket: to read the corpus and write its own error
# log (objectAdmin), and to look the bucket up in the first place
# (legacyBucketReader, which is what carries storage.buckets.get). Grant less
# than this and the import fails — first by trying to create a staging bucket
# of its own, then on the bucket lookup.
resource "google_storage_bucket_iam_member" "discoveryengine_object_admin" {
  bucket = google_storage_bucket.corpus.name
  role   = "roles/storage.objectAdmin"
  member = google_project_service_identity.discoveryengine.member
}

resource "google_storage_bucket_iam_member" "discoveryengine_bucket_reader" {
  bucket = google_storage_bucket.corpus.name
  role   = "roles/storage.legacyBucketReader"
  member = google_project_service_identity.discoveryengine.member
}

roles/storage.objectViewer is the intuitive grant and it is wrong twice over. Without write access the agent can’t drop its per-document error log, so it goes off to create a staging bucket of its own and dies on storage.buckets.create. And storage.buckets.get — needed just to look the bucket up — lives in legacyBucketReader, a role whose name actively discourages you from reaching for it.

The related fix is to tell the import where to put that error log in the first place, using a prefix in the bucket you already own:

labs/lab-vertex-ai-search-rag/cli/src/search/import.ts:

    const [operation] = await client.importDocuments({
      parent: branch,
      gcsSource: { inputUris: [metadataGcsUri], dataSchema: "document" },
      reconciliationMode: "INCREMENTAL",
      // Without an errorConfig, the service tries to create its own staging
      // bucket to hold the per-document error log, which needs project-wide
      // storage.buckets.create rights the service agent should not have.
      // Pointing it at a prefix in the bucket we already own avoids that.
      errorConfig: { gcsPrefix: errorGcsPrefix },
    });

Note dataSchema: "document". The design started out with "content", which points the import straight at raw files — and hashes the document IDs, so you lose the ability to say “this answer cited chunking-strategies.md” because you never chose that name. With "document" you hand it a metadata JSONL and keep control of the IDs:

labs/lab-vertex-ai-search-rag/cli/src/corpus/documents.ts:

/**
 * The `document` data schema: one JSON line per document, pointing at the real
 * file in Cloud Storage. This is what lets us choose the document ids instead of
 * having them hashed from the URI.
 */
export function buildMetadataJsonl(docs: CorpusDocument[], bucket: string): string {
  return `${docs
    .map((doc) =>
      JSON.stringify({
        id: doc.id,
        structData: { docId: doc.id, title: doc.title },
        content: { mimeType: MARKDOWN_MIME_TYPE, uri: corpusUri(bucket, doc) },
      }),
    )
    .join("\n")}\n`;
}

One more thing about ingestion that nobody warns you about: importDocuments returns as soon as the documents are accepted, not once they are queryable. Indexing takes minutes. The lab keeps that wait inside upload — polling listDocuments until all ten appear — precisely so it can’t turn verify into a flaky test that fails for reasons that have nothing to do with retrieval.

And a small piece of collateral damage worth knowing about if you are on a recent Node: the official Cloud Storage client is currently unusable.

labs/lab-vertex-ai-search-rag/cli/src/storage/upload.ts:

/**
 * The Cloud Storage client library still ships an auth stack built on
 * node-fetch 2, which breaks on current Node. A single-object upload is one
 * HTTP POST, so we make it ourselves with the auth the search client already uses.
 */
export function bucketWriter(auth: GoogleAuth, bucket: string): ObjectWriter {

@google-cloud/storage 8.0.1 transitively pins google-auth-library 9 → gaxios 6 → node-fetch 2, whose token refresh dies with ERR_STREAM_PREMATURE_CLOSE against oauth2.googleapis.com/token. The same refresh through google-auth-library 11 — which the Discovery Engine client already brings — works fine. There is no newer storage release, so the lab drops the dependency and does the upload as one POST to the JSON API.

The proof: facts that cannot be in the training data

Now the interesting part.

Every document in the corpus is a real explainer on a real RAG topic — chunking, embeddings, vector indexes, RAG vs fine-tuning, retrieval evaluation, grounding, reranking, prompt injection, MLOps, feature stores. And every one of them ends with a benchmark note citing exactly one number against a benchmark with an invented name:

labs/lab-vertex-ai-search-rag/corpus/chunking-strategies.md:

## Benchmark note

On the Frostvane-7 chunking benchmark, recursive splitting with a 200-token overlap scored 41.8 points, roughly nine points ahead of fixed-size splitting at the same chunk length.

*The Frostvane-7 chunking benchmark is fictional. It was invented for this lab so that an answer containing it could only have come from retrieving this document — no model could have learned it during pretraining.*

Frostvane-7. Halcyon-3. Marrowlight-12. These do not exist. They were invented for this lab, they appear nowhere else, and the disclaimer sits inside the document itself so nobody quoting the corpus out of context is misled.

That gives us the canary. If you ask a question only this corpus can answer and 41.8 comes back — cited to chunking-strategies.md, with a grounding score attached — the answer had to come from retrieval. There is nowhere else it could have come from.

But there’s a subtlety, and it’s the difference between a proof and a decoration:

labs/lab-vertex-ai-search-rag/cli/src/corpus/probes.ts:

/**
 * Each `fact` is the bare distinctive number from the document's benchmark
 * note, deliberately stripped of its unit word (e.g. "0.912", not "0.912
 * cosine similarity"). Two reasons:
 *
 * 1. It is robust. The number is the only part of the sentence the model
 *    cannot reword — units and word order are the model's choice, the digits
 *    are not. A model that paraphrases "0.912 cosine similarity" as "a
 *    cosine similarity of 0.912" still carries the number verbatim.
 * 2. It is the stronger evidence. The benchmark name appears in the question
 *    we ask, so a model could echo it back without retrieving anything. The
 *    number appears nowhere except inside that one document, so an answer
 *    containing it can only have come from retrieval.
 */
export const POSITIVE_PROBES: Probe[] = [
  {
    docId: "chunking-strategies",
    question:
      "What score did recursive splitting with a 200-token overlap reach on the Frostvane-7 chunking benchmark?",
    fact: "41.8",
  },
  // ...
];

Reason 1 was learned the hard way. The first suite failed a probe because the document says “achieved 0.912 cosine similarity” and the model answered “a cosine similarity of 0.912” — a correct answer, marked wrong by a brittle assertion. Match the number, never the phrasing.

Reason 2 is the one that makes the whole exercise honest. Asserting the benchmark name comes back would be theatre: the name is in the question, so a model can echo it without looking at anything. The digits are the part it cannot manufacture.

The control

A positive result on its own is still not a proof. You need the other half: the same question, with retrieval taken away.

labs/lab-vertex-ai-search-rag/cli/src/search/answer.ts:

/**
 * The control. Instead of letting the app search the corpus, hand the answer
 * generator one irrelevant passage. Anything it still says about the corpus
 * would have to come from pretraining — which is exactly what we want to rule out.
 */
export const UNRELATED_CONTEXT =
  "Tomatoes ripen faster when kept above 18 degrees Celsius and away from direct sunlight.";

export function unrelatedSearchSpec(): protos.google.cloud.discoveryengine.v1.AnswerQueryRequest.ISearchSpec {
  return {
    searchResultList: {
      searchResults: [
        {
          unstructuredDocumentInfo: {
            uri: "gs://example/unrelated.md",
            title: "Unrelated passage",
            documentContexts: [{ content: UNRELATED_CONTEXT }],
          },
        },
      ],
    },
  };
}

The answerQuery API lets you supply your own search results instead of letting it retrieve. So the control probes ask the exact same ten questions while feeding the answer generator a sentence about tomatoes. The check that runs against those answers is deliberately blunt:

labs/lab-vertex-ai-search-rag/cli/src/verify/checks.ts:

export function checkOmitsFact(name: string, result: AnswerResult, fact: string): Check {
  const passed = !result.text.includes(fact);
  const skipped = `skipped: [${result.skippedReasons.join(", ")}]`;
  const gotBack = `got back: ${excerpt(result.text)}`;
  return {
    name,
    passed,
    detail: passed
      ? `no sign of "${fact}" without retrieval (${skipped}, ${gotBack})`
      : `"${fact}" appeared without retrieval: ${excerpt(result.text)}`,
  };
}

Worth being precise about what this control does and doesn’t establish. It shows the number doesn’t come back when the corpus isn’t retrieved. It does not by itself prove the model has never seen the number, because a system that abstains under the control passes the check too. It is the right control anyway — it’s the half that rules out “the answer was in the weights all along”, which is the failure mode every RAG demo is silently vulnerable to.

Around those twenty probes the suite adds an abstention check (a question about Kubernetes pod disruption budgets, which the corpus never mentions) and a cross-document check (a question answerable only by combining two documents, which must cite both). Forty-five checks in total, all passing against a live deployment, with observed grounding scores between 0.924 and 0.994 against a 0.6 threshold.

The abstention result is worth a note: the corpus-free question came back with answerSkippedReasons: [OUT_OF_DOMAIN_QUERY_IGNORED] and zero citations, not the NO_RELEVANT_CONTENT the design expected. Both are honest refusals, so the check accepts any skip reason — but if you’re asserting on a specific string there, you’re asserting on something the API is free to change.

The bug that would have made all of this prove nothing

Here is the one worth the whole post.

The first live run produced beautiful answers. Grounding scores of 0.92 to 0.99. Correct invented numbers, every time. And zero citations on every single answer — so every checkCitesOnly assertion failed.

The obvious read is that citations were broken. They weren’t. The response shape was simply not the one the documentation’s examples show:

labs/lab-vertex-ai-search-rag/cli/src/search/shape.ts:

/**
 * A data store names its reference's source document differently depending on
 * which retrieval shape it returns: unstructured-document search nests it
 * under `unstructuredDocumentInfo`, chunk-based search under `chunkInfo`.
 * Both are read here so citations resolve regardless of which shape a given
 * data store uses.
 */
function referenceUri(reference: RawReference): string | null {
  return (
    reference.unstructuredDocumentInfo?.uri ?? reference.chunkInfo?.documentMetadata?.uri ?? null
  );
}

References came back under chunkInfo.documentMetadata.uri, and the shaping code only read unstructuredDocumentInfo.uri. One ?? away from working.

Sit with the failure mode for a second, because it is the entire argument for building the verification suite. Had the lab only checked “does the answer contain the fact” — which is what a reasonable person writes first — every check would have passed, the demo would have looked flawless, and the citation plumbing would have been silently returning nothing the whole time. It only surfaced because something was asserting on citations specifically. The proof caught the bug in the proof.

Reading the signals

Once citations resolve, ask --raw gives you the whole evidence chain:

labs/lab-vertex-ai-search-rag/README.md:

grounding score: 0.943
citations:
  - gs://vertex-search-rag-your-gcp-project-id/corpus/chunking-strategies.md
per-claim grounding:
  - 0.996 from gs://vertex-search-rag-your-gcp-project-id/corpus/chunking-strategies.md
  - 0.988 from gs://vertex-search-rag-your-gcp-project-id/corpus/chunking-strategies.md
  - 0.991 from gs://vertex-search-rag-your-gcp-project-id/corpus/chunking-strategies.md
  ...
  - 0.724 from gs://vertex-search-rag-your-gcp-project-id/corpus/chunking-strategies.md
  - 0.992 from gs://vertex-search-rag-your-gcp-project-id/corpus/chunking-strategies.md

Three distinct signals, and they answer different questions:

Add answerSkippedReasons for the case where the system correctly declines, and you have something most hand-rolled RAG stacks never bother to build: a per-answer audit trail you can assert on in CI.

Two operational things that will bite you

The answer quota is ten LLM calls per minute, per project (LlmRequestsPerMinutePerProject). A full verification run makes 22. Without backoff, verify dies a third of the way through:

labs/lab-vertex-ai-search-rag/cli/src/search/answer.ts:

/** gRPC status code for RESOURCE_EXHAUSTED — what the answer-generation quota reports. */
export const RESOURCE_EXHAUSTED_CODE = 8;
/** gRPC status code for DEADLINE_EXCEEDED — what an occasional slow answer call reports. */
export const DEADLINE_EXCEEDED_CODE = 4;
/** Total attempts at one answerQuery call, including the first — 3 retries beyond it. */
export const MAX_ANSWER_ATTEMPTS = 4;
/** First retry waits this long; each further retry doubles it. */
export const RETRY_BASE_DELAY_MS = 20_000;
/** The client's default 30s deadline is too tight for grounded answer generation. */
export const ANSWER_TIMEOUT_MS = 120_000;

That last constant is its own small lesson: the client’s default 30-second gRPC deadline is too tight for grounded answer generation, and a slow call killed a run before the timeout was raised to 120s. DEADLINE_EXCEEDED is treated as retryable alongside the quota error; everything else fails on the first attempt, because retrying a real error just wastes your time twice.

Deleting a data store reserves its ID for hours. terraform destroy returns as soon as the delete is accepted. Redeploy into the same project with the same names any time in the next couple of hours and you get:

Error: Error creating DataStore: googleapi: Error 400: DataStore projects/…/locations/global/collections/
default_collection/dataStores/vertex-search-rag-datastore is being deleted, please wait for deletion to
complete before recreating with the same ID. The deletion could take a couple of hours.

Either wait it out, or bump resource_prefix — which renames the bucket, data store and app together so the redeploy collides with nothing.

And do run the destroy. The data store and search app keep existing, and keep costing, until you delete them. A full run of this lab is well under a dollar; a data store you forgot about is not.

What this lab deliberately does not do

This is the simple base of the stack, on purpose. Every variable that isn’t “does retrieval work” has been removed:

Every one of those omissions is a lab in its own right, and that is the plan. This post is the baseline of a series that gets progressively deeper into Vertex AI Search and RAG generally — custom chunking and parsing, the ranking API as a second-pass reranker, grounding checks decoupled from retrieval, multiple data stores behind one app, and ACL-aware retrieval where the answer depends on who’s asking. Each of those complications is much easier to reason about when you already have a working, proven baseline to compare against — which is exactly what this one is for.

The technique, though, is the part that outlives any particular product. Put a fact in your corpus that the model cannot possibly know. Assert the answer contains it, cites the right source, and clears a grounding threshold. Then ask the same question with retrieval removed and assert the fact does not come back. That works against Vertex AI Search, against a hand-built pipeline, against whatever ships next quarter — and it is the only thing standing between a RAG system that retrieves and a RAG system that just sounds like it does.