~/tech-with-ugur

A living RAG corpus: syncing a Google Drive folder into Vertex AI Search

2026-09-09 aicloud

Run the companion lab

The previous lab in this series stood Vertex AI Search up over ten Markdown files in a bucket and spent almost all of its energy on one question: did retrieval actually happen, or is the model bluffing? It answered that. It also cheated, in one specific and deliberate way — the corpus was frozen. Upload once, ask questions, done.

Real corpora are not frozen. They live in a Drive folder that somebody edits on a Tuesday afternoon, reorganises on a Thursday, and archives half of in December. A RAG system over that corpus is judged on a different question entirely: did the index notice?

This lab makes a Google Drive shared drive the source of truth for the same ten documents and builds the sync loop that keeps a Vertex AI Search data store honest about them. Then it proves two things, in code, against live APIs:

That second one is the finding worth the price of admission, and it is not called out anywhere in Google’s Drive documentation. Every code block below is copied verbatim from the lab.

The obvious answer, and why it isn’t one

Vertex AI Search ships a first-party Google Drive connector. That should be the end of the post. Three things ruled it out:

Its fastest incremental sync is every three hours. Google’s connector documentation states that an incremental sync runs every three hours by default, three hours is also the shortest interval offered, and the only controls are pause and resume — there is no on-demand trigger and no streaming mode. A demo whose central claim is “watch the edit land” cannot be built on a three-hour floor.

It is console-only. There is no Terraform resource and no gcloud command for it. A lab that deploys from one script and leaves nothing to be clicked through would consist almost entirely of screenshots.

Folder scoping was removed. Administrator filters — folder and shared-drive scoping — are no longer supported for new Drive data stores, so the connector indexes an entire drive. This lab needs a corpus/ folder and an archive/ folder sitting side by side in the same shared drive, so that the move test has somewhere to move a folder to without that destination also being indexed. The connector cannot express that boundary.

(There is a fourth limitation people hit first — personal @gmail.com accounts have no Workspace customer ID and are not supported at all. That’s real, but it’s a footnote here: this lab targets Workspace tenants, where the other three still bite.)

So: walk the tree, export, stage, import, by hand. Slower to build, on-demand, scriptable, and scoped to exactly one folder.

The quota wall nobody warns you about

The first design instinct is a folder in someone’s My Drive, shared with a service account. That does not work, and the failure is not a permission error.

Service accounts have no storage quota and cannot own files. Anything a service account creates in a personal My Drive folder fails with storageQuotaExceeded regardless of how much free space the folder’s owner has, because the service account would have to own the new file and it cannot own anything. In a shared drive, the organisation owns everything, files count against pooled storage, and a service account can create freely.

So the corpus lives in a shared drive, and the reader’s manual setup is exactly two strings: paste the service account’s address into Drive’s sharing dialog, paste the shared drive’s ID into terraform.tfvars. That’s the entire cross-organisation link — and it genuinely is cross-organisation. The GCP project and the Workspace tenant can belong to completely unrelated orgs. There is no domain-wide delegation to configure, no OAuth client to register and have a Workspace admin trust. Every service account’s address ends in .iam.gserviceaccount.com, which makes it external to every Workspace domain by definition — including one created inside the very same organisation — so Drive has always treated it as an outside principal and the design costs nothing extra to work across the boundary.

A second, quieter benefit falls out of the shared drive: a file in a shared drive has exactly one parent. My Drive’s multi-parent model would have made “where does this document live” ambiguous, and the move logic below leans hard on it not being.

Content manager, not Editor

The service account gets added to the shared drive. At which access level is not a detail.

Drive’s UI renames the roles once you are inside a shared drive’s member list: Editor becomes Contributor, and Contributor cannot move a file or folder within a shared drive. It can create and edit content all day. Only Manager and Content manager can relocate something.

The verification suite moves a subfolder. Pick Editor and everything works — seeding works, syncing works, every answer comes back grounded and cited — right up until the move stage near the end of a twenty-minute run, where it fails in a way that has nothing to do with anything you would be looking at. Content manager is the lowest shared-drive role that includes move.

No key files: one account, two scopes

The service account is the only Drive identity in the CLI, and no key file is ever created for it. Terraform gives the operator Token Creator on the account and stops there:

labs/lab-vertex-ai-search-gdrive-sync/terraform/05_service_account.tf:

# This service account exists to be a Drive principal and nothing else. It holds
# no project IAM roles: no IAM role of any kind grants access to Drive content.
# Its access comes entirely from being added to the shared drive as Content
# manager, which is a manual step in the README.
resource "google_service_account" "corpus_sync" {
  project      = var.project_id
  account_id   = local.sync_sa_id
  display_name = "Corpus sync (Drive reader for the Vertex AI Search lab)"

  depends_on = [time_sleep.api_propagation]
}

# No key is ever created. The operator mints short-lived tokens for this account
# instead, which is why they need Token Creator on it.
resource "google_service_account_iam_member" "operator_can_impersonate" {
  service_account_id = google_service_account.corpus_sync.name
  role               = "roles/iam.serviceAccountTokenCreator"
  member             = var.operator_member
}

Read that comment again, because it is the single most useful sentence in the Terraform: no IAM role of any kind grants access to Drive content. Drive access is a Workspace-side grant. There is no roles/drive.reader waiting to be discovered. The share is the permission.

The CLI’s ADC then mints a short-lived token for that account, carrying only the scopes the running command needs:

labs/lab-vertex-ai-search-gdrive-sync/cli/src/drive/auth.ts:

/** Reading the corpus. `sync` holds this and nothing more, so it cannot write to Drive. */
export const DRIVE_READONLY_SCOPE = "https://www.googleapis.com/auth/drive.readonly";
/** Creating and moving. Only `seed` and the verify harness's deliberate mutations use this. */
export const DRIVE_WRITE_SCOPE = "https://www.googleapis.com/auth/drive";
/** The source credential's scope — enough to call generateAccessToken on the service account. */
const CLOUD_PLATFORM_SCOPE = "https://www.googleapis.com/auth/cloud-platform";

// ...

/**
 * No key file anywhere: the operator's ADC asks IAM Credentials for a
 * short-lived token belonging to the service account, carrying only the scopes
 * this command needs. The service account's Drive access comes from shared-drive
 * membership, not from any IAM role.
 */
async function impersonatedAuth(serviceAccount: string, scopes: string[]): Promise<Impersonated> {
  const sourceClient = await applicationAuth().getClient();
  return new Impersonated({ sourceClient, ...impersonationOptions(serviceAccount, scopes) });
}

The split between those two scopes is the part I’d actually keep from this lab. sync — the thing you would run on a schedule, unattended, forever — requests drive.readonly and nothing else. It is not policy that the sync loop can’t write to Drive; the token it holds structurally cannot. seed asks for the write scope because it has to create folders and Docs, and the verify harness asks for it because two of its three stages mutate the corpus on purpose.

Walking a shared drive

Drive’s query syntax has no recursive form. '<id>' in parents matches exactly one level, so a tree is a breadth-first walk with one listing per folder. That’s fine. The part that isn’t fine is the listing parameters:

labs/lab-vertex-ai-search-gdrive-sync/cli/src/drive/tree.ts:

/**
 * Every shared-drive listing needs all four of these. Omit any one and the API
 * returns an empty page instead of an error, which looks exactly like an empty
 * folder.
 */
export function childLister(drive: drive_v3.Drive, driveId: string): ChildLister {
  return async (parentId) => {
    const entries: DriveEntry[] = [];
    let pageToken: string | undefined;

    do {
      const { data } = await drive.files.list({
        q: `'${parentId}' in parents and trashed = false`,
        driveId,
        corpora: "drive",
        includeItemsFromAllDrives: true,
        supportsAllDrives: true,
        fields: LIST_FIELDS,
        pageSize: 100,
        pageToken,
      });
      // ...

driveId, corpora: "drive", includeItemsFromAllDrives, supportsAllDrives. Drop any one of them and the call succeeds and returns nothing. No error, no warning — an empty page, indistinguishable from an empty folder. Combine that with the other way this fails silently (a wrong or stale service account address in the sharing dialog also produces an empty listing rather than a permission error) and you have a sync loop that cheerfully reports success while indexing zero documents. sync therefore validates the fixture up front, before it walks anything, and staging refuses to proceed on an empty document set — labs/lab-vertex-ai-search-gdrive-sync/cli/src/storage/stage.ts:

    if (docs.length === 0) {
      throw new Error("Refusing to stage no documents: a FULL import would empty the data store.");
    }

Export has the mirror-image trap. files.export takes only mimeType — passing supportsAllDrives to it is an error, which is easy to get wrong when every neighbouring call demands it:

labs/lab-vertex-ai-search-gdrive-sync/cli/src/drive/export.ts:

/**
 * `files.export` takes only `mimeType` — no `supportsAllDrives`. Passing one is
 * an error, which is easy to get wrong given every other call needs it.
 * Exported content is capped at 10 MB, far above anything here.
 */
export async function exportDocAsMarkdown(drive: drive_v3.Drive, fileId: string): Promise<string> {
  const { data } = await drive.files.export(
    { fileId, mimeType: "text/markdown" },
    { responseType: "text" },
  );
  return typeof data === "string" ? data : String(data);
}

One small documentation contradiction worth recording, since the lab settled it empirically. seed creates each Google Doc by uploading the tracked Markdown file as the media body and letting Drive’s import conversion do the rest:

labs/lab-vertex-ai-search-gdrive-sync/cli/src/drive/seed.ts:

/**
 * The Doc is created by the service account but owned by the shared drive, which
 * is the only reason this works: a service account has no storage quota of its
 * own and cannot own files.
 */
export function driveDocUploader(drive: drive_v3.Drive, driveId: string): DocUploader {
  return {
    find: (parentId, name) => findByName(drive, driveId, parentId, name, DOC_MIME_TYPE),
    create: async (parentId, name, body) => {
      const { data } = await drive.files.create({
        supportsAllDrives: true,
        fields: "id",
        requestBody: { name, mimeType: DOC_MIME_TYPE, parents: [parentId] },
        media: { mimeType: SEED_MIME_TYPE, body },
      });
      // ...

Google’s uploads guide lists Word, ODT, HTML, RTF and plain text as the importable source formats for Docs, and omits Markdown entirely. Docs Help says a .md upload opens in Docs. The API sides with Docs Help — the conversion works, and the invented benchmark facts survive the Markdown → Doc → Markdown round trip intact, which the verification suite depends on because it matches strings like 41.8 points literally.

The document ID is the Drive file ID

This is the design decision the whole lab rests on, and it is one line in the metadata writer:

labs/lab-vertex-ai-search-gdrive-sync/cli/src/storage/stage.ts:

/**
 * The `document` data schema: one JSON line per document pointing at the staged
 * object. The id is the Drive file id, which is stable across renames and moves,
 * so reorganising the Drive folder never re-creates a document.
 */
export function buildMetadataJsonl(docs: StagedDocument[], bucket: string): string {
  return `${docs
    .map((doc) =>
      JSON.stringify({
        id: doc.driveFileId,
        structData: { driveFileId: doc.driveFileId, path: doc.path, title: doc.title },
        content: { mimeType: MARKDOWN_MIME_TYPE, uri: stagedUri(bucket, doc.driveFileId) },
      }),
    )
    .join("\n")}\n`;
}

Drive file IDs are stable for the life of a file, across renames and across moves. Key the index on the path or the filename instead and every rename is a delete plus an insert; every folder reorganisation re-ingests the whole corpus and invalidates every citation you’d previously handed a user. Key it on the file ID and a reorganisation costs nothing at all.

Which is exactly what the manifest diff is able to say, because it tracks the two dimensions independently:

labs/lab-vertex-ai-search-gdrive-sync/cli/src/drive/manifest.ts:

/**
 * The whole point of keying on the Drive file id: a move or a rename changes the
 * path, a real edit changes modifiedTime, and the two are independent. Only a
 * content change costs an ingest.
 */
export function diffManifests(previous: Manifest | null, next: Manifest): ManifestDiff {
  const before = new Map((previous?.entries ?? []).map((entry) => [entry.id, entry]));
  const diff: ManifestDiff = { added: [], contentChanged: [], moved: [], removed: [] };

  for (const entry of next.entries) {
    const old = before.get(entry.id);
    if (old === undefined) {
      diff.added.push(entry);
      continue;
    }
    if (old.modifiedTime !== entry.modifiedTime) {
      diff.contentChanged.push(entry);
    }
    if (old.path !== entry.path) {
      diff.moved.push({ entry, from: old.path });
    }
  }
  // ...

contentChanged and moved are separate buckets, not one “changed” bucket, and a document can land in either without the other.

The edit that has to show up

Freshness is the easy proof, and the lab automates it end to end: the verify harness edits a Doc in place through the Docs API — documents.batchUpdate with a replaceAllText request, the same thing a person typing in the browser would produce — re-syncs, and asks the question whose answer depends on that number.

labs/lab-vertex-ai-search-gdrive-sync/cli/src/verify/freshness.ts:

    checks.push(
      checkContainsFact("the new value is answerable", result, FRESHNESS_PROBE.replacement),
      checkOmitsFact("the superseded value is gone", result, FRESHNESS_PROBE.original),
      checkCitesDocument("the answer still cites the same document", result, expectedUri),
      checkExactly(
        "the document id did not change across the edit",
        [idOf(manifest, FRESHNESS_PROBE.docName)],
        [driveFileId],
      ),
    );
  } finally {
    // Put the document back so the run is repeatable, even if the checks above threw.

Four assertions, and the fourth is the one that separates “the index updated” from “the index grew a second copy”: same document ID before and after, same citation URI, new value present, old value gone. The corpus keeps the base lab’s invented benchmark facts precisely so that “the old value is gone” means something — a model cannot supply 41.8 points from pretraining, because 41.8 points on the Frostvane-7 chunking benchmark does not exist outside this lab.

Two details in that stage cost real debugging time. The restore runs in a finally, so a check that fails midway still leaves the corpus as it found it — without that, one bad run leaves the canary permanently swapped and the next run confusing. And the ask retries until the replacement value appears, because the document count is unchanged across an edit, which means the usual “wait for the expected number of documents” poll cannot tell you whether the new content is searchable yet.

The trap: a moved folder tells you nothing about its files

Here’s the part I’d want someone to read even if they never touch Vertex AI Search.

The intuitive way to build an incremental sync against Drive is changes.list: hold a page token, ask Drive what changed since, act on it. It’s the API Google provides for exactly this, and it’s cheap — no full tree walk, no re-export.

Move a folder out of your indexed root, and the changes feed reports the folder. It says nothing whatsoever about the files inside it. Drive’s change events are keyed to the object whose parents actually changed, and when you drag evaluation/ into archive/, the only object whose parents changed is evaluation/. The three documents underneath it did not move, by Drive’s reckoning. Their own parent is still evaluation/.

Meanwhile your index has three documents in it that are no longer in your corpus, and nothing in the feed will ever tell you so.

The lab asserts this rather than claiming it. The move stage moves the folder, reads the changes feed for exactly that interval, re-syncs with a full rebase, and checks all four facts at once:

labs/lab-vertex-ai-search-gdrive-sync/cli/src/verify/move.ts:

    checks.push(
      checkIdsAbsent("the moved documents left the data store", indexed, movedIds),
      checkCount(
        "the data store shrank by the moved folder",
        indexed.length,
        manifest.entries.length,
      ),
      checkIdsPresent(
        "the changes feed reported the folder",
        changes.filter((change) => change.isFolder).map((change) => change.fileId),
        [folder],
      ),
      checkIdsAbsent(
        "the changes feed said nothing about the files inside it",
        changes.map((change) => change.fileId),
        movedIds,
      ),

That third and fourth check are a pair: the feed did report the folder, and it did not report any of the files. Both, in the same interval, from the same token — so this isn’t a timing artefact.

And to be fair to the feed, the parameters do matter and are easy to get wrong in the same silent way the listing ones were:

labs/lab-vertex-ai-search-gdrive-sync/cli/src/drive/changes.ts:

/**
 * The feed is drive-wide, not folder-scoped, so callers must re-resolve which
 * changes fall inside the corpus. `restrictToMyDrive` must stay false: it "omits
 * ... shared files which have not been added to My Drive", and for a service
 * account the corpus is exactly that.
 */

That is the correct configuration, and it still doesn’t report the nested files — because there is nothing to report. The gap is in the model, not the parameters.

One nice type-level detail from the same area: corpora is a parameter of files.list alone. It is not a parameter of changes.list — in @googleapis/drive 22.0.0 the request type has no such property and passing it fails to compile, because driveId already does the scoping there. Two adjacent Drive APIs, near-identical parameter vocabularies, different sets.

Which is why the sync engine is a full re-walk

Given that, the sync loop stages the whole tree every time and imports with reconciliationMode: FULL:

labs/lab-vertex-ai-search-gdrive-sync/cli/src/search/import.ts:

/**
 * FULL rebases the data store against exactly what is staged, so documents that
 * left the Drive folder are removed. INCREMENTAL upserts by id and never removes
 * anything — which is why a deletion or a move-out is invisible to it.
 */
export async function importStaged(
  client: DocumentServiceClient,
  branch: string,
  metadataGcsUri: string,
  mode: "FULL" | "INCREMENTAL",
  logger: Logger,
): Promise<ImportOutcome> {
  try {
    logger.info({ branch, metadataGcsUri, mode }, "Importing documents...");

    const [operation] = await client.importDocuments({
      parent: branch,
      gcsSource: { inputUris: [metadataGcsUri], dataSchema: "document" },
      reconciliationMode: mode,
      // Without an error prefix the import goes looking for a staging bucket of
      // its own and fails on storage.buckets.create.
      errorConfig: { gcsPrefix: `${metadataGcsUri.replace(/\/[^/]+$/, "")}/errors` },
    });

FULL rebases: what isn’t staged gets removed. That single flag is what makes the move test pass, and it is doing the work the changes feed cannot. The lab exercises INCREMENTAL too — the freshness stage uses it, since an edit is a pure upsert by ID — so both modes are visible side by side with the property that distinguishes them made explicit: INCREMENTAL silently drops deletions.

Notice the shape of the tradeoff, though. The reason a full re-walk is affordable here is that the corpus is ten documents. The re-walk costs one Drive listing per folder and one export per Doc, every single sync, whether anything changed or not. At ten thousand documents across a few hundred folders that is a different conversation entirely — and the honest answer is that you would still not build a changes-feed-only sync, because of the trap above. You would build both: the feed as a cheap trigger for when to look, a scoped re-walk as the thing that decides what is true. The feed tells you something happened; it does not tell you what your index needs to know.

What this lab deliberately doesn’t do

Same discipline as the base lab — every variable that isn’t “does the sync loop stay correct” has been removed:

And the honest caveats about what was proven. The full sequence — deploy, seed, sync, verify, destroy — ran against a real GCP project and a real Workspace shared drive, and both headline claims held against live APIs. Not re-run, and therefore not claimed: seed’s idempotency against real Drive (it is unit-tested against fakes only), a second consecutive verify, and npm run changes across two invocations, which is the only way the persisted page token demonstrates its effect. Those are cheap to close on a redeploy and expensive to close on a first one.

The transferable part isn’t Vertex AI Search, and it isn’t Drive. It’s this: when your source of truth is a system that publishes a change feed, the feed tells you what it considers to have changed, which is a statement about its own data model — not about what your derived index now has wrong. Those two things overlap enough to be mistaken for each other, right up until somebody drags a folder.