~/tech-with-ugur

Talk to your Markdown: a streaming OpenAI voice RAG lab

2026-09-12 aiobservability

Run the companion lab

A small folder of Markdown should be enough to try a voice knowledge assistant. Start a connection, ask a question, inspect the passages behind the answer, and hear the reply. The interesting engineering work sits between those steps: deciding which component owns the question, the history, the retrieval, and the audio currently playing.

This companion lab connects a Next.js page, a Hono backend, and PostgreSQL with pgvector. Its default simulated mode runs without an API key. Live mode adds OpenAI embeddings, a Responses function-calling agent, and a Realtime WebRTC connection for transcription and spoken delivery.

The local application path has been tested. Paid inference, transcript accuracy, answer quality, audible latency, and voice intelligibility remain unverified. Keep that distinction in mind when trying the fictional handbook example.

Follow one question through the application

Clicking Start creates a browser UUID and acquires the microphone. In live mode the backend mints an ephemeral Realtime credential; the browser uses it to connect directly to OpenAI over WebRTC. The permanent API key stays in Hono. Only the frontend’s loopback port 3000 is published; the backend and database stay on the Compose network.

The Realtime connection produces a completed speech transcript. The page displays that recognized question and sends it to the backend as the query, together with the UUID. There is no intermediate voice-model rewrite of the question. Transcription itself can still misrecognize a word.

The backend loads completed history from PostgreSQL, forces a document-retrieval function call, and streams a final answer after retrieval. The page shows source passages and incoming text. Complete phrases enter a sequential Realtime speech queue while the rest of the answer is still being generated.

PathJobState and credentials
Browser ↔ RealtimeTranscribe microphone audio and speak supplied phrasesEphemeral credential; active voice connection
Hono ↔ embeddingsEmbed Markdown chunks and retrieval queriesBackend-only permanent key; vectors stored locally
Hono ↔ ResponsesSelect the retrieval query and generate an evidence-based answerBackend-only key; explicit history loaded from PostgreSQL
Hono ↔ PostgreSQLStore corpus, chats, turns, answers, and sourcesInternal database connection

The implementation uses a local Responses agent. Chat history belongs to PostgreSQL, and each request supplies bounded history with store:false. It does not use managed Agents sessions or provider conversation IDs.

Keep voice delivery separate from document answering

Realtime has no retrieval tools in this lab. Input turn detection also has automatic response creation disabled. A completed utterance therefore enters the application’s document-answer path instead of triggering an independent voice answer.

From labs/lab-openai-voice-rag/backend/src/provider/realtime.ts:

							turn_detection: {
								type: "server_vad",
								create_response: false,
								interrupt_response: false,
							},

The pinned live configuration uses gpt-4o-transcribe for input transcription, gpt-realtime-2.1 with voice marin for delivery, gpt-4.1-mini-2025-04-14 for the Responses agent, and text-embedding-3-small for 1,536-dimensional vectors. These are this lab’s configured models; the optional live check still needs actual account access.

Index Markdown without leaving stale chunks behind

The backend reads Markdown from the read-only tmp/documents mount. Startup applies checked-in Drizzle migrations, creates the vector extension, and refreshes the corpus. You can also refresh after editing files on the host.

Each document has a path identity, content hash, and embedding/chunking fingerprint. The fingerprint matters because an unchanged file still needs new vectors when switching between simulated and live embeddings, or when the chunking configuration changes. A content hash alone cannot tell you whether stored vectors belong to the current embedding space.

A refresh embeds new or changed documents, skips compatible unchanged documents, and removes deleted documents. Replacements and deletions commit atomically. A failed refresh preserves the previous corpus; failed startup regeneration prevents readiness. The reader does not need to reset the database to change modes.

This is a bounded local corpus: at most 256 KiB per Markdown file and 2 MiB in total, with 800-character chunks. Escaping symlinks are rejected. PostgreSQL data lives in the tmp/postgres bind directory and survives container recreation.

Retrieve the passage and its neighbors

The fictional handbook separates the amber valve description from its recovery code, ORCHID-47, across neighboring chunks. Finding the passage about the valve is useful only if retrieval also brings back the nearby instruction containing the code.

The backend embeds the query, validates the vector, and orders chunks by cosine distance using a parameterized Drizzle SQL expression. The default search takes five hits.

From labs/lab-openai-voice-rag/backend/src/corpus/retrieve.ts:

			const distance = sql<number>`${chunks.embedding} <=> ${JSON.stringify(vector)}::vector`;
			const hits = await tx
				.select()
				.from(chunks)
				.orderBy(distance, asc(chunks.id))
				.limit(topK);

For each hit, a second query includes the previous, matching, and next ordinal from the same file. The filename constraint prevents the neighbor window from pulling passages out of another document.

From labs/lab-openai-voice-rag/backend/src/corpus/retrieve.ts:

						...hits.map((hit) =>
							and(
								eq(chunks.filename, hit.filename),
								between(chunks.ordinal, hit.ordinal - 1, hit.ordinal + 1),
							),
						),
					),
				)
				.orderBy(asc(chunks.filename), asc(chunks.ordinal));

Overlapping windows return each matching database row once. Results are ordered by filename and chunk ordinal, then fitted into a 6,000-character context budget that includes source headers and separators. The returned sources carry IDs, filenames, ordinals, and the text actually included in context.

There is a useful tradeoff here: final context follows document order rather than global relevance rank, and the budget can truncate a later passage. Neighbor expansion restores nearby context; it does not guarantee that the answer is present or that the model will use it correctly. The simulated embeddings verify database behavior, not real semantic ranking quality.

Run a bounded two-round agent

The agent deliberately has a small execution path. Its first Responses request forces exactly one retrieve_documents call and disables parallel tool calls. The tool arguments are validated before local retrieval runs. Pre-retrieval model prose never becomes an answer delta in the browser.

From labs/lab-openai-voice-rag/backend/src/provider/openai.ts:

	const first = await client.responses.create(
		{
			...base,
			input,
			tools: [tool],
			tool_choice: { type: "function", name: tool.name },
			parallel_tool_calls: false,
		},
		{ signal },
	);
	const { call, output } = await retrievalOutput(first, signal);
	const { question, callId } = toolQuestion(call);
	const evidence = await retrieve(question, signal);
	signal.throwIfAborted();
	emit({ type: "sources", sources: evidence.sources });
	emit({ type: "status", stage: "answering" });

After retrieval, the application submits the evidence with the original function call ID. The second request has no tools and streams the final answer. There is no open-ended tool loop.

From labs/lab-openai-voice-rag/backend/src/provider/openai.ts:

	const finalInput: ResponseInput = [
		...input,
		...output,
		{
			type: "function_call_output",
			call_id: callId,
			output: JSON.stringify(evidence),
		},
	];
	const final = await client.responses.create(
		{ ...base, input: finalInput, tools: [], tool_choice: "none" },
		{ signal },
	);
	const answer = await finalText(final, signal, emit);
	return { answer, sources: evidence.sources };

The answer instructions request concise evidence-based output, citations, abstention when evidence is insufficient, and clarification for ambiguous entities. Those are behavioral requests. Retrieval is enforced by the application, but answer accuracy and faithful speech still require live evaluation.

A chat’s turns are serialized. Scheduling is bounded to 100 active chats and four active-plus-pending turns per chat; each turn has a 30-second deadline including queue time. The current question is limited to 2,000 characters. Completed history contains up to 20 whole turns and 40,000 characters, preserving user/assistant pairs rather than cutting through a turn.

The backend records the human question before execution and saves the complete assistant answer and sources before emitting final success. Failed and cancelled turns remain separate from completed history. Partial streamed text is never treated as a completed answer.

Stream text, then wait for actual playback

The HTTP response is newline-delimited JSON: searching status, sources, answering status, text deltas, and a final done after persistence. A stream ending without done is a failure. The Next.js proxy forwards the stream so the page can show the first delta before completion.

The frontend buffers complete phrases and sends them into a speech queue. Each phrase uses isolated supplied text with no conversation history and no tools. Citations stay on screen. Reading the phrase exactly is prompted model behavior, so this is not a guarantee of deterministic text-to-speech.

A subtle queue rule prevents overlapping replies: generation finishing is different from playback finishing. The queue records response.done, but advances only after the matching output_audio_buffer.stopped event.

From labs/lab-openai-voice-rag/frontend/src/voice/speech-queue.ts:

		if (e.type === "response.done" && response?.id === active.id) {
			active.generating = false;
			if (response.status !== "completed") {
				dispose();
				fail();
			}
			return;
		}
		if (
			e.type === "output_audio_buffer.stopped" &&
			e.response_id === active.id
		) {
			clearTimeout(active.timer);
			active = undefined;
			deliver();
		}

A new utterance interrupts the pending answer and clears active and queued speech while retaining the chat UUID. Stop aborts requests, releases microphone tracks, closes the voice connection, clears displayed results, and rejects late work. Another Start creates a fresh UUID; previous chat records remain in PostgreSQL.

A knowledge-request failure permits another question in the same connection. A connection failure requires Start again. Failed turns are not automatically retried, avoiding an invisible second paid request.

The teal robot orb measures incoming agent audio rather than the microphone. It stays still in simulated mode, and reduced-motion preferences replace movement with readable speaking feedback.

Try the keyless path first

You need Docker Compose, a browser, and a free localhost port 3000. From the lab directory:

From labs/lab-openai-voice-rag/README.md:

cp .env.example .env
docker compose up --build

Open http://localhost:3000, wait for readiness, click Start, and allow microphone access. Simulated mode accepts a typed question: What is the amber valve recovery code? Inspect the answer and source passages containing ORCHID-47.

The simulated transport acquires and releases a real microphone but does not recognize speech or produce audio. Its deterministic embeddings and scripted agent exercise the actual corpus, PostgreSQL queries, streaming protocol, and persisted turns without API charges.

Edit or add Markdown under tmp/documents, then run the refresh command:

From labs/lab-openai-voice-rag/README.md:

docker compose exec frontend npm run corpus:refresh

The JSON logs report added, changed, deleted, and unchanged counts. Delete a test document and refresh again to observe removal. Stop and Start to inspect the fresh-session behavior. The lab README contains the full test commands, persistence/reset steps, package pins, and optional live setup.

What verification establishes

The recorded final clean remote-checkout run passed 63 unit tests, 15 real-PostgreSQL integration tests, and eight browser tests, along with service typechecks, lint, unused-code checks, and the production build. Verification used isolated Compose ports and storage with explicit example-environment consumption rather than the reader’s default runtime. The paid live test was skipped.

Actual Next/Hono canary streams delivered sources and deltas before final completion. Database checks found two completed turns in one chat and one in a fresh chat, retained across container recreation. Controlled browser tests cover canonical transcript forwarding, early streaming and speech sequencing, interruption, recoverable failures, fresh sessions, and microphone cleanup.

These checks establish local contracts and persistence. They do not establish paid transcription accuracy, semantic retrieval quality, grounded live answers, or intelligible audio.

For live mode, the README explains setting the ignored local environment file to MODE=live with a backend-only application key. The optional prerecorded canary harness captures actual transcripts, NDJSON frames, speech events, and remote audio across a repeated chat and a fresh connection. A passing capture still needs someone to play the recordings and confirm intelligibility.

No measured live latency or cost is reported. Frame timestamps are observations rather than audible latency measurements. Numeric Realtime completion usage is captured when supplied; Embeddings and Responses usage must be collected separately from the provider dashboard. Missing usage remains pending.

Know what leaves the laptop and what stays

Live microphone audio goes directly to OpenAI Realtime. Document text and query text are sent for embeddings; questions, bounded history, and retrieved evidence are sent to Responses. Vectors being stored locally does not mean all inference stays local.

PostgreSQL retains document text, vectors, UUIDs, questions, completed answers and sources, plus failed/cancelled turn records. Backend operation logs include questions, filenames, counts, and timings, while withholding document contents, permanent keys, ephemeral credentials, and raw provider errors. Optional live artifacts retain transcripts and remote audio locally.

This is a local learning application without authentication or deployment configuration. Its useful pattern is explicit ownership: PostgreSQL owns durable chat history, Hono owns retrieval and answer generation, and the browser owns the active microphone and playback lifecycle. Keeping those boundaries visible makes the voice-to-document path easier to inspect and test.