Skip to content

Backfill embeddings

An embedding sink embeds records as they arrive. Records stored before it was running can be embedded afterwards, reading them back from the DatabaseSink’s tables.

Build the sinks as for live embedding, then pass the embedding sink to backfill_embeddings:

import os
from fastmcp import FastMCP
from fastmcp_feedback.instrumentation import (
DatabaseSink,
EmbeddingSink,
OpenAIEmbedder,
backfill_embeddings,
instrument,
)
app = FastMCP("Render Farm")
embedder = OpenAIEmbedder(
os.environ["EMBED_BASE_URL"],
api_key=os.environ.get("FFB_EMBED_API_KEY"),
model="mxbai-embed-large",
dim=1024,
)
db = DatabaseSink(os.environ["DATABASE_URL"], create_tables=True, embedding_dim=1024)
mw = instrument(app, [db, EmbeddingSink(embedder, db)])

The snippets from here on use top-level await. Run them inside your server’s event loop, an async function, or the python -m asyncio shell:

print(await backfill_embeddings(mw.embedding_sink))
print(await backfill_embeddings(mw.embedding_sink)) # a second run
{'feedback': 1, 'call_error': 1, 'event': 0, 'llm': 0, 'skipped': 0, 'errors': 0}
{'feedback': 0, 'call_error': 0, 'event': 0, 'llm': 0, 'skipped': 2, 'errors': 0}

It picks up the same four kinds of text as the live sink: calls with outcome error or soft_error, feedback.submitted events, llm.call events with a prompt or completion, and other events with any of the sink’s event_text_keys. Calls recorded before the outcome column existed (2026.09.27.4) have it NULL; those with an error_message count as errors.

Each text goes through the sink’s own extraction, redactor and max_chars, so it comes out with the same text, text_hash and source_id the live sink would give that record. That makes it idempotent: the second run above embeds nothing and reports the rows as skipped, and records arriving later dedupe against what was backfilled.

from datetime import datetime
counts = await backfill_embeddings(
mw.embedding_sink,
sources=["call_error"], # default: the sink's sources
since=datetime(2026, 9, 1), # naive means UTC
limit=500, # at most 500 texts, oldest first
dry_run=True, # count only; call no embedder
)

since and until select rows by timestamp, [since, until). With limit, running it again carries on where it stopped. Rows are read batch_size (500) at a time and embedded in chunks of the sink’s batch_size, so a large table never sits in memory. A failed chunk is logged and counted in errors, never raised; only a missing table or bad arguments raise.

If your feedback lives in your own tables rather than as feedback.submitted events, pass it to backfill_feedback as (ref, title, description) tuples. It builds the same text such an event would:

from fastmcp_feedback.instrumentation import backfill_feedback
rows = [("bug-7Q2X", "Render hangs", "Stuck at 99%"), ("bug-8K1D", "Crash", None)]
print(await backfill_feedback(mw.embedding_sink, rows))
{'feedback': 2, 'skipped': 0, 'errors': 0}

Close the sinks when you are done with await mw.aclose().

The same backfill runs from the command line over a database URL. The API key comes from the FFB_EMBED_API_KEY environment variable (or the variable that --api-key-env names), never from a flag:

Terminal window
python -m fastmcp_feedback.instrumentation.backfill \
--database-url "$DATABASE_URL" \
--since 2026-09-01 --sources call_error,feedback --dry-run

--database-url takes an async URL, such as postgresql+asyncpg://... or sqlite+aiosqlite:///calls.db. The command prints the counts as JSON:

{"dry_run": true, "feedback": 1, "call_error": 1, "skipped": 0, "errors": 0}

It exits 0 on success, 1 if any text failed to embed, and 2 on a setup error (bad arguments, a missing table, an unreachable database). --create-tables creates ffb_embeddings (and on PostgreSQL the vector extension) when it is missing; without it, a missing table is an error. All options are in the backfill reference.