Backfill embeddings
An embedding sink embeds records as they arrive. Records
stored before it was running can be embedded afterwards, reading them back from
the DatabaseSink’s tables.
From Python
Section titled “From Python”Build the sinks as for live embedding, then pass the embedding sink to
backfill_embeddings:
import os
from fastmcp import FastMCPfrom fastmcp_feedback.instrumentation import ( DatabaseSink, EmbeddingSink, OpenAIEmbedder, backfill_embeddings, instrument,)
app = FastMCP("Render Farm")embedder = OpenAIEmbedder( os.environ["EMBED_BASE_URL"], api_key=os.environ.get("FFB_EMBED_API_KEY"), model="mxbai-embed-large", dim=1024,)db = DatabaseSink(os.environ["DATABASE_URL"], create_tables=True, embedding_dim=1024)mw = instrument(app, [db, EmbeddingSink(embedder, db)])The snippets from here on use top-level await. Run them inside your server’s
event loop, an async function, or the python -m asyncio shell:
print(await backfill_embeddings(mw.embedding_sink))print(await backfill_embeddings(mw.embedding_sink)) # a second run{'feedback': 1, 'call_error': 1, 'event': 0, 'llm': 0, 'skipped': 0, 'errors': 0}{'feedback': 0, 'call_error': 0, 'event': 0, 'llm': 0, 'skipped': 2, 'errors': 0}It picks up the same four kinds of text as the live sink: calls with outcome
error or soft_error, feedback.submitted events, llm.call events with a
prompt or completion, and other events with any of the sink’s
event_text_keys. Calls recorded before the outcome column existed
(2026.09.27.4) have it NULL; those with an error_message count as errors.
Each text goes through the sink’s own extraction, redactor and max_chars, so it
comes out with the same text, text_hash and source_id the live sink would
give that record. That makes it idempotent: the second run above embeds nothing
and reports the rows as skipped, and records arriving later dedupe against
what was backfilled.
Narrow the run
Section titled “Narrow the run”from datetime import datetime
counts = await backfill_embeddings( mw.embedding_sink, sources=["call_error"], # default: the sink's sources since=datetime(2026, 9, 1), # naive means UTC limit=500, # at most 500 texts, oldest first dry_run=True, # count only; call no embedder)since and until select rows by timestamp, [since, until). With limit,
running it again carries on where it stopped. Rows are read batch_size (500)
at a time and embedded in chunks of the sink’s batch_size, so a large table
never sits in memory. A failed chunk is logged and counted in errors, never
raised; only a missing table or bad arguments raise.
Feedback kept in your own tables
Section titled “Feedback kept in your own tables”If your feedback lives in your own tables rather than as feedback.submitted
events, pass it to backfill_feedback as (ref, title, description) tuples. It
builds the same text such an event would:
from fastmcp_feedback.instrumentation import backfill_feedback
rows = [("bug-7Q2X", "Render hangs", "Stuck at 99%"), ("bug-8K1D", "Crash", None)]print(await backfill_feedback(mw.embedding_sink, rows)){'feedback': 2, 'skipped': 0, 'errors': 0}Close the sinks when you are done with await mw.aclose().
From a shell
Section titled “From a shell”The same backfill runs from the command line over a database URL. The API key
comes from the FFB_EMBED_API_KEY environment variable (or the variable that
--api-key-env names), never from a flag:
python -m fastmcp_feedback.instrumentation.backfill \ --database-url "$DATABASE_URL" \ --since 2026-09-01 --sources call_error,feedback --dry-runexport FFB_EMBED_API_KEY=... # your embedding endpoint's keypython -m fastmcp_feedback.instrumentation.backfill \ --database-url "$DATABASE_URL" \ --embed-base-url https://api.example.com/v1 \ --model mxbai-embed-large --dim 1024 --prefix ffb_ \ --create-tables--database-url takes an async URL, such as postgresql+asyncpg://... or
sqlite+aiosqlite:///calls.db. The command prints the counts as JSON:
{"dry_run": true, "feedback": 1, "call_error": 1, "skipped": 0, "errors": 0}It exits 0 on success, 1 if any text failed to embed, and 2 on a setup error
(bad arguments, a missing table, an unreachable database). --create-tables
creates ffb_embeddings (and on PostgreSQL the vector extension) when it is
missing; without it, a missing table is an error. All options are in the
backfill reference.