Skip to content

Backfill API and CLI

Both functions need the instrumentation,embeddings extras and are imported from fastmcp_feedback.instrumentation. How-to: Backfill embeddings.

await backfill_embeddings(embedding_sink, *, sources=None, since=None, until=None,
limit=None, batch_size=500, dry_run=False) -> dict[str, int]

Reads the tables of embedding_sink.database_sink and embeds what has no embedding yet:

SourceRows read
call_errorcalls with outcome error or soft_error, and calls with outcome NULL and an error_message (written before 2026.09.27.4)
feedbackfeedback.submitted events
llmllm.call events with a prompt or completion
eventother events with any of the sink’s event_text_keys
ParameterDescription
sourcesWhich of the above. Defaults to the sink’s sources.
since, untilOnly rows timestamped in [since, until). Naive datetimes are UTC.
limitEmbed at most this many texts, oldest first. A later run carries on.
batch_sizeRows read per query. Texts are embedded in chunks of the sink’s batch_size, each limited to its timeout.
dry_runCount what would be embedded; call no embedder and write nothing.

Returns texts embedded per source type plus skipped (already stored with the same hash) and errors (texts whose embed or store failed, logged). Raises ValueError for bad arguments and RuntimeError when a table it needs does not exist (including ffb_embeddings, unless the DatabaseSink has create_tables=True).

await backfill_feedback(embedding_sink, items, *, batch_size=500, dry_run=False) -> dict[str, int]

items is an iterable of (ref, title, description) tuples, where description may be None. Each becomes the text a feedback.submitted event with key ref would give, so feedback recorded later as an event dedupes against these rows. Returns {"feedback", "skipped", "errors"}.

Terminal window
python -m fastmcp_feedback.instrumentation.backfill --database-url URL [options]
OptionDefaultDescription
--database-urlrequiredAsync SQLAlchemy URL, such as postgresql+asyncpg://... or sqlite+aiosqlite:///calls.db.
--embed-base-urlrequired unless --dry-runOpenAI-compatible base URL up to the version.
--api-key-env NAMEFFB_EMBED_API_KEYEnvironment variable holding the API key. The key is never a flag.
--modelmxbai-embed-largeEmbedding model.
--dim1024Vector size.
--prefixffb_Table prefix.
--sourcesfeedback,call_error,event,llmComma separated.
--event-text-keyserror_excerpt,error,message,reasonComma separated.
--sinceISO 8601; naive means UTC.
--untilISO 8601, exclusive.
--limitEmbed at most this many texts.
--batch-size500Rows per query.
--max-chars4000Longest text embedded.
--create-tablesoffCreate ffb_embeddings (and on PostgreSQL the vector extension) if missing. Ignored with --dry-run.
--dry-runoffCount only.
-v, --verboseoffLog at INFO.

Prints the counts as one JSON object with a dry_run field. Exit status: 0 on success, 1 if any text failed to embed, 2 on a setup error (bad arguments, a missing table, an unreachable database).