Backfill API and CLI
Both functions need the instrumentation,embeddings extras and are imported from
fastmcp_feedback.instrumentation. How-to: Backfill embeddings.
backfill_embeddings
Section titled “backfill_embeddings”await backfill_embeddings(embedding_sink, *, sources=None, since=None, until=None, limit=None, batch_size=500, dry_run=False) -> dict[str, int]Reads the tables of embedding_sink.database_sink and embeds what has no
embedding yet:
| Source | Rows read |
|---|---|
call_error | calls with outcome error or soft_error, and calls with outcome NULL and an error_message (written before 2026.09.27.4) |
feedback | feedback.submitted events |
llm | llm.call events with a prompt or completion |
event | other events with any of the sink’s event_text_keys |
| Parameter | Description |
|---|---|
sources | Which of the above. Defaults to the sink’s sources. |
since, until | Only rows timestamped in [since, until). Naive datetimes are UTC. |
limit | Embed at most this many texts, oldest first. A later run carries on. |
batch_size | Rows read per query. Texts are embedded in chunks of the sink’s batch_size, each limited to its timeout. |
dry_run | Count what would be embedded; call no embedder and write nothing. |
Returns texts embedded per source type plus skipped (already stored with the
same hash) and errors (texts whose embed or store failed, logged). Raises
ValueError for bad arguments and RuntimeError when a table it needs does not
exist (including ffb_embeddings, unless the DatabaseSink has
create_tables=True).
backfill_feedback
Section titled “backfill_feedback”await backfill_feedback(embedding_sink, items, *, batch_size=500, dry_run=False) -> dict[str, int]items is an iterable of (ref, title, description) tuples, where
description may be None. Each becomes the text a feedback.submitted event
with key ref would give, so feedback recorded later as an event dedupes
against these rows. Returns {"feedback", "skipped", "errors"}.
Command line
Section titled “Command line”python -m fastmcp_feedback.instrumentation.backfill --database-url URL [options]| Option | Default | Description |
|---|---|---|
--database-url | required | Async SQLAlchemy URL, such as postgresql+asyncpg://... or sqlite+aiosqlite:///calls.db. |
--embed-base-url | required unless --dry-run | OpenAI-compatible base URL up to the version. |
--api-key-env NAME | FFB_EMBED_API_KEY | Environment variable holding the API key. The key is never a flag. |
--model | mxbai-embed-large | Embedding model. |
--dim | 1024 | Vector size. |
--prefix | ffb_ | Table prefix. |
--sources | feedback,call_error,event,llm | Comma separated. |
--event-text-keys | error_excerpt,error,message,reason | Comma separated. |
--since | ISO 8601; naive means UTC. | |
--until | ISO 8601, exclusive. | |
--limit | Embed at most this many texts. | |
--batch-size | 500 | Rows per query. |
--max-chars | 4000 | Longest text embedded. |
--create-tables | off | Create ffb_embeddings (and on PostgreSQL the vector extension) if missing. Ignored with --dry-run. |
--dry-run | off | Count only. |
-v, --verbose | off | Log at INFO. |
Prints the counts as one JSON object with a dry_run field. Exit status: 0 on
success, 1 if any text failed to embed, 2 on a setup error (bad arguments, a
missing table, an unreachable database).