Skip to content

Record LLM calls

When a tool calls a language model on the caller’s behalf, record each request with record_llm_call. It is an event of kind llm.call with a fixed set of attribute names, so token counts and latency can be queried the same way across servers.

Call it from inside the tool that made the request, and the event links to that call through call_id:

import time
from fastmcp import FastMCP
from fastmcp_feedback.instrumentation import MemorySink, instrument
app = FastMCP("Assistant")
sink = MemorySink()
mw = instrument(app, [sink])
@app.tool
async def summarize(text: str) -> str:
start = time.perf_counter()
summary = text[:40] # stand-in for a real model request
mw.record_llm_call(
model="claude-opus-4",
provider="anthropic",
duration_ms=(time.perf_counter() - start) * 1000,
input_tokens=1200,
output_tokens=90,
attrs={"cache_read_input_tokens": 1000},
)
return summary

To record a failure, pass the exception or message as error. ok then defaults to False, and an exception’s class name is stored as error_type:

@app.tool
async def translate(text: str) -> str:
try:
raise TimeoutError("model did not answer in 30 s") # stand-in for a failed request
except TimeoutError as exc:
mw.record_llm_call(model="claude-opus-4", provider="anthropic", error=exc)
return text
AttributeOTel GenAI equivalentNotes
modelgen_ai.request.modelAlways set
providergen_ai.provider.nameWhen given
duration_msgen_ai.client.operation.duration (in seconds there)When given
input_tokensgen_ai.usage.input_tokensWhen given
output_tokensgen_ai.usage.output_tokensWhen given
okDefaults to error is None
errorRedacted, cut to max_error_chars (500)
error_typeerror.typeWhen error is an exception

Anything in attrs is kept alongside, but these names win. key= works as for record_event, for example a chat turn id. Query them with attrs->>'model' on PostgreSQL or json_extract(attrs, '$.model') on SQLite.

Chat text can hold anything a user typed, so prompts and completions are discarded unless you opt in on the middleware:

import asyncio
from fastmcp import Client
app = FastMCP("Assistant")
sink = MemorySink()
mw = instrument(app, [sink], capture_llm_text=True)
@app.tool
async def ask(question: str) -> str:
answer = "The scene has 40 million polygons."
mw.record_llm_call(model="claude-opus-4", prompt=question, completion=answer)
return answer
async def main():
async with Client(app) as client:
await client.call_tool("ask", {"question": "Why does the render time out?"})
await mw.flush()
print(sink.events[-1].attrs)
asyncio.run(main())
{'prompt': 'Why does the render time out?', 'completion': 'The scene has 40 million polygons.', 'model': 'claude-opus-4', 'ok': True}

With capture_llm_text=True, both are redacted and cut to max_llm_text_chars, which defaults to 8 * max_error_chars (4000 characters). With the default capture_llm_text=False, both are dropped, including prompt and completion keys passed inside attrs.