Record LLM calls
When a tool calls a language model on the caller’s behalf, record each request
with record_llm_call. It is an event of kind llm.call
with a fixed set of attribute names, so token counts and latency can be queried
the same way across servers.
Record a request
Section titled “Record a request”Call it from inside the tool that made the request, and the event links to that
call through call_id:
import time
from fastmcp import FastMCPfrom fastmcp_feedback.instrumentation import MemorySink, instrument
app = FastMCP("Assistant")sink = MemorySink()mw = instrument(app, [sink])
@app.toolasync def summarize(text: str) -> str: start = time.perf_counter() summary = text[:40] # stand-in for a real model request mw.record_llm_call( model="claude-opus-4", provider="anthropic", duration_ms=(time.perf_counter() - start) * 1000, input_tokens=1200, output_tokens=90, attrs={"cache_read_input_tokens": 1000}, ) return summaryTo record a failure, pass the exception or message as error. ok then
defaults to False, and an exception’s class name is stored as error_type:
@app.toolasync def translate(text: str) -> str: try: raise TimeoutError("model did not answer in 30 s") # stand-in for a failed request except TimeoutError as exc: mw.record_llm_call(model="claude-opus-4", provider="anthropic", error=exc) return textAttributes
Section titled “Attributes”| Attribute | OTel GenAI equivalent | Notes |
|---|---|---|
model | gen_ai.request.model | Always set |
provider | gen_ai.provider.name | When given |
duration_ms | gen_ai.client.operation.duration (in seconds there) | When given |
input_tokens | gen_ai.usage.input_tokens | When given |
output_tokens | gen_ai.usage.output_tokens | When given |
ok | Defaults to error is None | |
error | Redacted, cut to max_error_chars (500) | |
error_type | error.type | When error is an exception |
Anything in attrs is kept alongside, but these names win. key= works as for
record_event, for example a chat turn id. Query them with
attrs->>'model' on PostgreSQL or json_extract(attrs, '$.model') on SQLite.
Store prompt and completion
Section titled “Store prompt and completion”Chat text can hold anything a user typed, so prompts and completions are discarded unless you opt in on the middleware:
import asyncio
from fastmcp import Client
app = FastMCP("Assistant")sink = MemorySink()mw = instrument(app, [sink], capture_llm_text=True)
@app.toolasync def ask(question: str) -> str: answer = "The scene has 40 million polygons." mw.record_llm_call(model="claude-opus-4", prompt=question, completion=answer) return answer
async def main(): async with Client(app) as client: await client.call_tool("ask", {"question": "Why does the render time out?"}) await mw.flush() print(sink.events[-1].attrs)
asyncio.run(main()){'prompt': 'Why does the render time out?', 'completion': 'The scene has 40 million polygons.', 'model': 'claude-opus-4', 'ok': True}With capture_llm_text=True, both are redacted and cut to
max_llm_text_chars, which defaults to 8 * max_error_chars (4000
characters). With the default capture_llm_text=False, both are dropped,
including prompt and completion keys passed inside attrs.