# Offline evaluation

Score conversations after the fact, or from a separate process — one eval job reporting metrics for many services, identifying each conversation by session id or trace id.

Offline evaluation means scoring somewhere other than the live session: a nightly job re-grading yesterday's conversations, a separate eval team running its own harness, a human reviewer the next morning. Often one process grades conversations from many services at once.

None of that needs a session scope. On each call you say which **service** the metric belongs to and which **conversation** it describes, and Brizz attaches it there.

## Name the service

Pass `service_name` / `serviceName` to file the metric under that service — the same name the service's own agent reports under. One eval process can report for any number of services this way; each metric lands on its own service's sessions and filters.

Leave it out and the metric goes to the `app_name` / `appName` you initialized Brizz with. An invalid service name (empty, over 256 bytes, or containing line breaks) drops the metric with a warning rather than filing it under the default.

## Identify the conversation

Pass a session id, a trace id, or both:

| You pass | Brizz attaches the metric to |
|---|---|
| `session_id` | That session. |
| `trace_id` | The session that trace belongs to, looked up when Brizz ingests the metric. |
| Both | The session. The trace id is kept for attribution. |

A trace id is useful when your eval dataset was built from traces and never stored the session id.

Three rules to know:

- **An explicit `session_id` always wins.** Brizz uses it as given and doesn't check it against the trace.
- **A trace id alone must resolve.** Brizz looks the trace up in the service you named, so the trace must belong to that service, must already have been ingested, and must still be within your data retention. If Brizz can't find it, the metric is dropped. Score after the conversation has landed, not while it's still in flight.
- **An explicit `trace_id` overrides a surrounding session scope.** Calling `record_metric(..., trace_id=t)` inside `start_session(...)` attaches the metric to the trace's session, not the surrounding one. It drops the surrounding [message id](/docs/instrument/message-ids.md) too, so a trace-only metric carries neither; with an explicit `session_id`, the surrounding message id is kept.

With neither a session id nor a trace id — and no surrounding session — the SDK warns and emits nothing. The SDK doesn't check a trace id's format; a malformed one is dropped when Brizz ingests the metric, like a trace it can't find.

## Backdate with `timestamp`

Pass `timestamp` to say *when the measured thing happened*, so the metric lands on the turn it describes rather than on the evaluation run. Without it, the metric is stamped with the time you recorded it.

## Re-scoring

Reporting the same metric for a session again supersedes the previous value — the newest report wins, even when it's backdated. A nightly job can re-grade conversations with a better judge and the scores update. See [Re-scoring and the latest value](/docs/platform/external-metrics.md#re-scoring-and-the-latest-value).

## Flush before a batch job exits

The SDK exports in the background. A short-lived job can exit before the last metrics leave the process, so shut Brizz down at the end: `Brizz.shutdown()` in Python (`await Brizz.ashutdown()` from async code), `await Brizz.shutdown()` in Node. Both flush whatever is still queued. Python also runs a bounded flush at interpreter exit, but calling it yourself makes the end of the job explicit.

## Example: a nightly eval job

One process grades rows from an eval dataset spanning several services. Each row carries the service, and whichever ids the dataset kept.

:::tabs
:::tab[Python]

```python
import os
from brizz import Brizz, record_metric

Brizz.initialize(
    api_key=os.environ.get("BRIZZ_API_KEY"),
    app_name="eval-runner",
)

for row in load_eval_rows():  # e.g. yesterday's sampled conversations
    score, rationale = judge.score(row.transcript)
    record_metric(
        "answer_quality",
        score,
        unit="score",
        min_value=0,
        max_value=1,
        polarity="positive",
        service_name=row.service,     # e.g. "support-agent", "billing-agent"
        session_id=row.session_id,    # may be None ...
        trace_id=row.trace_id,        # ... if the trace id resolves to a session
        timestamp=row.ended_at,
        comment=rationale,
        attributes={"evaluator": "gpt-4o", "rubric": "v2"},
    )

Brizz.shutdown()
```

:::tab[Node.js]

```typescript
import { Brizz, recordMetric } from '@brizz/sdk';

Brizz.initialize({
  apiKey: process.env.BRIZZ_API_KEY,
  appName: 'eval-runner',
});

for (const row of await loadEvalRows()) {
  const { score, rationale } = await judge.score(row.transcript);
  recordMetric({
    name: 'answer_quality',
    value: score,
    unit: 'score',
    minValue: 0,
    maxValue: 1,
    polarity: 'positive',
    serviceName: row.service,
    sessionId: row.sessionId ?? undefined,
    traceId: row.traceId ?? undefined,
    timestamp: row.endedAt,
    comment: rationale,
    attributes: { evaluator: 'gpt-4o', rubric: 'v2' },
  });
}

await Brizz.shutdown();
```

:::

## See also

- [Record metrics](/docs/instrument/record-metric.md) — fields, validation, and describing a metric once.
- [Online evaluation](/docs/instrument/record-metric/online-evaluation.md) — recording from inside the session.
- [External metrics](/docs/platform/external-metrics.md) — where the metrics show up once they land.

---

[All Brizz documentation](https://docs.brizz.ai/llms.txt) · [Full documentation (single file)](https://docs.brizz.ai/llms-full.txt)
