Documentation
Two halves: getting your own agent's runs into Pramana (this page's Quickstart), and reading what shows up once they're there (everything after it).
From a fresh signup to your first recorded trace, in about a minute.
Sign up at reliai.in — creates a new organization with you as its admin, no card required.
Create an API key. Settings → API keys → New key, role
engineer. This is what your agent code authenticates with — it's shown once, so
save it immediately (an environment variable, not committed to source control).
Install the SDK.
pip install pramana-sdk
Using Claude Code? Download the
Pramana integration skill,
drop it in your project's .claude/skills/pramana-integration/, and just ask
Claude to "add Pramana tracing to this agent" — it handles steps 4 onward itself, correctly,
including the multi-agent and tool-call cases below.
Instrument your agent. Three real lines, around code you already have — no rewrite, no wrapping every call by hand.
import uuid
import pramana
from pramana.sinks.http import HttpSink
pramana.init(
trace_id=str(uuid.uuid4()), # one per agent run
tenant_id="<your-tenant-id>", # Settings page, top of the page
sink=HttpSink(
"https://www.reliai.in/ingest/v1/events:batch",
tenant_id="<your-tenant-id>",
api_key="<your-api-key>", # or set PRAMANA_API_KEY and omit this
),
)
pramana.instrument(openai_client) # or an Anthropic client — same call either way
# ...run your agent exactly as you already do...
Every call through openai_client (or the Anthropic client) is now recorded —
the model, the prompt, the response, latency, token counts — without changing how you call it.
Watch it appear. Open the dashboard — your trace shows up within a few seconds of the run finishing.
Tool calls and inter-agent messages aren't auto-instrumented the way LLM
calls are — wrap them explicitly with interceptor.intercepted_call(...) (a tool
call) or pramana.send_message/receive_message (a handoff between two
of your own agents) if you want those to show up too. LLM calls alone already cover the most
common case.
Multiple agents in one run? Give each its own agent_id in
pramana.init(...), but the same trace_id — that's what groups them
into one Simulation view instead of five separate traces.
Three more pramana.init() arguments beyond the Quickstart's, and what your agent's
own shape means for two of them: streaming and async.
A function you supply that takes a payload and returns what should actually be stored:
pramana.init(
...,
redactor=lambda payload: strip_pii(payload),
)
It runs on everything recorded — the input, the model's response, every
stream chunk — before it is hashed or stored, on every mode: record, replay,
model-diff, and sandboxed batch alike. It also applies by default to a tool call you wrap
yourself with interceptor.intercepted_call(...), using whatever redactor is
configured on the trace — you don't have to pass it again at that call site.
Because it runs on responses too, not just prompts, it has to tolerate whatever shape your model or tool actually returns — the same function is called with an input dict, an output dict, and (for a streamed call) each chunk in turn. A redactor that only expects the input shape will be handed the output shape as well.
A set of key names to leave out of comparison — not storage:
pramana.init(
...,
exclude_keys={"request_id", "timestamp"},
)
This is a comparison control, not a privacy control. A key you list here is
dropped only from the hash that decides whether a call matches its recording — the value still
lands in storage exactly as it does without it. Its job is stopping a volatile field (a
timestamp, a request id your own code generates) from making every replay or model-diff
report a divergence that isn't real. Use redactor for anything that must not be
stored; use exclude_keys for anything that must not affect a comparison. They are
not substitutes for each other.
Your keys are added to a short list of vendor transport arguments (things like request headers) that Pramana already excludes by default — never a replacement for it. The default, with nothing passed, excludes nothing beyond that built-in list.
A free-form label naming which change a model-diff run is evidence for:
pramana.init(..., mode="model_diff", change_id="prompt-tone-v3")
# or: pramana model-diff <ref> --change-id=prompt-tone-v3 -- <command...>
Optional to run model-diff at all — nothing refuses without it. It matters once you build a
change-approval bundle: a bundle refuses to sign unless every
run it covers agrees on one change_id, because an approval record that can't say
which change it's evidence for isn't an approval record. Only meaningful in model-diff or
sandboxed batch mode — record and replay ignore it.
An AsyncOpenAI or AsyncAnthropic client is detected automatically —
instrument it exactly the same way as a sync one. Non-streaming async calls are fully
supported across record, replay and model-diff.
Async streaming is not supported, on either vendor, and fails loudly rather
than silently: calling with both stream=True and an async client raises
NotImplementedError at the call, naming the fix. If your agent streams and is
async, use a sync client for that call, or don't set stream=True.
Supported for OpenAI clients across record, replay, model-diff, and sandboxed batch — a streamed response is recorded and replayed chunk by chunk, and model-diff assembles the whole stream before comparing it. One narrower case: a streamed tool call cannot run inside sandboxed batch today (there's no recorded-chunk replay path for a tool yet) and refuses loudly rather than executing it for real — use single-trace model-diff for that call site instead, with tools live.
Anthropic streaming is not supported at all yet, sync or async — pass
stream=True to an instrumented Anthropic client and it raises
NotImplementedError before the call is made, rather than attempting it.
Both replay and model-diff accept more than one <ref>
— the same command runs once per trace, and one summary prints on top of each trace's own
report:
pramana model-diff loan-01 loan-02 loan-03 -- python your_agent.py
Two or more traces switches model-diff to sandboxed by default — see
the mode spectrum for what that means and why. A trace whose
run halts on a diverged tool call ends that trace's run early; the batch carries on to
the next one, and a halt does not fail the batch's exit code — only a trace that crashed
outright, producing no report at all, does that.
The dashboard's home screen.
Three numbers up top — total traces, how many diverged, and the divergence rate — computed from whatever you have permission to see. Below that, every recorded trace: a status badge (clean or diverged), how many agents took part, how many events, and when it started. Search by trace ID, filter to just the diverged ones, or sort by recency or size. Click a row to open it.
The header repeats the summary (agents, events, duration) plus — once an LLM call carries usage data — total tokens and estimated cost. Below it, a row of tabs; which ones you see adapts to the trace.
The default view for a single-agent trace. Reads top to bottom like a chat transcript: user and assistant turns as bubbles, tool calls as collapsible cards, and a caption under each assistant turn showing token counts and latency when that data was captured. A diverged call is outlined and expands to show exactly what changed.
The default view for a multi-agent trace — anything with two or more agents. Each agent gets its own horizontal lane; a scrubber plays back the run like a video, and curved connector lines show exactly which send matched which receive, even with several messages in flight across different agents. This animates a run that already happened; to actually re-execute your agent against it, see the replay command.
One row per event, in true causal order — the order the vector clocks reconstruct, not just each agent's own count. Reach for this when you want strict recorded order rather than grouped by agent or read as a conversation. Click any row to inspect its raw input/output.
LLM call Tool call Message State / system
Every point where a replay's live input didn't match what was originally recorded, collected in one place with a side-by-side expected/actual diff. This is Pramana's core promise made visible: not "the output changed," but the exact call site and the exact input that changed it.
Signed, exportable proof — see Evidence & verification below.
A separate screen from the trace list — not scoped to one trace by default. It aggregates every message ever sent between every pair of agents across your tenant's history: agents as circles (the most-connected one centered), message volume as line thickness, arrows for direction. Pick a specific trace from the selector at the top to see just that run's topology instead.
Every API key and every team member has exactly one role.
| Role | Sees raw payloads | Creates evidence bundles | Manages billing |
|---|---|---|---|
| Auditor | No — structure and divergences only | No | No |
| Engineer | Yes | Yes, if the plan allows it | No |
| Admin | Yes | Yes, if the plan allows it | Yes |
An auditor can confirm that a call diverged and exactly which call site — never the prompt or response text itself. By design: an auditor's job is confirming what changed and where, not reading production content they don't need.
An evidence bundle is a signed, tamper-evident export of a slice of a trace: the hash chain linking every event to the one before it, a Merkle root over that chain, and an Ed25519 signature over the whole thing.
pramana-verify
command-line tool — the identical check, usable without a browser at all.
A change-approval bundle is what model-diff produces evidence
for: signed proof that a comparison was run, under a stated configuration, and
concluded specific results. It has its own artifact_type
(model_diff_comparison) precisely so a verifier can never mistake it for the
production-run bundle above — the two make different claims and are checked differently.
What it attests: that a model-diff comparison ran against a named set of
traces, under one classifier version and one exclude_keys configuration, and
produced these classifications. Every model name observed during the comparison, when the
comparison started and finished, and — per trace — whether it halted, whether a control run
confirmed each finding, and how many call sites of each kind it saw.
What it does not attest: that these traces reflect production behavior — a
sandboxed run's tools were never called for real, and the bundle's own disclaimer says so in
plain language, not just in the artifact type. And it does not verify the contents of the
change itself. change_id is a label you supply when you run
model-diff — a string naming which change this is evidence for. The bundle attests that the
comparison happened and what it found; it has no way to check that the label describes the
change accurately, because nothing about the label is independently verifiable. Treat it the
way you'd treat a commit message: almost always accurate, never proof.
A trace can fail to make it into the bundle it was asked for — no rows recorded, or a run whose
bucket counts turned out incoherent — and when that happens the bundle is still signed over
whatever traces remain, never refused over one bad row. Every excluded trace is named, with a
reason, in the bundle's own excluded_traces field — inside the signed content, not
appended after, so an excluded trace can't be quietly dropped from a copy without invalidating
the signature. Three counts at the top — traces submitted, included, and excluded — reconcile
by construction, so the denominator for any call-site or cost math you do against it is never
ambiguous.
Built with a Python call today, not a CLI command or a dashboard button.
from pramana_evidence.model_diff_bundle import build_model_diff_bundle — pass it a
repository, a tenant id, the trace ids a model-diff batch covered, and a signing key. The
dashboard shows every model-diff run already (see CI
gate, model-diff & baselines); turning a batch of those runs into one signed bundle is
not yet a button there.
| Free | Pro | Enterprise | |
|---|---|---|---|
| Price | $0 | $99 / 30 days | Contact sales |
| Evidence export | ✕ | ✓ | ✓ |
| Seats (API keys + team) | 3 | 10 | Unlimited |
| Retention | 30 days | 365 days | Custom |
| Events / month | 50,000 | 2,000,000 | Custom |
Recording, replay, and every dashboard view work on every plan — evidence export is the compliance/audit feature Free doesn't include, and trying it there returns a clear "plan upgrade required" error rather than a generic failure. The events/month figure is a hard ceiling on ingest: exceed it and further writes are refused (a clear error, not silent data loss) until the next monthly period.
Upgrading is fully self-serve from the Billing page (admin only) — real payment, hosted by Cashfree. Pro is a one-time, 30-day purchase rather than an auto-renewing subscription; it isn't renewed automatically, and access reverts to Free once it expires.
Two SDK CLI commands, run locally or in a pipeline — no dashboard involved.
pramana replay <ref> [<ref>...] -- <command...>pramana model-diff <ref> [<ref>...] [--yes] [--change-id=<id>] -- <command...>
replay and model-diff look almost identical on a command line and
behave very differently. There are really three positions, not two:
replay — nothing executesmodel-diff, one trace — everything executesmodel-diff, two or more traces — sandboxed by default
Single-trace model-diff asks you to confirm before it runs and
refuses to run unattended unless you pass --yes. Run it
deliberately, against a trace whose tools are safe to re-execute. Never put it in a pipeline.
Batch mode asks too — model spend is still real — but with the side-effect warning dropped,
because the side effects are.
Sandboxing is a property of the batch, not a flag. One trace is always
live-tools; two or more is sandboxed unless you pass --live-tools, which opts the
whole batch back into executing everything for real and re-imposes the full warning. There is
no way to sandbox a single-trace run today.
When a sandboxed run reaches a tool call whose arguments no longer match the recording, it has three options and only one of them is honest.
It could replay the recorded output anyway — but the recording answers the question the agent asked last time, and this time it is asking a different one. Handing that back would fabricate a result and every step after it would be built on the fabrication. It could call the tool for real — but that breaks the one guarantee that makes running across a corpus safe at all. Or it can stop, and say exactly which call diverged and how.
So it stops. The halt is not a failure to finish — it is the finding. A tool whose arguments changed is a changed decision: the agent is about to do something different to the outside world, which is the thing you were looking for. It's reported as a root finding with the argument delta named, the same as any other:
[HALTED] issue_decision (05c7510c9e1a090e169626ddbe027f20#0) sandboxed run halted here — live arguments diverged: decision: 'denied' -> 'approved'
A halt ends that one trace's run early, and nothing else. The batch carries on to the next trace, the finding is saved and printed before the run unwinds, and a halted trace does not fail the batch's exit code — one routine halt in 500 traces must not turn model-diff back into a build gate. A trace that genuinely crashed, producing no report at all, is treated differently and does still count.
Streaming agents can't use sandboxed mode yet. A streaming call site raises
NotImplementedError: pramana mode 'sandbox' not implemented. Streaming works in
record, replay and single-trace model-diff; it is sandboxed batch specifically that has no
streaming path. If your agent streams, use single-trace model-diff and accept that tools
execute.
A pipeline step pointing at a raw trace id has a problem: the first time you deliberately change your agent's behavior, replay correctly goes red — and the only way to make it green again is to record a new trace and hand-edit the id in your pipeline config. Do that a few times and people delete the replay step instead.
So give the trace a name. Anywhere a command takes a trace id, it also takes a baseline name:
pramana baseline set checkout-flow <trace_id> # name it once
pramana replay checkout-flow -- python your_agent.py # what CI runs, forever
When a divergence turns out to be a change you meant to make, accept it with one command:
pramana rebaseline checkout-flow -- python your_agent.py
That runs your agent for real, records a fresh trace, and points checkout-flow at
it. Your pipeline config never changes. If the run fails, the name stays where it was — a
crashed agent is not a baseline.
Re-recording is unavoidable here, and worth understanding: replay compares what your code sends, not what comes back. A changed prompt has no recorded answer, and replay is blocked from asking for one, so no replay run can ever produce a new baseline for you. Accepting a change genuinely means running the agent again.
pramana baseline list shows every name and where it points. Names are lowercase
letters, digits, ., - and _, and deliberately cannot look
like a trace id — otherwise a name could shadow a real trace.
Only your API key — the same PRAMANA_API_KEY the SDK already uses to record. Both
commands read the recorded trace back over HTTPS, so there is no database credential to hand out
and nothing to install beyond the SDK itself:
export PRAMANA_API_KEY=... # engineer or admin key, from Settings
pramana replay <trace_id> -- python your_agent.py
Use an engineer or admin key, not an auditor one: replay needs
the recorded prompts and responses, and an auditor key deliberately never receives raw payload
content (see Roles). Running Pramana on your own infrastructure instead?
Set PRAMANA_PG to your own database and both commands read from it directly.
Both commands print their findings to your terminal, and every trace has a Replay tab in the dashboard carrying the same two commands pre-filled with that trace's id — so you don't have to come back here to look them up.
Model comparisons are also saved: that tab lists every model-diff run against the
trace, which model was compared, how many call sites changed, and a side-by-side recorded-vs-new
output for each one. That means a migration check isn't just terminal output one person saw —
the whole team can review it afterwards. (Replay itself runs on your machine against your own
code, so there is no button here that runs it for you.)
What each verdict means, and — just as important — what it does not mean. Every example below is real output from a 20-trace underwriting corpus, twelve instrumented call sites per trace.
Every run prints the same three parts: what the run was, two counts of what changed, then a line per call site. Here is a whole clean trace, unedited:
model-diff: CLOCK and RNG replayed from recording
model-diff: reordered tool calls are flagged separately, not counted as behavioural
model-diff: retries compared last-attempt-to-last-attempt; attempt count is not a behavioural signal
model-diff: classifier version 1.0.0
model-diff: 13 call site(s) compared
model-diff: without classification: 0 of 13 sites that ran changed (raw output ==)
model-diff: with classification: 0 behavioural · 0 consequent · 0 cosmetic · 12 unchanged · 1 superseded
model-diff: exclude_keys = (none)
model-diff: coverage is limited to instrumented calls — anything not wrapped by pramana is invisible to this report
[NO_CHANGE] identity_verify (732c5e8de713fa6b24c24d34d145e8d3#0)
[NO_CHANGE] sanctions_screen (1c8ebd76f556da3fd0fd56422b894949#0)
[NO_CHANGE] "Summarize: {'confidence': 0.97, 'document_valid'…" (827ed70057313df2fdeaf545eaf33597#0)
[NO_CHANGE] bureau_pull #2 (8f8f5dc5ffb2d9542027430a75865aec#1)
[NO_CHANGE] compute_dti (10d7b044915ba73c5a6953ffa33ca490#0)
...
...1 superseded retry attempt(s) collapsed — a later attempt at the same call site is the decision (1 no_change): bureau_pull
model-diff: final output — recorded: {'decision': 'approved', 'status': 'recorded'}
model-diff: — live: {'decision': 'approved', 'status': 'recorded'}
The two counts are the same 13 sites counted twice, deliberately. The first is what a plain equality check on the output would have told you. The second is what the classifier concluded. On a real corpus those two numbers are far apart, and the gap is the entire product: an LLM reworded 200 responses and changed no decisions.
Call sites are identified by where in your code the call is made, not by the
order it happened in. The #0 / #1 suffix is which time that same line
ran.
Each line also prints a label and, in parentheses, a hash: identity_verify
(732c5e8de713fa6b24c24d34d145e8d3#0). The hash is that call site's identity — stable
across runs, and what the counts above are keyed by. The label beside it is for a human, not a
lookup: a tool call is labelled from your tool metadata (identity_verify); an LLM
call with no name is labelled with the opening of its own prompt, because that's the thing you
can grep for in your source. Neither the hash nor the label records the source location
itself — the file and line the call was made from isn't captured today, so grepping the
label text is how you find the call site in your code, not a look-up Pramana does for you.
Every compared call site lands in exactly one of these. The counts are a strict partition — they sum to the number of sites compared, always, by construction — so a site is never counted under two headings.
NO_CHANGECOSMETICBEHAVIOURALCONSEQUENTSUPERSEDEDREORDEREDHALTEDUNSTABLEERROREDThe corpus these examples come from produced no REORDERED,
UNSTABLE or ERRORED sites, so there is no real output to show for
those three. They are described above rather than illustrated, rather than shown with invented
output.
This is the trace the whole corpus was built around: a prompt edit that also quietly loosened a risk threshold.
model-diff: 13 call site(s) compared
model-diff: without classification: 1 of 12 sites that ran changed (raw output ==)
model-diff: 1 halted before it could run
model-diff: with classification: 1 behavioural · 2 consequent · 0 cosmetic · 9 unchanged · 1 superseded
model-diff: (1 of the 1 halted site(s) above collapsed behind an earlier root cause on this trace, and are counted under consequent above, not halted)
...
[BEHAVIOURAL] risk_guardrail (2c66c0f22e99661e5d687433e53e5f6e#0) threshold: 50 -> 45
...2 finding(s) collapsed behind the root cause above (1 no_change, 1 halted): "You are a loan underwriter. Approve or deny base…", issue_decision
...1 superseded retry attempt(s) collapsed — a later attempt at the same call site is the decision (1 no_change): bureau_pull
model-diff: final output — run halted before completion; no final output produced
model-diff: (recorded output at the halted site: {'decision': 'approved', 'status': 'recorded'})
One root finding — threshold: 50 -> 45 — and two things collapsed behind it,
including the halt on issue_decision. That is the point of separating root from
consequent: the halt is real and serious, but it is not a second problem. There is one thing to
fix here, and the report says which one.
Note the disclosure line in the counts. A halt that is a known consequence of an earlier finding
is counted under consequent, not halted, so the halted count in the
classified line can legitimately differ from the raw one above it. Whenever the two disagree,
the report says so in that sentence rather than leaving you to notice.
A finding raises an obvious question: did the change cause this, or does the model just do this sometimes? So when a run produces a root finding, Pramana automatically re-runs the same trace once more against the original model and compares that too:
model-diff: re-running as a control against the original model, to rule out sampling noise
model-diff: control run — same classifier, same policies as the main run above
model-diff: 1 halted finding(s) in the main run — 1 confirmed, 0 explained by the control (model noise, not a real change)
It fires only on a root BEHAVIOURAL, ERRORED or HALTED
finding — never on a clean, cosmetic-only, reordered-only, or purely consequent run, because a
control on those has nothing to invalidate. That gating is what keeps a corpus-wide comparison
at roughly its original cost rather than double: only the traces that actually flagged pay for a
second run.
The verdict reads in opposite directions depending on the finding, and this is the one thing on this page most worth getting right:
BEHAVIOURAL and HALTEDUNSTABLE. The control not reproducing it means the
change is what caused it → confirmed.ERRORED
A missing control row is reported as inconclusive, never as "explained" — the
absence of a comparison is not evidence that a finding is noise.
Run more than one trace and you get one summary on top of every trace's own unchanged report. This is the real tail of the 20-trace run:
model-diff: 20 of 20 run(s) classified
model-diff: 260 call site(s) compared
model-diff: without classification: 4 of 257 sites that ran changed (raw output ==)
model-diff: 3 halted before they could run
model-diff: with classification: 2 behavioural · 4 consequent · 3 cosmetic · 230 unchanged · 1 halted · 20 superseded
model-diff: (2 of the 3 halted site(s) above collapsed behind an earlier root cause on the same trace, and are counted under consequent above, not halted)
model-diff: 3 of 20 run(s) had a changed decision — 3 halted on a diverged tool call, 0 diverged without halting
model-diff: root findings:
model-diff: loan-17-5a6eef: risk_guardrail threshold: 50 -> 45
model-diff: loan-18-b3e695: risk_guardrail threshold: 50 -> 45
model-diff: loan-19-233199: issue_decision sandboxed run halted here — live arguments diverged: decision: 'denied' -> 'approved'
Read it from the bottom. 3 of 20 runs had a changed decision, and the root findings are named — not just counted — because that last line is where most readers stop. Two traces had a risk threshold move from 50 to 45 under what was labelled a prompt-tone edit; one had its final decision flip on data that changed upstream.
The 20 superseded sites are the retried bureau call, once per trace. The
4 consequent are genuine knock-on effects on the two threshold traces. Those are
different things and they are counted separately — if they shared a bucket, this corpus would
report 24 downstream effects of 2 root causes, and 20 of them would be a retry that did nothing.
A trace that refused to classify at all is excluded from these counts and named separately, so the denominator always reconciles.
Team lists everyone on your tenant. An admin adds a teammate by email and role — this hands back a one-time temporary password to share directly, since there's no email service yet to send an invite link — or removes someone, which signs them out everywhere immediately.
Settings has your account summary (tenant, role, plan), a change-password form, and API key management: create a key for a given role, see every key ever issued (including revoked ones, for the audit trail), and revoke one immediately.
The four asked most often when someone is deciding whether this fits their system.
A flight recorder for AI agents. A three-line Python SDK records every non-deterministic decision your agent makes — each model call, tool call, clock read, random draw and agent-to-agent message — so any past run can be replayed exactly, and any range of it exported as a signed evidence record. It works with the OpenAI and Anthropic clients you already have.
The use it's built around is change approval: point it at conversations you already recorded, change the model or the prompt, and see which decisions changed — not which sentences.
Yes — that's the primary workflow. Pass as many recorded traces as you like; the same command runs against each, and one summary prints on top of each trace's own report:
pramana model-diff loan-01 loan-02 loan-03 --change-id=gpt-5-migration -- python your_agent.py
Two or more traces switches it to sandboxed mode by default. Your tools are never called; a recorded tool's output is replayed back to your agent for as long as the arguments your code is now sending still match the recording, and that trace stops the moment they diverge. So a whole corpus re-runs with zero side effects. Sandboxing is decided by the number of traces, not by a flag you have to remember — see the mode spectrum.
Model calls are still real in both modes, so a run still costs model-provider money. The report sorts every call site into one of nine outcomes and separates changes that altered a decision from changes that altered only wording — see Reading a model-diff report.
No. Replay needs the execution trace, not the transcript.
Pramana matches each recorded call to where in your code it was made, which is what lets it serve the right recorded answer back at the right point when your code runs again. A transcript of prompts and responses has no call sites, so there is nothing to replay against — and no way to know which line of your agent produced which message.
The practical consequence, stated plainly: your corpus starts the day you instrument, not retroactively. If evaluating a change against history matters to you, that's the argument for instrumenting sooner rather than later.
No — and it's deliberate about which half it does answer.
It tells you deterministically which changes were decisions and which were only wording: different tool arguments, a different policy outcome, a call that stopped happening. It reaches that without reading a word of prose — "cosmetic" is decided by elimination, not by resemblance. There is no similarity score and no threshold anywhere in the product. And when a run produces a root finding, that trace is automatically re-run against the original model to rule out ordinary model variance.
But whether a changed decision is a bad decision is still a human call. Semantic judgment of changed prose is deliberately not built.
pramana rebaseline; if it wasn't, you've just caught a bug before a customer did.checkout-flow rather than a UUID. Anywhere a command takes a trace id it takes a baseline name, so a CI pipeline never has to be edited when you accept an intentional behaviour change. pramana rebaseline records a fresh trace and moves the name to it.