Pramana Docs
← Back to app

Documentation

Everything you need to record, replay, and prove what your agent did.

Two halves: getting your own agent's runs into Pramana (this page's Quickstart), and reading what shows up once they're there (everything after it).

01Quickstart

From a fresh signup to your first recorded trace, in about a minute.

1

Sign up at reliai.in — creates a new organization with you as its admin, no card required.

2

Create an API key. Settings → API keys → New key, role engineer. This is what your agent code authenticates with — it's shown once, so save it immediately (an environment variable, not committed to source control).

3

Install the SDK.

pip install pramana-sdk

Using Claude Code? Download the Pramana integration skill, drop it in your project's .claude/skills/pramana-integration/, and just ask Claude to "add Pramana tracing to this agent" — it handles steps 4 onward itself, correctly, including the multi-agent and tool-call cases below.

4

Instrument your agent. Three real lines, around code you already have — no rewrite, no wrapping every call by hand.

import uuid
import pramana
from pramana.sinks.http import HttpSink

pramana.init(
    trace_id=str(uuid.uuid4()),          # one per agent run
    tenant_id="<your-tenant-id>",         # Settings page, top of the page
    sink=HttpSink(
        "https://www.reliai.in/ingest/v1/events:batch",
        tenant_id="<your-tenant-id>",
        api_key="<your-api-key>",         # or set PRAMANA_API_KEY and omit this
    ),
)
pramana.instrument(openai_client)  # or an Anthropic client — same call either way

# ...run your agent exactly as you already do...

Every call through openai_client (or the Anthropic client) is now recorded — the model, the prompt, the response, latency, token counts — without changing how you call it.

5

Watch it appear. Open the dashboard — your trace shows up within a few seconds of the run finishing.

Tool calls and inter-agent messages aren't auto-instrumented the way LLM calls are — wrap them explicitly with interceptor.intercepted_call(...) (a tool call) or pramana.send_message/receive_message (a handoff between two of your own agents) if you want those to show up too. LLM calls alone already cover the most common case.

Multiple agents in one run? Give each its own agent_id in pramana.init(...), but the same trace_id — that's what groups them into one Simulation view instead of five separate traces.

02Configuring what gets recorded

Three more pramana.init() arguments beyond the Quickstart's, and what your agent's own shape means for two of them: streaming and async.

redactor

A function you supply that takes a payload and returns what should actually be stored:

pramana.init(
    ...,
    redactor=lambda payload: strip_pii(payload),
)

It runs on everything recorded — the input, the model's response, every stream chunk — before it is hashed or stored, on every mode: record, replay, model-diff, and sandboxed batch alike. It also applies by default to a tool call you wrap yourself with interceptor.intercepted_call(...), using whatever redactor is configured on the trace — you don't have to pass it again at that call site.

Because it runs on responses too, not just prompts, it has to tolerate whatever shape your model or tool actually returns — the same function is called with an input dict, an output dict, and (for a streamed call) each chunk in turn. A redactor that only expects the input shape will be handed the output shape as well.

exclude_keys

A set of key names to leave out of comparison — not storage:

pramana.init(
    ...,
    exclude_keys={"request_id", "timestamp"},
)

This is a comparison control, not a privacy control. A key you list here is dropped only from the hash that decides whether a call matches its recording — the value still lands in storage exactly as it does without it. Its job is stopping a volatile field (a timestamp, a request id your own code generates) from making every replay or model-diff report a divergence that isn't real. Use redactor for anything that must not be stored; use exclude_keys for anything that must not affect a comparison. They are not substitutes for each other.

Your keys are added to a short list of vendor transport arguments (things like request headers) that Pramana already excludes by default — never a replacement for it. The default, with nothing passed, excludes nothing beyond that built-in list.

change_id

A free-form label naming which change a model-diff run is evidence for:

pramana.init(..., mode="model_diff", change_id="prompt-tone-v3")
# or: pramana model-diff <ref> --change-id=prompt-tone-v3 -- <command...>

Optional to run model-diff at all — nothing refuses without it. It matters once you build a change-approval bundle: a bundle refuses to sign unless every run it covers agrees on one change_id, because an approval record that can't say which change it's evidence for isn't an approval record. Only meaningful in model-diff or sandboxed batch mode — record and replay ignore it.

Async

An AsyncOpenAI or AsyncAnthropic client is detected automatically — instrument it exactly the same way as a sync one. Non-streaming async calls are fully supported across record, replay and model-diff.

Async streaming is not supported, on either vendor, and fails loudly rather than silently: calling with both stream=True and an async client raises NotImplementedError at the call, naming the fix. If your agent streams and is async, use a sync client for that call, or don't set stream=True.

Streaming

Supported for OpenAI clients across record, replay, model-diff, and sandboxed batch — a streamed response is recorded and replayed chunk by chunk, and model-diff assembles the whole stream before comparing it. One narrower case: a streamed tool call cannot run inside sandboxed batch today (there's no recorded-chunk replay path for a tool yet) and refuses loudly rather than executing it for real — use single-trace model-diff for that call site instead, with tools live.

Anthropic streaming is not supported at all yet, sync or async — pass stream=True to an instrumented Anthropic client and it raises NotImplementedError before the call is made, rather than attempting it.

Batch

Both replay and model-diff accept more than one <ref> — the same command runs once per trace, and one summary prints on top of each trace's own report:

pramana model-diff loan-01 loan-02 loan-03 -- python your_agent.py

Two or more traces switches model-diff to sandboxed by default — see the mode spectrum for what that means and why. A trace whose run halts on a diverged tool call ends that trace's run early; the batch carries on to the next one, and a halt does not fail the batch's exit code — only a trace that crashed outright, producing no report at all, does that.

03The traces list

The dashboard's home screen.

Three numbers up top — total traces, how many diverged, and the divergence rate — computed from whatever you have permission to see. Below that, every recorded trace: a status badge (clean or diverged), how many agents took part, how many events, and when it started. Search by trace ID, filter to just the diverged ones, or sort by recency or size. Click a row to open it.

04Inside a trace

The header repeats the summary (agents, events, duration) plus — once an LLM call carries usage data — total tokens and estimated cost. Below it, a row of tabs; which ones you see adapts to the trace.

Conversation

The default view for a single-agent trace. Reads top to bottom like a chat transcript: user and assistant turns as bubbles, tool calls as collapsible cards, and a caption under each assistant turn showing token counts and latency when that data was captured. A diverged call is outlined and expands to show exactly what changed.

Simulation

The default view for a multi-agent trace — anything with two or more agents. Each agent gets its own horizontal lane; a scrubber plays back the run like a video, and curved connector lines show exactly which send matched which receive, even with several messages in flight across different agents. This animates a run that already happened; to actually re-execute your agent against it, see the replay command.

Timeline

One row per event, in true causal order — the order the vector clocks reconstruct, not just each agent's own count. Reach for this when you want strict recorded order rather than grouped by agent or read as a conversation. Click any row to inspect its raw input/output.

LLM call Tool call Message State / system

Divergences

Every point where a replay's live input didn't match what was originally recorded, collected in one place with a side-by-side expected/actual diff. This is Pramana's core promise made visible: not "the output changed," but the exact call site and the exact input that changed it.

Evidence

Signed, exportable proof — see Evidence & verification below.

05Agent graph

A separate screen from the trace list — not scoped to one trace by default. It aggregates every message ever sent between every pair of agents across your tenant's history: agents as circles (the most-connected one centered), message volume as line thickness, arrows for direction. Pick a specific trace from the selector at the top to see just that run's topology instead.

06Roles

Every API key and every team member has exactly one role.

RoleSees raw payloadsCreates evidence bundlesManages billing
AuditorNo — structure and divergences onlyNoNo
EngineerYesYes, if the plan allows itNo
AdminYesYes, if the plan allows itYes

An auditor can confirm that a call diverged and exactly which call site — never the prompt or response text itself. By design: an auditor's job is confirming what changed and where, not reading production content they don't need.

07Evidence & verification

An evidence bundle is a signed, tamper-evident export of a slice of a trace: the hash chain linking every event to the one before it, a Merkle root over that chain, and an Ed25519 signature over the whole thing.

Export
Engineer/admin, and only if your plan includes it. Pick a range, create a bundle — it appears in the list immediately.
Verify
Paste the platform's public key (never one found inside the bundle itself — that would defeat the point) and click Verify. This recomputes the entire chain, root, and signature in your browser — an independent check, not a status Pramana's server merely asserts.
Download
The raw signed JSON, for offline verification with the standalone pramana-verify command-line tool — the identical check, usable without a browser at all.
Compliance report
The same bundle, restructured around EU AI Act Article 12's traceability fields — system identification, a timestamped event log, and the same integrity proof — for handing to an auditor rather than debugging a call site.

Change-approval bundles are a different artifact, not a smaller version of this one

A change-approval bundle is what model-diff produces evidence for: signed proof that a comparison was run, under a stated configuration, and concluded specific results. It has its own artifact_type (model_diff_comparison) precisely so a verifier can never mistake it for the production-run bundle above — the two make different claims and are checked differently.

What it attests: that a model-diff comparison ran against a named set of traces, under one classifier version and one exclude_keys configuration, and produced these classifications. Every model name observed during the comparison, when the comparison started and finished, and — per trace — whether it halted, whether a control run confirmed each finding, and how many call sites of each kind it saw.

What it does not attest: that these traces reflect production behavior — a sandboxed run's tools were never called for real, and the bundle's own disclaimer says so in plain language, not just in the artifact type. And it does not verify the contents of the change itself. change_id is a label you supply when you run model-diff — a string naming which change this is evidence for. The bundle attests that the comparison happened and what it found; it has no way to check that the label describes the change accurately, because nothing about the label is independently verifiable. Treat it the way you'd treat a commit message: almost always accurate, never proof.

A trace can fail to make it into the bundle it was asked for — no rows recorded, or a run whose bucket counts turned out incoherent — and when that happens the bundle is still signed over whatever traces remain, never refused over one bad row. Every excluded trace is named, with a reason, in the bundle's own excluded_traces field — inside the signed content, not appended after, so an excluded trace can't be quietly dropped from a copy without invalidating the signature. Three counts at the top — traces submitted, included, and excluded — reconcile by construction, so the denominator for any call-site or cost math you do against it is never ambiguous.

Built with a Python call today, not a CLI command or a dashboard button. from pramana_evidence.model_diff_bundle import build_model_diff_bundle — pass it a repository, a tenant id, the trace ids a model-diff batch covered, and a signing key. The dashboard shows every model-diff run already (see CI gate, model-diff & baselines); turning a batch of those runs into one signed bundle is not yet a button there.

08Plans

FreeProEnterprise
Price$0$99 / 30 daysContact sales
Evidence export✕✓✓
Seats (API keys + team)310Unlimited
Retention30 days365 daysCustom
Events / month50,0002,000,000Custom

Recording, replay, and every dashboard view work on every plan — evidence export is the compliance/audit feature Free doesn't include, and trying it there returns a clear "plan upgrade required" error rather than a generic failure. The events/month figure is a hard ceiling on ingest: exceed it and further writes are refused (a clear error, not silent data loss) until the next monthly period.

Upgrading is fully self-serve from the Billing page (admin only) — real payment, hosted by Cashfree. Pro is a one-time, 30-day purchase rather than an auto-renewing subscription; it isn't renewed automatically, and access reverts to Free once it expires.

09CI gate, model-diff & baselines

Two SDK CLI commands, run locally or in a pipeline — no dashboard involved.

pramana replay <ref> [<ref>...] -- <command...>
Replays a recorded trace against your command, bit-for-bit, no live network calls. Doubles as a CI regression gate: it exits non-zero if the replay diverges at all, even if the command itself exited 0 — wire it into a pipeline step to fail a build the moment an agent's real behavior drifts from a known-good recorded run.
pramana model-diff <ref> [<ref>...] [--yes] [--change-id=<id>] -- <command...>
Runs the same script with a different model swapped in, and diffs each recorded call site against what the original run produced — so you see exactly where a model or prompt change moves behavior, call by call, before you ship it. How much of your agent actually executes depends on which mode you're in — that's the next two sections, and it is the most important thing on this page.

Three modes, and which one you're in

replay and model-diff look almost identical on a command line and behave very differently. There are really three positions, not two:

replay — nothing executes
Sealed off. Every model answer comes back out of the recording, and outbound connections are blocked at the process level, below whatever HTTP client you use — so a recorded payment or email cannot fire a second time. Safe to run automatically on every pull request.
model-diff, one trace — everything executes
A real answer from a real model requires a real call, so this mode is not sealed off and cannot be. That means real charges from your model provider, and every tool your agent uses actually executes. If your agent issues refunds, this issues refunds.
model-diff, two or more traces — sandboxed by default
Passing more than one trace switches the default: your tools are never called. A recorded tool's output is replayed back to your agent for as long as the arguments your code is now passing still match what was recorded — and the moment they don't, that run stops. Model calls are still real (so this still costs money), but there are no side effects, however many traces you run. This is the mode built for pointing at a corpus of real production traffic.

Single-trace model-diff asks you to confirm before it runs and refuses to run unattended unless you pass --yes. Run it deliberately, against a trace whose tools are safe to re-execute. Never put it in a pipeline. Batch mode asks too — model spend is still real — but with the side-effect warning dropped, because the side effects are.

Sandboxing is a property of the batch, not a flag. One trace is always live-tools; two or more is sandboxed unless you pass --live-tools, which opts the whole batch back into executing everything for real and re-imposes the full warning. There is no way to sandbox a single-trace run today.

Why it stops instead of guessing

When a sandboxed run reaches a tool call whose arguments no longer match the recording, it has three options and only one of them is honest.

It could replay the recorded output anyway — but the recording answers the question the agent asked last time, and this time it is asking a different one. Handing that back would fabricate a result and every step after it would be built on the fabrication. It could call the tool for real — but that breaks the one guarantee that makes running across a corpus safe at all. Or it can stop, and say exactly which call diverged and how.

So it stops. The halt is not a failure to finish — it is the finding. A tool whose arguments changed is a changed decision: the agent is about to do something different to the outside world, which is the thing you were looking for. It's reported as a root finding with the argument delta named, the same as any other:

[HALTED] issue_decision (05c7510c9e1a090e169626ddbe027f20#0)  sandboxed run halted here — live arguments diverged: decision: 'denied' -> 'approved'

A halt ends that one trace's run early, and nothing else. The batch carries on to the next trace, the finding is saved and printed before the run unwinds, and a halted trace does not fail the batch's exit code — one routine halt in 500 traces must not turn model-diff back into a build gate. A trace that genuinely crashed, producing no report at all, is treated differently and does still count.

Streaming agents can't use sandboxed mode yet. A streaming call site raises NotImplementedError: pramana mode 'sandbox' not implemented. Streaming works in record, replay and single-trace model-diff; it is sandboxed batch specifically that has no streaming path. If your agent streams, use single-trace model-diff and accept that tools execute.

Naming a baseline, so CI never needs editing

A pipeline step pointing at a raw trace id has a problem: the first time you deliberately change your agent's behavior, replay correctly goes red — and the only way to make it green again is to record a new trace and hand-edit the id in your pipeline config. Do that a few times and people delete the replay step instead.

So give the trace a name. Anywhere a command takes a trace id, it also takes a baseline name:

pramana baseline set checkout-flow <trace_id>    # name it once
pramana replay checkout-flow -- python your_agent.py   # what CI runs, forever

When a divergence turns out to be a change you meant to make, accept it with one command:

pramana rebaseline checkout-flow -- python your_agent.py

That runs your agent for real, records a fresh trace, and points checkout-flow at it. Your pipeline config never changes. If the run fails, the name stays where it was — a crashed agent is not a baseline.

Re-recording is unavoidable here, and worth understanding: replay compares what your code sends, not what comes back. A changed prompt has no recorded answer, and replay is blocked from asking for one, so no replay run can ever produce a new baseline for you. Accepting a change genuinely means running the agent again.

pramana baseline list shows every name and where it points. Names are lowercase letters, digits, ., - and _, and deliberately cannot look like a trace id — otherwise a name could shadow a real trace.

What both commands need

Only your API key — the same PRAMANA_API_KEY the SDK already uses to record. Both commands read the recorded trace back over HTTPS, so there is no database credential to hand out and nothing to install beyond the SDK itself:

export PRAMANA_API_KEY=...        # engineer or admin key, from Settings
pramana replay <trace_id> -- python your_agent.py

Use an engineer or admin key, not an auditor one: replay needs the recorded prompts and responses, and an auditor key deliberately never receives raw payload content (see Roles). Running Pramana on your own infrastructure instead? Set PRAMANA_PG to your own database and both commands read from it directly.

Where the results show up

Both commands print their findings to your terminal, and every trace has a Replay tab in the dashboard carrying the same two commands pre-filled with that trace's id — so you don't have to come back here to look them up.

Model comparisons are also saved: that tab lists every model-diff run against the trace, which model was compared, how many call sites changed, and a side-by-side recorded-vs-new output for each one. That means a migration check isn't just terminal output one person saw — the whole team can review it afterwards. (Replay itself runs on your machine against your own code, so there is no button here that runs it for you.)

10Reading a model-diff report

What each verdict means, and — just as important — what it does not mean. Every example below is real output from a 20-trace underwriting corpus, twelve instrumented call sites per trace.

The shape of a report

Every run prints the same three parts: what the run was, two counts of what changed, then a line per call site. Here is a whole clean trace, unedited:

model-diff: CLOCK and RNG replayed from recording
model-diff: reordered tool calls are flagged separately, not counted as behavioural
model-diff: retries compared last-attempt-to-last-attempt; attempt count is not a behavioural signal
model-diff: classifier version 1.0.0
model-diff: 13 call site(s) compared
model-diff:   without classification: 0 of 13 sites that ran changed (raw output ==)
model-diff:   with classification:    0 behavioural · 0 consequent · 0 cosmetic · 12 unchanged · 1 superseded
model-diff: exclude_keys = (none)
model-diff: coverage is limited to instrumented calls — anything not wrapped by pramana is invisible to this report
  [NO_CHANGE] identity_verify (732c5e8de713fa6b24c24d34d145e8d3#0)
  [NO_CHANGE] sanctions_screen (1c8ebd76f556da3fd0fd56422b894949#0)
  [NO_CHANGE] "Summarize: {'confidence': 0.97, 'document_valid'…" (827ed70057313df2fdeaf545eaf33597#0)
  [NO_CHANGE] bureau_pull #2 (8f8f5dc5ffb2d9542027430a75865aec#1)
  [NO_CHANGE] compute_dti (10d7b044915ba73c5a6953ffa33ca490#0)
  ...
  ...1 superseded retry attempt(s) collapsed — a later attempt at the same call site is the decision (1 no_change): bureau_pull
model-diff: final output — recorded: {'decision': 'approved', 'status': 'recorded'}
model-diff:              — live:     {'decision': 'approved', 'status': 'recorded'}

The two counts are the same 13 sites counted twice, deliberately. The first is what a plain equality check on the output would have told you. The second is what the classifier concluded. On a real corpus those two numbers are far apart, and the gap is the entire product: an LLM reworded 200 responses and changed no decisions.

Call sites are identified by where in your code the call is made, not by the order it happened in. The #0 / #1 suffix is which time that same line ran.

Each line also prints a label and, in parentheses, a hash: identity_verify (732c5e8de713fa6b24c24d34d145e8d3#0). The hash is that call site's identity — stable across runs, and what the counts above are keyed by. The label beside it is for a human, not a lookup: a tool call is labelled from your tool metadata (identity_verify); an LLM call with no name is labelled with the opening of its own prompt, because that's the thing you can grep for in your source. Neither the hash nor the label records the source location itself — the file and line the call was made from isn't captured today, so grepping the label text is how you find the call site in your code, not a look-up Pramana does for you.

The nine verdicts

Every compared call site lands in exactly one of these. The counts are a strict partition — they sum to the number of sites compared, always, by construction — so a site is never counted under two headings.

NO_CHANGE
This site was reached, and what your code sent matched the recording. Does not mean the whole run was identical — only this one site.
COSMETIC
The model's text changed, but nothing downstream acted differently — no tool argument moved. A reworded answer, the same decision. Does not mean the difference is unimportant to a human reading the output; it means no instrumented decision changed because of it. If your users read that text, read it yourself.
BEHAVIOURAL
A real change in what the agent did — the arguments it passed to a tool or policy decision differ from the recording. This is the finding that matters. Does not mean the change is wrong. It means it is real, and yours to approve or reject.
CONSEQUENT
This site changed too, but downstream of an earlier finding on the same trace — the explanation for it is above it. Counted separately so ten knock-on effects of one root cause don't read as ten problems. Does not mean it's safe to ignore: a cascade can be much worse than its root. It means you have one thing to investigate, not ten.
SUPERSEDED
An earlier attempt at a call site that was retried, where a later attempt is the one that carried the decision. Retries are compared last-attempt-to-last-attempt, so the attempt count itself is not a behavioural signal. Does not mean a finding at all — a clean agent that retries anything produces these on every run. It has its own bucket precisely so it never inflates the finding count.
REORDERED
The same tool calls happened, with the same arguments, in a different order. Deliberately not counted as behavioural: a concurrent agent reorders its own work run to run, and calling that a regression would make the report cry wolf on every single trace. It is surfaced so you can see it — if order matters in your system, that judgment is yours, and the report will not make it for you.
HALTED
Sandboxed mode stopped here because the arguments your code is now passing no longer match the recording. See why it stops instead of guessing. Does not mean something broke. It means a tool call changed, which is a changed decision — the strongest finding the sandboxed mode can produce.
UNSTABLE
This site looked like a finding, but a control run against the original model produced its own divergence at the same place — so the model wanders here on its own and the difference is not attributable to your change. Does not mean the site is fine. It means this particular comparison cannot tell you anything about it, which is a different and more honest statement.
ERRORED
The call raised, where the recording returned a value. Does not mean a model regression by itself — it may be a flaky dependency, which is exactly what the control run is for, and it reads in the opposite direction from the others (below).

The corpus these examples come from produced no REORDERED, UNSTABLE or ERRORED sites, so there is no real output to show for those three. They are described above rather than illustrated, rather than shown with invented output.

Root and consequent, on one real trace

This is the trace the whole corpus was built around: a prompt edit that also quietly loosened a risk threshold.

model-diff: 13 call site(s) compared
model-diff:   without classification: 1 of 12 sites that ran changed (raw output ==)
model-diff:                           1 halted before it could run
model-diff:   with classification:    1 behavioural · 2 consequent · 0 cosmetic · 9 unchanged · 1 superseded
model-diff:   (1 of the 1 halted site(s) above collapsed behind an earlier root cause on this trace, and are counted under consequent above, not halted)
  ...
  [BEHAVIOURAL] risk_guardrail (2c66c0f22e99661e5d687433e53e5f6e#0)  threshold: 50 -> 45
  ...2 finding(s) collapsed behind the root cause above (1 no_change, 1 halted): "You are a loan underwriter. Approve or deny base…", issue_decision
  ...1 superseded retry attempt(s) collapsed — a later attempt at the same call site is the decision (1 no_change): bureau_pull
model-diff: final output — run halted before completion; no final output produced
model-diff:                (recorded output at the halted site: {'decision': 'approved', 'status': 'recorded'})

One root finding — threshold: 50 -> 45 — and two things collapsed behind it, including the halt on issue_decision. That is the point of separating root from consequent: the halt is real and serious, but it is not a second problem. There is one thing to fix here, and the report says which one.

Note the disclosure line in the counts. A halt that is a known consequence of an earlier finding is counted under consequent, not halted, so the halted count in the classified line can legitimately differ from the raw one above it. Whenever the two disagree, the report says so in that sentence rather than leaving you to notice.

The control run — what it is and when it fires

A finding raises an obvious question: did the change cause this, or does the model just do this sometimes? So when a run produces a root finding, Pramana automatically re-runs the same trace once more against the original model and compares that too:

model-diff: re-running as a control against the original model, to rule out sampling noise
model-diff: control run — same classifier, same policies as the main run above
model-diff: 1 halted finding(s) in the main run — 1 confirmed, 0 explained by the control (model noise, not a real change)

It fires only on a root BEHAVIOURAL, ERRORED or HALTED finding — never on a clean, cosmetic-only, reordered-only, or purely consequent run, because a control on those has nothing to invalidate. That gating is what keeps a corpus-wide comparison at roughly its original cost rather than double: only the traces that actually flagged pay for a second run.

The verdict reads in opposite directions depending on the finding, and this is the one thing on this page most worth getting right:

BEHAVIOURAL and HALTED
The control reproducing the divergence means the old model wandered here too → noise, reported as UNSTABLE. The control not reproducing it means the change is what caused it → confirmed.
ERRORED
Inverted. The control reproducing the failure means it is deterministic → a confirmed, real breakage. It is the control not reproducing it that makes the error flaky and unattributable to your change.

A missing control row is reported as inconclusive, never as "explained" — the absence of a comparison is not evidence that a finding is noise.

The batch aggregate

Run more than one trace and you get one summary on top of every trace's own unchanged report. This is the real tail of the 20-trace run:

model-diff: 20 of 20 run(s) classified
model-diff: 260 call site(s) compared
model-diff:   without classification: 4 of 257 sites that ran changed (raw output ==)
model-diff:                           3 halted before they could run
model-diff:   with classification:    2 behavioural · 4 consequent · 3 cosmetic · 230 unchanged · 1 halted · 20 superseded
model-diff:   (2 of the 3 halted site(s) above collapsed behind an earlier root cause on the same trace, and are counted under consequent above, not halted)
model-diff: 3 of 20 run(s) had a changed decision — 3 halted on a diverged tool call, 0 diverged without halting
model-diff:   root findings:
model-diff:     loan-17-5a6eef: risk_guardrail  threshold: 50 -> 45
model-diff:     loan-18-b3e695: risk_guardrail  threshold: 50 -> 45
model-diff:     loan-19-233199: issue_decision  sandboxed run halted here — live arguments diverged: decision: 'denied' -> 'approved'

Read it from the bottom. 3 of 20 runs had a changed decision, and the root findings are named — not just counted — because that last line is where most readers stop. Two traces had a risk threshold move from 50 to 45 under what was labelled a prompt-tone edit; one had its final decision flip on data that changed upstream.

The 20 superseded sites are the retried bureau call, once per trace. The 4 consequent are genuine knock-on effects on the two threshold traces. Those are different things and they are counted separately — if they shared a bucket, this corpus would report 24 downstream effects of 2 root causes, and 20 of them would be a retry that did nothing.

A trace that refused to classify at all is excluded from these counts and named separately, so the denominator always reconciles.

11Team & settings

Team lists everyone on your tenant. An admin adds a teammate by email and role — this hands back a one-time temporary password to share directly, since there's no email service yet to send an invite link — or removes someone, which signs them out everywhere immediately.

Settings has your account summary (tenant, role, plan), a change-password form, and API key management: create a key for a given role, see every key ever issued (including revoked ones, for the audit trail), and revoke one immediately.

12Common questions

The four asked most often when someone is deciding whether this fits their system.

What is Pramana?

A flight recorder for AI agents. A three-line Python SDK records every non-deterministic decision your agent makes — each model call, tool call, clock read, random draw and agent-to-agent message — so any past run can be replayed exactly, and any range of it exported as a signed evidence record. It works with the OpenAI and Anthropic clients you already have.

The use it's built around is change approval: point it at conversations you already recorded, change the model or the prompt, and see which decisions changed — not which sentences.

Can I test a model or prompt change against my past production runs?

Yes — that's the primary workflow. Pass as many recorded traces as you like; the same command runs against each, and one summary prints on top of each trace's own report:

pramana model-diff loan-01 loan-02 loan-03 --change-id=gpt-5-migration -- python your_agent.py

Two or more traces switches it to sandboxed mode by default. Your tools are never called; a recorded tool's output is replayed back to your agent for as long as the arguments your code is now sending still match the recording, and that trace stops the moment they diverge. So a whole corpus re-runs with zero side effects. Sandboxing is decided by the number of traces, not by a flag you have to remember — see the mode spectrum.

Model calls are still real in both modes, so a run still costs model-provider money. The report sorts every call site into one of nine outcomes and separates changes that altered a decision from changes that altered only wording — see Reading a model-diff report.

Can I import my existing conversation logs?

No. Replay needs the execution trace, not the transcript.

Pramana matches each recorded call to where in your code it was made, which is what lets it serve the right recorded answer back at the right point when your code runs again. A transcript of prompts and responses has no call sites, so there is nothing to replay against — and no way to know which line of your agent produced which message.

The practical consequence, stated plainly: your corpus starts the day you instrument, not retroactively. If evaluating a change against history matters to you, that's the argument for instrumenting sooner rather than later.

Does it tell me whether a change is bad?

No — and it's deliberate about which half it does answer.

It tells you deterministically which changes were decisions and which were only wording: different tool arguments, a different policy outcome, a call that stopped happening. It reaches that without reading a word of prose — "cosmetic" is decided by elimination, not by resemblance. There is no similarity score and no threshold anywhere in the product. And when a run produces a root finding, that trace is automatically re-run against the original model to rule out ordinary model variance.

But whether a changed decision is a bad decision is still a human call. Semantic judgment of changed prose is deliberately not built.

13Glossary

Trace
One recorded run of one agent, or several working together.
Event
One recorded decision within a trace: an LLM call, a tool call, a message send/receive, or a state commit.
Call site
The exact place in your code an LLM/tool call happens — used to line up a replay's live call with the one originally recorded there.
Divergence
A replay's live input didn't match what was recorded at a given call site. Not an error — the entire point of replay is surfacing exactly this. If the change was intended, accept it with pramana rebaseline; if it wasn't, you've just caught a bug before a customer did.
Evidence bundle
A signed, self-contained export of part of a trace. Checking one requires the platform's public key, which is not yet published at a stable URL — today you obtain it from us, which is why this is not described as independent verification. The check itself runs offline and needs no Pramana account.
Baseline
A movable name for a trace — checkout-flow rather than a UUID. Anywhere a command takes a trace id it takes a baseline name, so a CI pipeline never has to be edited when you accept an intentional behaviour change. pramana rebaseline records a fresh trace and moves the name to it.
Agent
A logical participant in a trace (e.g. "planner", "worker") — a multi-agent trace has more than one.