# Pramana > Pramana is a flight recorder for AI agents. It records every non-deterministic decision an agent > makes, replays any past run exactly, and exports a signed record of what happened. Built by > reliai.in. ## What it is, in one sentence Pramana records every non-deterministic decision your AI agent makes — each model call, tool call, clock read, random draw and agent-to-agent message — so any past run can be replayed exactly, and any range of it exported as a signed evidence record. Three verbs, in this order: **record → replay → prove.** ## What it is for: approving a change against real history Point it at the conversations you already recorded. Change the model or the prompt, and see which decisions changed — not which sentences. This is the use case Pramana is built around. Upgrading a model, editing a system prompt, or swapping a provider are all changes nobody can currently approve with evidence: the only options are a handful of eyeballed spot-checks or a benchmark that has nothing to do with your traffic. Pramana replays your own recorded production runs against the change and reports what actually moved. ## What a comparison reports - Every compared call site lands in exactly one of **nine outcomes**, and the counts are a strict partition — they always sum to the number of call sites compared. - The outcomes: unchanged, cosmetic, behavioural, consequent, superseded, reordered, halted, unstable, errored. - **"Cosmetic" is decided by elimination, never by reading the prose.** If every tool call is identical — same tool, same arguments, same order — then nothing the agent *did* changed, however differently it worded the answer. - **There is no similarity score and no threshold anywhere in the product.** No embeddings, no LLM-as-judge, no fuzzy matching. A decision either changed or it did not. - A root finding is reported separately from its knock-on effects, so one cause that produced ten downstream differences reads as one thing to investigate, not ten. - Retried calls are compared last-attempt-to-last-attempt; the retry count is never itself treated as a behaviour change. ## Running against a whole corpus safely Pass two or more recorded traces to `pramana model-diff` and it switches to **sandboxed mode by default**: - The tool calls you have wrapped are never executed. - A recorded tool's output is replayed back to your agent for as long as the arguments your code is now sending still match the recording. - The moment those arguments diverge, that trace stops — and the divergence is the finding, reported with the exact argument that changed. - So an entire corpus of historical traces can be re-run without firing your wrapped tools. - **Pramana only intercepts tool calls you have wrapped.** An unwrapped call is invisible to it and executes normally, once per trace. Before a batch runs, Pramana prints how many recorded tool call sites were wrapped and what that means multiplied across the batch. **Sandboxing is decided by the number of traces, not by a flag.** One trace is always live-tools; two or more is sandboxed unless explicitly opted out. Model calls are still real in both modes, so a run still costs model-provider money. ## The control run When a comparison produces a root finding, Pramana automatically re-runs that same trace once more against the **original** model and compares again. This separates a real change from ordinary model variance: if the original model produces the same divergence on its own, the finding is reported as noise rather than as a consequence of your change. Only flagged traces pay for the second run. ## What it cannot do - **You can only replay what Pramana recorded.** An existing database of conversation logs cannot be imported and replayed. Replay is keyed to *where in your code* each call was made, and a transcript has no call sites. Your corpus starts the day you instrument, not retroactively. - **It does not tell you whether a change is good or bad.** It tells you which changes were decisions and which were wording. Judging a changed decision is a human call. - **Python only.** No JavaScript or TypeScript SDK. - **The OpenAI and Anthropic clients are supported.** Azure OpenAI and `AnthropicBedrock` are routed to those adapters and covered by tests; neither has been exercised against live cloud credentials. Raw `boto3` is not supported. - **LangChain is tested and works.** LangGraph, CrewAI and LlamaIndex are untested — they may work, but that is not a claim being made. - **Async: non-streaming is supported; async streaming is not** and fails loudly rather than silently. Sync streaming is supported for OpenAI across record, replay and comparison; Anthropic streaming is not supported at all. - **A streamed tool call cannot run inside sandboxed mode** and refuses rather than executing. - **Model-diff is a report, not a build gate.** It deliberately does not fail a build. `pramana replay` is the command intended for CI, and it does exit non-zero on any divergence. - **The evidence public key is not yet published at a stable URL.** The verifier runs offline and needs no account, but today a third party still has to obtain the key from us — which is not the same as independent verification, and is not claimed as such. *(This limitation is removed from this file the moment the key is live at its well-known URL.)* - **No SOC 2, no SLA, no on-premise deployment, no paying customers yet.** On-premise, an SLA and SOC 2 are real commitments that begin when a named prospect requires them, not shipped features. ## Record, replay, prove - **Record** — a three-line Python SDK integration captures each decision as an event, matched by where in your code the call is made rather than by the order it happened in. Each step is chained into the record as it arrives. - **Replay** — re-runs your recorded trace against your own code, serving each recorded answer back at its own call site instead of making a live call. **What is guaranteed depends on the mode, and one mode's guarantee never carries over to the other.** In a plain replay, nothing leaves the process at all: outbound network access is blocked at the process level, below whatever HTTP client you use. In a sandboxed batch — two or more traces — the tool calls you have **wrapped** are served from the recording and never executed, the model **is** called for real because the new model is the thing you are testing, and a tool call you did **not** wrap is invisible to Pramana and will execute normally, once per trace. If your code now sends something different at a call site, that is reported as a divergence at that exact call site. - **Prove** — any range of a trace exports as a signed evidence bundle, signed at the moment of export. Once exported, nothing in it can be altered without detection. The bundle is checked with `pramana-verify`, a separate tool that runs offline and needs no Pramana account. A change-approval comparison exports as its own distinct artifact type, which attests that a comparison ran and what it concluded — never that the traces are production behaviour. ## Commands - `pramana replay [...] -- ` — no live calls; exits non-zero on divergence, so it works as a CI regression gate. - `pramana model-diff [...] [--change-id=] -- ` — one trace runs live-tools behind a confirmation; two or more runs sandboxed by default. - `pramana baseline set ` — give a recording a movable name so CI references the name, never a hardcoded trace id. - `pramana rebaseline -- ` — accept an intended change by recording a fresh trace and moving the name to it. ## Key pages - [Product & documentation](https://www.reliai.in/docs/) — quickstart, configuration, reading a comparison report, CI gate, evidence and verification, plans, glossary. - [`pramana-sdk` on PyPI](https://pypi.org/project/pramana-sdk/) — the recording SDK and CLI. - [`pramana-verify` on PyPI](https://pypi.org/project/pramana-verify/) — the standalone verifier. - [Sign up / sign in](https://www.reliai.in/) — free tier, no card required. ## Pricing Free ($0, 50,000 events/month, 30-day retention, 3 seats), Pro ($99 for 30 days, 2,000,000 events/month, 365-day retention, 10 seats), Enterprise (custom volume and retention, contact sales). Recording, replay, model-diff and the trace, timeline, conversation and multi-agent views all work on every plan, including Free. **What Free excludes is evidence export** — creating and downloading signed evidence bundles requires Pro. **Payment is INR-only today.** Checkout runs through an Indian payment provider that does not accept cross-border cards, so a US or EU customer cannot currently buy Pro even though the price is listed in dollars. Non-Indian customers use the free tier, or contact us directly.