Risk definition → model development

Risk is defined in words. Models are built on labelled data. Scenario Studio is the translation.

An expert describes a concern in plain English — or starts from a policy, advisory or enforcement finding. Studio turns it into reviewed scenarios with explicit expected responses, and the connected data to build and test against, in your schema. Experts direct the design. Engineers and modellers build the control. Every step leaves the evidence your model risk reviewers will ask for.

  1. 01

    Define your scope of controls

    Live in beta
    Decide which controls are in scope — and where their definition comes from.
    • Start from regulatory obligations: the advisories, master directions and rules you are held to
    • Or start from your own internal policy, procedures, or an expert's plain-English hypothesis
    • Scope is explicit — which typologies, channels, products and jurisdictions are in, and which are deliberately out
    • A source document informs the scenario; you review the interpretation and adopt it as yours
    Produces

    A scoped control definition, each control traced to the obligation, policy or hypothesis behind it

    Before the next step

    Your SMEs and compliance agree the scope — including what is deliberately out of it

  2. 02

    Build ground truth from that definition

    Live in betaTry it
    Your scoped definition becomes a reviewed specification — and then the connected data to test with.
    • Normal business history is built first, so the behaviour under examination can be told apart from it
    • Each control becomes a structured typology: what a pattern must, may and must not include
    • You choose the count and the mix — suspicious cases, and legitimate activity that looks like them
    • Connected evidence, not just rows: transactions, trades, communications, KYC documentation and risk profiles, agreeing on identities, dates, amounts and circumstances
    • A relationship graph and money-flow view make it reviewable by people who do not read SQL
    • One collection of scenarios is a dataset: TP and FP together, across every typology in scope
    • Development and evaluation cases are kept separate from the start, and the holdout is sealed here
    Produces

    A reviewed scenario pack — versioned spec, adopted expected responses, case counts, source references and the generated records

    Before the next step

    Live in beta: this is the step you can run today, before committing to anything downstream

    Try it

    Build your own scenario.

    Pick a typology. See the specification it becomes, and the scenarios generated from it — the ones that must alert, and the near-misses that must not.

    Live in betaIllustrative example
    The specificationUS · FinCEN

    Structuring — aggregated ACH below reporting threshold

    FinCEN advisory, red-flag indicators ¶14–17

    Must include
    • Three or more credits to one account within 7 days
    • Each credit below the applicable reporting threshold
    • Aggregate across the window at or above the threshold
    May include
    • Onward transfer within 48h of the final credit
    • Account opened within 90 days of first credit
    Must not include
    • A single credit at or above threshold — a reporting event, not structuring
    • Payroll or recurring credits from a known employer
    What it generates
    True positivemust alert

    Four ACH credits over five days from three unrelated originators into a nine-week-old account, consolidated out by wire on day six.

    Ground truth: Satisfies every must-clause plus two may-clauses.

    Near missmust not alert

    Four ACH credits of the same size and timing into the same account profile — all four from one payroll bureau against a registered employer relationship.

    Ground truth: Trips a must-not clause. An alert here is a tuning defect.

    This is a preview of the artifact. In Scenario Studio you edit every clause, set the thresholds and windows, tune the near-miss margin, and generate the full pack inside a synthetic bank.

    Build your own in Scenario Studio
  3. 03

    Measure your baseline

    Live in beta
    Run that data against the controls you already have — once, to fix the number everything later is measured against.
    • Runs against your existing controls exactly as they stand — rule-based, ML, or a mix of both
    • Two ways to run it: connect your model or scenario engine via API, or export the dataset into a dev environment and run your own test process
    • Requires a one-time schema configuration; after that, new scenarios reuse it
    • Then compare observed against expected — what fired, what did not, and what was never covered at all
    • Recall becomes computable: you can count what your controls miss, not only what they alerted on
    • Two different baselines, easily confused: the behavioural baseline is a customer's normal activity; this is the control-performance benchmark, a dated result for a stated test set
    Produces

    A dated control-performance benchmark: what your controls do today, on an agreed test set

    Before the next step

    Owner and validator agree the benchmark, its scope and its exclusions, as the reference for what follows

  4. 04

    Implement the control with eval-driven development

    Pilot
    Build the control against the dataset, not away from it — the expectation comes first, the implementation follows.
    • Compose rules, machine learning and agents into one workflow, each with an explicit job
    • Build on VERDICT, our eval-driven control engine — or bring your own model and evaluate it on exactly the same terms
    • The dataset from step 02 is the eval: TP and FP across every typology in scope
    • Tuned only on the development set — the sealed holdout stays untouched until it is tested
    • Conceptual soundness documented as you go: why this design, on what assumptions, and what it cannot do
    • Prompt, model, tools, retrieval sources and thresholds all versioned — for an agent, a prompt change is a model change
    Produces

    A control, versioned, with its design rationale and stated limitations — the developmental evidence MRM asks for, as a by-product

    Before the next step

    Build and validation stay separate — whoever builds it does not sign it off

    Steps 04 and 05 are a loop, not a hand-off. Implement, test, read the result, implement again — each pass hill-climbs against the same development set until the control clears the target you set.

  5. 05

    Test the control

    Pilot
    Same dataset, same scale as the benchmark. You know what you are getting before it reaches production.
    • Precision and recall, measured against the benchmark rather than asserted
    • Threshold and sensitivity analysis before you move anything
    • Bias across customer segments and geographies
    • Cost per alert, and per detected true positive
    • Determinism: same input, same decision, replayed
    • Reasoning stability: same decision, same stated reason
    • Evidence grounding: retrieved facts, not invented ones
    • Robustness: adversarial input, prompt injection, tool failure
    • Override rate: whether human review is real or nominal
    • A missed target sends you back to step 04 — that loop is the method, not a setback
    • Repeated feedback from a holdout erodes its independence: refresh it rather than treating it as permanently clean
    Produces

    A test report against the benchmark — the go / no-go artifact, on thresholds you set

    Before the next step

    Independent validation and approver sign-off. Polarisk measures; you set the pass mark.

  6. 06

    Monitor the control

    Roadmap
    The dataset that proved the control works becomes the regression suite that keeps it honest.
    • Drift detection by re-running the sealed dataset on a cadence
    • Guardrails and a kill switch with a named authority — who pulls it, on what trigger, and what it falls back to
    • Continuous quality assurance on live decisions, not only at release
    • Change triggers send you back to step 05: a prompt edit, a model version bump, a new channel, a threshold move — or your provider updating the foundation model underneath you
    • Findings flow into your issue tracker, and close only when the scenario passes on re-run
    • Honest limit: replaying a fixed suite catches regression on known cases. It does not tell you your customer population has changed — new risks need new scenarios
    Produces

    Monitoring reports, an issue register, and a replayable decision log

    Before the next step

    Re-validation on a defined cadence, and on every change trigger

The handoff

One reviewed scenario. Three people who need different things from it.

The translation usually fails as a handoff problem, not a modelling problem: an expert's judgment reaches the modeller as a spreadsheet of labels with no stated reasoning, and the disagreement surfaces months later. One artifact, inspectable three ways, is what stops that.

Risk owner / SME

Does this fit our business, and is this the response we would expect?

  • The case in plain English, with the money flow drawn
  • The evidence behind each expected response
  • What the scenario deliberately leaves out
Adopts it

Engineer / modeller

Can I build and test against this?

  • Connected records in your own schema, not a sample CSV
  • An explicit expected response per case, and why
  • Development and evaluation sets kept apart
Builds the control

Independent reviewer

Are the assumptions justified, and are the gaps explicit?

  • The spec version and the source paragraph behind it
  • Case counts, variation settings, declared coverage
  • What was adopted, by whom, and when
Challenges it

The artifact does not change between them — only the question being asked of it. That is what turns "our experts disagreed" into "we disagree about clause three": a disagreement you can version, resolve and evidence.

Controlled variation

The cases you cannot pull from your history.

Historical data holds what happened and was caught. It cannot hold the near-miss nobody logged, it cannot hold one condition changed while the rest stays fixed — and it cannot hold what your controls never saw, because an unmonitored channel produces no alerts, no cases and no labels.

Illustrative example
Suspiciousmust alert

A quiet consulting business starts forwarding unrelated incoming payments.

Tests a change that is inconsistent with the reviewed business profile.

Legitimate lookalikemust not alert

An established payroll provider receives client funds and pays employees, exactly as documented.

Tests whether the available context supports an ordinary business explanation.

Controlled variationyou choose

The same case with the timing, the counterparties, the prior history or the supporting documents changed.

Vary one factor at a time, or combine them deliberately — and pick how many of each.

Neither the business category nor a benign backstory decides the outcome. The expert states which observable facts justify the response — and whether the control can even see those facts at the moment it decides.

Governance at the core

Four threads that run through every step.

Not a review at the end. These are what connect one step to the next — each stage consumes the artifact the last one signed, so the model file writes itself as you build.

Accountability

A named owner, an independent validator and an approver at every gate. The bank owns the adopted interpretation — accountability does not transfer to a vendor.

Traceability

Every scenario, decision and finding traces back to the obligation paragraph or hypothesis it came from. A gap arrives as a reference, not as a number.

Change control

Anything that can change behaviour is versioned, and defined triggers force re-validation. For an agent, a prompt change is a model change.

Evidence

Each stage emits a dated, immutable artifact for the model file. Nothing rests on a claim that cannot be reproduced on demand.

Before you ask

The questions we would ask in your seat.

Stating the limits is not a hedge — it is what makes the rest of the claims worth anything. Every answer below is one we would rather you heard from us than found later.

  • We already have years of historical data. Why generate more?

    Volume is not the constraint — suitable examples with reliable labels are. Historical cases can be scarce for the exact behaviour you are testing, inconsistently tagged, and impossible to re-cut when the requirement changes. History also cannot contain what your controls never caught: an unmonitored channel produces no alerts and therefore no labels. Your history still matters, and can inform what normal looks like.

  • Why should I trust synthetic data to look like my bank?

    Schema matching proves compatibility, not realism, and we will not claim otherwise. Today the honest scope is: the reviewed scenario is realistic because your experts reviewed it, not because the background was fitted to your population. Calibrating that background from approved bank profiles or a scoped sample is a proposed extension we would scope and validate with you — not a live capability.

  • How can a synthetic case be “ground truth”?

    It means the designed case facts plus the expected response your experts reviewed and adopted, for one stated task. It is not a legal determination, and not an automatically correct reading of an advisory. A criminal backstory does not justify an alert if the control cannot observe the relevant evidence — which is why we state the expected response precisely, per control.

  • Could a control just learn to pass your scenarios?

    Yes, if development and evaluation are not separated — so they are, from step 02. Keep held-out cases outside the tuning loop, check generated records for shortcuts and label leakage, and test against real historical cases too. And after a holdout has fed back into development enough times, refresh it rather than pretending it is still independent.

  • Does a good score mean our AML programme is effective?

    No. It measures performance on a stated test population and task, and should always be reported with the scenario mix, sample size, scope, control version and exclusions. Recall here is the share of expected-positive cases detected — not the share of real-world laundering. A deliberately enriched test set is not an estimate of production alert volume or cost.

  • Does this replace model validation, or get us through an exam?

    Neither, and anyone promising that is selling you something. It produces evidence that supports your validation, independent testing and exam preparation — developmental evidence, outcomes analysis, a dated benchmark. Conceptual soundness, data quality, governance and operating effectiveness remain your validators' remit, and no regulator has endorsed this.

  • Do we need to adopt AI controls to get value?

    No. The first question can be about a rule you already run. Scenario-driven development is a method for specifying and testing expected behaviour; whether the control you end up building is a rule, a model or an agent is a separate decision — and the same evaluation basis lets you compare them fairly.

  • Does this need production access?

    Not to start. Scenario authoring needs none. Running against your controls needs a one-time schema configuration and either a supported endpoint or a dev environment to load into. No production data, no write access. Covering your production pipeline means testing that pipeline — schema compatibility does not prove it.

Start where you are.

A useful first engagement is one control question: suspicious cases, legitimate comparisons, controlled variations, reviewed and run. Most institutions begin at step 02 or 03 — that needs no integration and no production access. Success is a supported decision, and a confirmed pass is a real result.