For AML & financial crime compliance

Add AI to your controls without losing the ability to prove they work.

Scenario Studio manufactures the ground truth your monitoring has never had — the patterns that must alert, and the near-misses that must not. Effectiveness becomes a measured number on one scale, whether the control behind it is a rule, a model, or an agent.

Every typology is decomposed from a published source — FinCEN advisories, FATF typologies, RBI master directions — into what a pattern must include, may include and must not include. You review the decomposition, edit it, and adopt it as yours.

control_assurance_record · structuring_ach_004control: agentic triage · v3 · run 0c41Illustrative example
Recallfalse negatives foundBelow target91.7% · target ≥ 95%
Precisionalerts that were realMeets target78.6% · target ≥ 70%
Decision determinismsame verdict, replayed 5×Meets target96.0% · target ≥ 95%
Reasoning stabilitysame justification, replayed 5×Finding41.0% · target ≥ 80%
Clause attributionrationale cites the source ¶Meets target88.2% · target ≥ 85%
Two findings raised, neither blocks deployment. The verdict is repeatable; the justification for it is not yet — that's a governance finding, logged before it became an audit finding.
Point A to point B

You already know what a legacy control can tell you. This is what changes.

Nothing here asks you to rip out what runs today. A composite control stack keeps the rules that work and adds measured components around them — each one held to the same evidence standard, so the AI parts of the stack are never the ones you can't explain.

Today, attested

Coverage is attested. Someone maps transaction codes to typologies on a spreadsheet and signs it.

With Scenario Studio, measured

Coverage is tested. A named scenario either fires or it doesn't, per channel, before you sign anything.

Today, attested

Effectiveness is inferred from alert volume and the alert-to-SAR ratio — a measure of activity, not of what was missed.

With Scenario Studio, measured

Effectiveness is recall against ground truth — the one number a bank normally cannot compute, because you can't count what you never saw.

Today, attested

A new AI component stalls at model risk because there's nothing to evidence it against and no accepted standard to cite.

With Scenario Studio, measured

A new AI component ships with the same recall / precision / determinism / traceability record as the rule sitting next to it.

Today, attested

Independent testing starts from a blank scope and eight to twelve weeks of sampling.

With Scenario Studio, measured

Independent testing starts from measured coverage already in hand — scoped at what's actually weak, and reviewed with a record instead of a memo.

How institutions get there

A path with a floor at every step, not a leap.

We ship a baseline, then work with your team to make it yours: your schema, your adopted typologies, your thresholds. Each stage below is a real capability, not a slide — and we say plainly which ones are live today, which are running in pilot, and which are the road ahead.

STAGE 00

Baseline library

Live in beta
You do

Pick typologies from the library, or bring an advisory and we decompose it with you.

You get

A reviewable must / may / must-not specification for each typology, traceable to its source paragraph.

Needs

Nothing to integrate. Runs entirely on synthetic data in a segregated environment.

STAGE 01

Evidenced change

Live in beta
You do

Map your data schema once, generate a scenario pack in it, run it against your own monitoring.

You get

A dated coverage record for a specific product, channel or config change — before you sign off on it.

Needs

One schema mapping, done once. Your environment, your data — nothing leaves it.

STAGE 02

AI in the loop, measured

Pilot
You do

Point an agentic or model-based control at the same scenario packs used to test your rules.

You get

The composite record: detection, decision quality, explainability and cost, on one scale as the rules beside it.

Needs

A model endpoint we can invoke, and your model risk team's sign-off on the threshold — we measure, you set the pass mark.

STAGE 03

Continuous assurance

Roadmap
You do

Nothing manual — read-only connectivity to a non-production instance, reviewed once.

You get

Coverage that re-runs on a schedule, so a threshold change shows its effect before it reaches production, every time.

Needs

Your security review. This is the version that satisfies validation before and after deployment — it belongs past a pilot, not before one.

The standard, not a scorecard

One measurement, applied identically to a rule and to an agent.

Polarisk measures. The bank sets the pass mark — because a vendor that also sets the threshold has graded its own homework, and a reviewer will say so.

Detection

  • Recall against ground truth
  • Precision
  • False-positive suppression
  • F1

Decision quality

  • Evidence grounding
  • Clause attribution
  • Decision determinism (replayed)
  • Reasoning stability

Explainability & audit

  • Rationale present
  • Traceable to source paragraph
  • Replay reproducibility
  • Analyst override rate

Operating cost

  • Cost per alert
  • Median time to decision
  • Alerts per 1,000 transactions
  • Cost per detected true positive
MeasureThe question it answersWhy a reviewer accepts it
RecallWhat does the control miss?The Wolfsberg Group names false-negative assessment as an effectiveness measure and cautions against alert-to-SAR ratios alone. Ground truth is what makes it computable at all.
Scenario coverageWhat was never tested at all?FFIEC independent testing requires that the systems supporting BSA/AML be shown complete and accurate. This is that requirement, quantified per channel.
Decision determinism & reasoning stabilitySame input, same decision — and the same reason — on replay?Published work on LLM-judge auditing finds verdict agreement above 99% coexisting with reasoning stability far lower on some tasks. Reporting only the verdict would flatter the model; both numbers is what makes this an assurance record.
Clause attributionDoes the rationale cite the obligation it came from?Closes the chain from obligation to control: a finding arrives as a paragraph reference, not a bare number.

Illustrative example Every figure on this page is illustrative example data, labelled as such — a real coverage number belongs to your run, not ours.

Anatomy of one typology

An advisory is prose. A control is code. This is the thing in between.

Expanded here, static, no clicking required — a reader who studies this one block understands the whole product.

TYP-STR-004 · v1.2 · adopted

Structuring — aggregated ACH below reporting threshold

Source: FinCEN advisory, red-flag indicators ¶14–17
Must include
  • Three or more credits to one account within 7 days
  • Each credit below the applicable reporting threshold
  • Aggregate across the window ≥ threshold
  • Two or more distinct originating counterparties
May include
  • Onward transfer within 48h of the final credit
  • Account opened within 90 days of first credit
Must not include
  • A single credit at or above threshold — that's a reporting event, not structuring
  • Payroll or recurring credits from a known employer
  • Credits from an account under the same beneficial ownership
True positivemust alert

Four ACH credits over five days from three unrelated originators into a nine-week-old account, consolidated out by wire on day six.

Ground truth:Satisfies all four must-clauses plus one may-clause. Expected disposition: alert, escalate.
False positivemust not alert

Four ACH credits of similar size and timing into the same account profile — but all four originate from one payroll bureau against a registered employer relationship.

Ground truth:Trips a must-not clause deliberately. An alert here is a tuning defect, and this scenario is how you find it before an examiner does.

Deliberately generating what must not alert is the part incumbent tuning tools structurally can't do — they reason from your own alert history and can't test uncovered space. The interpretation is Polarisk's proposal; the adopted version is yours, versioned and traceable, because accountability for a third-party model does not transfer to the vendor.

Try it

Build your own scenario.

Pick a typology. See the specification it becomes, and the scenarios generated from it — the ones that must alert, and the near-misses that must not.

Live in betaIllustrative example
The specificationUS · FinCEN

Structuring — aggregated ACH below reporting threshold

FinCEN advisory, red-flag indicators ¶14–17

Must include
  • Three or more credits to one account within 7 days
  • Each credit below the applicable reporting threshold
  • Aggregate across the window at or above the threshold
May include
  • Onward transfer within 48h of the final credit
  • Account opened within 90 days of first credit
Must not include
  • A single credit at or above threshold — a reporting event, not structuring
  • Payroll or recurring credits from a known employer
What it generates
True positivemust alert

Four ACH credits over five days from three unrelated originators into a nine-week-old account, consolidated out by wire on day six.

Ground truth: Satisfies every must-clause plus two may-clauses.

Near missmust not alert

Four ACH credits of the same size and timing into the same account profile — all four from one payroll bureau against a registered employer relationship.

Ground truth: Trips a must-not clause. An alert here is a tuning defect.

This is a preview of the artifact. In Scenario Studio you edit every clause, set the thresholds and windows, tune the near-miss margin, and generate the full pack inside a synthetic bank.

Build your own in Scenario Studio
What ships, and what we build with you

The product you see is the baseline. The fit is the engagement.

We don't hand over a fixed tool and leave. Scenario Studio ships as a working baseline — a typology library, a specification format, a measurement standard — and the first weeks of an engagement are spent making it match your institution, not the other way round.

Ships as baseline

  • A library of typologies already decomposed from published advisories
  • A specification format — must / may / must-not — that any new typology follows
  • A synthetic bank environment, isolated from production, to generate test data in
  • The measurement standard: the same metric families on every control, rule or agent

We fit to you

  • Your transaction, account and customer schema, mapped once
  • Your adopted interpretation of each typology, reviewed and versioned by your team
  • The thresholds and pass marks your model risk function sets, not ours
  • The report renderings your board, auditors and issue tracker already expect

What this is

  • Internal evidence that strengthens your own independent testing function.
  • A way to test the negative space — behaviour your production data never contained because the channel was never monitored.
  • A reviewable interpretation of published typologies, versioned, traceable to source, adopted by you.
  • Run entirely on synthetic data, in a segregated environment, with no production access.

What this isn't

  • Not a substitute for model validation or for your independent testing function.
  • Not endorsed by any regulator — no supervisor has yet accepted synthetic data for AML model validation.
  • Not a source of novel typologies — generation reproduces what the adopted specification describes.
  • Not connected to your production monitoring system in beta.
How a first engagement runs

Scoped so the finding is never a problem you own alone.

A product that finds control gaps creates a real question for its buyer at the exact moment it works: who else can see this, and what happens next. Three things answer that up front, by design rather than as talking points.