An expert describes a concern in plain English — or starts from a policy, advisory or enforcement finding. Studio turns it into reviewed scenarios with explicit expected responses, and the connected data to build and test against, in your schema. Experts direct the design. Engineers and modellers build the control. Every step leaves the evidence your model risk reviewers will ask for.
A scoped control definition, each control traced to the obligation, policy or hypothesis behind it
Your SMEs and compliance agree the scope — including what is deliberately out of it
A reviewed scenario pack — versioned spec, adopted expected responses, case counts, source references and the generated records
Live in beta: this is the step you can run today, before committing to anything downstream
Pick a typology. See the specification it becomes, and the scenarios generated from it — the ones that must alert, and the near-misses that must not.
FinCEN advisory, red-flag indicators ¶14–17
Four ACH credits over five days from three unrelated originators into a nine-week-old account, consolidated out by wire on day six.
Ground truth: Satisfies every must-clause plus two may-clauses.
Four ACH credits of the same size and timing into the same account profile — all four from one payroll bureau against a registered employer relationship.
Ground truth: Trips a must-not clause. An alert here is a tuning defect.
A dated control-performance benchmark: what your controls do today, on an agreed test set
Owner and validator agree the benchmark, its scope and its exclusions, as the reference for what follows
A control, versioned, with its design rationale and stated limitations — the developmental evidence MRM asks for, as a by-product
Build and validation stay separate — whoever builds it does not sign it off
Steps 04 and 05 are a loop, not a hand-off. Implement, test, read the result, implement again — each pass hill-climbs against the same development set until the control clears the target you set.
A test report against the benchmark — the go / no-go artifact, on thresholds you set
Independent validation and approver sign-off. Polarisk measures; you set the pass mark.
Monitoring reports, an issue register, and a replayable decision log
Re-validation on a defined cadence, and on every change trigger
The translation usually fails as a handoff problem, not a modelling problem: an expert's judgment reaches the modeller as a spreadsheet of labels with no stated reasoning, and the disagreement surfaces months later. One artifact, inspectable three ways, is what stops that.
Does this fit our business, and is this the response we would expect?
Can I build and test against this?
Are the assumptions justified, and are the gaps explicit?
The artifact does not change between them — only the question being asked of it. That is what turns "our experts disagreed" into "we disagree about clause three": a disagreement you can version, resolve and evidence.
Historical data holds what happened and was caught. It cannot hold the near-miss nobody logged, it cannot hold one condition changed while the rest stays fixed — and it cannot hold what your controls never saw, because an unmonitored channel produces no alerts, no cases and no labels.
A quiet consulting business starts forwarding unrelated incoming payments.
Tests a change that is inconsistent with the reviewed business profile.
An established payroll provider receives client funds and pays employees, exactly as documented.
Tests whether the available context supports an ordinary business explanation.
The same case with the timing, the counterparties, the prior history or the supporting documents changed.
Vary one factor at a time, or combine them deliberately — and pick how many of each.
Neither the business category nor a benign backstory decides the outcome. The expert states which observable facts justify the response — and whether the control can even see those facts at the moment it decides.
Not a review at the end. These are what connect one step to the next — each stage consumes the artifact the last one signed, so the model file writes itself as you build.
A named owner, an independent validator and an approver at every gate. The bank owns the adopted interpretation — accountability does not transfer to a vendor.
Every scenario, decision and finding traces back to the obligation paragraph or hypothesis it came from. A gap arrives as a reference, not as a number.
Anything that can change behaviour is versioned, and defined triggers force re-validation. For an agent, a prompt change is a model change.
Each stage emits a dated, immutable artifact for the model file. Nothing rests on a claim that cannot be reproduced on demand.
Stating the limits is not a hedge — it is what makes the rest of the claims worth anything. Every answer below is one we would rather you heard from us than found later.
Volume is not the constraint — suitable examples with reliable labels are. Historical cases can be scarce for the exact behaviour you are testing, inconsistently tagged, and impossible to re-cut when the requirement changes. History also cannot contain what your controls never caught: an unmonitored channel produces no alerts and therefore no labels. Your history still matters, and can inform what normal looks like.
Schema matching proves compatibility, not realism, and we will not claim otherwise. Today the honest scope is: the reviewed scenario is realistic because your experts reviewed it, not because the background was fitted to your population. Calibrating that background from approved bank profiles or a scoped sample is a proposed extension we would scope and validate with you — not a live capability.
It means the designed case facts plus the expected response your experts reviewed and adopted, for one stated task. It is not a legal determination, and not an automatically correct reading of an advisory. A criminal backstory does not justify an alert if the control cannot observe the relevant evidence — which is why we state the expected response precisely, per control.
Yes, if development and evaluation are not separated — so they are, from step 02. Keep held-out cases outside the tuning loop, check generated records for shortcuts and label leakage, and test against real historical cases too. And after a holdout has fed back into development enough times, refresh it rather than pretending it is still independent.
No. It measures performance on a stated test population and task, and should always be reported with the scenario mix, sample size, scope, control version and exclusions. Recall here is the share of expected-positive cases detected — not the share of real-world laundering. A deliberately enriched test set is not an estimate of production alert volume or cost.
Neither, and anyone promising that is selling you something. It produces evidence that supports your validation, independent testing and exam preparation — developmental evidence, outcomes analysis, a dated benchmark. Conceptual soundness, data quality, governance and operating effectiveness remain your validators' remit, and no regulator has endorsed this.
No. The first question can be about a rule you already run. Scenario-driven development is a method for specifying and testing expected behaviour; whether the control you end up building is a rule, a model or an agent is a separate decision — and the same evaluation basis lets you compare them fairly.
Not to start. Scenario authoring needs none. Running against your controls needs a one-time schema configuration and either a supported endpoint or a dev environment to load into. No production data, no write access. Covering your production pipeline means testing that pipeline — schema compatibility does not prove it.
A useful first engagement is one control question: suspicious cases, legitimate comparisons, controlled variations, reviewed and run. Most institutions begin at step 02 or 03 — that needs no integration and no production access. Success is a supported decision, and a confirmed pass is a real result.