Defensible AI for financial crime investigations

They built the alarm. Nobody built the investigation.

Clarté investigates each alert, shows the evidence on both sides, says plainly what the data can't show, and produces decisions your team can replay and defend in an exam.

Built first for BSA/AML teams at community banks working under real examiner pressure. The ledger on the right is scored in code on this page, with the engine's own formula.

Evidence ledger · sample case

Riverside Auto Parts LLC
83 TXNS · BUSINESS · OCT 2025 – MAR 2026
—…
Detectorexculpatory −0+ suspiciouslog-odds

Three illustrative cases on synthetic data. Each bar is a detector's weighted contribution in log-odds; the posterior starts from a 3% prior and is capped at 99% by design. The demo shows 10 of the 40 detectors.

40Evidence detectors
49Typologies mapped · 44 active
17Specialized agents · 12 AI-backed
4,407Automated tests

The moment that matters

The examiner opens a case file and asks three questions.

Most compliance tools help you run alerts. Clarté is built for the conversation that happens afterwards.

"How did you decide this case was not suspicious?"
Without Clarté

"The score was low and the analyst reviewed it."

With Clarté

"40 detectors examined 83 transactions across six months. 31 had enough data to judge, and the nine that didn't are listed by name. Nothing pointed to suspicion and the challenger agreed. Here's the run you can replay, and the hash-verified audit trail."

"Why did this case take 45 days to resolve?"
Without Clarté

"We had a backlog."

With Clarté

"Flagged INVESTIGATE on day 1. Entered review on day 3. Additional history requested on day 12. Determination recorded on day 18. Every action timestamped and hash-chained."

"How do you know your model is working?"
Without Clarté

"Our vendor says it's validated."

With Clarté

"The model version changes with every scoring change. A standing blind sample measures how often officers agree with the machine, reported with its sample size. And one wrong close pauses automatic closing until we review it."

The investigator, end to end

Every case goes through the same eight steps. The evidence decides where it ends up.

This is what runs when a case arrives, with no one watching. The names below are the real stage names from the code, and each step writes what it did to the case's audit chain.

01Intake

The data is read strictly, and the gaps are counted.

Amounts, directions and timestamps are parsed without guessing. An ambiguous row is rejected with a reason. Counterparty names are tokenized before anything else sees them, after the risk attributes they carry have been derived.

Every upload reconciles: rows accepted and rejected, totals by direction and currency, date coverage, missing fields.
A date without a time of day is kept as a date. Timing detectors abstain rather than treat the day's payments as simultaneous.
Rule: unknown stays unknown. A missing field is never read as a reassuring one.
02Detect

Forty detectors weigh the evidence and say what they couldn't see.

Each detector reports evidence it observed, absence it confirmed, or that it couldn't look, with the data coverage it needed. Their weighted contributions add up in log-odds to one posterior, capped at 99% by design.

Floor rules override the score when one signal is too dangerous to average away.
A sensitivity range and a two-sided fragility check show whether removing one piece of evidence, in either direction, would flip the determination.
A decision-readiness check names the facts still missing and whether each can be obtained. A one-transaction case cannot be cleared on its score.
03Challenge

The result is attacked before anyone reads it.

A deterministic skeptic runs forty predicates against the engine's own findings. An independent challenger model scores the same case; if the two disagree by more than a tier, the case can't proceed until someone writes down why.

Consistency checks compare the case against the customer's declared profile, linked cases that share a counterparty, and the institution's own closed precedents.
The record says which checks ran and which couldn't. "No contradictions found" is never written when a check was skipped.
04Investigate

Four investigators work the case in parallel.

Counterparty intelligence, source of funds, pattern archaeology across the account's history, and an evidence-gap identifier that lists what an officer would still need to ask. Each writes analysis an investigator would otherwise have produced by hand.

Their output is validated against a schema and the matched typologies. A claim about a typology the engine didn't match is dropped, and the drop is recorded.
Rule: the investigators narrate. No model output ever decides control flow. Every decision below is a deterministic function of the engine's signals.
05Propose

The disposition is drafted and then checked against the evidence.

A disposition-rationale writer drafts the clearing or escalation reasoning. On an escalation, the SAR writer drafts the narrative. Both run only when the case has enough evidence to dispose; a case that stops with a question gets no draft.

Six grounding layers check every SAR draft: fact check against the transactions, pattern grounding, citation check, legal-conclusion blocker, investigation bounds, qualitative grounding.
A package-quality review scores the draft against the eight SAR-narrative requirements and flags any it fails.
06Package

The officer gets a briefing, not a score.

What's going on, in plain language with the numbers that matter. The evidence on both sides. The surviving signals and the ones set aside. The one check that would change the answer. Then the proposed action, in plain verbs.

Every line of the briefing is traced to the signal or agent output it came from.
The package records the model version, the analysis date and the fingerprint of everything it depended on. Any later action checks that fingerprint first.
07Route

The case lands in the one queue that matches what it needs.

Ready to close, ready for signature, decision needed, or evidence needed. The routing is a pure function with a dozen named rules, and the rule that fired is written on the case in plain English.

"Ready to close" requires an engine clear with nothing contradicting it, a stable result, sufficient evidence, and no floor rule involved.
"Ready for signature" requires an engine escalation, sufficient evidence, and a draft that passed all six grounding layers.
08Re-enter

New evidence sends the case back through all eight steps.

A customer's document arrives and its text is extracted. An officer resolves an information request, sends a case back with what's new, or reanalyzes it. Each of these queues the case again. The old package is superseded; nothing is decided on stale evidence.

Eight triggers re-queue a case: creation, a rescored evidence set, an RFI document attached, RFI responses logged, an RFI resolved, a send-back, a backfill, a manual retry.
Re-running any stored case reproduces its original result exactly, from the stored inputs and the model version that scored it.

The line the whole system is built on. Steps 02, 03, 07 and every decision are deterministic code over the evidence. Steps 04 and 05 use language models, and their output is narration that is validated and grounded before anyone reads it. The score, the routing, the auto-close and the tripwire never read a model's words.

Seventeen specialized agents sit behind these steps: twelve model-backed, five deterministic. Every one can be switched off per institution; the scoring core is version-controlled, not toggleable.

Status. Built and passing 4,407 automated tests, 2,419 of them against a real database. Not yet in production: staging validation comes before design-partner pilots.

Operator

The machine does the work. Your officer does quality control.

Operator runs the eight steps above on every case as it arrives, and again whenever new evidence lands. Then it works each queue the way an officer would, up to the line you set.

BUILT AND TESTED · NOT YET IN PRODUCTION
Ready to close

Clean, well-evidenced cases

With autonomy on, Operator closes them, writes the rationale and records the outcome. With it off, your officer confirms them one at a time or in batches of up to fifty.

Ready for signature

Escalations with a grounded SAR draft

The draft has passed all six grounding layers and the package-quality review. Your officer reads it inline and signs with one click.

Decision needed

One question decides it

Operator states the question and what each answer would mean. Your officer answers it on the case.

Evidence needed

Something is missing

Operator names the missing fact and drafts the request for it. A person sends it, and the reply re-enters the loop.

What never changes
A person signs every SAR.
Operator drafts and checks. It never files.
No automated customer contact.
Requests for information are drafted for a person to send.
Autonomy is your switch.
An admin turns automatic closing on or off with a recorded note. There is no waiting period and no hidden threshold.
How you know it's working
Blind quality control.
A standing sample goes to an officer who can't see the machine's answer. Agreement is reported with its sample size, and reads "not yet measured" until there is one.
A tripwire.
If Operator closes a case that a person later escalates, automatic closing pauses at once and stays paused until an admin reviews it.
Spot checks.
Every automatic close is listed and can be reopened with its history intact.
200 / dayDefault cap on investigations per institution, with a separate token budget.
5% · min 20Blind QC sample rate and monthly floor, rising to 20% for 200 cases after any model change.
PauseOne switch stops new work within a minute. In-flight runs finish and record.
Per agentEvery agent can be disabled per institution. A disabled SAR writer means no draft, so the case routes to a person.

Built-in safeguards

A system that knows when it doesn't know.

Most systems give you a score and leave you alone. Clarté tells you when to trust the score, when to question it, and when to stop and think.

!

The score says CLEAR, but one signal is screaming.

A floor rule overrides the score and forces review. Some signals are too dangerous to average away.

±

The result could easily go either way.

You see how fragile it is. When removing one piece of evidence, in either direction, would flip the determination, the case says so.

⇄

Two models disagree.

You can't proceed until you write down why. That rationale becomes part of the permanent audit record.

⊞

There's evidence on both sides.

Both are shown, explicitly. No hiding the green flags behind the red ones.

?

The data is too thin to say.

A missing country, a date with no time, three transactions where ninety days are needed: the check is marked "couldn't observe" and the gap is listed. Missing data never counts as reassurance.

Design-partner targets

What we're building toward.

Targets for a two-person BSA team handling about 300 alerts a month. These are objectives, not guaranteed outcomes. They will be measured during pilots and published in our calibration reports.

Time per case
45 min5 min

Review time per case, once routine cases are prepared or closed by Operator.

Monthly analyst hours
225 hrs25 hrs

200 hours back for actual investigation rather than alert triage.

Exam readiness
ScrambleAlways

Continuous readiness rather than a two-week fire drill before every exam.

This isn't headcount reduction. It's giving your team back the time to do the work that matters.

Founder

Three lines of defense. One founder who's worked all three.

I've sat across from examiners defending decisions I couldn't fully explain, because the tools didn't give me the language. I've written SARs at 11pm wondering if I caught everything. I've watched a two-person BSA team drown in 300 monthly alerts knowing that 285 of them were noise.

Most compliance tools are built by engineers who've never filed a SAR, or by consultants who've never written code. I've done both, and I've audited the tools other people built. That's why Clarté works the way it does.

"I know what the examiner asks because I've been the person the examiner asks. I built Clarté to give the answer I always wished I had."
First line: BSA analystCommunity bank

Investigated cases, filed SARs, survived exams. Learned what examiners actually ask, and what answers satisfy them.

Second line: financial crimes compliance advisorBig Four consulting

Led consent-order remediation and program reviews for global banks. Directed 40+ analysts across the US and APAC.

Third line: payments compliance auditMajor technology company

Audit across billions in payment volume. Built AI-assisted audit tools. Evaluated the systems other people trust.

CAMS certifiedACAMS

The industry-standard credential for AML professionals.

For model validators and risk committees

We publish everything.

Full methodology paper. Model risk alignment memo with mapped supervisory requirements. Known-limitations register. Validation framework with recalibration governance. The end-to-end review and our remediation record. Every parameter traceable, every limitation documented.

Methodology paperModel risk alignmentKnown limitationsValidation frameworkTest coverage · 4,407Review & remediation record

Available under NDA. Write to team@clartehq.com to request technical documentation.

Where we are today

A working build, a real founder, an open design-partner program.

Built onDirect BSA workflow experience

First-line analyst, second-line advisor, third-line auditor. Designed by someone who's been in the seat.

Open nowDesign-partner program

Selecting three to five community banks for free 60–90-day pilots. The first empirical calibration report follows once pilots produce real outcomes.

Built, not yet liveCurrent build and Operator

The rebuilt evidence engine, decision replay, Operator and blind QC are built and tested. Staging validation comes next; design partners see them first.

Available nowSample cases on this page

Three illustrative cases, scored the way the engine scores them. See the evidence ledger at the top of the page.

Under NDATechnical documentation

Methodology paper, model risk alignment memo, validation framework, limitations register. Available to validators on request.

See how your decisions would stand up in an exam.

Start with the sample cases, or request a design-partner walkthrough on your own data.

Request a design-partner walkthrough Back to the sample cases

We work with BSA teams who aren't just managing alert backlog. They're rethinking how decisions are made.