Compliance Intelligence Platform

run-20260821T172609Z-3489c0ed · created 2026-08-21T17:26:09Z

Three-minute analyst walkthrough

1. Clear result

Acme Galactik Holding Co

normalized to acme galactik holding co; no candidate met the versioned review threshold.

Inspect matching logic

2. Explainable review candidate

Contoso Maritime Grp → Contoso Maritime Group

score 95.2; tier strong_fuzzy; reasons normalized_sequence_similarity, country_overlap. Human review is still required.

Inspect source, score, and reason fields

3. Source unavailable

No verified snapshot → HTTP 503

The API refuses to turn missing data into a false clear result. This static page does not expose a public screening endpoint.

Inspect fail-closed contract tests

Follow the evidence into the FastAPI schemas, threat model, matching evaluation, and limitations.

12
entities screened
7
flagged for review
58%
flag rate
7
hits
4
exact-tier hits

Match-score distribution

123489–90: 1 hit92–93: 1 hit95–96: 1 hit99–100: 4 hitsweak ≥ 89strong ≥ 95exact ≥ 99.5859095100
Match-score distribution with the versioned review thresholds (evaluated-2026.08-v2).
Score distribution as a table
Score binHits
89–901
92–931
95–961
99–1004

Hits by source

SYNTHETIC_LISTSYNTHETIC_LIST: 7 hits7
Hits by source list.

Review queue

QueryMatched nameScoreTierSourceRecordReasons
Acme Galactic HoldingsAcme Galactic Holdings100.0exactSYNTHETIC_LISTFAKE-001normalized_sequence_similarity, country_overlap
Acme Galactic HoldingGalactic Acme Holding100.0exactSYNTHETIC_LISTFAKE-001normalized_sequence_similarity, token_order_normalized, country_overlap
Galactic Acme HoldingsAcme Galactic Holdings100.0exactSYNTHETIC_LISTFAKE-001normalized_sequence_similarity, token_order_normalized, country_overlap
Contoso Maritime GroupContoso Maritime Group100.0exactSYNTHETIC_LISTFAKE-002normalized_sequence_similarity, country_overlap
Contoso Maritime GrpContoso Maritime Group95.2strong_fuzzySYNTHETIC_LISTFAKE-002normalized_sequence_similarity, country_overlap
Acme Galactic HldgsAcme Galactic Holdings92.7weak_fuzzySYNTHETIC_LISTFAKE-001normalized_sequence_similarity, country_overlap
Acme Galactic Holdings IntlAcme Galactic Holdings89.8weak_fuzzySYNTHETIC_LISTFAKE-001normalized_sequence_similarity, country_overlap

Dataset freshness

SYNTHETIC_LIST
retrieved 2026-08-21T17:24:48Z
2 records · sha256 4a9ccd33f11a…
synthetic-fixture-4a9ccd33f11a
Synthetic demonstration data; not authoritative.

Questions a screen actually asks

Written for a recruiter or hiring manager with a few minutes, ahead of the engineer with an hour. Every answer links to the evidence it rests on, and every figure is read from the committed evaluation reports when this page is built.

What is this, in one sentence?

A sanctions-screening reference system: submit a name and get back either an explainable review lead against the OFAC and UN lists, or a refusal because the data behind the answer could not be trusted. It never issues a compliance determination.

Is it a real compliance tool? Can it clear someone?

No. It is an educational screening aid. Every output is a review lead for a qualified person, and a “clear” result means only that no name-similarity lead crossed the threshold in the loaded snapshots on that date. The disclaimer, the limits, and the human-review workflow are in limitations.md.

What role is this evidence for?

Backend and data engineering: ingestion with recorded provenance, a FastAPI service with contract tests, explainable fuzzy matching, versioned thresholds, evaluation that is regenerated in CI, analyst exports, and a container path. It is not evidence of legal or compliance credentials, and does not claim to be.

Is the data real?

The adapters pull the real, public OFAC SDN and UN Consolidated lists, and nothing from either is committed to the repository. Everything on this page and in the evaluation is fictional: synthetic rows carry an explicit flag, snapshots built from them get a synthetic- prefix, and the API refuses to screen against them unless told to. The data card records the isolation controls.

Was it built with AI?

Yes, as a collaboration between the author and Claude Code rather than a hand-off to either. The author set the problem, the data and safety boundaries, the stop conditions, and the rule that thresholds and reports are versioned and never silently changed, and reviewed and directed the work throughout; Claude Code did much of the implementation and testing under those rules. ROADMAP.md records the division of labour rather than leaving it to be inferred, including a finalization round and a read-only review of the code that found eight defects, each fixed with a regression test. The guardrail is mechanical rather than a promise: this page and the committed evaluation reports are regenerated in CI, and the build fails if either differs from what the code produces.

How good is the matching?

On the 65-case holdout split of a 163-case labelled set, precision is 1.000, recall 0.833, F1 0.909 at the recommended threshold of 89. 6 of the 9 misses are abbreviation aliases such as “CM Group” (6 of 6 alias cases in the split); the rest are 2 legal-suffix, 1 ambiguous-near-neighbor cases. The limitations name these as the known gaps. The labelled set is synthetic, single-annotator, and has been consumed for tuning, so the next scorer change needs fresh annotations. The full report is committed.

Some of the numbers are bad. Why publish them?

Because they are the measurement. Entity extraction scores a micro F1 of 0.536, and 0.000 on PERSON, because the small spaCy model fails on the ALL-CAPS, comma-inverted names sanctions notices use; place names score 0.882. In retrieval, over 24 labelled queries against 200 Federal Register notices, dense encoding alone is the weakest mode (Recall@10 0.546 against BM25’s 0.818), and hybrid ranking edges BM25 on MRR (0.853 against 0.839) while trailing it at Recall@10 (0.763). The model cards say which to use. Adjusting a number to look better would fail CI, because the reports are regenerated there.

What happens when the source data is missing or stale?

The API fails closed. With no verified snapshot loaded, screening returns HTTP 503 instead of an empty “clear”; a source file that parses to zero records is rejected at ingest; a name that normalizes to nothing is rejected with 422; and only the newest snapshot per source is served, with superseded files kept on disk for audit. The health endpoint reports snapshot age and turns degraded past the configured maximum. Each of these is a row in the threat model with a test behind it.

What can it not do?

It screens names only: identifiers such as passport or registration numbers are not matched. It covers two lists, not the EU or UK lists. The normalizer folds to ASCII, so names in Cyrillic, Arabic, or CJK are refused rather than silently cleared. There is no audit database of who screened what; runs are exported files. All of it is in limitations.md rather than discovered later.

Can I run it, and is there a hosted version?

It runs locally from the README quick start with a Python virtual environment, or as a container with Docker Compose bound to localhost. There is deliberately no hosted endpoint: a public free-text screening service creates abuse, privacy, freshness, and legal-presentation risks without adding evidence, so this page is the interactive artefact. The CI container job proves the two runtime contracts on every push: 503 without a verified snapshot, then an exact hit carrying snapshot provenance.

Are the numbers on this page hand-typed?

No. The dashboard above is rendered from committed run tables by the same code the CLI uses, and CI fails if the committed page differs from what the renderer produces. The figures in this FAQ are read from the committed evaluation reports when the page is built, and a test pins them to those reports. The matching report itself is regenerated in CI and compared with the committed copy. See the build script and its test.

Has anyone reviewed it?

A read-only code review on 10 September 2026, with findings reproduced against the committed OFAC and UN snapshots, found eight defects: an empty name scored a perfect match against records with Arabic-script aliases, a zero-record source file became a “verified” snapshot after which every screen returned clear, superseded snapshots were merged so delisted entities stayed flagged, the health check said “ok” with nothing loaded, and exports were open to spreadsheet formula injection, among others. Each is fixed with a regression test and listed in ROADMAP.md.

How long did it take?

The commit history is the record. The first commit is 10 August 2026 and the core system landed in the first two days; the walkthrough followed on 2 September and the review-and-finalization round on 10 September. Speed is not the point. The discipline is: thresholds are versioned and never silently changed, and every score-affecting change lands together with a regenerated report.

Is it finished, and is it maintained?

Finished, deliberately. ROADMAP.md marks it portfolio-ready and maintenance-only, schedules no engineering milestone, and states what would justify one: an observed weakness in the retained matching, extraction, or retrieval evaluation. It also carries stop conditions, including never exposing a public screening endpoint and never hand-authoring a result on this page.

Where should an engineer spend ten minutes?

engine.py for the scorer and the reason codes, test_api.py for the fail-closed contract, the matching report for what the scorer misses, and the threat model for how each failure mode maps to a control.