1. Clear result
Acme Galactik Holding Co
normalized to acme galactik holding co; no candidate met the versioned review threshold.
Inspect matching logicAcme Galactik Holding Co
normalized to acme galactik holding co; no candidate met the versioned review threshold.
Inspect matching logicContoso Maritime Grp → Contoso Maritime Group
score 95.2; tier strong_fuzzy; reasons normalized_sequence_similarity, country_overlap. Human review is still required.
Inspect source, score, and reason fieldsNo verified snapshot → HTTP 503
The API refuses to turn missing data into a false clear result. This static page does not expose a public screening endpoint.
Inspect fail-closed contract testsFollow the evidence into the FastAPI schemas, threat model, matching evaluation, and limitations.
| Score bin | Hits |
|---|---|
| 89–90 | 1 |
| 92–93 | 1 |
| 95–96 | 1 |
| 99–100 | 4 |
| Query | Matched name | Score | Tier | Source | Record | Reasons |
|---|---|---|---|---|---|---|
| Acme Galactic Holdings | Acme Galactic Holdings | 100.0 | exact | SYNTHETIC_LIST | FAKE-001 | normalized_sequence_similarity, country_overlap |
| Acme Galactic Holding | Galactic Acme Holding | 100.0 | exact | SYNTHETIC_LIST | FAKE-001 | normalized_sequence_similarity, token_order_normalized, country_overlap |
| Galactic Acme Holdings | Acme Galactic Holdings | 100.0 | exact | SYNTHETIC_LIST | FAKE-001 | normalized_sequence_similarity, token_order_normalized, country_overlap |
| Contoso Maritime Group | Contoso Maritime Group | 100.0 | exact | SYNTHETIC_LIST | FAKE-002 | normalized_sequence_similarity, country_overlap |
| Contoso Maritime Grp | Contoso Maritime Group | 95.2 | strong_fuzzy | SYNTHETIC_LIST | FAKE-002 | normalized_sequence_similarity, country_overlap |
| Acme Galactic Hldgs | Acme Galactic Holdings | 92.7 | weak_fuzzy | SYNTHETIC_LIST | FAKE-001 | normalized_sequence_similarity, country_overlap |
| Acme Galactic Holdings Intl | Acme Galactic Holdings | 89.8 | weak_fuzzy | SYNTHETIC_LIST | FAKE-001 | normalized_sequence_similarity, country_overlap |
Written for a recruiter or hiring manager with a few minutes, ahead of the engineer with an hour. Every answer links to the evidence it rests on, and every figure is read from the committed evaluation reports when this page is built.
A sanctions-screening reference system: submit a name and get back either an explainable review lead against the OFAC and UN lists, or a refusal because the data behind the answer could not be trusted. It never issues a compliance determination.
No. It is an educational screening aid. Every output is a review lead for a qualified person, and a “clear” result means only that no name-similarity lead crossed the threshold in the loaded snapshots on that date. The disclaimer, the limits, and the human-review workflow are in limitations.md.
Backend and data engineering: ingestion with recorded provenance, a FastAPI service with contract tests, explainable fuzzy matching, versioned thresholds, evaluation that is regenerated in CI, analyst exports, and a container path. It is not evidence of legal or compliance credentials, and does not claim to be.
The adapters pull the real, public OFAC SDN and UN Consolidated lists, and nothing from either is committed to the repository. Everything on this page and in the evaluation is fictional: synthetic rows carry an explicit flag, snapshots built from them get a synthetic- prefix, and the API refuses to screen against them unless told to. The data card records the isolation controls.
Yes, as a collaboration between the author and Claude Code rather than a hand-off to either. The author set the problem, the data and safety boundaries, the stop conditions, and the rule that thresholds and reports are versioned and never silently changed, and reviewed and directed the work throughout; Claude Code did much of the implementation and testing under those rules. ROADMAP.md records the division of labour rather than leaving it to be inferred, including a finalization round and a read-only review of the code that found eight defects, each fixed with a regression test. The guardrail is mechanical rather than a promise: this page and the committed evaluation reports are regenerated in CI, and the build fails if either differs from what the code produces.
On the 65-case holdout split of a 163-case labelled set, precision is 1.000, recall 0.833, F1 0.909 at the recommended threshold of 89. 6 of the 9 misses are abbreviation aliases such as “CM Group” (6 of 6 alias cases in the split); the rest are 2 legal-suffix, 1 ambiguous-near-neighbor cases. The limitations name these as the known gaps. The labelled set is synthetic, single-annotator, and has been consumed for tuning, so the next scorer change needs fresh annotations. The full report is committed.
Because they are the measurement. Entity extraction scores a micro F1 of 0.536, and 0.000 on PERSON, because the small spaCy model fails on the ALL-CAPS, comma-inverted names sanctions notices use; place names score 0.882. In retrieval, over 24 labelled queries against 200 Federal Register notices, dense encoding alone is the weakest mode (Recall@10 0.546 against BM25’s 0.818), and hybrid ranking edges BM25 on MRR (0.853 against 0.839) while trailing it at Recall@10 (0.763). The model cards say which to use. Adjusting a number to look better would fail CI, because the reports are regenerated there.
The API fails closed. With no verified snapshot loaded, screening returns HTTP 503 instead of an empty “clear”; a source file that parses to zero records is rejected at ingest; a name that normalizes to nothing is rejected with 422; and only the newest snapshot per source is served, with superseded files kept on disk for audit. The health endpoint reports snapshot age and turns degraded past the configured maximum. Each of these is a row in the threat model with a test behind it.
It screens names only: identifiers such as passport or registration numbers are not matched. It covers two lists, not the EU or UK lists. The normalizer folds to ASCII, so names in Cyrillic, Arabic, or CJK are refused rather than silently cleared. There is no audit database of who screened what; runs are exported files. All of it is in limitations.md rather than discovered later.
It runs locally from the README quick start with a Python virtual environment, or as a container with Docker Compose bound to localhost. There is deliberately no hosted endpoint: a public free-text screening service creates abuse, privacy, freshness, and legal-presentation risks without adding evidence, so this page is the interactive artefact. The CI container job proves the two runtime contracts on every push: 503 without a verified snapshot, then an exact hit carrying snapshot provenance.
No. The dashboard above is rendered from committed run tables by the same code the CLI uses, and CI fails if the committed page differs from what the renderer produces. The figures in this FAQ are read from the committed evaluation reports when the page is built, and a test pins them to those reports. The matching report itself is regenerated in CI and compared with the committed copy. See the build script and its test.
A read-only code review on 10 September 2026, with findings reproduced against the committed OFAC and UN snapshots, found eight defects: an empty name scored a perfect match against records with Arabic-script aliases, a zero-record source file became a “verified” snapshot after which every screen returned clear, superseded snapshots were merged so delisted entities stayed flagged, the health check said “ok” with nothing loaded, and exports were open to spreadsheet formula injection, among others. Each is fixed with a regression test and listed in ROADMAP.md.
The commit history is the record. The first commit is 10 August 2026 and the core system landed in the first two days; the walkthrough followed on 2 September and the review-and-finalization round on 10 September. Speed is not the point. The discipline is: thresholds are versioned and never silently changed, and every score-affecting change lands together with a regenerated report.
Finished, deliberately. ROADMAP.md marks it portfolio-ready and maintenance-only, schedules no engineering milestone, and states what would justify one: an observed weakness in the retained matching, extraction, or retrieval evaluation. It also carries stop conditions, including never exposing a public screening endpoint and never hand-authoring a result on this page.
engine.py for the scorer and the reason codes, test_api.py for the fail-closed contract, the matching report for what the scorer misses, and the threat model for how each failure mode maps to a control.