Portfolio project · synthetic data · not an ATC system

Airspace Conformance Platform

Four Python microservices connected by Kafka turn noisy simulated aircraft position reports into smoothed tracks and advisory safety alerts: predicted losses of separation, unmodelled manoeuvres, and emergency transponder codes.

60-second engineering tour

1. Watch the system

Replay a measured run from the real simulator, tracker, and conformance monitor. The traffic is synthetic; the recorded behavior is not mocked.

2. Inspect a boundary

Read the versioned message contracts and the idempotent persistence boundary.

3. Follow real integration proof

The broker integration test exercises Redpanda instead of replacing Kafka with a mock.

4. Read the negative result

See why a more complex probabilistic detector was not promoted, and why the simpler operating point remains defensible.

Representative evidence: separation logic, REST/WebSocket boundary, and end-to-end system test.

The system, replayed

T+00:00 4 aircraft

Not a video and not a mock-up. This is a recorded run of the real simulator, Kalman filter, and separation monitor, replayed in a canvas — the same log that produces the animation in the README, generated by scripts/make_demo.py. Two aircraft converge head-on at FL350; the advisory fires about five minutes before closest approach. Watch ACP303: it crosses the same point at the same moment, 4000 ft above, and is correctly ignored — a conflict needs both standards breached at once, and treating either alone as sufficient is the most common way to get this wrong.

Results

Conflict detection

Scored against simulator ground truth the detector never sees, while consuming only the noisy observation stream. The traffic is invented; the measurement is not.

MetricValueReading
Recall1.00Every real loss of separation was alerted before it happened. Sweeping manoeuvre density, recall holds — and is lowest when nothing manoeuvres at all
Precision0.56Two in five alerts concern a pair that never loses separation. Not good enough, and the reason is known
Median warning249 sJust over four minutes of lead time
Sample123 scenarios / 20.75 h39 real losses of separation
Precision of 0.56 is an operating point, not a defect. The median “false” alert is a pair that genuinely closed to 5.52 NM against a 5 NM standard, raised at 294 s against a 300 s lookahead ceiling. Halve the lookahead and precision reaches 0.87 — but lead time is the product, so the default stands. Quoting a precision figure without the lookahead it was measured at was the actual error.
That took two attempts to work out. The stated cause for four milestones was point-estimate thresholding, so the principled fix got built — probabilistic detection using the Kalman covariance. It won on the family it was tuned against and vanished on shifted traffic. The failed fix is what prompted measuring the alerts instead of theorising about them.

Trajectory prediction

A PyTorch model predicting the residual of dead reckoning, not the position. Horizontal error while turning, at a 60 s horizon:

SplitDead reckoningNeuralSkill
Unseen scenarios, same family2.412 NM1.220 NM+49.4%
Shifted family2.024 NM1.261 NM+37.7%
The neural net is not distinguishable from ridge regression on shifted traffic. A scenario-clustered bootstrap gives −12.7% to +5.5% at 60 s. Both beat dead reckoning decisively; what is uncertain is whether the extra capacity earns its place away from the training distribution. Shipping the linear model would be defensible, and the model card says so.

How it is built

feed

Plays a seeded scenario and emits noisy, dropout-prone surveillance reports at 1 Hz, keyed by aircraft address so Kafka orders each aircraft independently.

track

A constant-velocity Kalman filter per aircraft, with track initiation, coasting and termination. History to Postgres, live picture to Redis. The one service that shards.

conformance

Closest-point-of-approach geometry, trajectory conformance against an earlier prediction, and single-aircraft rules. Publishes alerts through a NEW→SUSTAINED→CLEARED lifecycle.

api

FastAPI REST and a WebSocket, plus the plan-view display. Never consumes Kafka and never writes — so a display refresh stays off the estimation path.

Two ideas the project is actually about

A filter chosen for its weakness. A constant-velocity Kalman filter lags during turns. That lag — the innovation — is published downstream and becomes the manoeuvre signal the conformance monitor thresholds on. The obvious upgrade to a constant-turn model would smooth away the quantity the system exists to notice.

A model that learns the correction, not the answer. The predictor outputs a residual on top of dead reckoning, so a broken, missing, or NaN-producing model degrades to physics rather than to nonsense — and that promise is executed in CI with the dependency genuinely uninstalled, not asserted in a comment.

Engineering

ConcernWhat is there
Testing700+ unit and contract tests at 82% branch coverage with no exclusions, 28 integration tests on real Redpanda/Postgres/Redis, 8 end-to-end under compose, and a measured latency budget
ContractsCommitted JSON Schemas and OpenAPI, drift-gated, plus a backward-compatibility diff against git. Schemathesis fuzzes the API against its own document
ObservabilityPrometheus metrics on all four services, a provisioned Grafana dashboard, and W3C trace context carried across Kafka into Jaeger
DevSecOpsbandit, pip-audit, gitleaks over full history, Trivy, and a syft SBOM. All eight third-party actions pinned to commit SHAs
DeploymentCompose and Kubernetes manifests, applied to a real kind cluster in CI. The image is built once and every consumer asserts it is the one that was tested
The most useful thing in the repository is what went wrong. Property tests found two geodesy bugs on their first run. An architecture test caught its own author importing one service from another. Fuzzing found the OpenAPI document lying, twice. A consumer-group rebalance could delete a live aircraft from the shared picture — invisible to 500 passing tests, because every one of them ran a single consumer. And the first CI run this project ever had found five more, including a scenario fingerprint that was not reproducible across operating systems. The full account.

What it is not

Not an air traffic control system, and never to be used as one. Advisory output only, synthetic data only, no certification of any kind — no DO-178C, no requirements traceability, no structural coverage analysis. It has never been connected to ADS-B or radar. It has no flight plans, so “non-conformance” can only mean the aircraft did not do what constant-velocity physics predicted, never what it was told.

The domain was chosen because it makes the engineering legible — why ordering matters per aircraft, why a stale track is worse than no track, why latency has a budget, why a false alert is expensive. Those are abstract in most demonstration projects and concrete here.

Read safety-notes.md for the operational boundary and future-work.md for what is deliberately not built, and why.

Run it

git clone https://github.com/nick-bellows/airspace-conformance-platform
cd airspace-conformance-platform
docker compose -f deploy/compose.yml up -d --build
# then open http://localhost:8000

Docker is the only prerequisite. No cloud account, no API keys, no data download — all traffic is generated locally from a committed seed.

Questions a screen actually asks

Written for a recruiter or hiring manager with a few minutes, ahead of the engineer with an hour. Every answer links to the evidence it rests on, and the figures quoted here are checked against retained evaluation output in CI.

What is this, in one sentence?

A system that watches simulated aircraft positions and warns a few minutes ahead when two are heading for a near miss. It is built the way a real one would be — four independent services passing messages through Kafka, a database, monitoring, automated testing, a deployment path — but it runs on a simulator written for the project, not on real air traffic.

Is it a real air traffic control system, or connected to one?

No, and it must never be used as one. Output is advisory only, every aircraft is synthetic, and there is no certification of any kind. It has never been connected to ADS-B or radar. The boundary is written down in safety-notes.md. Calling it “real-time air traffic control” would be a red flag in this industry, not a boast.

What role is this evidence for?

Backend and distributed-systems work: event streaming, idempotent consumers, versioned contracts, observability, a test pyramid that runs against real infrastructure, and security scanning in CI. It also shows engineering judgement — results published when they were unflattering, and claims withdrawn when the measurement disagreed. It is not evidence of aviation domain credentials or safety-certification experience, and does not claim to be.

Was it built with AI?

Yes, heavily, with Claude Code, and that is documented rather than left to be inferred. ai-assisted-development.md records what the assistance did, the guardrails, and what it got wrong. The dominant failure was not broken code, which types and tests catch, but confident, specific, false prose: a docstring claimed a projection was accurate to 0.1%, and measuring it gave 1.3%. The rule that came out of that is that no claim stands in prose unless something that runs checks it. Separately, no generative model runs inside the system; ADR 0008 explains why that is a safety question rather than a productivity one.

If AI wrote much of the code, what does the author actually understand?

Ask about the decisions rather than the code. Fourteen decision records give each significant choice and the alternatives it beat; how-it-was-built.md lists the defects found at each milestone and how; and interview-brief.md holds the answers the author expects to give in the room. The parts that cannot be pattern-matched are the negative results: a probabilistic detector that was built, measured, and not promoted, and a recall figure later shown to be flattered by the population it was measured on.

Is the data real?

No. All traffic comes from a seeded simulator, and nothing from a third party is in the repository. That was a choice (ADR 0002): the simulator emits ground truth that no pipeline component can see, so the detector is scored against what actually happened, which real surveillance data cannot offer because it carries no truth. The cost is the project’s biggest weakness, stated in limitations.md: every measurement inherits the assumptions of a simulator written by the same person.

How long did it take?

The commit history is the record. The first commit is 15 August 2026; milestones M0 through M6 landed on the 15th and 16th; the following three weeks were evaluation, external review and remediation, ending on 4 September. Speed is not the point. The verification discipline is: no number appears in the documentation without a committed script that reproduces it, and no check ships without being watched to fail once.

Can I run it, and is there a hosted version?

It runs locally with one prerequisite, Docker: the three-line quickstart above works from a clean clone with no accounts, keys, or downloads. There is deliberately no hosted instance — an internet-facing copy would add authentication, TLS, rate limiting and cost without improving the evidence, so the replay at the top of this page is the interactive artefact. Kubernetes manifests exist and are applied to a real kind cluster in CI, which proves they apply, not that a cluster is running anywhere.

Why is precision only 0.56?

Because the detector is asked to warn 300 seconds ahead, and constant-velocity error accumulates over that window. The median “false” alert is a pair that genuinely closed to 5.52 NM against a 5 NM standard. At a 120-second lookahead, precision is 0.87 at the same recall, and the shifted traffic family moves the same way. The default stays at 300 s because lead time is the product; ADR 0013 records the trade, including the argument it later withdrew.

Recall of 1.00 looks too good. Is it?

Treat it with suspicion; the repository does. Recall read 1.00 for the whole life of a real detector defect found by external review in September 2026: the vertical standard was checked only at the instant of horizontal closest approach. The number never moved because the scenario generator had never produced the geometry the detector was blind to. The fix is ADR 0014, and the lesson sits next to the number in limitations.md: a metric is only as strong as the adversarial quality of the population it was measured on.

Has anyone outside the author reviewed it?

Yes. Three external review rounds during the build found eighteen defects, and two independent LLM reviews with filesystem access (Codex and Cursor, 4 September 2026) each returned ADVANCE and each found what the earlier rounds had missed: the detector geometry defect above, a state-lifecycle defect, and a latency budget miss. That miss is disclosed rather than closed — the conflict scan runs at 342 ms p95 against its 250 ms budget, and the roadmap forbids fixing it by raising the budget. The full account is in ROADMAP.md.

Are the numbers on this page hand-typed?

No. test_portfolio_site.py reads the retained evaluation JSON and fails CI if the recall, precision, lead time, sample size, lookahead figure, false-alert geometry, or latency miss shown here diverge from it. The same test pins the ADR count to the directory, the pinned-action count to the workflow, and the test count to the README, because each of those had drifted before the guard existed.

Is it finished, and is it maintained?

Finished, deliberately. ROADMAP.md marks the repository portfolio-ready, schedules no further features, and records what would justify reopening it: a target role that makes one gap in future-work.md material and measurable. It also records two prohibitions, so that the numbers cannot later be made to look better by tuning the detector on the same scenario families or by moving a budget to meet a result.

Where should an engineer spend ten minutes?

separation.py for the geometry, ADR 0014 for the defect it had, test_messaging.py for integration against a real broker, and the “what went wrong” section of how-it-was-built.md. In that order it covers the code, the judgement, and the test depth.