Rugby Analytics Lakehouse
Rugby results can change. A useful analytics pipeline needs to remember what changed, preserve the source and keep its models consistent. This lakehouse brings that thinking to United Rugby Championship data.
The question
A final score tells you who won. I wanted to follow the longer story: how do fixtures, teams, player scoring and team performance change across United Rugby Championship seasons?
The current dataset spans 2021–22 to 2025–26: five seasons and 755 completed matches, reconciled after a backfill and a source repair. These are season snapshots, so a result can arrive late or be corrected. Keeping that history is part of the question, too.
A little context before kick-off: team wins include playoffs, so these are team records, not official league standings. Player analysis covers lineups and listed scoring events, not a complete picture of player performance. A missing result means “result unavailable”, not necessarily an upcoming fixture.
The engineering
From the final whistle to a useful answer.
- SOURCESeason JSONRugby match feeds, saved as local snapshots
- ORCHESTRATEAirflow in DockerUpload selected file → run ingestion → run dbt
- LANDDatabricks VolumeScan the folder for new snapshots
- PYSPARKBronzeKeep raw versions
- PYSPARKSilverModel fixtures & completed matches
- DBT COREGoldBuild analytics models & run quality tests
The job is to turn revisable source files into a consistent view of the game. Each layer has a specific responsibility: keep the evidence, model the matches, then make the results useful.
LOCAL JSONStart with a seasonA snapshot of the source, ready to load.
The input is a local season JSON file. It represents the season as the source knows it at that moment, rather than a stream of individual changes. Later results and corrections arrive as revised snapshots.
AIRFLOW · DOCKERCoordinate the runUpload a file. Run ingestion. Build and test.
The local, manually triggered DAG selects 2025–26 by default, with other seasons configurable. It validates and transforms the selected file into JSONL, uploads it, waits for the Databricks job, then runs dbt Core. A failed task stops the later steps.
DATABRICKS · MANAGED VOLUMELand it, then discover what’s newOne selected upload; a scan of the whole folder.
Landing filenames include the season, transform version and a content hash: a fingerprint of the snapshot. The Databricks job scans all matching files and checks the loaded-snapshot manifest. Known snapshots are skipped; changed content creates a new filename to ingest.
BRONZE · PYSPARK + DELTAKeep the original evidencePreserve raw versions when a match changes.
Match-level content hashes distinguish a correction from a repeat load. Identical content adds no new match version. A changed score or lineup creates a new version while the previous raw record stays in Bronze, making corrections traceable.
SILVER · PYSPARK + DELTAGive fixtures and results their own placeA fixture can exist before its result does.
Silver separates versioned fixtures from completed matches, player details and scoring events. Current views follow the latest published season snapshot, so late results and corrected records can replace the current view without erasing history. The snapshot manifest is published after its rows are written.
GOLD · DBT COREBuild useful models. Check the numbers.Facts, dimensions and analytics with explicit grains.
Nine dbt models turn current Silver data into reusable team, match and scoring views. The 46 tests check unique keys, relationships, fixture status, two team appearances per completed match and agreement between team and match score totals.
DASHBOARD + WEBSITE EXPORTPut the work in front of someoneTwo views of curated Gold data.
Read-only dashboard queries use Gold models. A separate exporter produces versioned JSON for the website, checks reconciliation and writes its manifest last so readers see a complete snapshot. The website reads that saved export, independently of Databricks. Refreshing the public snapshot remains a separate release step.
Repeat loads should be uneventful. Changed records should leave a trail.
What works today
All five seasons have been ingested and reconciled in Databricks Free Edition. The documented Gold build contains nine dbt models and 46 passing tests, covering keys, relationships, fixture status, team appearances and score totals.
- seasons reconciled
- 5
- dbt models built
- 9
- dbt tests passing
- 46
On 2 October 2026, a manually triggered local Airflow run in Docker uploaded the unchanged 2025–26 file, started the Databricks folder-ingestion job and completed dbt. The existing season snapshot was not duplicated.
That distinction matters: Airflow selects one local season file for upload by default; the Databricks job scans the whole landing folder for new snapshots. Dashboard queries and the website explorer then use saved exports of the curated Gold data.
Still to verify or deliver: a changed-source run through Airflow, automatic scheduling, Power BI reporting and automatic public website data refreshes. The successful orchestration run used an unchanged source file.
Beyond my laptop
The next test is a less glamorous one: can it keep working when I’m not there to press Run? Here’s how I’d take this working project towards a production platform.
Give the data a reliable way in
Automate source delivery and move orchestration to a hosted service, with a deliberate refresh schedule, retries and checks for late or missing files.
Make change safe to ship
Add CI/CD checks and promote tested changes through development, staging and production. Use managed identities or service credentials in a secret store, with access scoped to each environment.
Make problems easy to find
Monitor freshness, quality and failed runs, with actionable alerts. Extend lineage from source snapshot through Gold to each export, so a surprising number has a traceable explanation.
Plan for the awkward matchday
Document and rehearse recovery: replay a snapshot, backfill a season and roll back a bad release. Publish a governed, versioned website feed only after quality checks and reuse rights are settled.
These are deliberate next steps, not implemented features. Public redistribution of the source rugby data still needs appropriate reuse permission, or a source licensed for that use.
Source: Project README and implementation status. Project descriptions reflect documented implementation, not independently rerun results.
The game, from
another angle.
Pick a season. Follow a team.
See what the scoreline leaves out.
The explorer loads as you reach it.