- π§ Analytics Engineer β transforming raw data into reliable, well-structured pipelines and models
- π Passionate about clean data architecture, good documentation, and making data trustworthy
- π± Currently learning: dbt Β· Microsoft Fabric Β· Databricks Β· Python & Data Pipelines Β· Claude & AI
- π¬ Ask me about Analytics Engineering, SQL, dbt, Power BI and data modelling
- π Writing about data & tech on Shift with Jo
Transform & Model
Platforms
Tools
Portfolio case study β Databricks-native replacement for a manual, spreadsheet-and-email access-request process.
A self-service intake app, a guided admin approval app, and a live dashboard, cutting ~17.75 hours of manual admin work down to ~3.25 hours over a 21-day sample β an ~82% reduction β built on Databricks Asset Bundles, Unity Catalog, and Databricks Apps.
- Intake app: guided multi-step Streamlit form for new/edit access requests, live-sourced rule dropdowns, per-domain/person duplicate detection, and an edit-access lookup that resumes an existing or pending request
- Admin app: deterministic slash-command approval console (no LLM in the write path) β grant/deny, comments, ticket management, and full status/rule correction, every write stamped
reviewed_by/updated_at - Data layer: two Unity Catalog fact tables as the single source of truth, with an AI/BI (Lakeview) dashboard replacing a manual Power BI cross-check
- Everything as code: tables, jobs, apps, and dashboard defined and deployed via Databricks Asset Bundles
Python Streamlit Databricks Unity Catalog Delta Lake Databricks Asset Bundles
Capstone project β Data Engineering track, Data Girls Bootcamp 2026.
An end-to-end ETL pipeline that extracts a credit score dataset from Kaggle, cleans and transforms it with pandas, and loads the result into a Databricks Unity Catalog Volume via the Files API.
- Transform: handles corrupted/sentinel values, PII pseudonymization (hashing + reversible ID mapping), type casting, range checks, and outlier/missing-value treatment β with the target column always kept untouched
- Orchestration: Apache Airflow running in Docker, backed by a Postgres metadata database so DAG run history and the admin user survive container rebuilds
- Scheduling: a daily DAG with automatic retries and a failure callback that logs which task/run failed and where to find the logs
- Documentation: every data-quality and architectural decision is explained and justified in the README
Python pandas Airflow Docker Postgres Databricks
Capstone project β dbt/Analytics Engineering course.
A dbt project built on Snowflake from scratch: a medallion-architecture pipeline (raw β staging/bronze β silver) over real public aviation data (~72K airports, ~44K runways, and user comments from OurAirports, joined on the airport_ident ICAO code).
- Incremental models: built to process new/changed data efficiently rather than full-rebuild every run
- Snapshots: slowly changing dimensions (SCD Type 2) tracked via dbt's built-in snapshot feature
- Layered testing: generic, domain-specific, cross-model, and singular tests, extended with
dbt-expectations, plus test-failure persistence for debugging - Documentation:
doc()blocks, full model/column descriptions, and an interconnection overview tying the whole DAG together - Environment: managed with
uvfor reproducible Python/dbt tooling
dbt-core Snowflake dbt-expectations SQL uv
Three independent, self-contained data projects, cleaning, static reporting, and a live BI dashboard; each solving the same class of e-commerce/CRM problem with a different final delivery format.
A portfolio repo where every project lives in its own folder with no shared code, its own dataset, and its own README, so eacendently.
- Sales data cleaning: Streamlit app + CLI that diagnoses and fixes a CRM sales export, leaked status fields, flagged/contet capitalization and accentuation, mixed date formats, revenue as BR/US-formatted text, with genuinely missing or unparseable values left untouched and surfaced for manual review, never invented.
- Static analytics report: a CSV β JSON β HTML pipeline that computes e-commerce KPIs (revenue trends, category/brand bregaps, geographic distribution) into small intermediate JSON files and renders a single static report, keeping the ~3,000-row sales table out of memory/context all at once.
- Live analytics dashboard: a multi-page Streamlit dashboard over the same dataset served from a Supabase (Postgres) database read-only via the
anonkey β Sales, Price Positioning, and Customers pages each backed by a pure, unit-tested KPI module, with every formula and data-quality caveat documented.
Python pandas Streamlit Supabase pytest
Learning project β agentic databases and dbt orchestration with Cosmos + Airflow. An end-to-end pipeline over a synthetic Walmart retail dataset: data is seeded into an agentic Postgres, replicated into Databricks by CDC, modelled with dbt into a medallion architecture (bronze β silver β gold), and scheduled by Airflow. The orchestration layer is deliberately built twice, side by side, so the two approaches can be compared directly.
- Agentic source database: Ghost Postgres β disposable, forkable, and exposed over MCP, so schema creation and CSV loads were driven conversationally against a throwaway fork;
load_data.pyis kept alongside as the reproducible equivalent - Orchestration A β Astro CLI + Astronomer Cosmos: Cosmos parses the dbt graph and renders one Airflow task per model, snapshot, and test, so a failing bronze test blocks silver and every node is inspectable in the UI
- Orchestration B β plain
docker compose+BashOperator: the officialapache/airflowimage running onedbt buildper DAG β coarser granularity, far fewer moving parts
dbt Databricks Apache Airflow Astronomer Cosmos Astro CLI Postgres MCP Docker uv
A business owner asks a question over WhatsApp and gets back a number, the reason behind it, a recommendation and a chart. An AI agent writes its own SQL and chooses the chart itself.
- No hard-coded rules: two tools (
query_data,create_chart) and a description of the database. The agent works out the rest. - Read-only by permission: the database user can only
SELECTfrom two views - Serverless: a Netlify webhook replies to Meta right away and a background function does the analysis
Node.js PostgreSQL Supabase OpenAI loud API QuickChart


