A complete demo of MySQL → Debezium → Kafka → ClickHouse → Metabase + GrafanaMySQL → Debezium → RabbitMQ → ClickHouse → Metabase + Grafana, modelling 220 hospitals, daily diet boxes, and a retail e-commerce shop. Every layer of the modern data stack — change data capture, native streaming, dimensional modelling, business intelligence, observability — wired together and running on GKE, provisioned by Terraform and deployed via ArgoCD GitOps.
Hospital meals are a low-margin, high-trust business: a single missed SLA can cost a contract, a single cold-chain breach can cost a license. The data exists — it's just stuck in five systems no leader can read together.
Yesterday's CSV refresh tells you about a problem after the contract is already on the table.
OLTP, fleet GPS, cold-chain sensors, billing — five sources, zero correlation.
Diet box subscriptions and e-commerce live in different tools — leadership sees the trees, not the forest.
SREs need infra metrics; the COO needs revenue. One Grafana for both is a compromise nobody loves.
MediCater OPS streams every business event from your operational MySQL into a ClickHouse warehouse via Debezium and KafkaRabbitMQ — in seconds, not hours. Pre-aggregated KPIs power business dashboards in Metabase; raw infra metrics power on-call dashboards in Grafana. Both audiences win.
A classic medallion architecture, fully implemented and visible end-to-end. Every metric on every dashboard can be traced back through three explicit layers of the warehouse.
medicater.* ReplacingMergeTree mirrors · raw CDC shape · exact dedup with FINALmedicater_dw.* star schema · 11 conformed dims · 5 fact tables · dim_date · 3 dictionariesmedicater.kpi_* AggregatingMergeTree rollups · 6 KPI tables · v_*_kpi convenience views
Transactional state (orders, deliveries, contracts) flows through CDC for guaranteed ordering. Raw telemetry — GPS pings, cold-chain temperature, SLA micro-events — bypasses MySQL entirely and is consumed directly by ClickHouse's KafkaRabbitMQ engine. Same warehouse, two well-suited shapes.
Long-term contracts with hospitals. SLA-bound deliveries, breach penalties, contracted meal plans by tier.
Daily meal-kit deliveries to subscribers' homes. 10 plan SKUs from Slim 1200 to Sport 3000, monthly billing.
One-off retail orders via web, mobile, marketplace, and partner apps. Cards, BLIK, transfers, COD.
Debezium tails the binlog, Kafka transports, the Altinity sink lands rows into ClickHouse before the spinner ends.
Debezium tails the binlog, RabbitMQ transports, the ClickHouse RabbitMQ engine lands rows before the spinner ends.
High-volume telemetry (GPS, cold chain, SLA events) bypasses CDC and hits the ClickHouse Kafka engine directly.
High-volume telemetry (GPS, cold chain, SLA events) bypasses CDC and hits the ClickHouse RabbitMQ engine directly.
11 conformed dimensions, 5 fact tables, surrogate date_key, an incremental ETL container with checkpoints.
6 AggregatingMergeTree KPIs hand-tuned for dashboard latency. Sub-100 ms responses at billions of rows.
Vehicle utilisation, daily km, L/100km, fuel cost in PLN by fuel type, driver safety scores. All live.
Per-vehicle temperature curves with realistic spikes on door-open events. Violations flagged automatically.
Monthly billing, top clients, per-region revenue, penalty leakage by hospital tier — all stitched through the contract layer.
Prometheus scrapes ClickHouse, MySQL, Kafka, Connect. Five Grafana dashboards from CDC pipeline health to stack-wide SLO.
Prometheus scrapes ClickHouse, MySQL, RabbitMQ. Five Grafana dashboards from CDC pipeline health to stack-wide SLO.
Three threads continuously generate orders, telemetry, churn, and incidents with realistic seasonality and anomalies.
A Python script pre-builds 8 Metabase dashboards via REST API. No clicking required — open and present.
Every script can be re-run safely. Bootstrap, backfill, ETL, dashboard provisioning — all check before they write.
12 workloads on GKE. Terraform provisions the cluster and edge, ArgoCD syncs every Deployment from Git. No manual kubectl apply.
A single business event traces a clean line through the platform. Watch a hospital order travel from MySQL to the executive dashboard — under 5 seconds, every time.
Hospital ops dispatches a meal order. The simulator (or your real OLTP service) runs INSERT INTO meal_orders. MySQL writes the row and emits a binlog event.
The Debezium MySQL source connector tails the binlog, wraps the change in an envelope (op: c, before, after, source), and produces it to kafka_cdc.medicater.meal_orders.
The Debezium Server tails the binlog, wraps the change in a flattened JSON record, and publishes it to the rabbitmq_cdc exchange with routing key rabbitmq_cdc.medicater.meal_orders.
The ClickHouse sink consumes the topic, batches records, and inserts into the medicater.meal_orders ReplacingMergeTree mirror. _version is set from the source timestamp; deletes flip is_deleted.
ClickHouse's RabbitMQ engine table consumes the message, its materialized view transforms the JSON and inserts into the medicater.meal_orders ReplacingMergeTree mirror. Deletes flip is_deleted.
Each insert triggers mv_kpi_delivery_hourly, mv_kpi_hospital_daily, and the channel rollups. Aggregate state functions write to AggregatingMergeTree targets.
Once a minute, the etl container re-runs transform.sql: 11 dimension refreshes, 5 incremental fact loads keyed by etl_checkpoint. fact_delivery picks up the new row.
Metabase's executive panel re-runs its query against v_delivery_hourly_kpi on its 30-second refresh interval. The new order is visible. The COO sees on-time % update in real time.
Operations is a first-class concern, not an afterthought. Five SRE-focused dashboards cover the whole stack — every layer of the pipeline has a metric you can alert on.
The on-call's home base. INSERT/SELECT rate, batching efficiency, ClickHouse merger backlog, p50 latency, SELECT-to-INSERT ratio. Alert when batching drops or parts pile up.
Query rate, failed query rate, parts/tables/databases, background merge tasks, RSS, mark cache, uncompressed cache, file handles, data path utilisation.
Per-topic message rate, total cumulative offsets, lag by consumer group × topic. Alert when sink lag grows linearly.
Per-exchange publish rate, queue depth, consumer ack rate, unacked messages. Alert when queues grow linearly.
Slow query rate, threads connected/running, command-type breakdown, InnoDB row activity (insert/update/delete/read), buffer pool used/free/dirty.
Big green/red component status row (CH, MySQL, Kafka, Connect, Prometheus). End-to-end ingest rate (MySQL inserts/s vs ClickHouse inserts/s). Total Kafka lag. ClickHouse RSS. Backpressure indicators. The board to flash on the wall.
Big green/red component status row (CH, MySQL, RabbitMQ, Debezium, Prometheus). End-to-end ingest rate (MySQL inserts/s vs ClickHouse inserts/s). Queue depth. ClickHouse RSS. Backpressure indicators. The board to flash on the wall.
kubectl rollout restart deployment -n medicater <name> for a single component, or resync the app in ArgoCD to reapply everything from Git. Idempotent and safe.
Reset the topic offset: kafka-consumer-groups --reset-offsets --to-earliest --topic kafka_cdc.medicater.X --execute. Documented in CLAUDE.md.
Check queue depth: rabbitmqctl list_queues name messages. Purge if needed: rabbitmqctl purge_queue <name>. Documented in CLAUDE.md.
./scripts/verify-cdc.sh compares MySQL row counts against ClickHouse FINAL counts and reports any drift.
Every link below resolves to a live Service running on GKE right now. Click and present.
8 dashboards · 80+ cards · ready to demo
5 operational dashboards · folder "Operations"
CH · MySQL · Kafka · Connect scraping
CH · MySQL · RabbitMQ scraping
Debezium source · Altinity sink status
Exchanges · queues · connections · message rates