A million records a minute
A lakehouse on Apache Iceberg with Kafka and Flink stream processing, handling 1M+ records per minute in near real-time — plus a no-code builder so clients could wire their own ETL pipelines.
- Kafka
- Flink
- Iceberg
- ETL
Machine Learning Engineer · Johannesburg, South Africa · UTC+2
I build ML that survives contact with production.
Machine Learning Engineer with 7+ years building and deploying production ML, RAG/LLM and real-time data systems across banking, IoT and retail. Led a 10-person AI team, built streaming data platforms processing 1M+ records per minute, and shipped production RAG systems on Kubernetes. MEng in Computer Engineering (cum laude).
An open-source rules engine I built at Capitec, now productionised across business credit, retail credit and fraud detection — executing on 7M+ transactions a day. Presented at PyCon.
capitec/dsp-decision-engine ↗
A lakehouse on Apache Iceberg with Kafka and Flink stream processing, handling 1M+ records per minute in near real-time — plus a no-code builder so clients could wire their own ETL pipelines.
A production chatbot platform on LLMs, embeddings and vector search. Moved it off a lone EC2 box onto Kubernetes, rebuilt ingestion serverless, and made every answer traceable to a downloadable source.
Computer vision in places it has to actually work: planogram compliance at 97% accuracy with detection cut from 15 minutes to 5 seconds, and real-time person detection on bank branch CCTV.
Eight roles, seven years, one long thread of shipping models into production —the full timeline is here.