Clifton Brown Ommila — Nairobi, Kenya

Data scientist with a statistician's discipline.

I build models that hold up under scrutiny — from drought early-warning systems built on satellite data to fraud detectors that can explain their own decisions. Trained in biostatistics, working in development, finance, and climate data.

About

I believe every dataset has a story to tell. My passion lies in uncovering that story and translating it into insights that help organisations make confident, data-driven decisions.

With a degree in Biostatistics and training through the ALX Data Science programme, I bring statistical thinking, technical skills, and curiosity to solving real-world problems. I use Python, SQL, Excel, and Power BI to clean and analyse data, build predictive models, and communicate findings in ways that people can understand and act on.

I approach modelling with scalability and reproducibility in mind — from building structured data pipelines and validating inputs to writing modular code, documenting workflows, and using version control. Through my AgriRisk-AI project, I am applying these practices to agricultural and climate data in Kenya while strengthening my skills in machine learning engineering.

I enjoy working in a team where people share ideas, learn from one another, and bring different perspectives to a problem. I value clear communication and connecting technical work to the needs of the people who will use it.

I'm keen to connect with teams that see data as an opportunity to ask better questions, improve decisions, and create meaningful impact.

Tools I work in

Modeling & statistics

Python, R, scikit-learn

Data & pipelines

SQL, DBT

Geospatial

QGIS, Google Earth Engine

Visualization & BI

Power BI, Tableau, Excel

Skills

Selected projects

Climate & agriculture In progress

Crop yield failure & drought risk prediction

A county-season classifier that flags drought-driven crop failure risk using CHIRPS rainfall and MODIS NDVI data pulled via Google Earth Engine. Built as a logistic regression with leave-one-year-out cross-validation, so the model is judged on years it has never seen — the standard I'd want applied to my own claims. Deploying to Streamlit for field-facing use.

CHIRPS · MODIS NDVI · Google Earth Engine · logistic regression · Streamlit

Fintech / fraud In progress

Explainable fraud detection with a bias & trust report

A fraud model built on the CreditTransAct dataset — 15 million synthetic transactions across four customer segments with different fraud rates. The goal isn't just detection accuracy: every prediction carries SHAP-based reasoning, and the false-positive/false-negative rates are broken down by segment so the model's fairness is visible, not assumed. Documented with a full model card.

15M transactions · SHAP · segment-level fairness · model card

Recommender systems Done

MovieLens recommendation engine

A recommendation engine built on MovieLens rating data, IMDB enrichment, and genome tags — taken on specifically to close two gaps in my portfolio: cloud deployment and recommender systems. Served via FastAPI with an AWS deployment target.

MovieLens · FastAPI · AWS

Experience

Education & certifications