Data & AI · 2026
Iterative K-Means Clustering on MapReduce
K-Means clustering on a distributed Hadoop cluster with hand-written Python Mappers and Reducers via Hadoop Streaming, without Mahout or Spark.
What I did
- The Mapper assigns each point to the nearest centroid using Euclidean distance.
- The Reducer aggregates points and computes the new centroids.
- The Reducer output is fed back into the next MapReduce job to form the iterative loop.
Tech stack
- Python
- Hadoop
- Hadoop Streaming
- MapReduce
Related projects
PythonPySparkSpark Streaming
Data & AI · 2026
Real-Time Tweet Sentiment Analysis
An end-to-end big data pipeline that trains a sentiment classifier on 1.6M tweets with PySpark MLlib and serves live predictions with Kafka and Spark Structured Streaming.
Details →Next.jsHugging FaceTailwind CSS
Data & AI · 2025
AI Text Classifier
A Next.js app that classifies text against user-provided labels with Hugging Face's facebook/bart-large-mnli model, supporting Turkish and English.
Details →Apache PigPig LatinDocker Compose
Data & AI · 2026
Dockerized Apache Pig WordCount
Word counting and grouping with Apache Pig 0.17 in Docker, on data generated with Python.
Details →