All projects

Data & AI · 2026

Iterative K-Means Clustering on MapReduce

K-Means clustering on a distributed Hadoop cluster with hand-written Python Mappers and Reducers via Hadoop Streaming, without Mahout or Spark.

What I did

  • The Mapper assigns each point to the nearest centroid using Euclidean distance.
  • The Reducer aggregates points and computes the new centroids.
  • The Reducer output is fed back into the next MapReduce job to form the iterative loop.

Tech stack

  • Python
  • Hadoop
  • Hadoop Streaming
  • MapReduce

Related projects