K-means in R: How to visualize clusters without using "fviz_cluster" function after preprocessing data using PCA
21:44 17 May 2024

I am trying to write code in R that uses K-means clustering after preprocessing the data using PCA. I found the "fviz_cluster" function but it seems like the function uses the first two principal components as the X,Y axes. Are there any other way to visualize the clusters that doesn't uses the first two principal components as the X,Y variables but instead the all of the PCs as a result of the PCA transform (in this case the first 3 principal components)

# Load necessary libraries
library(ggplot2)
library(cluster)
library(GGally)
library(dplyr)
library(factoextra)

# Load the data
data <- read.csv("Hydro_lm.csv")

# Replace any missing data with NA
data[data == ""] <- NA

# Check for missing values
if (any(is.na(data))) {
  cat("Warning: There are missing values in the dataset. They have been replaced with NA.\n")
}

# Min-max normalization function
min_max_normalize <- function(x) {
  return((x - min(x, na.rm = TRUE)) / (max(x, na.rm = TRUE) - min(x, na.rm = TRUE)))
}

# Apply min-max normalization to each column
data_normalized <- as.data.frame(lapply(data, min_max_normalize))

# Perform PCA on the normalized data
pca_result <- prcomp(data_normalized, center = TRUE, scale. = FALSE)
summary(pca_result)

#keep first 3 PC
hydro_transform = as.data.frame(-pca_result$x[,1:3])
hydro_transform

# Decide on the number of clusters using Elbow Method on PCA results
fviz_nbclust(hydro_transform, kmeans, method = "wss")

#vizualize clusters
k = 4
kmeans_hydro = kmeans(hydro_transform, centers = k, nstart = 50)
kmeans_hydro
fviz_cluster(kmeans_hydro, data = hydro_transform)
print(kmeans_hydro)
r visualization k-means pca