Using countvectorizer, I extracted feature vectors from thousands of emails and saved it in a CSV file
dictionary = open (r'''C:\Users\User\Desktop\csmp3\stemmedDictionary.txt''',"r")
dic = list(set(dictionary.read().splitlines()))
cv = CountVectorizer(vocabulary = dic, binary = True)
#~PRESENCE FEATURE VECTOR~#
#TRAIN
pdt = open (r'''C:\Users\User\Desktop\csmp3\presence-dataset-training-stemmed.csv''',"w")
matWriter = csv.writer(pdt,delimiter = ',')
for i in range (1,2): #45252
processed_email = open(r'''C:\Users\User\Desktop\csmp3\processed\processed'''+str(i)+'''.txt''',"r")
presence_array = cv.transform(processed_email)
matWriter.writerow(presence_array)
processed_email.close()
pdt.close()
This is part of a Spam Filtering using Naive Bayes project and our data set is rather large. I'm hoping to use this sparse matrix for Bernoulli Naive Bayes' partial fit method. I just can't quite figure out how to load the sparse matrix from the file. I've already tried numpy.loadtxt but it gives me:
ValueError: could not convert string to float