I'm trying to populate a dataframe based on a class label and images in a folder.
I have a folder have over 10,000 images with the following name structure: ['leaflet_10000_1.jpg', 'leaflet_10000_2.jpg', 'leaflet_10001_1.jpg', 'leaflet_10001_2.jpg', 'leaflet_10002_1.jpg', 'leaflet_10002_2.jpg', 'leaflet_10003_1.jpg', 'leaflet_10003_2.jpg'
And an accompanying csv file of the structure:
ID,Location,Party,Representative/Candidate,Date
1000,Glasgow North,Liberal Democrats,,02-Apr-10
1001,Erith and Thamesmead,Labour Party,,02-Apr-10
I want to create a new csv file which has the paths of all the images for a said Party. I can separate a certain party from the full csv file using the commands:
df_ = df.loc[df["Party"] == "Labour Party"]
This will give me the party I am interested in, but how do I create a FULL list of all images associated with it.. from the image list shared above, it can be seen that ID 1001 has 2 images associated with it.. this is not a fixed number, some ID's have 3 to 5 images associated with them.
How do I get this new dataframe populated with all the required paths?
My thought process is to apply str.split(name, '_') on each file name and then search every ID against all the results but where to go from there?