Flatten a time series based on its correlation with another time series
22:33 13 Jul 2023

I would like to flatten a time series column based on its correlation with another column, I am trying a few approach and they are not working...

import pandas as pd
import numpy as np
import random
from datetime import datetime, timedelta

start_date = datetime(2023, 1, 1)
total_rows = 100

time_index = [start_date + timedelta(days=i) for i in range(total_rows)]

column1_values = np.random.randint(0, 100, total_rows)
column2_values = np.random.randint(0, 10, total_rows)

df = pd.DataFrame({'Time': time_index, 'Column1': column1_values, 'Column2': column2_values})

correlation = df['Column1'].corr(df['Column2'])

if abs(correlation) < 0.5:
    df['Column2'] = df['Column2'] - np.mean(df['Column2'])

# Approach 2
#    df['Column2'] = df['Column2']/np.max(df['Column2'])

Trying to make timeseries at column2 flat if no correlation exists.

python pandas dataframe time-series