Course
machine-learning-zoomcamp
Question
Is there alternative code that prevents the performance warning from this piece of code:
df_object = df[df.columns[df.dtypes == "object"]] for i in list(df_object.columns.values): df_object[i] = (df_object[i]).astype("str") for j in list(np.unique(df_object[i].values)): df["%s" % j] = (df_object[i] == j).astype("int")
Answer
First the reason for this warning is that the data is continually being added with lots of columns (df["%s" % j] = (df_object[i] == j).astype("int")). Every time a new column is added, the whole dataset is referenced, which requires lots of memory. I suggest you encode the categorical values into a list, or use the function DictVectorizer from sklearn.feature_extraction (you will learn more of this in lecture 3)
For the first option, here's the code:
df_object = df[df.columns[df.dtypes == "object"]].copy() df_num = df[np.array(list(df.columns[df.dtypes == "float"]) + list(df.columns[df.dtypes == "int"]))]
-->This separates the str values
modified_df_object = [] num_col = list(np.zeros (df_object.nunique().values.sum())) for i in range (len(df)): modified_df_object.append(num_col) print (np.array(modified_df_object).shape[1]) columns = [] for i in df_object.columns: for j in ((np.unique(df_object[i].values))): columns.append ("%s_%s" % (i, j)) modified_df_object = list(np.array(modified_df_object).T) row = 0 for i in df_object.columns: for j in ((np.unique(df_object[i].values))): modified_df_object[row] = list((df_object[i] == j).astype("int")) row = row +1 modified_df_object = list(np.array(modified_df_object).T)
-->This encode the data inside the df_object
df = pd.concat([modified_df_object, df_num], axis= 1) --> Merge two datasets together.
Checklist
Course
machine-learning-zoomcamp
Question
Is there alternative code that prevents the performance warning from this piece of code:
df_object = df[df.columns[df.dtypes == "object"]] for i in list(df_object.columns.values): df_object[i] = (df_object[i]).astype("str") for j in list(np.unique(df_object[i].values)): df["%s" % j] = (df_object[i] == j).astype("int")Answer
First the reason for this warning is that the data is continually being added with lots of columns (
df["%s" % j] = (df_object[i] == j).astype("int")). Every time a new column is added, the whole dataset is referenced, which requires lots of memory. I suggest you encode the categorical values into a list, or use the function DictVectorizer from sklearn.feature_extraction (you will learn more of this in lecture 3)For the first option, here's the code:
df_object = df[df.columns[df.dtypes == "object"]].copy() df_num = df[np.array(list(df.columns[df.dtypes == "float"]) + list(df.columns[df.dtypes == "int"]))]-->This separates the str values
modified_df_object = [] num_col = list(np.zeros (df_object.nunique().values.sum())) for i in range (len(df)): modified_df_object.append(num_col) print (np.array(modified_df_object).shape[1]) columns = [] for i in df_object.columns: for j in ((np.unique(df_object[i].values))): columns.append ("%s_%s" % (i, j)) modified_df_object = list(np.array(modified_df_object).T) row = 0 for i in df_object.columns: for j in ((np.unique(df_object[i].values))): modified_df_object[row] = list((df_object[i] == j).astype("int")) row = row +1 modified_df_object = list(np.array(modified_df_object).T)-->This encode the data inside the df_object
df = pd.concat([modified_df_object, df_num], axis= 1)--> Merge two datasets together.Checklist