printing the confusion matrix
LABELS = ['Normal', 'Fraud']
conf_matrix = confusion_matrix(yTest, yPred)
plt.figure(figsize =(12, 12))
sns.heatmap(conf_matrix, xticklabels = LABELS,
yticklabels = LABELS, annot = True, fmt ="d");
plt.title("Confusion matrix")
plt.ylabel('True class')
plt.xlabel('Predicted class')
plt.show()
6.3 Autism Prediction
Importing Libraries and Dataset
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sb
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import LabelEncoder, StandardScaler
from sklearn import metrics
from sklearn.svm import SVC
from xgboost import XGBClassifier
from sklearn.linear_model import LogisticRegression
from imblearn.over_sampling import RandomOverSampler
import warnings
warnings.filterwarnings('ignore')
The dataset will now be loaded into a pandas data frame, and the first five rows will be printed. (Collect data and store the file name as dats.csv)
df = pd.read_csv('train.csv') print(df.head())
Output
Now let’s check the size of the dataset,
df.shape
Output
(800, 22)
Let’s check which column of the dataset contains which type of data,
df.info()
Output
As per the above information regarding the data in each column we can observe that there are no null values,
df.describe().T
Output
Data Cleaning
Data obtained from primary sources, often referred to as raw data, requires extensive preprocessing before it can be used for analysis or modeling. This process, known as data cleaning, involves several essential steps, including:
Outlier Removal: Identifying and eliminating data points that deviate significantly from the dataset's overall pattern. v Null Value Imputation: Handling missing values by filling them with appropriate estimates or removing them to ensure data integrity. v Resolving Discrepancies: Addressing inconsistencies or errors in the data to maintain accuracy and reliability in analysis.
df['ethnicity'].value_counts()
Output
In the above two outputs we can observe some ambiguity that there are ‘?’, ‘others’, and ‘Others’ which all must be the same as they are unknown or we can say that null values have been substituted with some indicator.
df['relation'].value_counts()
Output
The same is the case with this column so, let’s clean this data, and along with this let’s convert ‘yes’ and ‘no’ to 0 and 1.
df = df.replace({'yes':1, 'no':0, '?':'Others', 'others':'Others'})
Now we have cleaned the data a bit to derive insights from it.
Exploratory Data Analysis
EDA is an approach to analyzing the data using visual techniques. It is used to discover trends, and patterns, or to check assumptions with the help of statistical summaries and graphical representations. Here we will see how to check the data imbalance and skewness of the data.
plt.pie(df['Class/ASD'].value_counts().values, autopct='%1.1f%%')
plt.show()
Output
The dataset we have is highly imbalanced. If we will train our model using this data then the model will face a hard time predicting the positive class which is our main objective here to predict whether a person has autism or not with high accuracy.
ints = []
objects = []
floats = []
for col in df.columns:
if df[col].dtype == int:
ints.append(col)
elif df[col].dtype == object:
objects.append(col)
else:
floats.append(col)
The ‘ID’ column will contain a unique value for each of the rows and for the column ‘Class/ASD’ we have already analyzed its distribution so, that is why they have been removed in the above code.