OER·harvester

← Back to the library
Zenodo PDF resource

Artificial intelligence (AI) with It's Applications

The book "Artificial Intelligence (AI) with It's Applications" provides a comprehensive insight into the field of AI, exploring its fundamental principles, modern applications, and future potential. It serves as a valuable resource for students, researchers, and professionals looking to understand AI’s role in shaping industries and everyday life. The book begins with an introduction to Artificial Intelligence , cov…

Licence
OPEN CC-BY-4.0
Authors
Dr. Dipikaben Umakant Thakar, Mrs. PL. Natchiammai, Dr. R. J. Kavitha…
Published
2025-03-18 · Zenodo
Language
eng
Length
66700 words
Type
narrative text
Open ↗ Download Open original ↗

Normalizing the features for stable and fast training.

scaler = StandardScaler()

X = scaler.fit_transform(X)

X_val = scaler.transform(X_val)

Now let’s train some state-of-the-art machine learning models and compare them which fit better with our data.

models = [LogisticRegression(), XGBClassifier(), SVC(kernel='rbf')]

for model in models:

model.fit(X, Y)

print(f'{model} : ')

print('Training Accuracy : ', metrics.roc_auc_score(Y, model.predict(X)))

print('Validation Accuracy : ', metrics.roc_auc_score(Y_val, model.predict(X_val)))

print()

LogisticRegression() :

Training Accuracy : 0.8664717348927876

Validation Accuracy : 0.782258064516129

XGBClassifier() :

Training Accuracy : 1.0

Validation Accuracy : 0.7491039426523298

SVC() :

Training Accuracy : 0.9405458089668616

Validation Accuracy : 0.8042114695340501

Model Evaluation

From the above accuracies, we can say that Logistic Regression and SVC() classifier perform better on the validation data with less difference between the validation and training data. Let’s plot the confusion matrix as well for the validation data using the Logistic Regression model.

from sklearn.metrics import ConfusionMatrixDisplay

ConfusionMatrixDisplay.from_estimator(models[0], X_val, Y_val)

plt.show()

Output

6.4 Wine Quality Prediction

Importing libraries and Dataset:

import numpy as np import pandas as pd import matplotlib.pyplot as plt import seaborn as sb

from sklearn.model_selection import train_test_split from sklearn.preprocessing import MinMaxScaler from sklearn import metrics from sklearn.svm import SVC from xgboost import XGBClassifier from sklearn.linear_model import LogisticRegression

import warnings warnings.filterwarnings('ignore')

Now let’s look at the first five rows of the dataset.

df = pd.read_csv('winequality.csv')

print(df.head())

Output

Let’s explore the type of data present in each of the columns present in the dataset.

df.info()

Output

Now we’ll explore the descriptive statistical measures of the dataset.

df.describe().T

Output

Exploratory Data Analysis

EDA is an approach to analysing the data using visual techniques. It is used to discover trends, and patterns, or to check assumptions with the help of statistical summaries and graphical representations. Now let’s check the number of null values in the dataset columns wise.

df.isnull().sum()

Let’s impute the missing values by means as the data present in the different columns are continuous values.

for col in df.columns:

if df[col].isnull().sum() > 0:

df[col] = df[col].fillna(df[col].mean())

df.isnull().sum().sum()

Output

Let’s draw the histogram to visualise the distribution of the data with continuous values in the columns of the dataset.

df.hist(bins=20, figsize=(10, 10))

plt.show()

Now let’s draw the count plot to visualise the number data for each quality of wine.

plt.bar(df['quality'], df['alcohol'])

plt.xlabel('quality')

plt.ylabel('alcohol')

plt.show()

Output

There are times the data provided to us contains redundant features they do not help with increasing the model’s performance that is why we remove them before using them to train our model.