OER·harvester

← Back to the library
Zenodo PDF resource

Artificial intelligence (AI) with It's Applications

The book "Artificial Intelligence (AI) with It's Applications" provides a comprehensive insight into the field of AI, exploring its fundamental principles, modern applications, and future potential. It serves as a valuable resource for students, researchers, and professionals looking to understand AI’s role in shaping industries and everyday life. The book begins with an introduction to Artificial Intelligence , cov…

Licence
OPEN CC-BY-4.0
Authors
Dr. Dipikaben Umakant Thakar, Mrs. PL. Natchiammai, Dr. R. J. Kavitha…
Published
2025-03-18 · Zenodo
Language
eng
Length
66700 words
Type
narrative text
Open ↗ Download Open original ↗

Convert 'object' columns to numerical if they represent numbers

for col in df.columns:

if df[col].dtype == 'object':

try:

df[col] = pd.to_numeric(df[col], errors='coerce') # Convert to numeric, replace non- convertibles with NaN

except:

pass # Skip columns that cannot be converted

plt.figure(figsize=(12, 12))

sb.heatmap(df.corr() > 0.7, annot=True, cbar=False)

plt.show()

From the above heat map we can conclude that the ‘total sulphur dioxide’ and ‘free sulphur dioxide‘ are highly correlated features so, we will remove them.

df = df.drop('total sulfur dioxide', axis=1)

Model Development

Let’s prepare our data for training and splitting it into training and validation data so, that we can select which model’s performance is best as per the use case. We will train some of the state of the art machine learning classification models and then select best out of them using validation data. df['best quality'] = [1 if x > 5 else 0 for x in df.quality]

We have a column with object data type as well let’s replace it with the 0 and 1 as there are only two categories.

df.replace({'white': 1, 'red': 0}, inplace=True)

After segregating features and the target variable from the dataset we will split it into 80:20 ratio for model selection.

features = features.fillna(features.mean())

features = df.drop(['quality', 'best quality'], axis=1)

target = df['best quality']

xtrain, xtest, ytrain, ytest = train_test_split(

features, target, test_size=0.2, random_state=40)