OER·harvester

← Back to the library
Zenodo PDF resource

Foundations of Artificial Intelligence & Machine Learning

Licence
OPEN CC-BY-4.0
Authors
Nidhi Sharma, Honey Singh, Ajay Sharma, Deepak Dagar
Published
2026-07-28 · Zenodo
Language
eng
Length
37166 words
Type
narrative text
Open ↗ Download Open original ↗
Need for Text Processing

Human language is naturally unstructured. Different people may express the same idea using different words and sentence structures. For example: AI is powerful. Artificial Intelligence is very powerful. Although both sentences express similar meanings, computers initially treat them as different patterns. Text processing helps normalize and organize such information for analysis. Tokenization Tokenization is one of the most basic NLP techniques. It involves dividing text into smaller units called tokens. These tokens may represent:

  • words
  • sentences
  • phrases For example: Machine Learning is transforming technology. may be divided into: Machine | Learning | is | transforming | technology Tokenization helps systems process language step by step.

Figure 8.1: Tokenization Process

The figure illustrates how tokenization breaks textual information into smaller units that can be processed by NLP systems. Stop Word Removal Certain words occur very frequently in language but contribute little meaning during analysis.

Examples include:

  • is
  • the
  • and
  • of These words are called stop words.

Removing stop words helps reduce unnecessary processing and improves analysis efficiency. Stemming and Lemmatization Words often appear in multiple forms. For example:

  • playing

  • played

  • plays all originate from the root word “play.” Stemming and lemmatization techniques reduce words to their base forms, helping systems treat related words similarly. This improves consistency in language analysis. Text Vectorization Machine Learning models cannot directly understand textual information. Therefore, text must be converted into numerical form. This process is known as vectorization. Common methods include:

  • Bag of Words

  • TF-IDF

  • word embeddings

These techniques represent words mathematically so that machine learning algorithms can process language data. Bag of Words Model The Bag of Words approach represents text based on word frequency. Instead of focusing on grammar or sequence, the model simply counts how often words appear within documents. Although simple, this technique is useful for:

  • text classification
  • spam detection
  • sentiment analysis