Human language is naturally unstructured. Different people may express the same idea using different words and sentence structures. For example: AI is powerful. Artificial Intelligence is very powerful. Although both sentences express similar meanings, computers initially treat them as different patterns. Text processing helps normalize and organize such information for analysis. Tokenization Tokenization is one of the most basic NLP techniques. It involves dividing text into smaller units called tokens. These tokens may represent:
words
sentences
phrases
For example: Machine Learning is transforming technology. may be divided into: Machine | Learning | is | transforming | technology Tokenization helps systems process language step by step.
Figure 8.1: Tokenization Process
The figure illustrates how tokenization breaks textual information into smaller units that can be processed by NLP systems. Stop Word Removal Certain words occur very frequently in language but contribute little meaning during analysis.
Examples include:
is
the
and
of
These words are called stop words.
Removing stop words helps reduce unnecessary processing and improves analysis efficiency. Stemming and Lemmatization Words often appear in multiple forms. For example:
playing
played
plays
all originate from the root word “play.” Stemming and lemmatization techniques reduce words to their base forms, helping systems treat related words similarly. This improves consistency in language analysis. Text Vectorization Machine Learning models cannot directly understand textual information. Therefore, text must be converted into numerical form. This process is known as vectorization. Common methods include:
Bag of Words
TF-IDF
word embeddings
These techniques represent words mathematically so that machine learning algorithms can process language data. Bag of Words Model The Bag of Words approach represents text based on word frequency. Instead of focusing on grammar or sequence, the model simply counts how often words appear within documents. Although simple, this technique is useful for:
DCMI Metadata Terms. Dublin Core has no element that separates the original file from the text extracted out of it, and none for LOM's educational characterisation. Both survive here as provenance statements and in the record itself, not in the projection.
the standard ↗
The groups below are this library's, for reading. DCMI Terms itself has no categories; each term keeps its standard name.
Works this one cites, when the source declares them as relations. What its text links to and its reference list cites is inferred, and stands under it apart.
Retrieved from Zenodo on 2026-10-09 in response to the search string “("artificial intelligence" OR "machine learning" OR "generative AI" OR "deep learning" OR "reinforcement learning" OR "large language model" OR "AI") AND ("AI concepts" OR "types of AI" OR "AI fundamentals" OR "recognizing AI" OR "recognising AI" OR "general versus narrow AI" OR "narrow AI" OR "general AI" OR "machine intelligence" OR "AI strengths and weaknesses" OR "traditional software" OR "rule-based systems" OR "introduction to AI" OR "introduction to artificial intelligence" OR "artificial intelligence introduction" OR "AI primer" OR "foundations of artificial intelligence" OR "overview of AI" OR "understanding AI" OR "history of AI" OR "AI essentials" OR "AI terminology" OR "metaphors for AI" OR "AI fundamental concepts" OR "AI key concepts" OR "philosophy of AI" OR "critical AI literacy")”. Zenodo served the resource and is not asserted to be its publisher or author.
Text extracted from pdf to Markdown by pdf-inspector; the original is retained unchanged beside it.
Where it was collected from, what was converted, and what container it came out of — the custody statements that would otherwise be mistaken for authorship.
IEEE 1484.12.1 Learning Object Metadata. LOM has no element for an SPDX identifier or a licence URI, so both are written into 6.3 Rights.Description. Flattening this record into simple Dublin Core would lose more again, which is why the two projections exist side by side rather than one being generated from the other.
the standard ↗
Role, entity and date per declared contribution. A role outside LOM's vocabulary is reported in the entry's description instead.
3 Meta-metadata
4/4
Identifier3.1
this engine
resource_id
URI: tag:aim-pro.eu,2026:oer/2237c62a2a9c/record
The identifier of this metadata record — the resource's own, with /record after it, because the record is a description of the resource and not the resource.
Contribute3.2
this engine
source
AIM-PRO WP3 OER harvester (Zenodo) — creator
Who generated this record and when — a statement about the record, not about the resource.
Yes unless the licence reserves nothing — attribution is a restriction. The conditions after the dash are the licence gate's reading; the export carries LOM's bare term.
Markdown extracted from the original by pdf-inspector
The container a file was found inside, and the Markdown extracted from the original. What a lab requires, and the lab a component belongs to, are inferred and stand apart.
8 Annotation
0/3
Entity8.1
not collected — this library does not fill it
Comments on the resource's educational use, by whoever made them. The platform's review grades competencies, which are classification (9), and writes no comment here.
Date8.2
not collected — this library does not fill it
Description8.3
not collected — this library does not fill it
9 Classification
0/4
Purpose9.1
not collected — this library does not fill it
Empty in the record: no source declares a competency. The alignment reads the resource for them and stands beside the record, never in it, and a taxon path derived from the search string that found it would be a claim about the query.
Taxon path9.2
not collected — this library does not fill it
Where the competency framework goes. Empty in the record for the reason above.
Description9.3
not collected — this library does not fill it
Keyword9.4
not collected — this library does not fill it
What could not be established5
Where the source's metadata could not be carried
over as it was — missing, contradictory, with no matching term in the
standard, restructured, or taken from the repository — and what was done
instead. Without these notes, an empty element would look like something
the harvester missed.
Status
Field
Why
Not available
description
the source published no description or abstract
Not available
subjects
the source published no keywords
Not available
rights_holder
no rights holder is named at source; the licence is recorded without one rather than attributed to the platform that served it
Not available
publisher
the source named no publisher of the work; where it was collected from is recorded as collection provenance instead, which is a different claim
Not available
educational
the source declared no educational metadata — no resource type, audience, context, difficulty or learning time. Nothing here estimates them