Large datasets sometimes contain unnecessary or highly correlated features that contribute little to prediction quality. High-dimensional data may create problems such as:
slow model training
increased memory usage
overfitting
difficulty in visualization
Dimensionality reduction helps simplify data and makes Machine Learning systems more efficient.
Figure 6.2: High-Dimensional Data Reduction
The figure illustrates how dimensionality reduction techniques simplify complex datasets by reducing the number of variables while preserving important information. Principal Component Analysis (PCA) Principal Component Analysis is one of the most widely used dimensionality reduction techniques. PCA transforms complex datasets into smaller sets of important variables called principal components. These components capture the maximum possible information from the original data. The technique helps:
reduce complexity
improve visualization
speed up training
remove redundant information
PCA is commonly used in facial recognition systems, image compression, and pattern recognition applications. Advantages of Dimensionality Reduction Dimensionality reduction improves Machine Learning efficiency in several ways. It helps:
reduce computational cost
simplify datasets
improve training speed
reduce overfitting
Smaller datasets are also easier to visualize and interpret. Challenges in Dimensionality Reduction Reducing dimensions may sometimes result in information loss. If important features are removed accidentally, model performance may decrease. Selecting the correct reduction technique therefore requires careful analysis of the dataset and problem type.
6.3 Association Rule Learning
Association Rule Learning is an unsupervised learning technique used to discover relationships and hidden connections between items in large datasets. The main purpose of this method is to identify patterns showing how different items or events are related to each other. This technique is commonly used in business analytics, especially in market basket analysis, where companies study customer purchasing behavior.
For example, a supermarket may discover that customers who purchase bread and butter often also buy milk. Such hidden relationships help businesses improve product placement, marketing strategies, and recommendation systems. Association Rule Learning focuses on identifying frequent item combinations within datasets. Understanding Association Rules Association rules are generally represented in the form: If A occurs, then B is likely to occur. For example: If a customer buys a laptop, they may also purchase a mouse. These relationships are not fixed rules but probability-based patterns identified from historical data. Association Rule Learning helps organizations understand customer behavior and decision patterns more effectively.
Figure 6.3: Association Between Purchased Items
The figure illustrates how association rule learning identifies relationships between items frequently purchased together in transactional datasets. Apriori Algorithm The Apriori Algorithm is one of the most commonly used association rule learning methods. It identifies frequent item combinations by analyzing transaction datasets repeatedly and selecting patterns that occur frequently. The algorithm works on the principle that:
if an item combination appears frequently,
its smaller subsets are also likely to appear frequently.
Apriori is widely used in:
DCMI Metadata Terms. Dublin Core has no element that separates the original file from the text extracted out of it, and none for LOM's educational characterisation. Both survive here as provenance statements and in the record itself, not in the projection.
the standard ↗
The groups below are this library's, for reading. DCMI Terms itself has no categories; each term keeps its standard name.
Works this one cites, when the source declares them as relations. What its text links to and its reference list cites is inferred, and stands under it apart.
Retrieved from Zenodo on 2026-10-09 in response to the search string “("artificial intelligence" OR "machine learning" OR "generative AI" OR "deep learning" OR "reinforcement learning" OR "large language model" OR "AI") AND ("AI concepts" OR "types of AI" OR "AI fundamentals" OR "recognizing AI" OR "recognising AI" OR "general versus narrow AI" OR "narrow AI" OR "general AI" OR "machine intelligence" OR "AI strengths and weaknesses" OR "traditional software" OR "rule-based systems" OR "introduction to AI" OR "introduction to artificial intelligence" OR "artificial intelligence introduction" OR "AI primer" OR "foundations of artificial intelligence" OR "overview of AI" OR "understanding AI" OR "history of AI" OR "AI essentials" OR "AI terminology" OR "metaphors for AI" OR "AI fundamental concepts" OR "AI key concepts" OR "philosophy of AI" OR "critical AI literacy")”. Zenodo served the resource and is not asserted to be its publisher or author.
Text extracted from pdf to Markdown by pdf-inspector; the original is retained unchanged beside it.
Where it was collected from, what was converted, and what container it came out of — the custody statements that would otherwise be mistaken for authorship.
IEEE 1484.12.1 Learning Object Metadata. LOM has no element for an SPDX identifier or a licence URI, so both are written into 6.3 Rights.Description. Flattening this record into simple Dublin Core would lose more again, which is why the two projections exist side by side rather than one being generated from the other.
the standard ↗
Role, entity and date per declared contribution. A role outside LOM's vocabulary is reported in the entry's description instead.
3 Meta-metadata
4/4
Identifier3.1
this engine
resource_id
URI: tag:aim-pro.eu,2026:oer/2237c62a2a9c/record
The identifier of this metadata record — the resource's own, with /record after it, because the record is a description of the resource and not the resource.
Contribute3.2
this engine
source
AIM-PRO WP3 OER harvester (Zenodo) — creator
Who generated this record and when — a statement about the record, not about the resource.
Yes unless the licence reserves nothing — attribution is a restriction. The conditions after the dash are the licence gate's reading; the export carries LOM's bare term.
Markdown extracted from the original by pdf-inspector
The container a file was found inside, and the Markdown extracted from the original. What a lab requires, and the lab a component belongs to, are inferred and stand apart.
8 Annotation
0/3
Entity8.1
not collected — this library does not fill it
Comments on the resource's educational use, by whoever made them. The platform's review grades competencies, which are classification (9), and writes no comment here.
Date8.2
not collected — this library does not fill it
Description8.3
not collected — this library does not fill it
9 Classification
0/4
Purpose9.1
not collected — this library does not fill it
Empty in the record: no source declares a competency. The alignment reads the resource for them and stands beside the record, never in it, and a taxon path derived from the search string that found it would be a claim about the query.
Taxon path9.2
not collected — this library does not fill it
Where the competency framework goes. Empty in the record for the reason above.
Description9.3
not collected — this library does not fill it
Keyword9.4
not collected — this library does not fill it
What could not be established5
Where the source's metadata could not be carried
over as it was — missing, contradictory, with no matching term in the
standard, restructured, or taken from the repository — and what was done
instead. Without these notes, an empty element would look like something
the harvester missed.
Status
Field
Why
Not available
description
the source published no description or abstract
Not available
subjects
the source published no keywords
Not available
rights_holder
no rights holder is named at source; the licence is recorded without one rather than attributed to the platform that served it
Not available
publisher
the source named no publisher of the work; where it was collected from is recorded as collection provenance instead, which is a different claim
Not available
educational
the source declared no educational metadata — no resource type, audience, context, difficulty or learning time. Nothing here estimates them