OER·harvester

← Back to the library
arXiv HTML resource

Foundations of Artificial Intelligence Frameworks: Notion and Limits of AGI

Within the limited scope of this paper, we argue that artificial general intelligence cannot emerge from current neural network paradigms regardless of scale, nor is such an approach healthy for the field at present. Drawing on various notions, discussions, present-day developments and observations, current debates and critiques, experiments, and so on in between philosophy, including the Chinese Room Argument and G…

Licence
OPEN CC-BY-4.0
Authors
Khanh Gia Bui
Published
2025-11-23 · arXiv
Language
en
Length
36705 words
Type
narrative text

Cites 147 works

inferred
Open original ↗

2 Understanding Artificial General Intelligence (AGI)

We have understood the general notion of artificial intelligence, in one form or another. It is then naturally that we extend such conversation to the notion of Artificial General Intelligence, or AGI in short. In essence, what does an AGI constitute? The AGI notion relies partially on the concept of fragmented intelligence. This was first apparent by the apparatus of the Turing test (Turing (1950)), that suggest quantifying different human capabilities for determining a machine’s ability to be ‘human’, and as if the machine can surpass human in such range. This view is supported in a different form in Gardner (1983) book The Theory of Multiple Intelligences, of which again posit that intelligence exists in different forms, and not a singular object of quantification. Given such, Artificial General Intelligence posits that we are able to construct, and would be able to construct, an AGI with all of such capabilities that one can consider to be human, with consciousness, with intelligent learning capabilities, with thoughts, and so on. A smaller camp, yet vocal, posits further that the structure of LLM, large language models, or agentic AI (Sapkota et al. (2025); Derouiche et al. (2025); Schneider (2025); Wei and others (2025); Raza et al. (2025)), will be able to achieve this goal.

Nevertheless, the issue that plagues such notion is that the term itself is not fully understood, nor there exists any given consensus on what is the acceptable form that constitute the baseline definition of an AGI. Technically speaking, AGI can be attributed to the fact that many AI systems are constructed in fragments, of which for example, computer vision, language processing, signal processing, classification analysis, robotic spatial movements’ extrapolation, and else, all of which then if can be combined into one, would inherently make a human-like form of intelligence. In between such, LLM, per its role as the language processor, is deemed to stand in between such. However, the definition in such term itself has its own fallacy. Suppose that the AGI $\mathsf{AGI}_{x}$ has multitude of ability that is inherently of its own domain and environment - computers, operating systems, software framework, etc, with full understanding and exploratory sense of such, but has no spatial movement, no sensors and the like that can attribute it of features of which human exhibits in natural sense. Would that disqualify it as an artificial intelligence, or just have to reclassify it into a different environment?

For now, we need to formalize what is AGI actually saying, in context. Or rather, a definition that is informal per its natural topic. What can be intrinsically defined to be AGI. Based on our current understanding, AGI can be defined using the basis of the fragmented intelligence theory, and by the previous statement of AI. Simply speak, of Conjecture 1.1 and Conjecture 1.2, fragmented intelligence begins with considering all construct of which fits certain limited set of qualities, one or more, but distinct. Then, AGI is the plateau that there exists such construct that can fit all qualities of its quantified notion.

Definition 2.1 (Artificial general intelligence)**.**

Based on Conjectures 1.1 and 1.2 and the fragmented intelligence scheme, define a quality basis $\mathsf{AI}_{n}=\{A_{1},\dots,A_{n}\}$ of arbitrary given qualities specified of an artificial intelligence construct. Then, a construct $\mathsf{AGI}\in\mathsf{AI}_{n}$ is called an artificial general intelligence, if $\mathsf{AGI}$ can be represented as

$$ \mathsf{AGI}=\alpha_{1}A_{1}+\alpha_{2}A_{2}+\dots+\alpha_{n}A_{n},\quad\alpha_{i}>0,i=1,\dots,n $$

That is, $\mathsf{AGI}$ expresses every given qualities of the quality basis, to a given evaluation degree $\alpha_{i}$ associated with each quality. Thus, we say $\mathsf{AGI}$ spans the entire quality space.

Here, the definition is arbitrary on purpose, as for many evaluation metrics, testing setup, or different quality that is regarded of AGI, for example, the updated version of the Turing test versus the traditional one, we resolve to such definition by the action of generality. Of course, this begs the question if certain quality basis $\mathsf{AI}_{n}^{i}$ is better than others, and the answer is yes, but subjectively — since the qualities themselves usually cannot be quantified exactly. Nevertheless, we posit that such construct can exist, from the construct’s structural system itself, that when mapped of its operation onto the arbitrary basis of choice, yields such resultant observation. There then exists the fundamental problem of such definition — for such artificial intelligence construct to exist, within the fragmented theory of intelligence, then there must then exist internal connections between its components, and the expression oward separated qualities’ evaluations. Because such processing connector also retains its own quality specification, in such framework, aside from being a hidden, internal component of the model (of which then handle such connection arbitrarily and cannot be evaluated simply), it must then be of certain construct that is present within the basis. Such is usually reserved for large language model, as it is on the basis of natural language qualities, and is naturally the candidate in such regard as the main focus of core component for determining the feasibility of an AGI. Such framework is then called, the agentic AI framework. By certain arbitrary consideration, we can say current AGI, of certain qualities such as the Turing test, is already an AGI, within the above definition. Nevertheless, by such basis, we can set the bar so low that such AGI is unbearably naive, or set the bar too high such that the qualities in consideration is considered an infeasible event by ‘superhuman standard’. Or by a list of all details, measurable set of qualities with overlapping, yet does not and could not usually take into account of the internal structures of the model and the inherent natural abstraction, or emergence of whatever definition being the arbitrary emergence, is for such given construct. In a sense, we already have AGI. The problem is how effective it is. It is perhaps natural that we re-question the anecdotal concept of LLM, and of the basis, and of the entirety of connector framework, and extrapolate from such — what can be different from LLM and such current understanding, and what is inherently wrong and right with it?

This view is debated fiercely, as we have introduced it before, and also because of its intrinsic nature of the field itself. For the example, the arbitrary of such consideration is a problem as there would imply the problem of inconsistency in between criteria, qualities, and so on, of which has been surprisingly apparent as many times standards and qualities have been modified and pushed forward in the case of determining ”What is then even A(G)I?”. However, there are also voices aside from such camp, telling us that current structures of artificial intelligence understanding, machine learning anecdotal knowledge, are enough to construct such AGI. Some further argue that the current framework is complete, of which provide exponential growth of such toward AGI in such matter. What is the correct answer to such question? Of this paper, the current stance is no. We then have to present logics and evidences to back such answer up.

2.1 The fallacy of defining intelligence

We mentioned the notion of the criterion of intelligence. However, what should we define it? How should we know to even evaluate it, is a very hard question even that we did not (or unable to) fully realize yet, then what we want to do with it? This question is where a lot of things in the artificial intelligence research was based upon. For example, the (Total) Turing Test in which outlines possible outlook for intelligence, for capabilities that then defines the fields in which we are having nowadays, for example, computer vision for the capability of visual perception, natural language processing (NLP) for the capacity of language, and more. We also have various conceptual criterions in which people have been suggesting about the model of the intelligent being, for example, various set of criterions that outlines and includes even consciousness, some suggest behavioural conditions, some goes for the exhibition of chain of thoughts, and some even goes further than that, which is perhaps irrelevant aside from mentioned for example. Overall, it is perhaps a mess.

We still do not know what to come of criteria, or rather, in the quest of producing intelligence, we base ourselves onto it too much. As a species capable of intelligence and more sophisticated notion, we have the basis, and the advantage of being able to examine ourselves. By that, eventually, as the highest example of intelligent being, we use ourselves as standard, for examples, of psychology, neurological behaviourism, neuroscience, applied onto the quest of going for artificial intelligence. Hence, there exists the total Turing test, and there exists the conflicts between various definitions and criterion of artificial intelligence. A mistake perhaps has been made, doubtfully so that one did not realize of such. While it is said that AI researcher has been working on, or at least researching on the general notion of artificial intelligence principles, it is, in fact, not so much of a principle, as we do not realize, yet, that what we are doing is still the act of mimicking ourselves - creating a plane by replicating a bird. By phenomenologically absorb and construct architectures, models on the higher-level surface of what artificial intelligence constitute, the deeper construct is still non-existent. By copying the apparent capabilities of human and related intelligence being, biological rather than not, the core of which those behaviours occur, and facilitate the organs and observations made is perhaps, manifested, simply does not exist. Ironically, while being too strict, wrongfully abhorrent to the fallacy of themselves, and too resistant to changes, symbolic approach got one of the right thing. If there exists intelligence, then it must be universal by virtue. That is, you cannot argue that alien from another universe is not intelligent, because they do not satisfy one of the criteria of the Turing test, just because such notion does not exist in such universe.

Historically, it does not prevent people in the field of artificial intelligence to search for intelligence in their bold claims, of which their arguments make clear that they are phenomenologically mimicking the attributes of which is of human, without the substances underneath:

Once (my intelligent system) OSCAR is fully functional, the argument from analogy will lead us inexorably to attribute thoughts and feelings to OSCAR with precisely the same credentials with which we attribute them to human beings. Philosophical arguments to the contrary will be passé. (Pollock, 1995)

From such view, I objectively don’t think we should, or we could define artificial intelligence, at least of this particular stage that we are in. Philosophically, being an armchair philosopher would not help in pursuing such notion, yet again because we are arguing on the basis of our own existence, and not the subject’s matter viewpoint. There are problems related to it, also, of such that the mind and consciousness is arguably debatable in every given sense, of which no one seems to agree on the mundane notion that intelligence and consciousness come from chemical and the weird ’quantum effect’ that would be then believed to be. And, truth to be told, we are not even endorsing such direction. In actuality, we don’t even know what is intelligent, and also don’t even know what can be of artificially made rather than matching mathematics.

On the flip side, computationally and neuroscientifically, the lack of formal treatment and overall encompassing knowledge conjunctions plague the construction and foremost attempt to do anything, simply because too many things have been said yet none can unify them together. Such is also to say different directions and different methodologies being conducted, yet they are so distinctively separated to be unable to conform one to another, despite them taking on the same object. Furthermore, there are a lot of assumptions given in computational theory, and the overall application thereof. As for anything, assumptions can be broken, and reinforced, for whatever it is being inconsistent as a virtue.

It is wise to remember that, for now with neuroscience being not advanced enough and in a perhaps different direction from what can be seen, while certainly for empirical science we can utilize neuroscience’s knowledge, we should not take in the philosophical arguments and ’idea’, including computational theory of mind. For empirical neuroscience, it is also not the fully-encompassing field that observe the brain from every angle, and observe consciousness of everything if ever, at least of the present. And, for the philosophical and idealistic view, only one thing can be said about such being ”the lines on the map is made up”.

2.2 The statistical model critique

AI, as of current, have shifted its structural formations and logical acumen back to mathematics. That is, right now, artificial intelligence looks like nothing but the thing it is originated from, but rather statistical models on data. This view is iterated in several literatures, of which we list of Penrose and Severino (1997); Peyré (2025) and Kutyniok (2022). While not downplaying the role of mathematics in its application and advancements of the field, and successful creations such as many models have been created, one question remains - is it actually artificial intelligence, or just a statistical model trained and probabilistically interpreted to mimic certain aspect or tasks of which is considered intelligent? This is also the stance that Searle (1980a) argued against, of which produced the long-standing Chinese Room Argument. We can summarize the argument simply as followed.

CRA is based on a thought-experiment in which Searle himself stars. He is inside a room; outside the room are native Chinese speakers who don’t know that Searle is inside it. Searle-in-the-box, like Searle-in-real-life, doesn’t know any Chinese, but is fluent in English. The Chinese speakers send cards into the room through a slot; on these cards are written questions in Chinese. The box, courtesy of Searle’s secret work therein, returns cards to the native Chinese speakers as output. Searle’s output is produced by consulting a rulebook: this book is a lookup table that tells him what Chinese to produce based on what is sent in. To Searle, the Chinese is all just a bunch of — to use Searle’s language — squiggle-squoggles. (Searl, 1980s)

It is notable to point out that Searle’s argument against the current advancement of AI using CRA is particular still effective even in the modern current landscape of artificial intelligence. In fact, this voice resonates with a large pool of people, whether because of the fear of losing identity, or from analytical assessment of current models. Large Language Models and advanced models often fail miserably when changing context or changing setting, losing information or hallucinating and so on for a large spectrum of situations, for imperfections that lies outside its intended encoding. Coincidentally, this also fit the old-age argument that Descartes made on the notion of machine and human, or simply the Descartes’s argument (Descartes (1950)). Descartes’s argument start with the justification of the apparent reactive behaviours observed by human themselves, who at the time, was largely considered to be the only species capable of advanced rational thoughts and processes.

(I)f someone touched it (= the machine) in a particular place, it would ask what one wishes to say to it, or if it were touched somewhere else, it would cry out that it was being hurt, and so on. But it could not arrange words in different ways to reply to the meaning of everything that is said in its presence, as even the most unintelligent human beings can do. [Descartes, 1700]

Here, Descartes argues that in order for human-like robots to acquire intelligence, they have to gain a universal capability to accurately react to any unknown situation that may happen in the environment. However, what machines can do is no more than to respond to a single situation one-on-one via a specific organ, hence, they cannot be considered to have a universal capability that even unintelligent human beings can enjoy.

Continuing, Descartes argues that those machines do not act on their knowledge, but the disposition of organs.

For whereas reason is a universal instrument that can be used in all kinds of situations, these organs need a specific disposition for every particular action. It follows that it is morally impossible for a machine to have enough different dispositions to make it act in every human situation in the same way as our reason makes us act.

The argument is quite clear. Human is universal of the environment. Whereas machine is no more than a combination of abilities that are applicable only to certain situation that the creator could imagine when they built the automated machine. This simply posit, and be relevant of our observations on the current machine learning models and large language models. While they are large and as a result, of great range of capability, they are inevitably hard-coded with what the designer wishes them to do. They are not intelligent, in a sense, so to speak of their capability as to be artificial intelligence of its true ‘general form’. Hence, of this debate, we can simply say, to the disdain of the empiricists themselves, that artificial intelligence of the current structure simply is not distinguishable from the statistical model view.

The same plagues the new-found sensation of agentic AI, of which is believed to obtain general intelligence by cutting and wrapping different small specialized-AI structures together, for example, connecting an LLM to a computer vision system or image recognition for scanning documents. However, by default, such connection is inherently shallow, as there exists only a black-box connection of input-output, process resultant in between those components together, which makes it similar to CRA and Descartes’s argument in question.

2.3 The intelligence model

Along the same line as the statistical model critique, we now change to the perspective of the current architecture of choice. Assuming the current knowledge of artificial intelligence and thereof, within respect to machine learning and other development fields, can we say that we understand, or at least can construct artificial intelligence, and hence the true form called AGI? The answer is both yes and no.

Current understanding of AI, Russell and Norvig (2009), as specified, focus on the self-reflection of the human researchers on themselves. Such reflection are often surface-level, for example, behaviours of which intelligent choice might occur rather than not, situation of which there exists patterns in which the mind choose to operate, and so on. However, this in particular, face heads-on with a problem that even now cannot be explained or go through - the domain problem, or the frame problem. The frame problem is the problem that an AI cannot autonomously distinguish important factors from unimportant ones when it tries to cope with something in a certain situation. The problem arises, for example, when we let AI robots operate in the real world. This problem was proposed by John McCarthy and Patrick J. Hayes in 1969, of which is considered a philosophical problem that cannot be merely reduced to a technical problem. Historically, the problem is narrowly defined for the field of logic-based artificial intelligence. But it was taken up in an embellished and modified form by philosophers of mind, and given a wider interpretation, and hence, is since then applicable to almost all formal system that wishes to call themselves artificial intelligence. We will not cover all of it here, at least for now. For authoritative literature, it is recommended to refer to the Stanford’s article on the frame problem Shanahan (2016), and other literatures Gryz (2013); Seager (2010); Briggs (2014). However, it is indeed a dilemma of which both symbolic Ai and the kind of statistical, variable AI of current form cannot proceed. Even with the structure of neural network and modern deep learning, symbolic encoding, expert systems, current AI structures are what called interpolator in the purest form, and not extrapolator. This problem is inherently similar to how we would interpret the problem of out-of-bound cases are, as seen in Bahng et al. (2022); Xu et al. (2020); Traber et al. (2020); Hüllermeier and Waegeman (2021); Amini et al. (2021); Sensoy et al. (2024); Jose et al. (2022) of a wide range of such notion. Additionally, there are many problems that would not have answer and thereof, for example, the problem of hallucination apparent in typical LLM settings, symbolic grounding problem in Harnad (1990b), and so on. The problem is not with identifying the problem and observing it — the problem lies mostly in the form of inexplainable phenomena. To do such, we have to inquire further into what is being done in the current theory landscape.

2.3.1 Artificial Intelligence Theory

The current artificial intelligence theory can be branched into several aspects, as seen in Jackson (2019); Goertzel and Pennachin (2006); Chowdhary (2020); Russell and Norvig (2009), and more. Those considerations of the field are what we can call as both phenomenological inspired systems, and general modelling theory. We can list a few of them, for example, Shortliffe (1976)’s MYCIN structures, Newell and Simon (1976) works on symbol systems, Collins and Quillian (1969) semantic memory, Davis et al. (1977) knowledge-base representation, and Lindsay et al. (1993) DENDRAL expert system for working functionals of chemical compounds. There core driven such developments are the artificial intelligence idea focus on the reasoning process of the brain, using either pure logic (Logicism), Expert Systems, Non-monotonic Logic, Planning Systems, Argumentation, Semantic Network/DL, and Modal/Temporal Logic; more can be found in Liang et al. (2025).

Theory like knowledge representation, agent structures, algorithmic searches, decision theory, rule-based or relational network-based reasoning theories, first-order logic, specific domain of thoughts of computationalism, symbolic AI, or neurosymbolic unifications, while proved to be effective in a sense, those theories do not resolve the origin problem for those behaviours, and furthermore is restricted in the same restrictions that we stated as above, of AI models in generality. Logic in the construct of symbolic approach deliberately catches surface logic, representations, knowledges, properties and attributes, and assigning strict logical conclusion together for the logic itself. This falls into the range of Descartes’s argument, in which one’s machine deliberately can only follow and operate of its designer’s configuration, nothing more and nothing less. The restriction of a decision tree can also be founded to be limited in such case, for there are limitless consideration and fluctuated information in a given setting. Furthermore, a decision network on itself does not have any meaning. For example, the programming language PROLOG, which is prominent of such framework design (for reference, see Clocksin and Mellish (2003); Colmerauer et al. (1972); Colmerauer and Roussel (1993) for details on the standard of PROLOG), different facts are encoded manually by atom, a unit of logical fact in the database, and interpreting symbolic tasks being the goal, of which PROLOG then traverse the symbolic graph encoded to find the solution, of which also return, in the same sense, an encoded answer.

(a) Resultant process.

(a) Resultant process.

(b) Program.

(b) Program.

(a) Resultant process.

The program, while succeed, only works in its own environment. It is simply a program with encoded ruleset, of configured system of objects, and would rather be classified as an algorithmic program than artificial intelligence. Such parent-child relationship only makes sense in the eye of the designer, yet questionably non-apparent of the program itself. It is similar to the CRA, in which symbol manipulation and fact encoding only get it thus far, without any sufficient notion of understanding, for the arbitrary definition of such term is defined. Furthermore, applications using such system is severely limited at scale, for manual addition of facts, expert systems (from expert knowledge and so on), of which add into the rigid mimicry, and many more design initiatives.

When the focus switch to another constructive structure of connectionism, of which is utilized in statistical and data-driven frameworks, models like probabilistic logic, statistical relational learning, Markov/Bayesian process network and logic, Causal Inferences, Deep Neural Networks (DNN), and so on, they also struct the same disadvantage as discussed of the majority of symbolic-based program, sometimes worse. In terms of mathematical modelling, (2024), such models in statistical framework and on are interpretations of a phenomenological model, or black-box model, in which data and observations are encoded as statistical points for interpolation, either soft (for $\epsilon>0$, the disparity between the wrong answer and the correct one is not required to be reduced to zero) or hard (the condition explicitly said to reduce to absolute 0, or point-fit), without any knowledge of the underlying mechanical details. Hence, they are usually nothing but statistical fit of particular problems, for example, classifications, regressions, predictions, and so on of numerical data encoding. Sometimes, the encoding range and model structure is extended so hard, that the model seems too real, as seen of LLM, yet is not real, by the observable statistical-inherent errors and limitations. A limited observation can be said, as in Song et al. (2025).

The flaw of the current artificial intelligence theory can simply be attributed to looking at the wrong way, of which while proved insightful, is misleadingly taken as the way forward of making an intelligent construct. Approaches and constructions listed above relies on the basic idea of mathematical modelling, the modelling theory, yet such theory is underdeveloped, hence the true nature of the modelling scheme, of the generality on the theory of the representation language of which represents and is used for descriptive construction and analysis of models, often not realized. Such is also seen of many theories that is of apparent importance to artificial intelligence advancement, yet is rather forgotten in pursuit of the state-of-the-art phenomenological models. Different conceptual understanding, while making sense and is rigorous of logic and foundation, often found itself in trouble of interpreting the strong argument about origin of such system, once ask of the designer to make the model learn of the concept by itself. As of now, only in the sense of mathematical reduction, data points reduction that the notion of learning is then realized, and yet such is insufficient to be called intelligent, but intelligent algorithm. There exists no unifying or general theory of which is concerned of the topic of encoding, representation, architectural design and implications, operation theory, percolation theory and systematic emergence and thereof, which makes any inquiry into the topic harder than it should be. Current theories also use their limited systems to make bold claims, for example of manifesting consciousness, emergent behaviours, models as ‘living’ while such criteria of living is not fully realized[^3], making it inconsistent, and created a skew trend toward practical, empirical construction but no theoretical understanding. Again, it does not prevent us from making use of it, for applications on such first ever dynamic system modelling is enabled, opening a wide variety of different structural encoding and thereof. But even then, the theory of such system itself is insufficient, of limited depth, or coherence that is typically seen of a matured field.

Other than such, treatments of empiricism on such artificial intelligence theory, either by reconstruction of the brain for neurophysiology and neuroscience, empirical constructions and continual improves, such as current model development of AI systems, remains deeply shrouded in unanswered questions, inexplainable behaviours, constrained results or returns, high computational cost of inefficiency, scaling issues, and many more. Of such, it seems questionable to continue pushing toward that end, as the theory has exhausted a fair lot of its development, without gaining much insight or exploratory evidences into the deeper questions about AI, to ever reach the node of AGI. On the philosophical side, many has argued against the notion of computationalism, whether computers can actually be used as the foundation to create intelligence form, and understandably yet unsurprisingly the current theoretical treatment cannot have an answer toward such argument, other than keep working on it until a roadblock is hit.

2.3.2 The learning theory

In all of artificial intelligence theory foundation, one if not the most important aspect, and hence field of research that now dominated the field, is the theory of learning. Inherently, from the onset, the particular intelligent behaviour that one can instantly attribute to a construct to be called intelligent, is the ability for it to learn. The notion of learning is difficult to define, hence current theory seek to determine the mechanical equivalent in mathematical form instead, of which is expressed using the current machine learning theory, Valiant (1984a); Wolpert (1996); Shalev-Shwartz et al. (2010). In between the learning theory, there exists the fundamental main side, one of which is concerned with pure logical encoding, or rather learning machines of symbolic logic with applications in a wide range of different fields, as seen in Newell and Simon (1956); Newell et al. (1959); Winograd (1972); Michalski (1969), Minsky (1961); Minsky (1968), and Quinlan (1986). On the other side, is the looser end of the learning theory, which seeks phenomenological learning through enough inferences onto data, or rather, information-based learning, attributing learning to pattern recognitions on observables of specific problem set, as seen of Solomonoff (1964a); Solomonoff (1964b); Gold (1967); Newell et al. (1958); Cover (1965), and its formal treatment in theory by Vapnik and Chervonenkis (1968); Vapnik (1999); Vapnik and Chervonenkis (1971); Valiant (1984b); Angluin (1988); Angluin (1989); Board and Pitt (1992); Hajek and Raginsky (2021); Mohri et al. (2012); Shalev-Shwartz and Ben-David (2014).

Learning theory is regarded strongly as one of the most important part of a model in the modern modelling landscape. In essence, it is considered detrimental to be able of lifting up the problem of which previous versions and generations of symbolic and rule-based AI, much because of their rigidity, stumbled upon. Hence, the learning theory often makes use of the notion of statistical phenomenological approximation, and different methods to construct models and expression that take into accounts information, deviations, noisy settings, and thus can be applied more generally and of specific setting accurate of observations.

These two are fundamentally different approach, yet of the same goal of learning, or adaptation of reasoning and unseen situation of which then previous knowledge can be applied in certain way, or by extrapolate new information using a copula of approach, either by observation or by deductions. Despite their success, symbolic learning is plagued with its own originality — it cannot be scaled effectively, it is too rigid for such understanding to be made, and it is inefficient in handling truth or logic outside its rigid domain. Statistical learning, or rather learning by data in its pure form, relies on the model of black-box modelling approach, usually not totally by can be seen of as the grey-box modelling, which has its own problem about interpreting the concept wrongly, surface-level phenomenological approach which leads to insufficiency of mechanical details that underlies the concept that it learns, the reliance on data and large scale observation sets, while being scalable remains deeply mystery, of which cannot be explained (so-called inexplainable AI/DL), unanswerable questions regarding the different phenomena and computational process, results and performances, additionally structural interpretation, and so on. Furthermore, the theoretical treatment itself is fragmented, with many theories competing for interpretation and explanation in the wide spectrum, and with many different non-unified outlook of the same system dynamic, for example, McAllester (1999); Alquier (2008); Haussler et al. (1996); Jeon and Roy (2025); Abreu et al. (2025), and so on. For the large majority of current theoretical treatment of machine learning, it is stuck interpreting a fixed domain of result and setting, reliance on interpretation of the statistical parameterized model $\theta$, and using learning strictly in the domain of optimization theory only. Newer frontier in architectural researches, such as modern deep learning Marcus (2018); Goodfellow et al. (2016); Demuth et al. (2014), proves tremendously difficult to fully understanding using such theory, as for many of unanswerable questions were born from such architecture alone. The ergonomic and logistic of a unified theory of learning, either both symbolic (experts’ voice of unifying both phenomenological learning with symbolic knowledge) has become apparently so large in the eye of practitioners that some inherently try to create new theory of interpreting the learning framework, or push further in some directions regarding such, of which then again, mimic the 1970s-1980s age of artificial intelligence discovery, where many interpretation, exotic AI systems, and prototypes exist far and large.

It is undoubtedly successful of advancements and technological leap that the current theories landscape has provided with regard to artificial intelligence knowledge, for example, of figuring out a system better in some regard, human counterparts. Yet, it is clear to see that such theory is not capable of handling various degree of new problems, of new consideration, and of new standards of which has been hidden from the main argument of AI development, the question of interpretability, the question of originality, many questions of behavioural concepts, such as the question on ”common sense learning” (while it is foolish to study this at this stage). It is clear that we either need a unification of all such learning concept, and a new development of a more powerful learning framework, or to jump the bandwagon entirely; or something else.

The main critique also lies in the heart of the current implementation of the machine and its learning process as of date. Indeed, one can see a learning process as not the machine itself, but as two separated entities: the pair $(\mathcal{A},\mathcal{M})$ of the algorithm $\mathcal{A}$ for the process of learning, and the model $\mathcal{M}$ itself. While being considered such, it is in fact concerning, at least in certain part of the dialogue between the camp of learning theorist, and others, is the claim of which the machine or model itself is already learning, while in fact it is not.

Figure 2: The conceptual framework of a learning agent (Stanford Encyclopedia of Philosophy (2018)). The agent itself in such framework then indeed, internally ‘understand’ the metric notion of mistake, knowing and then fixing, typical process of which a learning system can be considered.

Figure 2: The conceptual framework of a learning agent (Stanford Encyclopedia of Philosophy (2018)). The agent itself in such framework then indeed, internally ‘understand’ the metric notion of mistake, knowing and then fixing, typical process of which a learning system can be considered.

As such, for those argued that machine can already learn and think, it is instead a very sophisticated pattern matching and statistical inferencing process. In typical situation, model in inference typically can be considered as post-dynamic, since they have been through the process of active construction and optimization, and thus at such point remains static for accurate prediction. Of the argument of which it can ‘somehow store memory’, such can be said of the regime in which RNNs or LSTM, of which takes into account of the encoding, historic and compressed memory influence between passing, and further on of the same principle as TNN (Transformer). The lag and processing speed can be clearly shown with lengthened discussion, as evidence. Furthermore, the process of learning itself is external, meaning that even the basic structural sophistry of which we consider the learning agent is already insufficient with respect to current ML system and models. Instead, what we are having is in fact a toolset, of which effective in making such model to a certain degree of prediction and statistical inference, but intrinsically cannot learn, and intrinsically black-box.

2.3.3 Architectural Insufficiency

Architectural insufficiency is the aspect in which we consider, the architectural design that is used in creating and constructing the construct that is evaluated for intelligence. Within its development history, AI has received numerous structural designs, some of which more successful than others. A fair share of our issues have been directed to the symbolic, logicism camp, so for now we would like to inquire on the matter of the statistical camp — machine learning models, and deep learning framework.

While classical frameworks have been noted of their own limitations and thereof, modern neural network structures, as deep learning architectures are much richer and much more interesting of such concern. For now, let us elaborate with the points on the architectural insufficiency observed of neural network, with respect to a provisional view toward a construct that is ideal of intelligence.

In essence, the current neural structures are concerned of the operation of neurons together in a unit-wise fashion. This started way earlier, with precursor constructions notable as Rosenblatt (1958); McCulloch and Pitts (1943), and more. In modern land scape, scaling such system is often done in layers, of which make up of the many smaller neurons organized and sequentially activated, i.e. feedforward, where individual layers sequentially feed its outputs to the inputs of the next layer. This is usually called layer connections, in some cases, and hence if two layers are connected such that each neuron $n_{j}\in L_{2}$ is connected to every neuron $n_{i}\in L_{1}$, we say the network is fully connected. Such operation is then repeated, and its learning behaviours controlled by the algorithm of backpropagation, of which propagate gradient errors network-wide and perform numerical correction, such that there then exists a correct solution to the given problem of the optimal configuration of the neural network. Then, at such stage, one simply fix the neural network, and let it run on the problem landscape with great accuracy. The most basic form of such architecture is the multilayer, fully-connected neural network, or multilayer perceptron (MLP), for $K$-layer, $m^{(0)}=d$ and $m^{(K)}=1$ for $m$ the width of the network, and $d$ the shape of the input vector.

Definition 2.2 (Standard multilayer network, Zhang et al. (2023))**.**

We define a $K$-layer fully-connected deep neural network with real-valued output. Let $m^{(0)}=d$ and $m^{(K)}=1$ for $m$ the width of the network, and $d$ the shape of the input vector. We then recursively define:

$$ x_{j}^{(0)} =x_{j}\quad(j=1,\dots,m^{(0)}), \tag{2} $$

$$ x_{j}^{(k)} =h\left(\sum^{m^{(k-1)}}_{j^{\prime}=1}\theta_{j,j^{\prime}}^{(k)}x_{j^{\prime}}^{(k-1)}+b_{j}^{(k)}\right)\quad(j=1,\dots,m^{(k)}),\quad k=1,2,\dots,K-1 \tag{3} $$

$$ f(x) =x_{1}^{(K)}=\sum^{m^{(K-1)}}_{j=1}u_{j}x_{j}^{(K-1)} \tag{4} $$

where the model parameters can be represented by $w=\{[u_{j},\theta_{j,j^{\prime}}^{(k)},b_{j}^{(k)}]:j,j^{\prime},k\}$ with $m^{(k)}$ being the number of hidden units at layer $k$; $\theta\in\mathbb{R}^{m}$ the weight of the neuron.

Such structural definition, of which use the unit abstraction of neuronal units, and the dynamic consideration of connections between parallel layer computational, allows for a variety of different architectural designs and application of the framework itself. Some of such includes the aforementioned neural networks in charge of revolutionary works in NLP, Graph Neural Network (Oono and Suzuki (2020); Scarselli et al. (2009); Hamilton ()) for graph-like structural data, physics ODE-based neural network like Chen et al. (2019) and so on, with numerous applications.

The first problem that come with this type of standard architecture is the expressive problem. Standardly speaking, both classical modelling, symbolic computation, and continuous/discrete Markov chain process[^4], relies on the perspective of function construction/approximation, that is, we assume that, there always exists, of the best-case scenario and the internal machine assumption, the concept $c\in\mathcal{C}$ can be expressed of the function space, that is, a collection of function that can be approximated, reconstructed, and so on. This ranges from symbolic process function, transition functions, time-evolution function given a system state $S$ of variables and parameters, or simply variables-related functional relationship. As such, structures and solutions on such space relies on reconstruction and approximation on various metric $d(f,c)$ that gauges the ability to approximate the learning objective in such numerical approximation landscape. Expressibility is then required, as the question of ”what kind of function class can encompass the entire concept space, and can there be a universal approximator?” for such various structures. For neural networks defined by definition 2.2, such is usually guaranteed, over certain class of functions, by the Universal Approximation Theorem, first proved by Cybenko (1989) for sigmoidal network, Hornik et al. (1989) for general activation function class, non-Weierstrass (polynomial) activation in Leshno et al. (1993), and partially, of probabilistic version on $L^{p}$ space in Hornik (1991). There are many statements and frameworks of such UAT is considered, though, UAT on neural network is considered in two different ways, as for two different architectural build-up of the network - arbitrary depth (layer), or arbitrary width (layer’s neuron). However, in general, it states:

Theorem 2.1 (Universal approximation theorem — UAT)**.**

Single hidden layer $\Sigma\Pi$ feedforward networks can approximate any measurable function arbitrarily well regardless of the activation function $\Psi$, the dimension of the input space $r$, and the input space environment $\mu$. That is, for every squashing function $\Psi:R\to[0,1]$ of which is non-decreasing and $\lim_{\lambda\to-\infty}\Psi(\lambda)=0$, every $r$ and every probability measure $\mu$ on $(R^{r},B^{r})$, both $\Sigma\Pi^{r}(\Psi)$ and $\Sigma^{r}(\Psi)$ are uniformly dense on compacta in $C^{r}$ and $\rho_{\mu}$-dense in $M^{r}$.

a perhaps simpler statement can be as followed.

Theorem 2.2 (Universal approximation theorem, simplified)**.**

For a class of functions $\mathcal{F}$ and a compact set $S\subset\mathbb{R}^{d}$, if for every continuous function $g$ on $S$ and for any $\epsilon>0$, there exists $f\in\mathcal{F}$ such that

$$ ||f-g||_{\infty}=\max_{\mathbf{x}\in S}|f(\mathbf{x}-g(\mathbf{x})|\leq\epsilon $$

Then, the class of functions $\mathcal{F}$ is a universal approximator of all continuous functions on $S$. We then indict that $N(w,d)$ of the neural network structure, derived from definition 2.2 is such universal approximator on the set of all continuous function on $[a,b]$ of arbitrary measure $\mu$.

While not determining correct optimal procedure that can lead to such result, or arbitrary properties of fitness, stability and so on that is useful for operational process, UAT still determines a weak existence theorem — such that in said domain, there exists the fundamental optimal model that can reconstruct partially the structure of the concept objective of the approximation task. Here, we still notice that we assume the observables of facts come in the form of function, and only function. Nevertheless, in particular, UAT allows for the guarantee of its expressive power over such class of function as the target. As such, many developments, of which derives from the class of optimization technique on function encoded space — numerical encoding on $\mathbb{R}$, landscape optimization such as gradient descent, coordinate descent (Luo and Tseng (1992); Tseng (2001); Friedman et al. (2007); Wright (2015); Tseng and Yun (2009)), extension of such into the process of backpropagation (Rumelhart et al. (1986) and older historical development). Because our argument is about architectural insufficiency, one then can ask what is the drawback in such framework?

First and foremost, UAT and the functional structure deliberately choose and restrict itself around smooth, continuous, well-behaved function. Not taking the impossible domain lies beyond the compact subset, many functions or encoded function cannot be determined, or can be approximated partially to a given arbitrary $\epsilon>0$, yet not satisfactory. Variations of UAT usually can be considered weak, with some stronger theorems deliberately require the use of stronger assumptions and more specialized constraints. Any given different setting, structures, data type and the encoding space of such concept, for example, as graph theoretical data, requires to be encoded or embedded into the continuous differentiable pipeline of a neural network to be utilized, and in the process making the structure suboptimal in certain standards. This result in the structure of Graph Neural Network, or GNN, of which use an encoding scheme, and a typical block aggregator over local graph information, for example, as

$$ m_{v}^{(k)} =\psi^{(k)}\!\left(\{\,M^{(k)}(h_{v}^{(k)},\,h_{u}^{(k)},\,e_{uv})\mid u\in\mathcal{N}(v)\,\}\right), \tag{6} $$

$$ h_{v}^{(k+1)} =U^{(k)}\!\left(h_{v}^{(k)},\,m_{v}^{(k)}\right), $$

Where $e_{uv}$ represents the edge features, if any, $h^{(k)}_{v}$ is the embedding of node $v$ at layer $k$, $\mathcal{N}(v)$ denotes the neighbours of $v$, $M^{(k)}$ denotes the message function producing messages from neighbours, $U^{(k)}$ denotes the update function, and $\psi^{(k)}(\cdot)$ denotes the message aggregation up to layer $k$ or such node in consideration. Even though GNN of such ‘rigorous’ in a sense framework can be used in a variety of sessions, being restricted to such neural network framework brings extraordinary drawbacks. For example, the expressive power on the native graph data is at best, equivalent to 1-dimensional Weisfeiler-Leman (1-WL) test. Such message passing scheme often fails because of the bottleneck problem intrinsic within the framework requirement and complex backpropagation adaptation, as seen in Alon and Yahav (2021). Further problems like structural information losses, scalability problems, trainability issues, over-smoothing, ontological structural mismatch, and so on, as seen in Maron et al. (2019). Internally of the neural network framework, such problem is not restricted to such adaptation, which requires both complex restructuring and architectural mapping, but also problems like the parity or memory-sample problem in statistical query model (Blum et al. (2003); Garg et al. (2021)), planted clique problem (Barak et al. (2016)), worst-case analysis indication of neural network being NP-complete or $\exists\mathbb{R}$-complete (Blum and Rivest (1992); Abrahamsen et al. (2021)) for low-level, 3-node, 2-layer configuration, computational problem and its expressibility (Eldan and Shamir (2016); Telgarsky (2016); Siegelmann and Sontag (1991); Weiss et al. (2018)). Perhaps more interestingly, is the actuated and controversial No Free Lunch theorem from Wolpert and Macready (1997), of which basically states that a ”structural-less”, or naively defined structural copulative information, construction on the encoding scheme, is equally bad as with any given model of arbitrary sense, and also the learning procedure associated with such. Setting this up with respect to a particular set of preset inductive bias, or structural assumptions thereof, scaling exorbitantly both of components, parameters, and repeated structures like typical deep learning structure, is not feasible and reduce its success to uncertainty, as the stronger, relative-NFL theorem takes effect — given arbitrary structural information as the origin, there then exists, particularly, certain point of which without further information introduced to the setting, the supposed ‘information’ can be reiterated as misleading, thus average out the model’s evaluation and performance.

Aside from some archetype of learning, for example, online learning of which keep dynamically update the model in a sense, generally, neural network and such type of parameterized models are static. They act, just as CRA indicted, to be an algorithmic machine only, of which fits to a description, to a task, to a given degree of accuracy within certain manually designed encoded space, manually designed objective, limited or no interpretation of the scheme itself (classification machines think in numerics and matching errors, thus would not know what a cat is, but only know that an object named 0x0AB4 must be evaluated in some way that can be optimized and reduced). The operating stream itself of the neural network is limited, as we can say of the feed-forward scheme. The extension that neural network is configured to is limited, and lest to say that the neural network scheme itself is underdeveloped, ultimately relies on the simplest possible modelling structure to then give rise to a multitude of different architectures of date. Computational optimization and layer-optimization, practically constrained structures further hampers down such understanding of the neural network itself, as nobody knows what happen in the neural network, less to try adapting it to different situations.