Review status

Primary-source, structure, and link review has been completed by the author. No independent specialist in machine learning, philosophy of mind, or Qur’anic exegesis has yet reviewed this dossier, and it is not an academically peer-reviewed publication.

External review pending

Direct answer

The current evidence supports neither “language models are only surface parrots” nor “fluent language proves human-like understanding.” Predictive training can produce internal representations that track structured properties of a task or domain, and in some bounded experiments those representations have shown causal relevance to model behaviour. Yet decodability is not the same as causal use; causal use in a game is not the same as open-world grounding; grounding is not intention; and none of these results by themselves establish phenomenal consciousness or moral responsibility. Qur’an 2:30–33 contributes a different kind of claim: Adam’s knowledge is taught, creaturely knowledge is limited, and knowledge appears within a scene of vicegerency and responsibility. The passage does not describe Transformer architecture or machine-learning training.

What predictive training can showInternal structure can emerge even when the explicit objective is next-token prediction.
What a probe can showInformation may be decodable from activations; causal use requires stronger evidence.
What multimodality addsImages, sensor state, and action provide stronger world-linked evidence than text alone.
What remains unresolvedIntention, responsibility, and phenomenal consciousness are not settled by linguistic fluency.

The question this dossier actually answers

“Does AI understand?” is too compressed to be scientifically or philosophically useful. It can mean at least six different things: whether a system distinguishes symbols, whether it encodes stable relations, whether it recombines rules in new cases, whether its symbols connect to the world, whether internal state is used causally in reasoning or action, and whether the system has intention or responsibility. A seventh issue—phenomenal consciousness, whether there is something it is like to be the system—does not follow automatically from the first six.

The central question here is therefore narrower: what must be demonstrated, layer by layer, before successful use of names or language can justify progressively stronger claims about meaning, reference, world modelling, intention, or responsibility? A second question is kept separate: what does the Qur’anic scene of teaching Adam the names add to the intellectual and ethical framing without being converted into a modern engineering theory?

A glossary before the argument

TermMeaning in this dossierCommon overreach avoided
TokenA computational unit used by a language model.It is not automatically identical to a Qur’anic or philosophical “name.”
RepresentationA model state from which some property may be encoded or decoded.Decodability alone does not prove causal use.
GroundingA relation between symbols and perception, action, external reference, or socially anchored use.It is not equivalent to consciousness.
World modelAn internal structure that tracks relevant state or regularities of an environment.A bounded game-state model is not a complete model of the open world.
IntentionA stronger claim about goal-directed agency and ownership of action.First-person language is not independent evidence of intention.
ConsciousnessPhenomenal or subjective experience.No behavioural benchmark in this dossier is treated as a decisive consciousness test.

Why the Qur’anic unit is 2:30–33, not verse 31 alone

The teaching of the names appears after the declaration of a khalifah on earth, the angels’ question about corruption and bloodshed, and the response, “Indeed, I know what you do not know.” It is followed by the presentation or disclosure associated with the names, the angels’ statement that they possess no knowledge except what God has taught them, and Adam’s act of informing them.[1] The sequence matters. The text does not present naming as an isolated vocabulary exercise.

This does not mean that every theological implication has one uncontested formulation. It means that the passage’s argumentative structure includes knowledge, limitation, human earthly vocation, and the moral risk already raised by the angels’ question. Any modern comparison that extracts “names” while discarding the rest of the unit is methodologically weak.

What the classical exegetes actually disagree about

Al-Tabari: reporting a broad field of views is not the same as preferring all of them

Al-Tabari reports several positions on what Adam was taught, including names of things, names connected to the angels, and names connected to Adam’s descendants. The important corrective is that his own preferred reading is narrower than the common modern summary “Tabari said Adam was taught the names of everything.” He gives weight to a reading tied to Adam’s descendants and the angels, in part because of the plural pronoun in ʿaraḍahum, while acknowledging that a broader generic reading remains linguistically possible.[2]

The methodological lesson is simple: a view that an exegete reports is not necessarily the view he prefers, and linguistic possibility is not the same as an exclusive interpretation.

Al-Qurtubi: name, naming, and referent are not identical

Al-Qurtubi records interpretive diversity and spends substantial effort on the relation between the name, the act of naming, and what the name refers to. He gives substantial weight to a broad reading because of “all of them,” while still preserving narrower positions and grammatical questions about what exactly was presented.[3] This is useful for conceptual discipline: the expression is not the external thing, and the ability to produce a label is not automatically complete knowledge of the referent.

That distinction is not presented here as a medieval theory of embeddings or symbol grounding. It is a terminological guardrail that prevents a modern technical term from being silently substituted for an exegetical one.

Al-Razi: properties and realities matter more than labels alone

Al-Razi develops a reading in which the names may involve attributes, properties, and realities of things, and he argues that knowledge of such properties carries greater significance than merely possessing verbal labels.[4] This creates a genuine historical point of contact with a modern question: is successful label use enough, or must a system encode relations and properties that travel with the label?

The contact is analogical, not identical. Al-Razi was not proposing a cognitive benchmark, and his exegetical argument cannot be used as empirical evidence about neural networks.

Ibn Ashur: naming, expression, and the beginning of transmissible knowledge

Ibn Ashur explicitly connects the teaching of the names to the preceding statement that God knows what the angels do not know. He treats Adam’s receptivity to teaching and the capacity for naming and expression as part of the wisdom of vicegerency, and links these capacities to the beginnings of sciences, laws, and the transmission of mental content.[5] He also argues that the lesson of the passage does not depend on resolving every disputed detail about the exact inventory of names or languages.

Abu Hayyan: Adam’s answer follows teaching; it is not independent knowledge

Abu Hayyan discusses whether the relevant object is names, referents, or persons, and records multiple possibilities for scope and method. A central point is that teaching precedes Adam’s disclosure; Adam’s successful response is therefore the effect of teaching, not proof of autonomous or uncaused knowledge.[6]

ExegeteUseful contribution hereWhat this dossier does not attribute to him
Al-TabariInterpretive plurality and a narrower preferred reading than common summaries suggest.A modern theory of concepts or machine learning.
Al-QurtubiDistinction among name, naming, and referent.Modern reference theory as such.
Al-RaziProperties and realities can matter more than labels alone.An empirical cognitive model.
Ibn AshurNaming and expression inside the wisdom of vicegerency and transmissible knowledge.A theory of Transformer training.
Abu HayyanCreaturely epistemic limitation and the causal priority of teaching in the narrative.A claim about machine consciousness.

The religious-scientific boundary

The comparison is legitimate only after both domains have been stated independently. Qur’an 2:30–33 can be read as a passage about taught knowledge, limitation, human earthly vocation, and responsibility. Machine-learning papers can test representation, robustness, grounding, and causal intervention. Neither domain is allowed to borrow the other’s evidentiary authority.

Accordingly, this dossier does not claim that the Qur’an predicted artificial intelligence, that “teaching the names” is equivalent to training a language model, that tokens are Qur’anic names, or that an embedding is a complete concept. It also does not argue that a machine cannot understand merely because it is not human. Each technical claim must stand or fall on technical evidence.

A predictive objective does not fully describe what is learned

The Transformer architecture introduced in 2017 replaced recurrence-heavy sequence processing with attention mechanisms and demonstrated strong performance in machine translation.[10] Nothing in that paper is a theory of consciousness. “Attention” is an engineering term, not proof that the model attends in the phenomenological sense.

Autoregressive language models are trained to improve next-token prediction. GPT-3 showed that scaling such training could yield broad in-context task performance without task-specific weight updates for every new prompt.[11] But the loss function specifies what optimization rewards, not an exhaustive list of all internal structures that can emerge as useful means to that end. A sequence model may discover compressed regularities because they help prediction.

The correct conclusion is therefore two-sided. A simple predictive objective does not justify saying that every learned representation is superficial. Yet broad behavioural performance does not justify saying that the system has human-like meaning, intention, or awareness.

Three levels of evidence for an internal representation

  1. Decodability: an external classifier can recover a property from model activations.
  2. Stability and transfer: the representation persists across paraphrases, entities, contexts, or datasets.
  3. Causal use: changing the proposed representation changes behaviour in the direction predicted by the interpretation.

The first level is the weakest. Activations can contain traces of information that the model does not actually use. The third level is stronger because it tests whether the proposed internal variable participates in computation. Even then, a causal result remains local to a model, task, and intervention design.

Othello-GPT: the strongest bounded example in this dossier

Li and colleagues trained an eight-layer GPT-style model on sequences of Othello moves without giving it a visual board state or an explicit symbolic encoding of the rules. In the synthetic setting the model became highly accurate at predicting legal moves. More importantly, nonlinear probes recovered board-state information from intermediate activations, and interventions designed to alter the encoded board state shifted the model’s move predictions toward the counterfactual state.[14]

This supports a specific claim: in a closed synthetic domain, sequence prediction can induce an internal representation of task state that plays a causal role in behaviour. That is stronger than “the model memorized strings.” It is still far weaker than “the model understands the open world.” Othello has a finite board, explicit rules, and a tightly specified state space. Nothing in the experiment establishes subjective experience, a desire to win, or moral agency.

Space and time in language-model activations

Gurnee and Tegmark examined spatial and temporal information in Llama-2 models across several geographic and historical domains. They found linearly decodable representations of space and time that showed some stability across prompt formulations and entity types.[15] The result is significant because it suggests that language models can organize disparate entities along structured real-world dimensions.

But decodability is not enough to say that the decoded coordinates are always the cause of the model’s answer. Nor does a spatial coordinate amount to a complete world model. The cautious formulation is that these are structured, recoverable components that may contribute to world modelling.

Why probes need control tasks

Hewitt and Liang showed why probing results can be overinterpreted. A powerful probe may succeed because the target property is cleanly represented in the model, or because the probe itself has enough capacity to memorize associations. They proposed control tasks with randomized labels and a selectivity criterion: a good probe should perform well on the real linguistic task while remaining poor at the random control task.[16]

The lesson is not “probes are useless.” It is that probe accuracy alone does not establish the algorithm used by the model. Stronger evidence combines controlled probes, transfer tests, and intervention.

From correlation to causal abstraction

Geiger and colleagues provide a formal framework for asking whether a high-level interpretation is a causally faithful abstraction of low-level neural computation. Their work unifies a family of mechanistic-interpretability techniques around intervention-based tests of correspondence.[17] This is methodologically important because it raises the bar: a concept label attached by the researcher is not automatically a true description of the model’s mechanism.

A representation claim becomes stronger when the proposed variable behaves as the interpretation predicts under counterfactual interventions. But even a successful intervention does not escape domain boundaries. A variable that causally tracks board state in Othello is evidence about Othello computation, not a general proof of semantic understanding.

Compositional generalization is neither absent nor guaranteed

A robust concept should survive more than familiar wording. Ordered CommonGen tests whether models can use specified concepts in a required order, rather than merely include them somewhere in a fluent sentence. Sakai and colleagues reported persistent order biases and incomplete instruction following across a broad set of models.[18]

Ismayilzada and colleagues tested morphological generalization in Turkish and Finnish, including application of patterns to novel roots and increasingly complex combinations. Performance showed systematic weaknesses on novel compositions and greater complexity.[19]

These results do not imply a universal percentage of “understanding.” They show something more useful: success on familiar combinations cannot be assumed to transfer to novel compositions, and benchmark design must deliberately separate memorized or familiar surface structure from rule-like generalization.

High scores can depend on irrelevant form

Zhao and colleagues tested content-preserving perturbations such as option length, problem format, and irrelevant noun substitutions. In their experiments, some models’ scores shifted sharply under changes that should not alter the underlying answer.[20] Such sensitivity is evidence that benchmark performance can be contaminated by shortcuts or formatting preferences.

The strongest interpretation is not “the model has no abstract representation.” A model can contain meaningful structure and still use brittle policies or shortcuts. The point is that a high score should not be treated as a pure measure of understanding until invariant-preserving perturbations have been tested.

An apparent sudden ability may partly be a metric effect

Schaeffer and colleagues demonstrated that discontinuous or nonlinear evaluation metrics can make smooth improvements appear as sudden emergent jumps. Their argument does not show that all emergent abilities are illusory. It shows that the shape of a benchmark curve is not by itself evidence that a qualitatively new cognitive mechanism suddenly appeared.[21]

Claims such as “understanding emerged at scale” therefore require more than a threshold crossing. The metric, partial-credit behaviour, error distribution, and transfer properties all matter.

The grounding dispute: text carries the world, but indirectly

Bender and Koller argue for a sharp distinction between linguistic form and meaning tied to communicative intention and external reference.[12] Bisk and colleagues extend the concern to physical and social experience, arguing that language understanding should be studied with grounding beyond text.[13] Harnad’s symbol-grounding problem gives the classic formulation of the difficulty of defining symbols only through other symbols.[9]

The strongest counterpoint is that text is not random form. Human-produced language carries regularities inherited from embodied and social interaction with the world. Piantadosi and Hill argue that conceptual-role relations among internal states can support genuine aspects of meaning, so the architecture or objective alone cannot settle the issue.[24]

This dossier therefore avoids both extremes. Text-only learning can acquire deep relational structure because text is a record of world-connected human practice. Yet a text-only model receives those connections through a mediated channel and cannot independently correct every relation by perception or action.

What multimodal and embodied systems add

PaLM-E combines language with images, continuous state estimates, and robotic task data. The system demonstrated transfer across language, vision, and embodied planning tasks in the settings reported by its authors.[22] This is stronger grounding evidence than text alone because words can be constrained by visual state and action consequences.

It still does not establish intention or consciousness. A robotic policy in a bounded environment is not equivalent to human embodiment, developmental history, pain, social membership, or moral accountability. Grounding is an evidential layer, not a shortcut to personhood.

Formal linguistic competence and functional competence

Mahowald and colleagues distinguish formal linguistic competence—patterns and structural properties of language—from functional competence involved in using language with world knowledge, social reasoning, goals, and broader cognition.[23] The distinction helps explain how a system can be remarkably strong at syntax, translation, or textual transformation while remaining uneven in tasks requiring stable world-state tracking or social inference.

Mitchell and Krakauer similarly argue that human-like understanding should not be inferred from benchmark performance without examining the capacities and representations that support behaviour.[25]

The Name–World–Responsibility Matrix

The following framework is an original analytical proposal by Ahmed Alhafiz. It is not a validated consciousness scale and it is not presented as an established scientific taxonomy. Its function is to expose what kind of evidence is required before moving from a weaker claim to a stronger one.

LayerCore questionEvidence that helpsWhat it does not establish
1. Symbol discriminationDoes the system distinguish forms, units, and positions?Controlled recognition, transformation, and prediction under surface changes.Meaning or external reference.
2. Relational structureDoes it encode stable relations among internal states?Representations that remain decodable across contexts and entities.That the relation is used causally or grounded externally.
3. Composition and generalizationCan it recombine relations in genuinely new cases?Novel roots, orders, combinations, and out-of-distribution transfer.Open-world understanding.
4. Reference and groundingDo symbols connect to perception, action, objects, or socially anchored use?Multimodal and embodied tests with corrective world feedback.Intention or subjective experience.
5. Causal world-model useDoes changing an internal state change behaviour as the proposed world variable predicts?Counterfactual intervention, activation patching, and causal-abstraction tests.General agency or consciousness.
6. Intention and responsibilityDoes the system own persistent goals and qualify as a bearer of responsibility?No agreed benchmark currently settles this package.First-person language is not sufficient evidence.

Phenomenal consciousness is not a seventh rung that automatically follows from layer six. It is a separate unresolved question requiring additional theory and evidence. Functional access to information, self-report, planning, or global availability should not be silently converted into proof of subjective experience.

A ten-part adversarial protocol

  1. Name swap: replace familiar entity names with arbitrary labels while preserving relations.
  2. Paraphrase and order: change wording and sequence while preserving the proposition.
  3. Irrelevant information: add distractors that should not change the answer.
  4. Option-format controls: vary answer length and placement.
  5. Novel composition: apply a learned rule to new roots, combinations, or structures.
  6. Text versus perception: separate what can be solved from text from what requires sensory evidence.
  7. Probe controls: compare linguistic probes with randomized control tasks and report selectivity.
  8. Internal intervention: alter the proposed representation and test the predicted counterfactual output.
  9. Environment transfer: move the task to a new distribution or setting.
  10. Calibration: evaluate confidence and uncertainty, not accuracy alone.

No score from these ten tests would amount to a single “understanding percentage.” The protocol is designed to distinguish memorization from transfer, decodability from causality, and fluent behaviour from stronger claims about reference or agency.

Five strong objections—and what remains after them

Objection 1: behaviour is the only evidence we ever have

We infer human understanding largely from behaviour, so demanding internal evidence from machines can appear unfair. The response is not to reject behaviour. It is to design behaviour tests that discriminate competing explanations and to use internal access where it is available. Humans also come with a long embodied, developmental, and social history that is part of our evidence for agency and responsibility.

Objection 2: humans learn many things from language alone

Correct. People know distant cities, extinct civilizations, and abstract science largely through testimony and language. Direct sensory contact with every referent is not a sensible requirement for meaning. The real distinction is that human language learning occurs inside an embodied and social system that can often return to perception, action, or trusted communities for correction. Text-only models inherit those traces indirectly.

Objection 3: causal interventions already prove world models

In bounded cases such as Othello, they provide strong evidence for a task-relevant internal state model. The matrix is designed to recognize this evidence, not dismiss it. The remaining question is scope: a board-state model is not automatically a model of open-ended physical and social reality.

Objection 4: multimodality solves grounding

It addresses an important part of grounding by linking words to images, sensors, and actions. It does not settle all grounding questions, and it does not by itself establish persistent self-generated goals, moral agency, or phenomenal consciousness.

Objection 5: the six layers are themselves philosophical choices

Yes. The matrix is an authorial synthesis, not a discovered law of cognition. It can be criticized, reordered, or subdivided. Its value lies in making hidden inferential jumps explicit and generating harder tests—not in claiming to be the final definition of understanding.

Turing and Searle: what their arguments can and cannot settle

Turing proposed replacing the vague question “Can machines think?” with an observable imitation game.[7] This remains a powerful methodological move: turn metaphysical language into testable behavioural criteria where possible. But conversational indistinguishability is not direct measurement of consciousness.

Searle’s Chinese Room argues that implementing a formal program is not sufficient for intrinsic intentionality or understanding.[8] It is an influential philosophical argument, not an experimental demonstration that machine understanding is impossible. The present dossier therefore treats it as a challenge to sufficiency claims, not as a laboratory result.

Why responsibility remains human unless stronger evidence appears

Language models can generate plans, explanations, requests, apologies, and first-person statements. None of those linguistic forms by themselves identify the bearer of responsibility. Current systems are trained, deployed, constrained, and assigned goals through human and institutional decisions. A statement such as “the model decided” can be operationally convenient while still hiding the chain of human choices around design, data, deployment, permissions, and reliance.

This is where the Qur’anic framing adds an ethical question without pretending to answer a technical one. In the passage, knowledge is not detached from creaturely limitation or the larger question of human action on earth. In the author’s forthcoming book Qul Siru fi al-Ard, the thematic path moves from name to concept, relation, and model, then from language and writing to civilizational memory, technology, and human responsibility. That manuscript is the origin of the question here, not evidence for the machine-learning claims. The empirical claims stand on the technical literature cited below.

Compact claim map

IDClaimConfidenceMain boundary
C01Qur’an 2:30–33 frames teaching the names within vicegerency and creaturely epistemic limits.HighNot a complete modern cognitive theory.
C02Classical exegetes differ over the scope and nature of the names.HighDiversity does not validate every modern projection.
C04A predictive objective does not by itself determine every internal structure learned.HighDetectable structure does not prove grounding.
C06Structured game, space, and time representations can be decoded in bounded settings.MediumDecodability does not prove causal use.
C07Causal intervention is stronger evidence than correlation alone.HighStill local to model and task.
C08High benchmark scores can be brittle under content-preserving changes.HighOne failure does not erase all abstraction.
C09Compositional generalization is partial and task-dependent.HighNo universal percentage applies.
C10Multimodal perception and action strengthen grounding evidence without settling intention.MediumBounded robotics is not human embodiment.
C13The six-layer matrix is an analytical proposal for separating evidence types.OpenNot a validated scale or consciousness test.
C14The Qur’an–AI connection here is epistemic and ethical, not an identity between revelation and Transformer architecture.HighComparison is allowed only after domain boundaries are explicit.

The complete bilingual claim graph, including support, qualification, confidence, and limitations for all 14 claims, is available in the machine-readable claim ledger.

What this dossier does not claim

  • It does not claim that the Qur’an predicted artificial intelligence.
  • It does not equate teaching Adam the names with machine-learning training.
  • It does not claim that language models understand nothing at all.
  • It does not claim that an embedding is a complete concept.
  • It does not claim that probe success establishes causal use.
  • It does not claim that a bounded world model establishes consciousness.
  • It does not treat a model’s self-report as independent evidence of subjective experience.
  • It does not present the Name–World–Responsibility Matrix as a validated scientific scale.
  • It does not use Ahmed Alhafiz’s book as empirical evidence for any technical claim.

Conclusion

The best-supported position is not a slogan. Predictive language training can generate internal structures that go beyond literal string memorization. Some are stable enough to decode, and in bounded settings some can be manipulated in ways that causally affect behaviour. At the same time, modern models remain uneven under composition, distribution shift, irrelevant formatting changes, and tasks that require robust world-linked correction.

The resulting hierarchy matters. Symbol use is not automatically reference. Decodable structure is not automatically causal use. Causal use in a bounded task is not automatically open-world understanding. Grounding through sensors and action is not automatically intention. Intention, even if one day supported by stronger evidence, would still not automatically settle phenomenal consciousness or moral responsibility.

Qur’an 2:30–33 adds a distinct intellectual frame: knowledge is taught, creaturely knowledge has limits, and human knowledge appears in a story already concerned with action on earth. The comparison becomes useful precisely when it resists identity. Revelation is not an engineering manual, and a neural network is not an exegetical category. The point of bringing them into one dossier is to sharpen the questions we ask of knowledge, capability, and responsibility—not to make one domain impersonate the other.

External review status

No independent external specialist review has yet been completed. The page is published with its sources, confidence boundaries, and correction state visible so that claims can be inspected. It must not be represented as a peer-reviewed academic paper.

Suggested citationAlhafiz, Ahmed. “Teaching the Names and AI: What Counts as Understanding?” Official website of Ahmed Alhafiz, 3 September 2026. https://ahmedalhafiz.com/en/articles/teaching-names-ai-understanding/

References

  1. Qur’an 2:30–33. The complete textual unit used in this dossier.
  2. Al-Tabari, Jamiʿ al-Bayan, commentary on Qur’an 2:31. Reported positions and Tabari’s preferred reading.
  3. Al-Qurtubi, commentary on Qur’an 2:31. Name, naming, referent, and interpretive scope.
  4. Fakhr al-Din al-Razi, commentary on Qur’an 2:31. Properties, attributes, and realities versus labels alone.
  5. Ibn Ashur, al-Tahrir wa-l-Tanwir, commentary on Qur’an 2:31. Naming, expression, vicegerency, and transmissible knowledge.
  6. Abu Hayyan, al-Bahr al-Muhit, commentary on Qur’an 2:31. Scope, referents, and the priority of teaching.
  7. Alan Turing, “Computing Machinery and Intelligence,” Mind, 1950. The imitation-game proposal.
  8. John Searle, “Minds, Brains, and Programs,” 1980. The Chinese Room argument.
  9. Stevan Harnad, “The Symbol Grounding Problem,” 1990. Classic grounding formulation.
  10. Vaswani et al., “Attention Is All You Need,” NeurIPS 2017. Transformer architecture.
  11. Brown et al., “Language Models are Few-Shot Learners,” NeurIPS 2020. Autoregressive scaling and in-context task performance.
  12. Bender & Koller, “Climbing towards NLU,” ACL 2020. Form, meaning, communicative intent, and reference.
  13. Bisk et al., “Experience Grounds Language,” EMNLP 2020. Physical and social grounding research agenda.
  14. Li et al., “Emergent World Representations,” arXiv:2210.13382. Othello state representations and intervention evidence.
  15. Gurnee & Tegmark, “Language Models Represent Space and Time,” ICLR 2024. Decodable spatial and temporal structure.
  16. Hewitt & Liang, “Designing and Interpreting Probes with Control Tasks,” EMNLP-IJCNLP 2019. Probe selectivity and controls.
  17. Geiger et al., “Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability,” JMLR 2025. Intervention-based causal faithfulness.
  18. Sakai et al., ACL 2025, Ordered CommonGen. Ordered concept composition and instruction-following failures.
  19. Ismayilzada et al., NAACL 2025. Morphological compositional generalization in Turkish and Finnish.
  20. Zhao et al., EMNLP 2025. Content-preserving perturbations and shortcut sensitivity.
  21. Schaeffer et al., “Are Emergent Abilities of Large Language Models a Mirage?” NeurIPS 2023. Metric-induced apparent emergence.
  22. Driess et al., “PaLM-E: An Embodied Multimodal Language Model,” 2023. Vision, continuous state, language, and robotic planning.
  23. Mahowald et al., “Dissociating language and thought in large language models,” Trends in Cognitive Sciences, 2024. Formal versus functional competence.
  24. Piantadosi & Hill, “Meaning without reference in large language models,” 2022. Conceptual-role counterargument.
  25. Mitchell & Krakauer, “The debate over understanding in AI's large language models,” PNAS 2023. Competing conceptions of understanding.
  26. Engels et al., “Not All Language Model Features Are Linear,” 2024 preprint. Nonlinear feature geometry and limited intervention evidence.
  27. Gurnee et al., “Verbalizable Representations Form a Global Workspace in Language Models,” 2026 preprint. Frontier functional evidence; excluded from the central consciousness conclusion.
  28. Lake, Salakhutdinov & Tenenbaum, “Human-level concept learning through probabilistic program induction,” Science 2015. Human compositional concept-learning comparison.

Primary texts and research papers are used according to their own domains. The Qur’anic and exegetical sources do not serve as empirical AI evidence, and the machine-learning papers do not establish theological claims. The unpublished manuscript is a thematic origin for the question, not an evidentiary source. See the research and correction method.