Africa’s AI Almanac

Introducing the Grounding Series: Understanding Foundation Models

9 min read

Part of Foundation Models

jr korpa 9XngoIpxcEo unsplash

In August 2021, Stanford researchers named a shift already under way in AI. What a foundation model is, why the name was contested, and why the argument reached the statute

This is the first article in the Africa Foundation Models Project's Grounding Series. Before examining what AI means for African policy, regulation and governance, the series steps back to the technology itself asking what foundation models are, what distinguishes them from the AI systems that came before, and how their construction and deployment shape the questions that follow. This article begins with the term itself, where it came from, and why the argument about it reached statute.

In August 2021, more than a hundred researchers put a name to a shift already under way in AI, the foundation model. Researchers were building increasingly large models trained on vast and varied datasets, intended for adaptation to a wide range of uses. What was missing was a common way to describe the approach. The researchers proposed one.

Their definition was concise. A foundation model is a model trained on broad data, generally using self-supervision at scale, that can be adapted through methods such as fine-tuning to a wide range of downstream tasks. Three ideas sit within this definition, and each marks a departure from earlier practice including the breadth of the training data, the way the model learns from it, and its adaptability afterwards.

  1. Broad data

A foundation model is not trained with a single task in mind. Its training data is deliberately broad, drawing on many sources and formats including text from the open web, code from public repositories, books, transcripts, images and their captions, and other publicly or commercially available material.

The aim is not a dataset that makes the model good at one job. It is to expose the model to enough different kinds of information, language and pattern that it develops capabilities useful across many tasks. Breadth, rather than fitness for a particular purpose, is the defining feature, and it determines what exists at the end of training: a general-purpose system rather than a finished tool.

  1. Self-supervision at scale

The second part of the definition concerns how the model learns. For much of the history of machine learning, training relied on human-labelled examples. People identified which emails were spam, marked which X-rays showed a fracture, or classified whether a sentence expressed positive or negative sentiment. That approach remains useful where the task and the desired outcome are clearly defined, but it requires substantial quantities of labelled data along with the time and expertise to produce it.

Self-supervised learning takes a different route. Rather than requiring each example to be labelled by a person, the model derives a training signal from the structure of the data itself. A word can be hidden from a sentence and the model is asked to predict what belongs there, with the surrounding text supplying the context from which it learns.

Since individual examples do not have to be labelled by hand, this can be applied across very large datasets, and the scale of pretraining expands with the availability of suitable data and the computing resources to process it. It is important to note that human work does not disappear from the process. It moves into deciding what data is collected and filtered, and into a later stage of training that shapes how the model behaves. This is further discussed in subsequent articles. .

  1. Adaptation, and the split it creates

Adaptation can take several forms. In fine-tuning, the model undergoes additional training on a smaller, task-specific dataset. Alternatively the underlying model is left largely unchanged and instructions or examples are supplied as part of the input, allowing it to perform a new task without further training at all.

This is a significant departure from older practice, in which one model was generally built for one task. A sentiment classifier trained on labelled product reviews could classify sentiment and little else; extracting clauses from contracts meant starting again with a new dataset, a new model and another round of labelling.

Foundation models split that process into two stages separated by cost, time and often institution. Pretraining produces the general model and consumes enormous quantities of data, computing power and money. Adaptation turns that model into a working application and is comparatively cheap. The foundation model is therefore not the finished application; it is the system from which many different applications are built.

Both halves of the split were demonstrated within about two years. BERT, released by Google researchers in 2018, was pretrained on unlabelled text and then fine-tuned to state-of-the-art results across eleven language tasks by adding a single output layer. GPT-3, published by OpenAI in 2020, went further. Its authors reported that for all tasks the model was applied "without any gradient updates or fine-tuning", specifying tasks through text instructions and a small number of demonstrations. Adaptation, at that point, no longer necessarily required training. It could require writing a prompt.

The asymmetry between the two stages matters well beyond the technical architecture. Very few organisations can perform the first stage. Very many can perform the second. African universities, hospitals, banks, courts and startups therefore encounter this technology largely at the adaptation layer, building applications on models whose training data, design choices and failure modes were determined elsewhere. Subsequent articles look at some of the models being used in the continent.

Why the word "foundation"

The authors were explicit that the term was meant to fill a conceptual gap. Existing terminology, they argued, captured the technical dimension of these systems but not the significance of the shift, and not in a way accessible to readers outside machine learning. "Large language model" was too narrow, because the same approach was emerging in vision, audio, code and biology. "Pretrained model" described a method without saying what the method had changed.

A reflection published by the Stanford centre in October 2021 records how many alternatives were considered and rejected, among them base model, inframodel, platform model, polymodel and pluripotent model, before a vote settled on foundation model.

The report's own explanation is that the word specifies the role these models play which is that, a foundation model is itself incomplete, but serves as the common basis from which many task-specific models are built through adaptation. The authors also chose it to signal the importance of architectural stability, on the reasoning that badly built foundations invite disaster while well-executed ones provide reliable bedrock for what follows.

That metaphor carries a warning the authors made explicit. Writing in 2021, they said they did not fully understand the nature or quality of the foundation these models provide, and could not say whether it was trustworthy. A foundation is load-bearing, which means faults in it can affect everything built above. What follows when one model sits beneath many applications will be examined later.

The dispute over the name

The term was contested almost immediately, from two directions. The first objection was that it overstates what these systems are. Jitendra Malik of UC Berkeley accepted that large language models had real practical use but rejected the idea that they could serve as a foundation for AI more broadly. As they lack sensorimotor grounding, they learn from text without direct connection to physical experience, he called them "castles in the air". Writing in The Gradient, Gary Marcus and Ernest Davis made a related argument. Relabelling pretrained language models does not make them the foundation that AI actually needs, which in their view would require causal reasoning and robust world knowledge these systems did not possess.

The second objection concerned what happens when a technology is renamed. Meredith Whittaker, then at the AI Now Institute, was reported as characterising the move as an attempt at erasure of the older term, arguing that a fresh label detaches a technology from the criticism accumulated around its predecessor and gives it a cleaner start in public discourse.

The report's answer to both objections is in the term itself. "Foundation" was chosen to signal incompleteness, not quality and that calling something a foundation does not assert that it is a good one.

Why the argument matters outside the academy

A naming dispute among researchers would matter little had the term stayed in academic papers. However, it did not. In its June 2023 negotiating position on the AI Act, the European Parliament introduced "foundation model" as a distinct legal category. The final legislation dropped the term and regulates "general-purpose AI models" instead. In the United States, Executive Order 14110 of October 2023 defined a "dual-use foundation model" by reference to broad training data, self-supervision, a parameter count in at least the tens of billions and applicability across many contexts, together with a capability test tied to national security. That order was revoked in January 2025.

The result is not merely a disagreement about words. Different instruments use different terms for broadly overlapping sets of systems, and apply different tests to decide what falls inside each category. For anyone drafting, applying or complying with AI rules, that creates a practical difficulty. What counts as a foundation model under one framework may not count under another, and the boundary of the legal category is not where the boundary of the technical one sits.

That is why the argument over a name matters. The words used to describe these systems do not only organise technical ideas. They determine what is regulated, who falls within the rules, and where responsibility sits. The next article turns to that problem of scope.

 

Sources

  1. Bommasani et al., On the Opportunities and Risks of Foundation Models, Stanford CRFM, August 2021 - definition, authorship, and the naming rationale at section 1.1.1. Report landing page.

  2. Center for Research on Foundation Models, Reflections on Foundation Models, Stanford HAI, 18 October 2021 rejected name candidates and the vote. CRFM copy.

  3. Jitendra Malik, Response on "Foundation Models", CRFM commentary, 18 October 2021 - the sensorimotor-grounding objection.

  4. Gary Marcus and Ernest Davis, "Has AI found a new Foundation?", The Gradient.

  5. Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, 2018. ACL Anthology version.

  6. Brown et al., Language Models are Few-Shot Learners, 2020.

  7. Executive Order 14110, Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence, 30 October 2023 -  full text.

  8. Regulation (EU) 2024/1689 (the AI Act), Article 3(63) - https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng . For context on the shift from the Parliament's June 2023 position to the adopted text, see Stanford CRFM, Foundation Models under the EU AI Act, and the Center for Democracy and Technology's brief on general-purpose AI models.

  9. The instrument revoking Executive Order 14110, January 2025 - https://www.whitehouse.gov/presidential-actions/2025/01/removing-barriers-to-american-leadership-in-artificial-intelligence/ 

 

Comments 0

Comments

Join the conversation. Comments are reviewed before they appear.

No comments yet — be the first to share a thought.

Leave a comment

Related reading