AI, Explained Plainly

How a chatbot writes an answer, one piece at a time

It starts by cutting text into pieces

Before a language model reads anything, the text gets split into chunks. A common word can be one chunk, and a rare word gets broken into several smaller ones. GPT-3 splits text with a version of a method that translation researchers adapted for this job in 2016: start with single letters, find the pair that shows up together most often, merge that pair into one new chunk, and keep repeating until there's a fixed number of chunks. Rico Sennrich, Barry Haddow and Alexandra Birch adapted it from a data compression technique, and the method is called byte pair encoding. Each chunk is called a token. In the 2020 paper describing GPT-3, the authors put the average at about 0.7 words per token.

Training is practice at guessing the next token

GPT-3 was trained on 300 billion tokens, drawn mostly from a filtered copy of web pages, plus books and Wikipedia. During training, the model reads text and keeps trying to predict the next token, and training adjusts the model's internal numbers so its predictions get closer to the text it's reading. GPT-3 has 175 billion of those trainable numbers, and they're called parameters.

The researchers who later turned GPT-3 into an instruction-following model describe the goal of that first stage plainly: predicting the next token on a web page from the internet.

Writing an answer is the same guess, repeated

When you type a question, the model predicts one token, adds it to the end of the text, and then predicts the next token using everything so far, including the token it just wrote. It does that again and again until the answer is finished. The 2017 paper that introduced the transformer, the design these models are built on, describes the model consuming the symbols it already generated as input when generating the next one. That one-token-at-a-time process is called autoregressive.

There's a limit to how much text the model can take into account at once. For GPT-3 it was 2,048 tokens, and that limit is called the context window.

A chatbot gets a second round of training

In 2022, researchers at OpenAI wrote that making a model bigger doesn't, on its own, make it better at doing what the user asked. To train that in, they hired about 40 contractors to write example answers and then rank different answers from the model from best to worst. They trained the model to produce the kind of answer those raters preferred, and called the result InstructGPT. Raters preferred answers from a 1.3 billion parameter InstructGPT model over answers from the 175 billion parameter GPT-3, even though it's more than 100 times smaller.

The same paper reports that InstructGPT can still make up facts. On tasks where the answer should only use information given in the question, like summarizing a document, it added information that wasn't there in 21 percent of cases, compared with 41 percent for GPT-3.

What this means when you use one

Neither training stage involves looking anything up. The first rewards text that fits what usually comes next, and the second rewards answers that raters liked. So a chatbot's answer reads smoothly whether or not it's correct, because smooth, plausible text is what both stages trained it to produce. Some AI tools add a separate step that searches the web before answering, which gives the model current pages to work from.

In practice, use a chatbot to draft, outline and rephrase, and check any specific number, date, name, price or policy in its answer against a source you trust before it goes to a customer.

Next step

If you'd like to see a tool that adds that search step, DataPsych's AI Content Strategist searches what's trending in your niche when you ask, then writes hooks and scripts around it. Its page lays out what it does and what it costs: datapsych.net/ai-strategist.

Sources