Introduction
The main goal of generative AI shows up clearly in one of the first AI-generated images to win an art competition. In August 2022, Jason M. Allen, a Colorado game designer, entered “Théâtre D’opéra Spatial” in the digital art category at the Colorado State Fair. The piece took first prize.
Allen made the image with Midjourney. He typed at least 624 prompts before the output matched the picture in his head. In 2023, the US Copyright Office refused to register the work. The Office ruled the human contribution too small to count as authorship.
The ruling treated Midjourney as the source of most of the image. Midjourney didn’t copy an existing painting to make it. The model built a new picture from patterns learned in training, which is the main goal of generative AI. This article explains how large language models fit the goal, what tokens are, and where training data comes from.
What Are Generative AI and LLMs?
Generative AI and LLMs are two levels of the same technology. Beginners often mix up the two terms.
What is Generative AI?
Generative AI is a type of artificial intelligence able to produce new content, including text, images, audio, video, and code. A generative model studies large volumes of examples and learns the statistical patterns within them. The model then creates new material following the same patterns. Marketing copy, product photos, and working software are all created this way.
What is a Large Language Model (LLM)?
A large language model is a generative AI system trained on text. An LLM learns from billions of sentences. When someone uses it, the LLM predicts the next piece of text in a sequence, one piece at a time. This is how ChatGPT, Claude, and Gemini answer a question or draft an email.
The word “large” refers to the parameter count. Parameters are the internal values a model adjusts during training, and current LLMs carry billions of them.
How Are Generative AI and LLMs Related?
Every LLM is generative AI. The reverse doesn’t hold, since plenty of image and audio systems carry no language model at all. Generative AI is the parent category. The branches below it split by output type.
Language models handle text. Diffusion models cover images and video, while speech models cover audio. Coding assistants belong to the language branch, because a model treats code as one more kind of text. That’s why prompt engineering habits carry over from writing prose to writing software.

What is the Main Goal of Generative AI?
The main goal of generative AI is to create new, original content such as text, images, audio, code, or video by learning patterns from existing data. A generative model builds an output piece by piece. Each piece follows from what usually comes next in the training material. The model never looks up an answer in a database.
What is the Primary Goal of a Generative AI Model?
The primary goal of a generative AI model is to produce new outputs that could plausibly belong to the training set. Train a model on cat photos, and the model should make a cat photo nobody has taken yet. A copy of a photo from the training set counts as a failure.
Why Does Generative AI Invent Things?
Generative AI invents things because the model aims for plausible output rather than true output. A 2023 federal court case, Mata v. Avianca, shows the difference.
Steven Schwartz, a New York personal injury lawyer, used ChatGPT to research the case. Schwartz wrote a brief citing six court decisions. His colleague Peter LoDuca signed and filed the brief. None of the six decisions existed.
ChatGPT had invented the case names, the quotes, and the rulings. Judge P. Kevin Castel later fined Schwartz, LoDuca, and their firm $5,000.
The invented cases weren’t a bug. They followed from what the model is built to do. Schwartz asked for supporting case law, and ChatGPT produced text shaped like case law. The shape was the target, so whether each citation was real never entered the calculation. The model is tuned for plausibility, which leaves it no way to tell a real citation from an invented one.
Checking is a separate job. Retrieval systems, external tools, and human review handle accuracy. None of the three comes free with the model. Even the model with the best benchmark scores will still invent details. So, the useful question for any new deployment is where wrong output gets caught.

Generative AI vs Traditional AI Goals
Traditional AI answers questions about data already in hand, while generative AI adds new material to the pile. The table compares the two goals across five dimensions.
|
Dimension |
Traditional (Discriminative) AI |
Generative AI |
|---|---|---|
|
What the Model Learns |
Boundaries between categories |
The structure of the data |
|
What the Model Produces |
A label, score, or prediction |
New text, images, audio, or code |
|
Typical Task |
Spam filtering, credit scoring |
Drafting, summarizing, image creation |
|
Right Answer Exists |
Usually yes |
Usually no |
|
Failure Mode |
Wrong label |
Plausible invention |
The two approaches work well together. A support tool might use a discriminative model to route each ticket and a generative model to draft the reply.
What is the Key Feature of Generative AI?
The key feature of generative AI is the ability to produce new content rather than only analyze existing data. A spam filter labels an email, while a generative model writes a new one.
So, what is a key feature of generative AI beyond raw creation? The answer is plain-language control. A user describes the outcome in ordinary words. The system handles the rest, so nobody needs a query language or a design tool.
AI integration projects often start with plain-language control. A natural language front end is one of the easiest features to add to existing software.
Plain-language control now works across formats too. Current models accept an image and return text, or accept text and return audio, inside one system. This mix of input and output types is called multimodality.
Context awareness works alongside multimodality. A model can hold a long document in working memory. The model then answers questions from the document rather than from general knowledge.
Adaptability is the last of the main features. The same base model behaves differently after fine-tuning on company documents. A change of wording in the prompt shifts the model’s behavior too, so one model can serve many jobs.
Each of these features comes with limits. Generative AI doesn’t verify facts. A model also knows nothing about events after its training cutoff. The model reaches a live database only when engineers connect one. The model remembers yesterday’s conversation only when the product stores the history.
Which of the Following Best Describes Generative AI?
Generative AI is best described as a system able to create new text, images, audio, code, or video by learning patterns from existing data. Certification exams and interview screens often ask the question in multiple-choice form, like the one below.
- A system able to sort incoming email into spam and inbox
- A system able to create new text, images, audio, code, or video by learning patterns from existing data
- A robot arm on a production line following fixed instructions
- A dashboard showing sales trends from last quarter
Option 2 is correct. Sorting email is classification, where the model picks a label from a fixed set. A robot arm on fixed instructions is rule-based automation, closer to software bots than to a generative model. A sales dashboard is analytics, summarizing numbers already recorded. Only option 2 describes a system producing material nobody wrote.
Which Combination of Tools Constitutes Generative AI?
Generative AI is built from a combination of machine learning, deep learning, neural networks (especially transformers), large training datasets, and heavy computing power. No single component produces a working system alone. The table shows what each one contributes.
|
Component |
What the Component Contributes |
|---|---|
|
Machine Learning |
The method of learning rules from examples |
|
Deep Learning |
Stacked layers, so subtle patterns survive |
|
Transformers |
Attention across a whole sequence at once |
|
Training Data |
The examples every pattern comes from |
|
Computing Power |
GPU clusters and weeks of training time |
Eight Google researchers published the transformer architecture in 2017. Attention is the central idea inside the transformer. Attention lets a model weigh every word in a sequence against every other word, which makes modern text generation practical.
Scale matters as much as design. Train the same architecture on a small corpus, and the result is a weak model. That’s why the table includes training data and computing power alongside the algorithms.
How is a Generative AI Model Trained?
A generative AI model is trained in three stages. Pretraining comes first. Engineers feed enormous volumes of raw text through the model until the model absorbs the structure of language.
Fine-tuning then narrows the model’s behavior with smaller, curated datasets. Human feedback comes last. People rank candidate answers, and the rankings nudge the model toward useful, safe replies.
Fine-tuning is only as good as the curated datasets behind it. Building a clean dataset is a data engineering problem before it’s a modeling problem. Reliable data pipelines decide how clean the dataset ends up.

What is a Token in Generative AI?
A token is a small unit of text that a model reads and generates. A single token can be a short word, a fragment of a longer word, or a punctuation mark. A model never sees raw text. The model reads its input as a sequence of tokens. The model writes its output as tokens too, which then convert back into readable text.
How Tokenization Works
Tokenization splits text into pieces drawn from a fixed vocabulary. Common words usually survive whole, while rare words, names, and code get chopped into fragments. This means an unusual term costs several tokens where an everyday word costs one.
OpenAI’s guidance on token counts puts a single token at roughly four characters of English, or about three-quarters of a word. At that rate, 100 tokens run to about 75 words. For a longer example, the same guide counts the US Declaration of Independence at 1,695 tokens.
Language changes the math. In OpenAI’s example, the Spanish phrase “Cómo estás” takes five tokens for 10 characters, so non-English text often costs more to process. Both counts come from older GPT tokenizers. Newer tokenizers split the same text into fewer pieces.

Why Do Tokens Matter?
Tokens matter because context windows, prices, and output caps are all measured in them.
A context window sets how much a model can read at once. A 200,000-token window holds roughly 150,000 words of input and conversation history combined. API providers charge per token and count input and output separately. Output length carries a cap as well, which is why a long answer sometimes stops mid-sentence.
Where Does Generative AI Get Its Data?
Generative AI gets data from public web pages, books and articles, licensed datasets, code repositories, human feedback, and whatever a user supplies at the moment of use. Model builders blend the sources and filter heavily. Most builders keep the exact recipe private. The table lists each source, what the source supplies, and its main concern.
|
Source |
What the Source Supplies |
Main Concern |
|---|---|---|
|
Public Web Pages |
Breadth of topics and language |
Quality and bias |
|
Books and Articles |
Long-form structure, edited prose |
Copyright |
|
Licensed Datasets |
News archives, stock libraries |
Cost and coverage |
|
Code Repositories |
Programming syntax and conventions |
License terms |
|
Human Feedback |
Rankings steering tone and refusals |
Rater bias |
|
User Input |
The prompt and files sent at runtime |
Privacy and retention |
For public web pages, most models start with Common Crawl. The nonprofit has crawled the open web since 2008. Its archive now holds more than 10 petabytes. Common Crawl publishes a new crawl of more than two billion pages roughly once a month, free to download.
Raw crawls arrive messy and duplicated, so data collection and cleaning absorb most of the effort long before training starts.
Data work dominates inside companies too, even when they build on a hosted model. Private company knowledge reaches the model through retrieval rather than training. As a result, data extraction from contracts, tickets, and PDFs becomes the real project.

Data Privacy and Copyright Concerns
The source list raises two concerns, copyright and privacy. Publishers, authors, and artists argue that training on copyrighted work needs a license. Courts in the US and UK are still working through the argument.
Privacy is the second concern, since a public chatbot may keep anything a user types into it. Companies with regulated records need data management rules for prompts and uploads, not just for databases and warehouses.
What is Regenerative AI and How Does the Term Differ?
Regenerative AI is not an established category in computer science. People using the phrase usually mean generative AI. The similar-sounding word is an easy mix-up.
A second group uses the phrase on purpose. For this group, regenerative AI means a system that keeps improving after deployment through live feedback, rather than through a scheduled retrain. The usage appears in vendor writing and trade articles. No research field, standards body, or benchmark carries the name.
A third use comes from outside computing. “Regenerative” already describes fields like medicine and agriculture, so AI used in medicine or agriculture sometimes picks up the adjective.
The table compares the two terms.
|
Dimension |
Generative AI |
Regenerative AI |
|---|---|---|
|
Standing |
Established technical category |
Informal, no agreed definition |
|
Main Aim |
Produce new content |
Keep improving after deployment |
|
Where the Term Appears |
Research papers and products |
Vendor blogs and trade articles |
|
Technical Basis |
Generative models such as LLMs |
Online and reinforcement learning |
The two labels aren’t rivals. Generative AI names a real technology, while regenerative AI names a wish for self-improving software. Today, engineering teams build self-improving behavior with AI agents, feedback loops, and scheduled retraining.
Real-World Examples of Generative AI
Generative AI already sits inside tools most people use at work. Support chatbots answer customer questions in plain language. The chatbots pass hard cases to a person.
Creative and technical tools use the same models. Image generators produce concept art and product shots from a written description. Coding assistants finish functions and explain unfamiliar code. Writing tools turn a brief into outlines, emails, and product descriptions at volume.
Inside companies, the same models sit behind document summarizers and meeting notes. AI automation flows also use them to decide how to route an invoice or a support ticket.
Conclusion
Jason Allen’s prize-winning image and the six invented cases in Mata v. Avianca came from the same behavior. Both models produced new material that fit the patterns in their training data. Creation is the goal of generative AI, so accuracy has to come from somewhere else.
Everything above points to one design choice for any real deployment: a separate layer for checking. Retrieval, evaluation, and guardrails do the checking. Each of these pieces sits inside the scope of a generative AI engagement at Data Prism.
Book a Free 30-Minute Meeting
Discover how our services can support your goals — no strings attached. Schedule your free 30-minute consultation today and let's explore the possibilities.
Book a Free Call