Home
Getting Started
Study Support
Integrity & Policy
Resources
Start Here Use AI Responsibly

Understanding How AI Works

Learn the fundamentals of large language models to use them more effectively and understand their capabilities and limitations.

Why Understanding AI Matters for Students

Understanding how AI tools work helps you use them more effectively, recognise their limitations, and make better decisions about when and how to apply them to your studies. You do not need to be a computer scientist to benefit from this knowledge.

What Is a Large Language Model?

Generative language models learn patterns from training material and use supplied context to generate outputs. A model is one part of an AI system: the application around it may also provide documents, retrieval, tools or web search. See AI Platforms for current access links.

Language models operate over tokens and generate outputs probabilistically from learned patterns and supplied context. They estimate possible next tokens and select among them as generation proceeds. Producing plausible language does not mean that the model has established whether a claim is true.

๐Ÿง  What They Are

  • Pattern-matching systems trained on vast amounts of text
  • Statistical models that estimate probabilities for token sequences
  • Tools that can process and generate human-like text
  • Systems that learn relationships between concepts and words

โŒ Limits to Remember

  • Fluent language is not proof of human understanding
  • A model alone is not a verified database of facts
  • Some systems can additionally retrieve documents or search the web
  • Tool and source access do not guarantee a correct answer

How LLMs Are Trained

Stage 1: Pre-Training

During initial training, a model learns patterns from large collections of material. The sources and methods vary between models. Training can encode useful relationships as well as errors, omissions and bias.

Key Point: The model learns from examples in its training data, not from a structured database of facts. This is why it can sometimes "know" information but present it incorrectly.

Stage 2: Further Training

Many models undergo further training to shape their responses. Methods can include examples, preference feedback and reinforcement learning; the combination varies. These processes aim to improve behaviour, but do not guarantee accuracy or suitability for your task.

Why This Matters: An assistant can produce helpful-sounding, detailed language while still making factual or reasoning errors. Evaluate its claims rather than its tone.

Key Concepts Every Student Should Understand

๐Ÿ“ Tokens

Language models represent text as small units called tokens. A token might be a word, part of a word, or punctuation. During generation, the model uses the context available to it to select subsequent tokens.

Student Example: A token can represent part of a word, a whole word or punctuation. The exact split depends on the model and language; token counts are not fixed word counts.

๐Ÿงต Context Window

The context window limits what a model can consider during generation. Depending on the system, this may include instructions, conversation text and selected source passages. Uploading a document does not guarantee that all of it is included.

Practical Impact: Long conversations or documents may be shortened, summarised or retrieved in selected passages. Check important responses against the original material rather than assuming the system considered everything.

๐ŸŽฒ Temperature

Where a system exposes it, temperature is a setting that affects variability in token selection. Available settings and their effects differ between systems.

When to Care: Less variable output is not necessarily more accurate. Low temperature, clearer prompting and repeated agreement between models do not establish factual accuracy.

๐Ÿ”ฎ Probabilistic Generation

A model generates tokens using probability estimates informed by learned patterns and context. Selection need not always choose the most probable token, so repeated requests can produce different responses.

Why This Matters: Different responses may include different errors. Agreement between responses is also not independent evidence: check important claims against appropriate original or credible sources.

How AI Generates Responses

A simplified view of language generation is below. Some systems may also call tools or retrieve material before or during this process.

1

Input Processing

Your message is broken into tokens and converted into numbers (embeddings) that the model can process.

2

Pattern Matching

The model analyses patterns in your input and relates them to patterns it learnt during training.

3

Token Prediction

The model estimates probabilities for possible next tokens from learned patterns and the context available to it. Generation settings influence which token is selected.

4

Iterative Generation

Generation continues with selected tokens becoming part of the context, until a stopping condition or limit is reached.

Models, sources and tools

An AI system may search supplied documents, retrieve passages from a collection, use a calculator or other tool, or search the web. These can help you investigate a question and trace candidate evidence.

Access to tools or sources does not guarantee a correct answer. The system may select irrelevant material, misread a source or draw an unsupported conclusion. Read and check the evidence yourself using the detailed verification guidance.

โš ๏ธ What This Means for Your Studies

Next: Limitations & Best Practices โ†’ โ† Back to Home