Level 0 · Understand — the model
See the machine
What a token is, where the memory lives, why it’s fluent even when it’s wrong — and every surface from chat to the API.
01 How it actually works
It predicts the next token
Claude writes one small piece at a time, each time picking the most likely next piece. That one trick, at scale, is the whole engine.
-
Your wordsPlus everything earlier in the chat
- split
-
TokensSplit into small chunks
- read all
-
Predict oneThe likeliest next tokenappend · repeat
- repeat
-
The replyFluent by design — not checked
Prediction at scale
Reasoning, code and judgment all come out of next-token prediction.
Always plausible
It always writes a fluent continuation — even when it doesn’t know.
No lookup by default
It isn’t searching anything unless a tool, like web search, actually runs.
02 The unit of everything
Everything is tokens
A token is a chunk of text — often part of a word. Claude reads, remembers and bills in tokens, not words.
Cost and speed
Both scale with tokens — what goes in and what comes out.
It all counts
Files, tool results and the reply itself are tokens too.
Long pastes aren’t free
A big paste fills memory and slows every turn after it.
6 tokens · 4 words. Your message, your files, the tool results and the reply are all counted this way — in and out, every turn.
Going in
What Claude reads
Your message, files, pastes, tool results
Billed per token
Fills the window
Coming out
What Claude writes
The reply — every word it writes
Billed per token, too
Fills the window, too
03 Where the memory lives
Its memory is the window
The context window is everything Claude can see right now. It’s finite, and it’s re-read on every turn.
Re-read every turn
Each reply works from the whole window, top to bottom.
Full means forgetting
Fill it and the chat hits its limit or gets summarized — early detail is lost; long chats drift.
One task, one chat
No memory between chats by default (unless it’s in a Project or memory is on).
Each turn is re-read every time. Keep adding — the window only holds so much, and early detail is what gets lost. (The app’s system prompt always stays.)
04 Why it’s fluent when it’s wrong
Fluent isn’t true
With no facts to work from, Claude predicts a plausible answer — confident, specific and sometimes invented. That’s a hallucination.
Knowledge has a cutoff
Training stops at a date. Anything newer is a guess.
Ground it
Paste the facts, attach the file, or turn on web search.
Verify outside the chat
Check facts at the source — never ask it to vouch for itself.
When will our new vendor’s contract clear legal review?
Reviews like these typically take 2–3 weeks — yours should be signed off around March 14.
⚠ Confident, specific — and invented. It can’t see your legal queue, so it predicted a plausible answer instead of a true one.
Legal’s tracker says: “In review, est. 8–10 weeks, submitted Jan 20.” When will it clear?
Per the tracker, 8–10 weeks from the Jan 20 submission puts it late March to early April. I can’t tighten that without a newer status.
✓ Same model — now it reasons over facts you handed it, and flags what it still doesn’t know.
05 What Claude reads
Instructions stack up
Every reply is built from layers, broad to specific. Write the durable ones once; spend your message on the task.
- 1
System prompt
Set by the app. You don’t see or change it.
- 2
Standing instructions
A Project’s instructions, or a
CLAUDE.md(which can importAGENTS.md) — written once. - 3
Your message
This turn’s task. Keep it about the work, not the background.
CLAUDE.md / AGENTS.md at Levels 3–4.-
System promptset by the app
-
Standing instructionsa Project, or a
CLAUDE.mdyou · once -
Your messagee.g. “Draft the weekly summary…”
06 The capability surface
Same model, more reach
The ladder isn’t smarter models. It’s the same mind with more reach at each rung — tap one to see what it adds.
Search and files ground it
Current facts and your own documents beat the training data.
Connectors let it act
Reading and acting on your real tools starts at Level 2.
Agents take whole jobs
Cowork and Code run multi-step work — Levels 3 and 4.
The raw model, reasoning over what you paste.
Unlocks Draft, explain, or think a decision through.
07 One more dial
Pick the model for the stakes
Match the setting to the task, not to habit. A quick rewrite and a board decision deserve different dials.
Model tier
Which model answers
Lighter, faster models — or the most capable one
Hard, high-stakes, multi-step work
Quick rewrites, summaries, simple questions
Thinking budget
How long it reasons first
Room to reason before it answers: slower, often sharper
Plans, numbers, tricky trade-offs
Short drafts and simple lookups
The API
No app around it
The same models with no app around them, paid per token
Building Claude into a product or workflow
Anything a chat window already does
08 Do it now
Five small tests, in real Claude
-
Step 1 of 5
See your words as tokens
Once you see them, pricing, speed and limits make sense.
You should see: Your own sentence, split into the units Claude actually reads.
paste into ClaudeExplain what a token is, using this exact sentence as the worked example: "The quarterly forecast shifted." Show roughly how it splits into tokens, and why token count (not word count) is what drives cost and memory.
-
Step 2 of 5
Feel the context window
Its “memory” is just this conversation — test it.
You should see: An honest recap — and the limits of it — straight from the machine.
paste into ClaudeWithout scrolling up (you can't): what were the first and most recent things I said in this conversation? Then explain what your context window is and what happens to a chat that outgrows it.
-
Step 3 of 5
Find the cutoff
Know where knowledge ends, and when to demand search.
You should see: The boundary between knowledge and guesswork, stated plainly.
paste into ClaudeWhat is your knowledge cutoff? Give one kind of question I should never trust you on without web search, and one where the cutoff barely matters. Explain the difference.
-
Step 4 of 5
Catch a hallucination
See the failure once, safely — the verify reflex sticks.
You should see: Fluency and accuracy pulled apart in front of you.
paste into ClaudeWithout using web search: give me one precise-sounding statistic about my industry. Then critique your own answer — which parts might be fabricated, and how exactly would I check them outside this chat?
-
Step 5 of 5
Map your loadout
Every chat has different reach — know yours.
You should see: Your real capability surface, not the theoretical one.
paste into ClaudeList the tools and abilities you actually have in this chat right now — files, web search, connectors, artifacts, anything else. For each one: a single thing it makes possible that a plain chat can't do.
One step = one paste into Claude (the mobile app is perfect). Each step you mark done climbs a floor; progress saves on this device.
09 Decide
Where do you start climbing?
You’ve seen the machine. What do you want to get better at first?
Pick the closest match — you’ll get a verdict.
10 Capstone
Make Claude your tutor
- Paste the prompt into a fresh chat.
- Answer its check question after each concept.
- Ask for a real example from your own work wherever one feels abstract.
You're my tutor on how Claude actually works. Teach me, one concept at a time, checking my understanding before moving on: 1. Tokens — what they are, why they're the unit of cost and memory. 2. The context window — what you hold, what happens when it fills, why long chats drift, and why there's no memory between chats by default. 3. Training and the knowledge cutoff — and why ungrounded answers can be confidently wrong (hallucination). 4. The instruction stack — system prompt vs standing instructions (Project instructions, CLAUDE.md / AGENTS.md) vs my message. 5. The capability surface — chat, search, files, artifacts, projects, connectors (MCP), skills/slash commands, Cowork, Claude Code, and the API — one line each on what it adds. 6. Model tiers and thinking budgets — how to match the model to the task. Use concrete examples from how I'm talking to you right now. After each concept, ask me one quick question to confirm I've got it before continuing.
Wrap-up · check yourself
You’ve seen the machine
Remember
You’re ready when tokens, window, cutoff and grounding feel like tools, not jargon — and a confident answer makes you ask “grounded in what?” before you act on it.