The Machine You're Actually Talking To
By the end of this module you can say, in one sentence and without hand-waving, what happens when you send a message and something comes back. You can name four kinds of work it does well and four it is built badly for, and you can tell the two apart before you start rather than after you have wasted an afternoon. You finish with a scored list of five things you do every week and one task chosen to carry through the rest of the course.
┌───────────────────────────┐
│ everything you gave it │
│ this message + history │
│ + files + instructions │
└─────────────┬─────────────┘
│
▼
┌───────────────────────────┐
│ the most plausible │
│ continuation of that │
└─────────────┬─────────────┘
│
┌────────┴────────┐
▼ ▼
┌───────────┐ ┌───────────────┐
│ true │ │ merely │
│ │ │ plausible │
└─────┬─────┘ └───────┬───────┘
│ │
└────────┬─────────┘
▼
┌──────────────────┐
│ reads the same │
│ either way │
└──────────────────┘Nothing inside the machine separates the left branch from the right. You are the separator.
- State in one sentence what the machine is doing when it produces an answer.
- Explain why it invents citations, why it agrees with you, and why it miscounts, using that one sentence and nothing else.
- Name the four kinds of work it is genuinely strong at, with an example of each from your own week.
- Name the four kinds of work it is structurally bad at, and say what goes wrong in each case.
- Apply the task test to any piece of work: does this need knowledge it has, or knowledge only you have?
- Recite the course spine sentence from memory.
ai-map.md — a one-page map of five tasks you do every week, each scored against what the machine is good and bad at, with one task marked as the one you will carry through the next eleven modules.
Here is the honest version. It is more interesting than the dismissive one.
When you send a message, the machine takes everything sitting in front of it — your message, the conversation so far, any files you attached, any standing instructions you set up — and produces the most plausible continuation of that material. Not the truest continuation. The most plausible one. It writes a word, then reconsiders everything including the word it just wrote, then writes the next one. That is the whole loop.
You will have heard this described as "just autocomplete", usually by someone who wants to end the conversation. Ignore that. The phone keyboard that suggests "meeting" after "see you at the" is doing something in the same family, in the way a paper plane and an airliner are both aerodynamics. What this machine has read is enormous, what it has absorbed from that reading is deep enough to hold arguments, structures, formats, tones, and the shape of how ideas usually fit together — and the continuation it produces is often better than the one you would have written. Calling that autocomplete is not a warning, it is a way of not looking.
But notice what "most plausible continuation" does not include. It does not include a step where the machine stops, consults a store of established facts, and confirms that what it is about to say is so. There is no such step. There is no second system standing behind the first, checking its work. The fluency and the correctness come out of the same process, at the same moment, with nothing separating them.
Key insight: it produces the most plausible continuation of what it has been given. Plausible is usually true. It has nothing inside it that checks which one you got.
Once you hold that one sentence, you stop needing a list of quirks to memorise. Every strange behaviour you have heard about is the same fact wearing a different coat.
Why it invents citations. You ask for sources on a topic. In everything it has read, a paragraph like the one it is writing is followed by a citation that looks a certain way: an author with a plausible name, a journal with a plausible title, a year, a volume, a page range. It produces the most plausible continuation, and the most plausible continuation of that paragraph is a citation shaped exactly like a real one. Whether that particular paper exists is a different question, and nothing in the process asks it.
Why it agrees with you. You push back and say "are you sure? I thought it was the other way round." Now look at what is in front of it: a conversation in which one participant has expressed doubt. The plausible continuation of a person expressing doubt is the other person accommodating. So it accommodates. This is not politeness or a desire to please you. It is the same mechanism, running on text you just changed.
Why context changes everything. The continuation is a continuation of what it was given. Change what it was given and you change the answer, not by persuading it but by handing it a different thing to continue. This is why the same question produces a decent answer in one window and an excellent one in another. That is the whole subject of Module 03, and it is the lever with the most travel in it.
Why it cannot count. Counting is not a plausibility problem. There is no way to write out the eleventh word of a paragraph by producing the sort of thing that usually comes next. Counting requires actually counting, step by step, holding a running total. Text does not usually contain the counting, only the result, so what it produces is a number of the right size and shape. Ask how many times a word appears in a document and you will get a confident, specific, frequently wrong integer.
There is a pattern to the strengths, and it is the same pattern each time: the answer is already implied by what you handed over.
Transformation. You give it text and ask for the same content in a different shape — long to short, notes to prose, prose to a table, formal to plain, English to Spanish. The material is present. The job is rearrangement. This is where it is at its most reliable, and it is also the least demoed, because nothing about it looks like magic.
Drafting. A blank page is the hardest object in professional life and the easiest thing to hand over. It will produce a complete first version of an email, a plan, an agenda, a paragraph you have been avoiding since Tuesday. The draft is worth having even when it is wrong, because reacting to a wrong draft is faster than starting.
Explanation. Take a thing you do not understand and ask for it at a different level, with an analogy, in half the length, without the acronyms. Explanations of well-covered subjects are dense in what it has read, and the plausible continuation of "explain X simply" is a good explanation of X.
Extraction. Pull the dates out of the contract, the decisions out of the transcript, the addresses out of the mess. You supply the document, so the facts are in the room. This is the highest-volume real job most people have, and Module 08 is entirely about doing it properly.
These are not gaps waiting for the next version. They come from the same mechanism the strengths come from.
Facts it was not given. If the answer is not in your message, your files or your instructions, it produces the most plausible version of that answer instead. Specific figures, someone's job title, what a particular clause of a particular contract says: all of these come back fluent and sometimes fabricated.
Arithmetic and counting. Covered above. Treat any number it produced by reasoning in prose as unverified until you check it, and give it a calculator or a spreadsheet when one is available.
Recency. It learned from text collected up to a point, and after that point it knows nothing unless you or a web-access feature hand it the news. Ask about something from last month and, without a source in front of it, it will produce a plausible last month.
Anything about you. Your team, your deadlines, your customer, your quality bar, your history with this document. It has never encountered any of it. There is no way for the machine to guess your context correctly, and it will not say that it is guessing.
Warning: the four weaknesses do not announce themselves. A fabricated case citation, a miscounted total and a real quotation from a real source all arrive in the same confident register, in the same paragraph, with the same punctuation.
You have spent your whole life reading text written by people, and you have an instinct calibrated on that: writing that is clear, well-organised and specific is usually writing by someone who knew what they were talking about. Vagueness signals uncertainty. Confidence signals competence. That instinct has served you well and it is now actively dangerous, because this machine produces the register of competence whether or not the content is right. Fluency was never a signal about truth. It was a signal about the author, and the author has changed.
The practical consequence is that you cannot use your feeling of being convinced as a check. When something reads well, that tells you it reads well. Verification has to be a separate action you take on purpose, which is why Module 04 makes you write your checks down before you run the task.
The most common way to start is to put one question to three services and read the answers side by side, looking for the one that sounds best. That test measures the wrong thing. The question is unspecified, so what gets ranked is house style: which one opens with a summary, which one reaches for bullets, which one hedges. None of that predicts which will be useful on Thursday with your own material in front of it. At this stage the gap between the leading tools is much smaller than the gap between a vague request and a good one, and the ranking changes every few months anyway, so whatever you conclude tonight expires.
The skills in this course transfer. A brief with five parts works everywhere. A context file works everywhere. A verification habit works everywhere. Pick one mainstream tool, learn its actual surfaces, and get good. Module 01 is that decision, made once, in about ten minutes.
Before you hand anything over, ask one question: does this need knowledge it has, or knowledge only I have?
If the work needs general knowledge — how a decent apology email is structured, what a project plan usually contains, how to explain compound interest — it has that, and you can hand it the job with a light brief. If it needs specific knowledge — what your director actually said in that meeting, what your renewal date is, why this customer is already annoyed — then it does not have it, and you have two options: supply the knowledge, or do it yourself. What you must not do is hand over a task that needs your knowledge without supplying it and then judge the output. That is not a test of the tool. That is a test of whether it can guess, and it will always guess.
When this test is the wrong tool. Do not run it on small things. If the task takes two minutes by hand, scoring it costs more than doing it, and you will end up with an elaborate map of trivia. The test earns its keep on work that is recurring, or long, or done badly under time pressure. Run it on those five weekly tasks, not on every email.
Learn this. It is the order of the next eleven modules, and Module 11 asks you to recite it.
Point it at an outcome, give it your context, check what comes back, keep what worked — and only then let it act.
Read it left to right. Each phrase is a part of the course, and the parts are in the order that people who get real value out of this actually learned them. Note where it ends. Letting it act is last, on purpose, even though it is the part that demos best and the part most courses open with. Everything before the dash is what makes the part after the dash safe.
Fifteen minutes, a blank document, your own week. No account required for the first four steps.
- Every one of your five entries names a specific recurring task, not a category of work.
- At least one entry names a specific missing fact — a number, a name, a document — rather than saying "it doesn't know my context".
- At least one entry is honest that the task is mostly knowledge only you have.
- Exactly one task is starred, and you could say in one sentence why that one.
- The file exists on your computer as
ai-map.md, not in a chat window.
Use this after you have done the lab by hand, not instead of it. Its job is to argue with your scoring: it will push back on tasks you were optimistic about and flag ones you dismissed too quickly. Paste it into any mainstream chat tool with your five tasks in the slot. Read the output as a second opinion from something that does not know your job, because that is exactly what it is.
<task> Audit a list of recurring work tasks for how well each one fits what a large language model is actually good at, and push back where the person scoring them has been too optimistic or too dismissive. </task> <context> The person's work in one line: [describe what you do, who for, and the kind of material you handle day to day] </context> <input> Their five recurring weekly tasks, one per line: [paste the five tasks from your lab, one per line, each specific enough that a colleague would recognise it] </input> <instructions> For each task, decide which of these it mostly is: transformation, drafting, explanation, extraction, or none of them. Then name every weakness it would run into from this list: facts the model was not given, arithmetic or counting, recency, or knowledge specific to this person and their organisation. For each weakness you name, state the exact missing item — the specific document, number, name or decision that the model would not have. Then answer, for that task: does it need general knowledge, knowledge only this person has, or both. Finally, rank all five by how much value there is in handing the task over, and say which single task you would start with and why. </instructions> <output_format> A markdown table with one row per task and these columns: Task, Strength type, Weaknesses, Exact missing input, Whose knowledge, Start here (yes/no). After the table, three short paragraphs under the headings "Where you were too optimistic", "Where you were too dismissive", and "Start with this one". No preamble before the table and no summary after the paragraphs. </output_format> <rules> Do not invent details about the person's job, tools, industry or colleagues beyond what the context line states. If a task is too vague to score, say so in its row instead of guessing what it means. Do not use percentages, scores out of ten, or any number that is not a rank from 1 to 5. Do not recommend a specific product or service. If a task looks like it needs judgement, relationships or accountability that a person must hold, say so plainly rather than finding a way to include it. </rules>
Take your five tasks to a colleague who does similar work and read them out with your scoring. Ask them one question about each: what would someone need to know to do this that is not written down anywhere? The answers are the missing context, and the task with the shortest answer is the one to start with. This is slower than the prompt and usually better, because your colleague knows the things the machine cannot.
You now have a scored map of your own week and one starred task. Keep the file somewhere you will find it again, because three later modules ask for it. Module 01 creates the folder this file will live in for the rest of the course, and moves it there.
Save as: ai-map.md — moved into your kit in Module 01, read again in Module 10.
Say these out loud. If either one stalls, reread the section named beside it.
- I can say in one sentence what the machine is doing when it answers me. (What It Is Doing When It Answers)
- I can name one task on my list I should not give it, and why. (The Task Test)
- I can recite the spine sentence. (The Sentence the Rest of the Course Hangs On)
Module 01, Your Working Setup, has you choose one tool, create the ai-kit/ folder, move ai-map.md into it, and write ai-kit/about-me.md — the file every module after this one reads.