AI readiness quiz
Last updated: 2026-10-01. Brief, not yet built. From Sheila, 2026-10-01.
The idea
A short quiz that tells a person what they need to understand and be able to do with AI to be hired by a company that already runs on it. Tandem is the model for that company, but the quiz never says Tandem. The point it makes: AI knowledge and skills are becoming as basic as typing. You would not list typing on a resume. Soon you will not list "uses AI" either. You will just be expected to.
The quiz ends with a plan to start: three WTF is AI products matched to the gaps it found, and one prompt, free, so the plan begins before any purchase.
Who it is for
- The career path first. Someone employed or looking, who suspects AI is now part of the job description.
- The business path second. An owner who hires and wants to know what to look for, or who wants to measure herself.
- Not for employers as the primary reader. Open question whether a "send this to your team" version comes later.
Versions
| Version | Where | What it is |
|---|---|---|
| v1.1 | WTFISAI-APP/mockups/quiz/v1.1/ | Ten scenario questions, five dimensions, one obviously best answer per question. Kept for comparison. Sheila, 2026-10-01: "the first three sections are very obvious correct answers." |
| v1.2 | WTFISAI-APP/mockups/quiz/v1.2/ | Twenty-seven questions, five minutes. A different evaluation method, below. Current. Updated 2026-10-02 with eight more questions from Sheila and a fix to the best-and-worst interaction (no auto-advance, markers inline). |
Both run locally with python3 -m http.server 4324 --directory mockups/quiz from the WTFISAI-APP folder, then http://localhost:4324.
The evaluation method, v1.2
The problem with v1.1: a multiple-choice question with one clearly right answer measures whether you can spot the right answer, which is a test-taking skill. It does not measure the skill, and it flatters everyone. v1.2 uses six instruments, each chosen because it is harder to game and because it measures a different thing.
| Instrument | Questions | What it measures | Why it is hard to game |
|---|---|---|---|
| Role first | 2 | Role type, and how AI shows up at work today | Not scored. Sets the benchmark, the product fit, and the wording of the free prompt |
| Felt (agree or disagree) | 8 | Feels prepared. Can name how the job changes. Wants to learn and has time. Takes answers at face value. Could explain an LLM. Could explain NLP. Believes there is more than prompt engineering. Feels confident bringing subject matter expertise to AI work beyond a chatbot | Never scored right or wrong. Prepared and job-impact become one number, "felt", shown next to "shown". The LLM and NLP statements are compared with the literacy score and produce a flag either way. The last two produce the "translation, not ability" flag |
| Confidence-weighted knowledge | 11 | Literacy: LLMs versus machine learning, what an LLM is doing when it answers, "I understand how you feel" versus consciousness, what a script is, applied AI versus the science, what a repo is, what NLP is, what a knowledge base is, two true-or-false statements about AI slop (it cannot be fixed and comes from low-quality LLMs: false; only technical people can fix it: false), who owns AI output | After each answer: sure, fairly sure, guessing, or "I don't know" up front. Sure and wrong scores zero. "I don't know" scores above guessing wrong. The number of sure-and-wrong answers is reported on its own as the overconfidence line. This is how the quiz assesses the taker's relationship with facts |
| Best and worst | 2 | Judgment and governance in situations with four plausible moves | Picking the best is half the score. Picking the worst is the other half. Standard in hiring assessments because the distractors are all defensible |
| Fabrication spotting | 1 | Pushback on facts delivered by AI | A four-line AI output. One line carries an invented statistic with an invented source. "Which line would you check first?" |
| Behavioral frequency | 2 | Experimentation and real use | Counts in the last 30 days and hours in a typical week, not intentions |
| Transferable strengths | 1 | What the taker already brings | Pick up to three things people come to you for. Each maps to an AI skill with a different name. Not scored. Opens the result on a strength |
Scoring. Four dimensions get a bar: Literacy, Judgment, Governance, Practice. Each is a percentage of the points available in the questions that had a defensible answer. "Shown" is the average of the four. "Felt" comes from the prepared and job-impact statements. Three flags can appear under the bars: sure and wrong N times, "you said you take answers at face value and the test agreed", and "you said you rarely do and the invented statistic still got through".
Topics added in v1.2 from Sheila, 2026-10-01: AI governance, data and IP security, GitHub and repos, applied AI versus the science of AI, what "writing a script" means, LLMs and machine learning, how an LLM decides versus consciousness, feelings about AI preparation, whether they understand how their job will be affected, whether they want to learn new skills, and experimentation level.
Open for discussion:
- Is 27 questions too many for a shared quiz? Five minutes is the edge of what people finish unpaid. Each instrument could lose one or two questions and the method would hold. Candidates to cut first: one of the two experimentation counts, one slop statement, and the repo question for non-technical roles.
- The seven knowledge questions have defensible answers, but the copyright one is US-specific and the "applied versus science" one is still soft. Both could use a lawyer and a researcher respectively for one minute.
- Whether the role answer should change the questions themselves, not just the benchmark and the prompt. A technical person should probably get a harder literacy set.
- Whether "felt versus shown" is the right headline, or too confronting for the cautiously curious. The copy currently treats every gap shape as common and fixable.
What it measures, v1.1
Five dimensions, two questions each, ten questions. Every question is a scenario, not a definition.
| Dimension | What a strong answer shows | Sample question |
|---|---|---|
| Context | Knows that the quality of the answer comes from what you give the model | "You need a first draft of a client proposal. What do you paste in first?" |
| Tool choice | Can pick between the main tools for a task, and knows when not to use one | "A spreadsheet of 400 leads needs cleaning. Which tool, and why?" |
| Judgment | Checks outputs, knows what hallucination looks like, knows what not to share | "The model cites a statute. What do you do before it goes in the memo?" |
| Workflow | Has made AI part of a repeated task, not a one-off | "Name a task you do every week. Where does AI sit in it today?" |
| Building | Has made something: a custom prompt, a skill, a small automation, a vibecoded page | "Have you built a tool, however small, that someone else used?" |
Scoring is per dimension, zero to four. The result is not a single number. It is a shape: where she is strong and where the gap is.
The result page
- Her shape. Five bars, plain words under each. No grades, no "beginner". The honesty rule from the plan applies: say what the answers showed, nothing more.
- The line. "Here is what hiring looks like at a company that runs on AI." Three or four expectations, written as a job post would write them, without naming any company.
- Three products. Matched by the two weakest dimensions, from the Library by id, with real prices. Free first where a free item fits, then a $29 playbook, then the 101. Same rule as the plan's gap section.
- One free prompt. Tailored to her weakest dimension, copy-ready, so she can start today in whatever AI she uses. This is the prompt-builder pattern from the GTM doc: answer questions, get the exact prompt for your situation.
- Email to keep it. See the flow below.
The flow, proposed 2026-10-01
Sheila's shape of it: the quiz is free, the results come by email, the email links to the prompt and the suggested items, signing up for the app is free, and a first-time discount code lands after the quiz. Recommended version, with one change:
| Step | What she sees | Why |
|---|---|---|
| 1. Ten questions | No email asked. Two minutes | Asking for an email before value is where quizzes die |
| 2. Her shape, on screen | The five bars and the plain words, right away | Show the result before asking for anything. The screen is also what gets screenshotted and shared |
| 3. "Email me my plan" | One field. The button says what the email contains: the prompt, three items, a code | This is the capture. The shape is free to see, the plan is free to keep, both need the address |
| 4. The email | Her shape. The free prompt, in full, copy-ready. The three items with prices. The code QUIZ10, 10% off a first purchase, 14 days. One button: Open my plan | Everything of value is in the email, so it earns its place in the inbox. The button is a Supabase magic link. Clicking it creates the account and signs her in. That is the "sign up for free", with no second form |
| 5. In the app | Her quiz shape is on the profile. The plan's gap section already knows her two weakest dimensions. The three items are on the Library with the code pre-applied at Checkout | The quiz and the intake share answers. She never types the same thing twice |
The one change from the original idea: the email is the sign-up, not a link to a sign-up. One click instead of an email, a link, a form, and a code. Account creation is silent and free. iOS gets the same link once the app exists, opened through a universal link.
The code. A Stripe promotion code, first-time transaction only (a Stripe setting, nothing to build), 10% off, expires 14 days after it is sent. On a $29 playbook that is $2.90 off, on the 101 it is $50. Open for Sheila: whether it applies to everything or playbooks only. Recommendation: everything, since the person who buys the 101 off a quiz is the person we most want.
What it costs to build beyond the quiz: one transactional email template in Resend, one Worker route that sends it and mints the magic link with the Supabase admin API, one promotion code in Stripe, and a quiz_results column on the profile. About a day.
Where it lives
- In the app at
app.wtfisai.co/quiz, signed out allowed, so it can be shared as a link. - The website gets a card and a CTA: "How ready are you, really?"
- Results write to the profile when she signs in, so the plan can use them. The two new intake questions and the quiz overlap on purpose. One answer feeds both.
Build notes
- Static questions in a manifest file, like the Library. No model call needed to score. One model call, optional, to write the free prompt from her answers, using the same plan agent.
- Reuses the Library card component for the three products and the plan's stat styling for the five bars.
- Share card: her shape as an image with the tagline. Nothing personal in it beyond the five bars.
- Estimated build: two days after the Oct 10 launch, once the Library names are final.
Open
- The name. Working titles: "The AI readiness check", "Are you hireable in the AI economy?", "The typing test".
- The exact ten questions and their scoring. Sheila writes the first draft, Claude Code tightens.
- Whether the business path gets a different question set or the same ten with different result copy.