WTF is AI BuildersInternal

Product · updated 2026-10-01

AI readiness quiz

Last updated: 2026-10-01. Brief, not yet built. From Sheila, 2026-10-01.

The idea

A short quiz that tells a person what they need to understand and be able to do with AI to be hired by a company that already runs on it. Tandem is the model for that company, but the quiz never says Tandem. The point it makes: AI knowledge and skills are becoming as basic as typing. You would not list typing on a resume. Soon you will not list "uses AI" either. You will just be expected to.

The quiz ends with a plan to start: three WTF is AI products matched to the gaps it found, and one prompt, free, so the plan begins before any purchase.

Who it is for

Versions

VersionWhereWhat it is
v1.1WTFISAI-APP/mockups/quiz/v1.1/Ten scenario questions, five dimensions, one obviously best answer per question. Kept for comparison. Sheila, 2026-10-01: "the first three sections are very obvious correct answers."
v1.2WTFISAI-APP/mockups/quiz/v1.2/Twenty-seven questions, five minutes. A different evaluation method, below. Current. Updated 2026-10-02 with eight more questions from Sheila and a fix to the best-and-worst interaction (no auto-advance, markers inline).

Both run locally with python3 -m http.server 4324 --directory mockups/quiz from the WTFISAI-APP folder, then http://localhost:4324.

The evaluation method, v1.2

The problem with v1.1: a multiple-choice question with one clearly right answer measures whether you can spot the right answer, which is a test-taking skill. It does not measure the skill, and it flatters everyone. v1.2 uses six instruments, each chosen because it is harder to game and because it measures a different thing.

InstrumentQuestionsWhat it measuresWhy it is hard to game
Role first2Role type, and how AI shows up at work todayNot scored. Sets the benchmark, the product fit, and the wording of the free prompt
Felt (agree or disagree)8Feels prepared. Can name how the job changes. Wants to learn and has time. Takes answers at face value. Could explain an LLM. Could explain NLP. Believes there is more than prompt engineering. Feels confident bringing subject matter expertise to AI work beyond a chatbotNever scored right or wrong. Prepared and job-impact become one number, "felt", shown next to "shown". The LLM and NLP statements are compared with the literacy score and produce a flag either way. The last two produce the "translation, not ability" flag
Confidence-weighted knowledge11Literacy: LLMs versus machine learning, what an LLM is doing when it answers, "I understand how you feel" versus consciousness, what a script is, applied AI versus the science, what a repo is, what NLP is, what a knowledge base is, two true-or-false statements about AI slop (it cannot be fixed and comes from low-quality LLMs: false; only technical people can fix it: false), who owns AI outputAfter each answer: sure, fairly sure, guessing, or "I don't know" up front. Sure and wrong scores zero. "I don't know" scores above guessing wrong. The number of sure-and-wrong answers is reported on its own as the overconfidence line. This is how the quiz assesses the taker's relationship with facts
Best and worst2Judgment and governance in situations with four plausible movesPicking the best is half the score. Picking the worst is the other half. Standard in hiring assessments because the distractors are all defensible
Fabrication spotting1Pushback on facts delivered by AIA four-line AI output. One line carries an invented statistic with an invented source. "Which line would you check first?"
Behavioral frequency2Experimentation and real useCounts in the last 30 days and hours in a typical week, not intentions
Transferable strengths1What the taker already bringsPick up to three things people come to you for. Each maps to an AI skill with a different name. Not scored. Opens the result on a strength

Scoring. Four dimensions get a bar: Literacy, Judgment, Governance, Practice. Each is a percentage of the points available in the questions that had a defensible answer. "Shown" is the average of the four. "Felt" comes from the prepared and job-impact statements. Three flags can appear under the bars: sure and wrong N times, "you said you take answers at face value and the test agreed", and "you said you rarely do and the invented statistic still got through".

Topics added in v1.2 from Sheila, 2026-10-01: AI governance, data and IP security, GitHub and repos, applied AI versus the science of AI, what "writing a script" means, LLMs and machine learning, how an LLM decides versus consciousness, feelings about AI preparation, whether they understand how their job will be affected, whether they want to learn new skills, and experimentation level.

Open for discussion:

What it measures, v1.1

Five dimensions, two questions each, ten questions. Every question is a scenario, not a definition.

DimensionWhat a strong answer showsSample question
ContextKnows that the quality of the answer comes from what you give the model"You need a first draft of a client proposal. What do you paste in first?"
Tool choiceCan pick between the main tools for a task, and knows when not to use one"A spreadsheet of 400 leads needs cleaning. Which tool, and why?"
JudgmentChecks outputs, knows what hallucination looks like, knows what not to share"The model cites a statute. What do you do before it goes in the memo?"
WorkflowHas made AI part of a repeated task, not a one-off"Name a task you do every week. Where does AI sit in it today?"
BuildingHas made something: a custom prompt, a skill, a small automation, a vibecoded page"Have you built a tool, however small, that someone else used?"

Scoring is per dimension, zero to four. The result is not a single number. It is a shape: where she is strong and where the gap is.

The result page

  1. Her shape. Five bars, plain words under each. No grades, no "beginner". The honesty rule from the plan applies: say what the answers showed, nothing more.
  2. The line. "Here is what hiring looks like at a company that runs on AI." Three or four expectations, written as a job post would write them, without naming any company.
  3. Three products. Matched by the two weakest dimensions, from the Library by id, with real prices. Free first where a free item fits, then a $29 playbook, then the 101. Same rule as the plan's gap section.
  4. One free prompt. Tailored to her weakest dimension, copy-ready, so she can start today in whatever AI she uses. This is the prompt-builder pattern from the GTM doc: answer questions, get the exact prompt for your situation.
  5. Email to keep it. See the flow below.

The flow, proposed 2026-10-01

Sheila's shape of it: the quiz is free, the results come by email, the email links to the prompt and the suggested items, signing up for the app is free, and a first-time discount code lands after the quiz. Recommended version, with one change:

StepWhat she seesWhy
1. Ten questionsNo email asked. Two minutesAsking for an email before value is where quizzes die
2. Her shape, on screenThe five bars and the plain words, right awayShow the result before asking for anything. The screen is also what gets screenshotted and shared
3. "Email me my plan"One field. The button says what the email contains: the prompt, three items, a codeThis is the capture. The shape is free to see, the plan is free to keep, both need the address
4. The emailHer shape. The free prompt, in full, copy-ready. The three items with prices. The code QUIZ10, 10% off a first purchase, 14 days. One button: Open my planEverything of value is in the email, so it earns its place in the inbox. The button is a Supabase magic link. Clicking it creates the account and signs her in. That is the "sign up for free", with no second form
5. In the appHer quiz shape is on the profile. The plan's gap section already knows her two weakest dimensions. The three items are on the Library with the code pre-applied at CheckoutThe quiz and the intake share answers. She never types the same thing twice

The one change from the original idea: the email is the sign-up, not a link to a sign-up. One click instead of an email, a link, a form, and a code. Account creation is silent and free. iOS gets the same link once the app exists, opened through a universal link.

The code. A Stripe promotion code, first-time transaction only (a Stripe setting, nothing to build), 10% off, expires 14 days after it is sent. On a $29 playbook that is $2.90 off, on the 101 it is $50. Open for Sheila: whether it applies to everything or playbooks only. Recommendation: everything, since the person who buys the 101 off a quiz is the person we most want.

What it costs to build beyond the quiz: one transactional email template in Resend, one Worker route that sends it and mints the magic link with the Supabase admin API, one promotion code in Stripe, and a quiz_results column on the profile. About a day.

Where it lives

Build notes

Open