Assessment design

SkillMedia

Lets your agent write and review quiz and test questions that actually measure the skill they claim to measure.

Instructions available. Your AI can read the instructions. Execution depends on the setup they require.

Add ahel to your AI once: Claude, ChatGPT, Cursor, Claude Code or Codex. Then ask it to use this.

Then ask your AI: use the Assessment design skill

About this skill

Builds questions and assessments that measure what they claim, matching item type to the thing being assessed, writing distractors that diagnose, and telling a question that measures learning from one that measures reading speed or test-taking. Use this to write or review assessment items, design a

What this skill tells your AI

The instructions your AI receives, as published by cbrock84/headcount in verticals/education/skills/education/assessment-design/SKILL.md and read by ahel’s review.

Every question measures something. The work is making sure it measures the thing you meant, because the alternatives — reading speed, familiarity with the format, willingness to guess — are always available and are usually easier for the student to use.

Name the inference before writing the item

An assessment is an argument: the student did this, therefore they can do that. The argument is where assessments fail, and it fails silently.

Write down the claim first — "can decompose a two-digit number into tens and ones" — then ask what performance would be evidence for it, and what performance would be evidence against. An item that a student who lacks the skill can still get right is not evidence. An item that a student who has the skill can still get wrong, for reasons unrelated to it, is worse: it produces a false negative that gets acted on.

The most common unrelated reason is reading. Any item whose stem is harder to read than the skill is to perform has quietly become a reading assessment.

Match the item type to the claim

  • Selected response (multiple choice, matching, true/false) is efficient and can only ever provide evidence of recognition. A student who can recognize the correct answer cannot be assumed to produce it.
  • Constructed response shows the path, which is what makes partial understanding visible. It costs scoring time and needs a rubric written before the responses arrive, not after.
  • Performance tasks are the only honest evidence for anything described as applying, investigating or designing — most NGSS performance expectations and every C3 inquiry, for instance, cannot be assessed by selected response at all.

Mismatch is the usual failure: a standard describing explanation assessed by a multiple-choice item, which measures whether the student can pick an explanation someone else wrote.

Distractors are the diagnostic

In a well-built multiple-choice item, each wrong answer is the result of a specific, predictable error. Then the pattern of wrong answers says what to reteach, and the item earns its place.

Distractors that are merely wrong — a random number, an obviously absurd option — turn a four-option item into a two-option one and tell you nothing beyond right or wrong.

The mechanical tells of a weak item, all of which students learn to exploit long before they learn the content:

  • The longest or most qualified option is correct.
  • One option is grammatically inconsistent with the stem.
  • "All of the above" appears, and is usually correct.
  • Two options are synonyms, so neither can be right.
  • The correct answer repeats wording from the stem.

One item is not evidence

A single item carries noise — a misread word, a slip, a lucky guess — that swamps the signal for any individual student. Inferring mastery from one response is the most common measurement error in classroom material, and the most consequential, because it gets recorded.

Several items per claim, varied in surface form so that recognition of the format is not what is being measured. If a claim is worth recording against a student's name, it is worth three items.

Spacing matters as much as quantity: performance on a skill immediately after it is taught measures something closer to short-term recall than to learning. The same item a fortnight later measures more.

Formative and summative are different products

Formative assessment exists to change what happens next, which means it must be quick, frequent, low-stakes, and read immediately. An assessment that takes a week to score cannot be formative whatever it is called.

Summative assessment exists to record a judgment, which means it must be defensible: enough items, a rubric written in advance, and conditions that were the same for everyone.

Material sold as practice is usually formative in function and summative in appearance, which is worth stating plainly to a buyer. education:learning-materials-design covers the practice sequence this sits inside.

Sources

references/sources.md in this skill lists the outside authorities that settle the questions here — what each one is authoritative for, and what you may do with it. Check them before answering on anything they cover, and cite what you used. Most are free to read and not free to reproduce; the use note on each is binding.

Never

  • Write an item before naming the claim it is evidence for.
  • Assess an explanation standard with a recognition item.
  • Build distractors that are wrong without being diagnostic.
  • Record mastery from a single response.
  • Make the stem harder to read than the skill is to perform.

Signals

GitHub stars
2k
Forks
262
Last commit
Sep 2026
Advanced
Item type
skill
Key
assessment-design
Source
github.com/cbrock84/headcount