Pre/post assessment

Kirkpatrick level 2 (learning). Administer before Session 1 and after Session 3. Use different scenario examples at post-test (same constructs). Participant-held codes link responses without names.

Reference: Kirkpatrick, J. D., & Kirkpatrick, W. K. (2016). Kirkpatrick’s Four Levels of Training Evaluation. ATD Press.

See also: session evaluation form (level 1 reaction) and 30-day transfer follow-up (level 3 behaviour, self-reported). Level 4 (organizational results) is outside this kit’s scope.

Part A: Performance scenarios (primary measure)

Score each scenario with the four-point skills rubric (maximum 4 points per scenario). Primary proficiency indicator: total rubric score on the exit scenario set.

Pre-test scenario (Session 1 construct)

Synthetic case excerpt: A draft SMS says “Road B is open. All entrances are step-free. Go there now.” The bulletin says “Road B status unconfirmed. Entrance E only is step-free. Await authorized alert.”

Rubric row Participant evidence (score 0 or 1)
Input/use decision
Source verification
Bias or omission
Human control

Pre-test total (0 to 4):

Post-test scenario (Session 3 construct)

Synthetic brief: A partner offers an AI analytics platform for a 2,400-household registration dataset with names and vulnerability flags. Your organization has no AI governance framework.

Rubric row Participant evidence (score 0 or 1)
Input/use decision
Source verification
Bias or omission
Human control

Post-test total (0 to 4):

Organizational proficiency target (example): 4/4 on post-test scenario, or improvement of 2+ points from pre to post paired samples.

Part B: Self-efficacy (secondary measure)

Use a 5-point scale: 1 = not at all confident, 5 = very confident.

  1. I can decide when AI is appropriate for a work task.
  2. I can check an AI output against a trusted source.
  3. I can protect sensitive data when using digital tools.
  4. I can explain my AI use decision to a colleague or manager.
  5. I can identify when to escalate to privacy, protection, or IT.

Report mean and distribution. Treat as secondary to rubric scores. A rise of 0.5 points or more may indicate increased confidence, but it is not proof of workplace behaviour change.

Part C: Knowledge items (supplementary, 12 items)

Optional multiple-choice check. Score 1 point per correct item. Use different stems at post-test.

Sample pre-test items

  1. Data: A colleague wants to paste beneficiary names into a free online chatbot to draft a letter. What is the best first response?
      1. Stop and check whether the tool is approved for personal data ✓
      1. Help them write a better prompt
      1. Anonymize only the surnames
      1. Proceed if the chat is deleted afterward
  2. Verification: An AI summary says “78% of households lack clean water.” What should you do before using this number?
      1. Round it to 80% for clarity
      1. Trust it if it sounds reasonable
      1. Find the original source and check date ✓
      1. Add a footnote saying “AI generated”
  3. Translation: A French SMS alert says roads are open when the English source says they are closed. What is the highest-priority risk?
      1. Spelling errors
      1. Wrong tone
      1. Character limit
      1. Dropped or reversed negation affecting safety ✓
  4. Human control: Who is accountable when AI supports a decision about aid eligibility?
      1. The AI vendor
      1. A named organizational decision-maker ✓
      1. The person who typed the prompt
      1. No one if the AI confidence score is high
  5. Tool: Your organization has not approved any generative AI tools. A staff member uses one for internal meeting notes. Best classification?
      1. Amber ✓
      1. Green
      1. Red
      1. Green if notes are deleted
  6. Protection: A chatbot tells a beneficiary where to collect cash assistance. The location is wrong. Primary harm?
      1. Embarrassment
      1. Extra travel cost only
      1. Brand damage
      1. Exclusion or safety risk from acting on false information ✓
  7. Purpose: AI is proposed to draft a donor report from last year’s files. First question?
      1. Which model is best?
      1. Can we finish by Friday?
      1. What is the manual alternative and is AI necessary? ✓
      1. Who has the longest prompt guide?
  8. Governance: A pilot AI tool will classify incoming calls. Minimum governance step?
      1. Buy the cheapest licence
      1. Define success, failure threshold, human override, and exit plan ✓
      1. Train all staff on prompting
      1. Publish a press release
  9. Sharing: A partner asks you to upload a shared spreadsheet to an AI tool they provide. You should:
      1. Upload if the partner insists
      1. Upload only rows without names
      1. Refuse all partner tools
      1. Check your organization’s approval and data-sharing agreement ✓
  10. Bias: An AI-generated needs summary mentions only urban areas. What principle applies?
      1. Harmful omission and exclusion ✓
      1. Efficiency
      1. Translation quality
      1. Cost recovery
  11. Monitoring: When should you stop using an approved AI tool?
      1. Never, if it was approved once
      1. Only when the licence expires
      1. When it fails a set threshold, terms change, or harm is reported ✓
      1. When a newer model launches
  12. Accountability: Classroom role play about consultation with affected people is:
      1. A substitute for real consultation
      1. Legally sufficient
      1. Practice only; it does not replace consultation with affected people ✓
      1. The same as a community meeting

Post-test

Replace examples (e.g., procurement memo, protection note, vendor pitch) while testing the same constructs. Swap distractors. Do not reuse identical stems.

Scoring guide

Section Scoring Report Role
Performance scenarios 0 to 4 per scenario Mean pre and post; paired gain Primary
Self-efficacy 1 to 5 per item Mean per item pre and post Secondary
Knowledge (optional) 0 to 12 Mean, median, % at 10+ Supplementary

How to read results with small samples

Privacy

Store anonymous codes only. Do not collect names on the same form as scenario answers unless your organization has a separate consent process.