duo-efficacy-measurement
Prove your product actually works — with instruments you did not author, on axes that do not substitute for each other, and in a way a skeptic outside your company can check. Covers the four-axis efficacy framework (enjoyment, mastery, real-task application, external standard), borrowing the incumbent's benchmark instead of inventing a metric, authoring an item pool independently of the content you ship, thin-sampling many users to grade your content rather than your users, embedding a pre-test and post-test as a gate, adaptive assessment, partial-credit scoring, and reading calibration off the score distribution. Use when asking how do I show this actually teaches anything, what do I measure besides engagement, should I build my own test or borrow one, how many questions do I need, where does a pre-test go, why is my assessment too easy, or should I report one score.
- Version
- 2.0.0
- License
- MIT
Pinned to revision 78072c9528fb, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/duo-efficacy-measurement/SKILL.md
- skills/duo-efficacy-measurement/references/adaptive-is-shorter-not-easier.md
- skills/duo-efficacy-measurement/references/borrow-the-instrument-you-do-not-control.md
- skills/duo-efficacy-measurement/references/build-the-item-pool-independently-of-the-content.md
- skills/duo-efficacy-measurement/references/check-the-distribution-not-the-score.md
- skills/duo-efficacy-measurement/references/embed-a-pre-test-and-post-test-in-the-product.md
- skills/duo-efficacy-measurement/references/four-axes-that-do-not-substitute.md
- skills/duo-efficacy-measurement/references/partial-credit-beats-binary-scoring.md
- skills/duo-efficacy-measurement/references/thin-sample-many-people-to-grade-the-curriculum.md
Every link opens the file at its source, pinned to the revision this page describes.