How many repetitions, and what evidence of transfer, should count as "mastered" for pencil control, decoding, and counting?
The question
"What does early-childhood-education research say about appropriate independent-mastery thresholds (repetition count, cross-context transfer) for skills like pencil control, decoding, and counting?"
Context: Scribble Works packs currently choose repetition count and advance timing by product guess. This brief looks for sourced numbers and turns them into a default rule.
What we already know (from the vault)
- The skill-tree proposal already rejects a universal "3 plays" or "80% correct" rule. Its pilot heuristic is "independent success on more than one occasion, plus a different example/context, followed by parent confirmation", and it labels that heuristic as unvalidated. It also states "one repeated worksheet cannot establish transfer" and "repeated play alone never promotes mastery" ([[skill-tree-proposal]]).
- The pre-v2 printables research gives age benchmarks for pre-writing strokes from occupational therapy (OT) sources: vertical line mastered around 3, horizontal line and circle around 3, cross 3.5-4, square 4, diagonals 4.5, X and triangle 5. Imitation comes before copying ([[2026-08-30-printables-inspiration-research]]). Those are age norms, not repetition thresholds.
- Cohort feedback exists, but no pack has been tested against any mastery criterion yet ([[2026-09-05-cohort-feedback-synthesis]], [[2026-09-01-cohort-week1-kit]]).
What the web says
Accuracy criteria (applied behavior analysis, ABA). This is the only body of work that has directly tested "what threshold predicts retention."
- Fuller & Fienup (2018) taught literacy skills to 50%, 80%, and 90% criteria and probed weekly for 3-4 weeks. The 90% criterion produced the best maintenance (Fuller & Fienup 2018).
- Richling et al. (2019) compared 60-100% criteria, each held across three consecutive sessions, with four children aged 6-9 with developmental disabilities. Only skills mastered at 100% across three sessions maintained several weeks after teaching stopped (Richling et al. 2019, JABA). The same group's descriptive review found that the field's common default sits near 80% and that maintenance is rarely measured (Richling, Williams & Carr 2019 descriptive analysis; PDF returned 403, so the exact modal criterion is from prior knowledge and not re-verified). Follow-up work compares how many sessions the criterion must hold for (mastery criterion frequency comparison; Pitts 2021). Across this literature, "80% once" is the weak end: at least 90% held across more than one session maintains better. Caveat: small samples, mostly children with developmental disabilities, and one-to-one teaching.
Handwriting and pencil control (OT dosage).
- Hoy, Egan & Feder (2011, Canadian Journal of Occupational Therapy), systematic review: whatever the approach, interventions without actual handwriting practice, and interventions with fewer than 20 practice sessions, were ineffective (Hoy et al. 2011). This is a dosage floor for a skill family (legibility), not a per-letter threshold. It measures over weeks, not within one sitting.
Decoding and orthographic learning.
- Nation, Angell & Castles (2007): 8-9-year-olds read novel words 1, 2, or 4 times. Learning rose with exposures, and some one-trial learning lasted. Context (meaningful text vs isolated word) made no difference (Nation et al. 2007). Share's earlier self-teaching work, in Hebrew-reading second graders, found orthographic learning after about 4 successful decodings (from prior knowledge, not re-fetched this run). These numbers are for children who can already decode. They answer "how many times must a decoded word be seen" (low single digits), not "how many drills make a 4-year-old a decoder."
- The National Reading Panel (2000) found systematic phonics effective, with larger effects in kindergarten and first grade (overall effect size about 0.41; from prior knowledge). It set no per-skill repetition count. Its evidence supports ordered, explicit sequences, not a specific number of repetitions.
Counting (early numeracy).
- Gelman & Gallistel's counting principles and Wynn's "Give-a-Number" research describe cardinality as a stage change ("knower levels", with most children becoming cardinal-principle knowers around 3.5-4.5). Clements & Sarama's learning trajectories (Building Blocks) advance children by observed behavior at each trajectory level and use formative assessment, not a repetition count (prior knowledge; not fetched this run). The transfer test is built into the construct: a child who can recite "1, 2, 3, 4, 5" but cannot give you 5 blocks when asked has not mastered cardinality.
Transfer and spacing (general).
- Stokes & Baer (1977), "train sufficient exemplars": generalization has to be programmed with several examples. It does not happen by default. Vlach, Sandhofer & Kornell (2008) found that 3-year-olds generalized categories better after spaced presentations than after massed presentations (both prior knowledge). Bloom-style mastery learning uses 80-90% on formative checks followed by correctives (Guskey's summaries), but it was built for school-age units, not preschool motor or number skills.
Precision teaching / fluency aims: frequency aims (correct responses per minute) are well developed for school-age academics. I found no validated per-minute aims for 3-5-year-olds on paper tasks. This is a gap in the literature. Timed drills also fit poorly with a calm, paper-first product for preschoolers.
Convergences and contradictions
- Convergence: the vault proposal's "more than one occasion plus a different example" matches both the ABA maintenance findings (multiple sessions beat single-session criteria) and the generalization literature (several exemplars, spaced). The vault reached the right shape without the citations. This brief supplies them.
- Partial tension: the vault says "do not use a universal 80%-correct rule." The ABA evidence agrees but goes further: if an accuracy bar is used at all, 80% is too low for retention, and 90% or higher is the supported floor.
- No consensus: no study sets a per-skill repetition count for typically developing preschoolers on paper tasks. The numbers above come from clinical populations (ABA), older readers (orthographic learning), or intervention-level dosage (OT). Everything below that applies them to preschoolers is extrapolation.
Synthesis for RDCO
The research does not give a magic repetition number. It gives three rules that fit together. (1) The accuracy bar should be about 90% or higher, not 80%. (2) The bar has to hold on more than one separate occasion. (3) Mastery only counts once the child succeeds on an example they have not practiced. Repetition count comes out of these rules: it is not the rule itself. For a paper product, "90%" becomes "9 of 10 items right, or all but one," and the parent does the scoring. So the criterion must be something a parent can judge in 10 seconds at the kitchen table.
Proposed default (confidence labels are Ray's, not the literature's):
- Pack contents per objective: about 3 short practice sessions of 8-10 items each, where each session uses different exemplars (different words, objects, or stroke contexts). Add one "new example" check page. Medium confidence: the multiple-session and multiple-exemplar structure is well supported. The exact number 3 is borrowed from ABA convention.
- Advance criterion: independent (no hand-over-hand help, no verbal prompting) at 9 of 10 or better on 2 separate days at least 24 hours apart, plus success on the new-example page. Medium confidence: the 90% bar and multiple sessions come from Fuller & Fienup and Richling. The 24-hour gap comes from spacing research. Two days rather than three is a product trade-off against parent effort.
- Maintenance check: put one mastered item back into a pack 1-2 weeks later. If it fails, return to practice without taking away the badge. Medium-high confidence: maintenance probes are standard, and the no-erasure part matches the existing skill-tree rule.
- Per-domain overrides:
- Pencil control: do not expect mastery in a single pack. Plan for 20 or more short practice sessions across a skill family over weeks (Hoy et al.). Judge each stroke by a quality rubric that a parent can check (for example, "stays within the thick guide, closes the circle") and by moving from tracing to copying to drawing from memory. Do not use accuracy percentage. Respect the age norms already in the vault. Medium confidence on the dosage floor; low on the rubric specifics.
- Decoding: the unit of mastery is a letter-sound set, not a word list. Transfer means reading untaught consonant-vowel-consonant (CVC) words built from mastered sounds. A child reading only the practiced words has memorized them, not decoded them. About 4 successful decodings per word helps fix the word as a sight word, but that is a separate, later goal. Medium confidence.
- Counting: the transfer test is the cardinality question ("give me 5", "how many?") across different objects, arrangements, and set sizes. It is not rote counting on a worksheet. Medium-high confidence; this is how the construct is defined.
The product implication: packs should be designed around varied exemplars plus a new-example check, not around more copies of the same page. That shapes the AI generation step directly. The generator should vary exemplars within an objective and hold back one novel item for the check page, instead of repeating identical items. The claim "our advancement rule follows the published mastery-criterion research" is defensible to parents, as long as we frame it as research-informed, not research-validated for preschoolers on paper.
Why this is in the vault
It replaces the unsourced repetition and advance guesses in the Scribble Works skill-tree / mastery design ([[skill-tree-proposal]]) with a cited default. It also gives the early-childhood educator review, which that proposal already calls for, a concrete rule to accept or reject.
Open follow-ups
- For typically developing 3-5-year-olds, is there any study comparing 1-session vs multi-session mastery criteria, or does all of the evidence come from clinical ABA samples?
- How reliable is a parent's judgment of independent correctness on preschool paper tasks compared with a trained observer (inter-rater agreement)? Our whole criterion depends on it.
- What legibility or quality rubrics exist for pre-writing strokes at ages 3-5 (for example, the Beery VMI scoring criteria) that could be simplified into a parent checklist?
Related
- [[skill-tree-proposal]]
- [[2026-08-30-printables-inspiration-research]]
- [[2026-09-05-cohort-feedback-synthesis]]
- [[2026-09-01-cohort-week1-kit]]
- [[DESIGN-scribble-works]]
- [[2026-09-11-scribble-works-proxy-demand-signals]]
Sources
- Vault: ~/rdco-vault/01-projects/printables-product/planning/2026-09-12-skill-tree/skill-tree-proposal.md
- Vault: ~/rdco-vault/01-projects/life/printables/2026-08-30-printables-inspiration-research.md
- Vault: ~/rdco-vault/01-projects/printables-product/reviews/2026-09-05-cohort-feedback-synthesis.md
- Vault: ~/rdco-vault/01-projects/printables-product/2026-09-01-cohort-week1-kit.md
- Vault: ~/rdco-vault/02-sops/DESIGN-scribble-works.md
- Vault: ~/rdco-vault/06-reference/research/2026-09-11-scribble-works-proxy-demand-signals.md
- Fuller & Fienup 2018, Behavior Analysis in Practice: https://www.researchgate.net/publication/321581517_A_Preliminary_Analysis_of_Mastery_Criterion_Level_Effects_on_Response_Maintenance
- Richling et al. 2019, Journal of Applied Behavior Analysis: https://onlinelibrary.wiley.com/doi/10.1002/jaba.580 (403 on fetch; findings from search abstract)
- Richling, Williams & Carr 2019 descriptive analysis: https://www.researchgate.net/profile/Sarah-Richling/publication/333472428 (403, flagged)
- Mastery criterion frequency comparison: https://www.researchgate.net/publication/355674095_A_preliminary_comparison_of_mastery_criterion_frequency_values_Effects_on_acquisition_and_maintenance
- Pitts 2021, Behavioral Interventions: https://onlinelibrary.wiley.com/doi/full/10.1002/bin.1778
- Hoy, Egan & Feder 2011, Canadian Journal of Occupational Therapy: https://journals.sagepub.com/doi/abs/10.2182/cjot.2011.78.1.3
- Nation, Angell & Castles 2007, Journal of Experimental Child Psychology: https://pubmed.ncbi.nlm.nih.gov/16904123/
- Prior knowledge, not re-fetched this run (verify before quoting externally): Share 1999 self-teaching; National Reading Panel 2000; Gelman & Gallistel 1978; Wynn 1990/1992; Clements & Sarama learning trajectories; Stokes & Baer 1977; Vlach, Sandhofer & Kornell 2008; Guskey on Bloom mastery learning.