Category: General

  • Why OTP’s need better assessments-update on the FMSAT

    Why OTP’s need better assessments-update on the FMSAT

    Measuring Neuromotor Lateralization Across the Lifespan: Progress on the FMSAT Norming Project (& Why OTP’s need better assessments)

    How a one-minute screener for hand dominance, modern Rasch psychometrics, and the shift to value-based care are converging to change how occupational therapists, educators, and clinicians measure what we do.

    Published May 24, 2026 by Stephanie Seymore Wick, MSOT, OT/L · Founder, Learning Charms and O.T. Wizard

    A short progress note before we go deeper

    Every week I send out an update to the therapists, educators, and clinicians contributing to the FMSAT norming project. The updates focus on data, leaderboards, and the age bands we still need to fill. This week the data deserves a deeper look than a weekly email can carry. The findings touch on questions that go well beyond a single screener, including how our profession measures what we do, how those measurements connect to insurance reimbursement, and why occupational therapy salaries have not kept pace with the cost of becoming an OT.

    If you are a pediatric occupational therapist who has ever had to defend a Beery score at an IEP meeting that did not match the child sitting in front of you, this post is for you. If you are an educator looking for a fast, fair screening tool to identify students who may benefit from earlier support, this is for you. If you are an adult-focused OT, a neurology specialist, or a clinician working with progressive motor conditions, this is for you too. And if you are an OT, OTA, or therapy leader thinking about the coming shift to value-based care and what it means for your practice, this is also for you.

    Here is what is in this post:

    • What the FMSAT is, what it actually measures, and why it works across the lifespan
    • How the FMSAT can be used as a one-time screener, a progress monitoring tool, and a lifelong tracking instrument
    • What Rasch analysis is, in plain English, and why it is considered the gold standard in modern assessment
    • The first Rasch results on the FMPRS, our companion rating scale
    • What ecological validity means and why your favorite assessments may be missing it
    • Why the FMSAT itself uses Classical Test Theory and the FMPRS uses Rasch
    • Why value-based care is about to make all of this matter more than it ever has before
    • How the O.T. Wizard platform, soon to be rebranded as MyTherapyWizard, fits into the bigger picture

    A note on terminology before we go further

    Throughout this post you will see three terms that are sometimes used interchangeably but mean different things in occupational therapy practice and in psychometrics. Getting the distinction right matters because the FMSAT is one of those terms and not the others.

    A screener is a brief, low-burden tool designed to flag people who may benefit from further evaluation. Screeners take a minute or two, can often be administered by people without specialized clinical training, and produce a simple signal that says proceed to next step or no further action needed. Screeners do not diagnose, do not establish eligibility, and do not produce a comprehensive clinical picture. Examples in healthcare include the M-CHAT for autism, the PHQ-9 for depression, and vision and hearing screens in schools.

    An assessment is a more comprehensive structured evaluation that produces detailed information sufficient to support diagnosis, eligibility decisions, and treatment planning. Assessments typically take 30 to 90 minutes, require trained administrators, and produce multiple subscores or domain scores. Examples in pediatric OT include the Beery-Buktenica Developmental Test of Visual-Motor Integration, the Peabody Developmental Motor Scales, the Bruininks-Oseretsky Test of Motor Proficiency, the Sensory Processing Measure, and the Pediatric Evaluation of Disability Inventory Computer Adaptive Test (PediCAT).

    An evaluation is the broader clinical process that uses one or more screeners and assessments together with observation, interview, and chart review to produce a full clinical picture and recommendations.

    The FMSAT (Fine Motor Speed and Accuracy Test) is a screener. It is not designed to replace any assessment in your toolkit. It is designed to do the screener job well, which means flagging test takers who may benefit from a comprehensive evaluation when their lateralization or fine motor speed scores fall outside expected ranges. The assessments that follow a positive FMSAT screen will be whatever your clinical reasoning and your setting indicate. The point of a good screener is to make sure the right people get to those assessments faster than they would have otherwise.

    The FMPRS (Fine Motor Participation Rating Scale), on the other hand, is functioning more like a brief assessment instrument during the validation phase of this project, because its job is to produce the multi-domain rating data needed to establish ecological validity for the FMSAT. After validation is complete, the FMPRS will not typically be used alongside the FMSAT in routine clinical screening.

    What the FMSAT measures and why it works across the lifespan

    The FMSAT, which stands for Fine Motor Speed and Accuracy Test, is a one-minute screener. The test taker is given a bubble-popping worksheet and asked to pop as many bubbles as they can in 30 seconds with one hand, then 30 seconds with the other hand. The score for each hand is the number of bubbles popped. That is it. No rater training, no expensive test kits, no proprietary materials beyond a printed worksheet.

    The name says fine motor speed and accuracy, and that is what each hand score reflects. But the construct the instrument was designed to measure, and the reason the two hands are tested separately, is neuromotor lateralization. Lateralization is the degree to which a person has committed one side of the brain, and therefore one hand, to specialized motor work. Strong lateralization means the dominant hand performs precision tasks fluently while the non-dominant hand serves as a stabilizer. Weak or absent lateralization means the two hands perform more similarly, hand preference is inconsistent, or the person switches hands mid-task. Fine motor speed is the metric we use to detect that pattern, because timed performance under demand reveals the lateralization signal more clearly than untimed observation does.

    This matters because lateralization is not just a developmental milestone of early childhood. It is a lifelong neuromotor property that emerges during the preschool years, consolidates through school age, holds through adulthood, and can shift or erode in response to neurological changes later in life. The FMSAT was designed to capture that lateralization signal at any age, which is why our normative dataset spans from preschool through older adulthood rather than stopping at age 12 or 17.

    What makes the FMSAT useful is not the bubble popping itself. It is what the two hand scores together tell us about how well a person has consolidated hand dominance, and how that pattern compares to typical lateralization for their age. The dominant hand score reflects fine motor capability. The difference between the two hands reflects lateralization. Both pieces of information matter for clinical and educational practice, and neither one is captured well by the standardized assessments most occupational therapists currently use.

    How the FMSAT can be used: one-time screener, progress monitoring, and lifelong tracking

    Because the FMSAT is fast, standardized, and produces a numeric score, it can support several clinical and educational use cases that most current assessments cannot. Three are worth naming explicitly.

    Use case 1: One-time screening for hand dominance and fine motor concerns

    This is the most familiar use case. A pediatric occupational therapist, school OT, or educator administers the FMSAT to a student who has been flagged for fine motor concerns, handwriting difficulty, or unclear hand dominance. The score, compared to age-appropriate normative bands, gives a quick objective marker of whether the student’s performance and lateralization fall within expected ranges. This is the use case the validation manuscript will focus on first.

    Use case 2: Response to Intervention and progress monitoring

    Because the FMSAT takes one minute and produces a numeric score, it can be re-administered at intervals to track whether a person is making measurable progress over time. This is exactly the kind of brief repeatable measurement that Response to Intervention frameworks and Multi-Tiered Systems of Support models require. A school OT could administer the FMSAT at the start of an intervention block, midway through, and at the end, and have objective data showing whether the lateralization or fine motor speed measure is changing in response to the intervention. A clinic-based OT could do the same across a course of therapy. The same minimum-detectable-change thresholds that value-based payment models are increasingly requiring would apply directly. Once the normative dataset is finalized, the FMSAT becomes one of the few fine motor measures fast enough to use for routine progress monitoring.

    Use case 3: Educator-administered screening in school settings

    The FMSAT requires no rater training and no clinical interpretation to administer. A teacher, paraprofessional, school nurse, or interventionist can give the test and record the scores. With enough normative data in place, the score itself does the screening work and a teacher does not need to be an OT to identify a student who falls below the expected band. This opens the door to universal screening at the classroom or grade level, the same way schools currently screen vision and hearing. Students flagged by the screener can then be referred to occupational therapy for full evaluation. This is a long-term vision rather than an immediate use case, but it is exactly the kind of MTSS Tier 1 screening application that schools have been asking for and that occupational therapy has not yet had the tools to support.

    Use case 4: Lifespan monitoring, including progressive neuromotor conditions

    This is the use case that emerges directly from norming the FMSAT across all ages rather than only in pediatrics. Lateralization can shift across the lifespan in response to neurological events and conditions. A person recovering from stroke may show changed dominant-hand performance. A person living with multiple sclerosis, Parkinson’s disease, or another progressive neuromotor condition may show gradual erosion of the lateralization signal as the condition advances. A person experiencing age-related changes in motor control may show compression of the dominant-hand advantage. The FMSAT, administered periodically, could detect those shifts earlier and more objectively than self-report or general clinical observation. With enough normative data across all ages, the same one-minute test that screens a five-year-old for emerging hand dominance could monitor a 55-year-old neurologist for early signs of motor change in the years after a diagnosis. That is the broader clinical reach the lifespan normative dataset makes possible.

    None of these use cases require a separate test. They all use the same one-minute bubble-popping task. What changes is who is administering it, how often, and what the score is being compared against. That flexibility is one of the reasons we are investing in the validation work the way we are.

    What is Rasch analysis and why is it considered the gold standard?

    Many occupational therapy assessments you have used in your career, including those most commonly used to demonstrate progress in pediatric settings, were developed using Classical Test Theory, often shortened to CTT. The Beery-Buktenica Developmental Test of Visual-Motor Integration, the Bruininks-Oseretsky Test of Motor Proficiency, the Peabody Developmental Motor Scales, and the Sensory Processing Measure are CTT-based instruments. CTT produces a total score that is then compared to a normative sample. The score tells you where a test taker falls relative to peers, but it has some real limitations.

    The biggest limitation is that CTT treats every item on a test as if it were equally difficult, and every point on the score scale as if it represented the same amount of skill. A test taker who scores 84 versus one who scores 89 may differ by a meaningful amount of skill, or may differ by almost nothing, depending on where on the scale those scores fall and which specific items they got right. CTT cannot tell you which.

    This is true for both screeners and assessments built under CTT, though the limitation is more consequential for assessments because assessments are doing more of the clinical decision-making work. A screener producing a CTT score is still useful as a yes/maybe/no signal. An assessment producing CTT scores is being asked to support diagnosis, eligibility, and treatment planning decisions on the same imprecise measurement scale.

    Rasch analysis, developed by Danish mathematician Georg Rasch in the 1960s, takes a fundamentally different approach. Rasch places each test item and each person on the same interval scale, called the logit scale. This means the distance between a score of 5 logits and 10 logits represents the same amount of skill change as the distance between 20 and 25. Rasch also calibrates each item individually, telling you which items are easy, which are hard, and whether each item is actually pulling its weight in measuring the construct.

    Rasch is to assessment what a ruler is to measurement. CTT scores tell you a child is somewhere in the middle. Rasch scores tell you exactly where, on a scale where the units are equal.

    Rasch analysis is considered the gold standard for modern instrument development because it is the framework used by some of the most respected and most defensible pediatric assessments in the field. The PEDI-CAT, the AMPS or Assessment of Motor and Process Skills, the School Function Assessment, and the WeeFIM are all Rasch-based or Item Response Theory based instruments. These are the tools that produce data insurance companies and researchers trust. If you have ever wondered why some assessments seem to have stronger research backing than others, the framework behind them is usually a big part of the answer.

    Behind the scenes: the first Rasch checkup on the FMPRS

    The FMPRS, or Fine Motor Participation Rating Scale, is the 12-item observer-rated companion to the FMSAT. It captures four dimensions of fine motor function across real-world tasks: fine motor speed, fine motor precision, laterality and bilateral differentiation, and participation in daily roles. Our Rasch consultant, Angie, ran the first calibration on 139 paired FMSAT and FMPRS records earlier this month. The results were strong for a first calibration on a relatively small sample.

    Finding one: the item difficulty order matches the developmental theory

    Rasch ordered the 12 items from easiest to hardest based on how raters actually responded to them. The Laterality items, which ask about hand dominance and bilateral hand use, came out as the easiest. The Participation items, which ask about sustained engagement in fine motor tasks throughout the day, came out as the hardest. This is exactly what we would predict developmentally. Hand dominance consolidates earlier in childhood than sustained occupational engagement, so on a rating scale measuring fine motor function across the developmental arc, laterality items should be easier to endorse than participation items. The data confirmed the theory.

    Finding two: the instrument separates clinical and typical test takers cleanly

    On the Rasch-derived person measure, typical preschoolers scored more than two logits higher than clinical preschoolers. In plain terms, the typical preschooler scored higher than approximately 98 percent of the clinical sample. This is what is called a known-groups validity effect, and a Cohen’s d effect size above 2.0 is unusually strong. Most pediatric assessments are pleased to show known-groups effects in the 0.5 to 0.8 range. The FMPRS is producing separation that is roughly three times stronger. As the normative dataset grows beyond preschool ages, we expect the same separation pattern to extend across older age bands and into adult populations where clinical and typical comparison groups can be defined.

    Finding three: the FMPRS and the FMSAT are picking up the same underlying construct

    The Rasch-derived person measure on the FMPRS correlated with FMSAT dominant hand scores at r = 0.575. That correlation tells us that two completely different methods of measurement, a one-minute performance task and a 12-item observer rating scale, are picking up the same underlying construct. This is exactly the kind of cross-method convergence a strong validation manuscript needs.

    What still needs work

    Five items in the FMPRS came back as overfitting in the Rasch model, which means they are too internally redundant. They are not bad items. They are just not adding as much new information as they could. Those items will be revised for the next version of the FMPRS based on what the calibration showed us. This kind of iterative refinement is how serious instrument development works, and it is exactly why we are doing this calibration now, before the manuscript is finalized.

    Ever wondered why Beery, PDMS-3, or BOT-2 scores do not match real-world function?

    Here is a question every occupational therapist has wrestled with at some point in their career. Why do scores from the Beery-Buktenica VMI, the Peabody Developmental Motor Scales, or the Bruininks-Oseretsky Test sometimes fail to line up with what you actually see in the classroom, at home, on the playground, or in adult daily life?

    You are not imagining it. A student can score below average on a tabletop visual motor task and still write legibly, manage their lunchbox, and participate fully in PE. Another can score in the average range and still struggle every single day to keep up with handwriting demands or self-care routines. The same pattern shows up in adult assessments. The mismatch is real, and it has a name in the psychometric literature. It is called an ecological validity gap.

    What ecological validity actually means

    Ecological validity is a psychometric term that answers a simple question: does this assessment measure something that actually matters in real life? An instrument with strong ecological validity produces scores that connect to how a person functions in their daily environment. An instrument with weak ecological validity produces scores that connect mainly to how a person performs on the test itself, with limited evidence that the score predicts daily function.

    Many of the assessments OTs use most often were designed to measure isolated motor performance, not real-world participation. They are good at what they measure. They were just never built to answer the participation question. When a school-based therapist is asked to defend a Beery standard score at an IEP meeting, or when a clinic-based therapist is asked to justify medical necessity to an insurer using a PDMS-3 score, the underlying problem is often that the score is being asked to do something the assessment was not designed to do.

    Why the FMPRS is essential to FMSAT validation, even though clinicians will not use both in practice

    Here is a question that comes up almost every time I explain this project. If the FMSAT is a screener, why pair it with the FMPRS at all? Won’t clinicians be expected to do both?

    The answer is no. The FMPRS is doing critical work right now, during validation, so that the FMSAT will not need to be paired with it in clinical practice later. The whole value of the FMSAT as a screener depends on it being a one-minute, paper-and-pencil, standalone score. Asking clinicians to also complete a 12-item rating scale every time they administered the screener would defeat the entire point of having a screener in the first place.

    What the validation work establishes, and what every paired FMSAT and FMPRS submission helps establish, is that when the FMSAT score is elevated or compressed in a particular way, it is reflecting something that shows up in the test taker’s real-world functional life. Once that link is established and published in a peer-reviewed manuscript, the FMSAT score on its own carries that ecological meaning forward. Clinicians using the FMSAT in practice will be able to point to the published validation evidence rather than having to demonstrate the connection every time.

    This is the same approach used to validate other widely accepted screeners. The Modified Checklist for Autism in Toddlers, known as the M-CHAT, was validated against full ADOS and ADI-R diagnostic batteries. Pediatricians using the M-CHAT today do not run an ADOS alongside it. The validation work was done once and the screener now stands on its own. The PHQ-9 depression screener was validated against structured psychiatric interviews. Primary care providers use it on its own today. The pattern is consistent across well-validated screeners. Validate against a richer companion measure once, publish the validity evidence, then use the screener on its own.

    Why the FMSAT itself uses Classical Test Theory and the FMPRS uses Rasch

    This is a question that any sharp reader will be asking by now. If Rasch is the gold standard, why is the FMSAT being calibrated using Classical Test Theory rather than Rasch?

    The answer comes down to the measurement structure of each instrument. The FMSAT produces raw bubble counts on a zero to 80 scale for each hand. That kind of continuous count data is well-suited to CTT-style descriptive statistics, percentile norms, and known-groups validity comparisons. These are the analyses the FMSAT validation manuscript will lean on, and they are the analyses that produce the percentile bands and severity cutoffs clinicians actually use at the point of care.

    Rasch is the right framework for the FMPRS because the FMPRS uses ordered category responses on a four-point scale, and each item can be at a different difficulty level on the same underlying trait. That is exactly the kind of measurement structure Rasch was designed to handle. The two instruments are built differently on purpose, and each one is being analyzed using the framework that fits its measurement structure.

    The Rasch work on the FMPRS gives the manuscript the modern psychometric backbone peer reviewers expect. The CTT work on the FMSAT keeps the screener simple and interpretable for the clinicians who will actually use it. Both pieces matter, and together they create a defensible validation argument.

    Why value-based care is about to make all of this matter much more than it has before

    If you have been practicing for more than a few years, you already know that occupational therapy reimbursement has been under pressure for a long time. The 2026 Medicare Physician Fee Schedule final rule from the Centers for Medicare and Medicaid Services, released October 31, 2025, continued a trend of flat or declining payment rates for outpatient OT. According to OT Potential’s 2026 reimbursement analysis, the proposed 1% decrease to OT and PT relative value units for 2026 came after a 0% increase in 2025 and a 3% decrease in 2024. The trajectory is real and most OTs feel it directly in their paychecks.

    What is changing right now, and what most clinicians have not fully internalized, is that the structure of the entire payment system is shifting underneath us. The shift is called value-based care.

    What value-based care actually means

    Value-based care is a payment model where providers and health systems are reimbursed based on the outcomes their patients achieve, not the volume of services delivered. Under traditional fee-for-service payment, an OT bills for each visit and gets paid for each visit. Under value-based care, payment is increasingly tied to whether the patient demonstrated measurable functional improvement against established benchmarks.

    CMS has been driving this shift through several specific programs. The Quality Payment Program, the Merit-based Incentive Payment System known as MIPS, and an expanding suite of Alternative Payment Models are all moving rehabilitation services toward outcomes-based reimbursement. The 2026 payment updates included a 0.75% increase for qualified APM participants and a 0.25% increase for everyone else, an early but clear signal that participating in alternative payment models will increasingly be where the financial upside is.

    The era of writing patient made progress toward goals in a discharge note and being reimbursed for it is ending. The era of demonstrating measurable functional change against defensible benchmarks is beginning.

    Why occupational therapy is structurally underprepared for this shift

    Here is the connection most clinicians have not drawn explicitly. The shift to value-based care requires outcomes data, and outcomes data is only as good as the assessments and progress-monitoring instruments producing it. Most of the assessments occupational therapists use to demonstrate progress, and most of the screeners they use to identify who needs services in the first place, were developed under CTT frameworks that produce raw scores, percentile bands, and standard scores. None of those formats give insurers what value-based payment models actually require.

    Insurers under value-based care want interval-level evidence of measurable functional change against established minimum-detectable-change thresholds. When a third-party reviewer asks whether a patient made meaningful progress, the answer they want is not, the patient’s standard score improved from 84 to 89. The answer they want is, the patient’s interval-level fine motor measure shifted by 0.45 logits, which exceeds the minimum detectable change threshold of 0.30 logits established in the calibration sample. One of those answers is opinion-vulnerable. The other is data.

    Closing this evidence gap requires investment at every level of the measurement pipeline: better screeners that identify who needs services earlier and more accurately, better assessments that produce the diagnostic and eligibility data on a defensible measurement scale, and better outcomes instruments that document functional change over an episode of care. The FMSAT and the FMPRS sit at the screener and ecological validity ends of that pipeline. They are one contribution among many that the profession needs.

    This evidence gap is one of the underrecognized reasons our profession has struggled to make the reimbursement case at the level of physical therapy or speech-language pathology. Both adjacent professions have invested more heavily in Rasch-calibrated, IRT-based assessment development over the last 20 years. The PEDI-CAT, the AM-PAC, and similar tools represent what that investment looks like. Occupational therapy has far fewer Rasch-calibrated tools across the screener, assessment, and outcomes layers, and that thinness in our measurement infrastructure shows up downstream as flatter reimbursement, narrower coverage policies, and ultimately compensation that has not kept pace with the cost of the training required to enter the field.

    How the O.T. Wizard platform fits into this picture

    The FMSAT and the FMPRS are not standalone projects. They are pieces of a larger evidence infrastructure being built into the O.T. Wizard platform, which is being rebranded as MyTherapyWizard.

    O.T. Wizard is a digital evaluation and outcomes platform built from the ground up on modern psychometric standards. The FMSAT lives inside the platform as a fast, defensible screener with clear research foundations. The FMPRS lives alongside it as the ecological validity companion during validation. The broader platform houses structured evaluation templates designed to produce the kind of data that holds up under value-based payment scrutiny. The architecture is PHI-free and operates under a 1EdTech-approved legal framework, which means the platform itself functions as a passive-accrual research engine. Every paired evaluation contributes to the dataset that makes the next generation of assessments stronger.

    The rebrand to MyTherapyWizard reflects the platform’s expanding scope beyond occupational therapy into a multi-discipline space for pediatric therapy professionals. The underlying mission stays the same. Build the measurement infrastructure our profession needs to move forward, in step with where reimbursement is going rather than chasing it after the fact.

    What you can do

    Our profession needs more Rasch-calibrated assessments. It needs more validated rating scales with strong ecological validity. It needs more normative datasets large enough to defend in peer review. And it needs more clinicians and educators willing to contribute the data that makes all of that possible.

    If you are an occupational therapist, an educator, a clinician working with adult or geriatric populations, or anyone interested in supporting the development of evidence-based assessment tools, here is how to get involved:

    • Request to be on the Norming Tryout Team and Contribute FMSAT data. The screener takes one minute per test taker. If you administer it after a session or screening you would have run anyway, the marginal time cost is essentially zero.
    • Complete the FMPRS when you can. The paired data is what makes the validation manuscript possible. Every paired submission directly strengthens the published evidence base our profession will use.
    • Look for the bands we need most. As of this week, the most urgent recruitment gaps are adolescents ages 12 to 17, both clinical and typical, three-year-olds in both groups, adults age 50 and older, and left-dominant test takers at every age.
    • Share this work with colleagues. The bigger and more representative the dataset, the stronger the eventual screener will be for the people you serve.

    Every paired submission you contribute is a small but real piece of building the measurement infrastructure our profession needs. Building a Rasch-calibrated rating scale and a CTT-validated performance screener together, on a normative dataset large enough to defend in peer review and broad enough to span the lifespan, is exactly the foundational psychometric work the field has needed for years. If we want occupational therapy to be reimbursed at the level our training and clinical expertise warrant under the new value-based payment models, we have to produce the kind of evidence other professions have already produced.

    That work does not happen in conference panels or position papers. It happens in datasets, calibrations, and validation manuscripts. It happens in projects like this one. And the people producing it are not academics in distant labs. They are clinicians and educators like you who choose to spend a few minutes on a Tuesday afternoon contributing to something larger than a single evaluation.

    About this project

    The FMSAT, Fine Motor Speed and Accuracy Test, is a one-minute screener for neuromotor lateralization and fine motor speed, currently in active normative data collection across the lifespan toward a peer-reviewed validation manuscript. The FMPRS, Fine Motor Participation Rating Scale, is the 12-item observer-rated companion used to establish ecological validity during the validation phase. Both instruments are part of the O.T. Wizard platform, rebranding to MyTherapyWizard. Pearl IRB Not Human Subjects Research determination on file (ID 2026-0154).

    Related topics Neuromotor lateralization assessment, hand dominance evaluation across the lifespan, Rasch analysis in rehabilitation, fine motor screening for educators, Response to Intervention RTI fine motor measures, MTSS Tier 1 and Tier 2 screening, ecological validity in occupational therapy, value-based care for outpatient therapy, Medicare Physician Fee Schedule 2026, evidence-based occupational therapy practice, school-based occupational therapy, OT reimbursement, alternative payment models for rehabilitation, progressive neuromotor condition monitoring, multiple sclerosis fine motor tracking, stroke rehabilitation outcomes measurement, lifespan motor assessment

  • When Behavior Eats Occupation: ABA’s Expansion

    When Behavior Eats Occupation: ABA’s Expansion

    A $639 million projection and a 3% pay cut

    On April 27, 2026, North Carolina Health News published a piece that should be required reading for every pediatric occupational therapist, physical therapist, and speech-language pathologist in the state. The headline: NC moves to rein in soaring autism therapy costs amid fraud concerns.

    The numbers were staggering. But the official NC DHHS policy paper Ensuring Person-Centered Care for Children with Autism Spectrum Disorder in the NC Medicaid Program, released for community feedback in late 2025, tells the underlying story even more clearly. NC Medicaid spending on Research-Based Behavioral Health Treatment (RB-BHT), which is overwhelmingly Applied Behavior Analysis (ABA), grew from $121.7 million in State Fiscal Year 2022 to $329.4 million in SFY 2024, a 171 percent increase in two years. The state’s own actuarial projection for SFY 2026 is $639 million. That is a 425 percent increase in four years, in one service line, in one state.

    In 2024, the NC General Assembly authorized a 15 percent rate increase for ABA across all seven RB-BHT CPT codes (97151 through 97157). In the same window, NC Medicaid pediatric occupational, physical, and speech therapy rates received no equivalent increase. They had not received a meaningful increase in nearly 20 years.

    Then, on October 1, 2025, NC Medicaid announced a 3 percent rate reduction across multiple service lines due to funding shortfalls. That reduction was applied to ABA, but only after ABA’s recent 15 percent raise. The same 3 percent was applied to pediatric OT, PT, and SLP, on a fee schedule that had not moved in two decades. The OT, PT, and SLP cut was subsequently paused after legal challenges and provider pushback, but the signal it sent is the point. The state was prepared to cut three established, board-credentialed, medically licensed pediatric therapy disciplines while a fourth service line, delivered overwhelmingly by paraprofessionals without medical credentialing, was on a trajectory to consume more than $600 million of the Medicaid budget in a single year.

    There is a question buried in those numbers that the rest of this post will try to answer. Medicaid coverage is statutorily anchored to medical necessity. If medical necessity is the threshold, how does a service delivered primarily by individuals with 40 hours of generalist training, no required college, no fieldwork, and no state healthcare licensure receive the lion’s share of the spend, while licensed medical professionals with master’s and doctoral degrees, board examinations, supervised fieldwork, and state licensure receive a rate cut?

    This post is not about whether ABA helps any individual child. It does, for some. This is about scale, scope, credentialing, and what happens to the developmentally rigorous, board-credentialed pediatric therapy disciplines when one adjacent discipline absorbs 30 to 40 hours per week of a child’s life, expands faster than its evidence base supports, and triggers fraud investigations that will eventually wash back across all of pediatric therapy.


    A brief history of ABA

    ABA’s foundations are in B.F. Skinner’s operant conditioning. The application to autism came from O. Ivar Lovaas at UCLA, whose 1987 paper is, to this day, the citation that anchors most of the field’s claims to insurers, school systems, and state legislatures.

    Lovaas reported that 47 percent of children who received 40 hours per week of intensive behavioral intervention for two to three years achieved “normal” intellectual and educational functioning, compared to 2 percent in the control group. That single study is where the 40-hour-per-week prescription standard came from. It is also where the language of “recovery” entered autism intervention discourse.

    The methodological problems with Lovaas 1987 are well documented and were acknowledged by Lovaas and his collaborators themselves: non-random group assignment, a sample functioning at a higher cognitive level than typical for autistic children at the time, unblinded outcome assessment, and an outcome definition built around IQ scores and mainstream classroom placement rather than quality of life. The original protocol also included aversives such as slaps and electric shock, which the field has since disavowed but which were central to the methods that produced the cited outcomes.

    Despite these problems, Lovaas 1987 became the evidence base cited in every state autism insurance mandate passed between 2007 and 2019. By the time more methodologically rigorous follow-up studies emerged, the reimbursement infrastructure was already built.


    What the research actually shows

    The strongest naturalistic dataset on ABA outcomes in the United States is the Department of Defense’s Autism Care Demonstration, which has tracked roughly 16,000 TRICARE-eligible children since 2014. The 2020 DoD Annual Report to Congress reached a remarkable conclusion for a discipline universally described as evidence-based: in their words, the current format of the demonstration project and the delivery of ABA services was not working for most TRICARE beneficiaries. Of the children studied over a one-year window, 76 percent showed no improvement on the standardized outcome measure, 16 percent improved, and 9 percent got worse. The report also stated, and this is the finding that should trouble anyone billing 30-plus hours per week, that “the number of hours rendered does not appear to impact outcomes.” The dose-response curve that justifies high-intensity ABA prescribing did not appear in the data.

    TRICARE still has not approved ABA as a basic medical benefit. It has been covered only under demonstration project structure for over a decade because it does not meet TRICARE’s hierarchy of evidence standard for proven medical effectiveness.

    In late 2025, the National Academies of Sciences, Engineering, and Medicine (NASEM) published a report commissioned by Congress to evaluate the demonstration. NASEM concluded that comprehensive ABA-based interventions are evidence-based practices and recommended that DHA cover ABA as a basic TRICARE benefit. But the same report explicitly criticized the way ABA is currently delivered: rigid hour prescriptions, mandatory assessments that do not inform treatment, restrictive setting requirements, and outcome measures that do not capture what families actually care about. The NASEM finding is essentially that the principles have evidence, but the delivery system the U.S. has built around them is misaligned with what the evidence supports.

    The honest summary: ABA principles, applied skillfully and in moderation by qualified clinicians, have empirical support for skill acquisition. The 40-hour-per-week dosage standard, the universal application to every autistic child, and long-term quality-of-life outcomes do not have the evidence base the field claims when speaking to payers.


    The fraud problem

    The NC Health News investigation, and the NC DHHS policy paper that followed, showed in publicly accessible numbers what happens when a payment stream grows 425 percent in four years with weak oversight.

    NC Health News obtained spending data through a public records request and found that 80 of the 200-plus ABA providers participating in NC Medicaid received at least $1 million in reimbursement in 2025. Payments to individual companies ranged as high as $64.91 million for a single Utah-based provider with 11 NC facilities. The Private Equity Stakeholder Project reported that 15 private equity-backed ABA companies operate more than 130 facilities in NC, making the state one of the most saturated PE-backed ABA markets in the country.

    NC Attorney General Jeff Jackson confirmed at an April 2026 House Select Committee on Oversight and Reform hearing that his office is conducting ongoing investigations into ABA billing in the state, including improper payments and “phantom billing,” which is the practice of submitting claims for therapy sessions that never took place or billing for more hours than were delivered. North Carolina is not alone. Federal prosecutors in Minnesota charged a defendant last year in what they described as the first criminal case tied to a sprawling ABA fraud scheme involving shell companies and millions in fraudulent Medicaid claims.

    The four-hat problem

    The NC DHHS policy paper named, in the state’s own words, the structural fraud vector that has been hiding in plain sight. From Action 8 of the policy paper:

    Some providers have reported to NCDHHS that these requirements in the RB-BHT Clinical Coverage Policy are insufficiently clear on which provider types may make an ASD diagnosis, referral to RB-BHT, or referrals for other ASD services… As a result, providers that do not offer RB-BHT sometimes refer an individual to an RB-BHT provider to make an ASD diagnosis, which raises conflict-of-interest concerns. In practice, the same provider may currently function as the diagnosing provider, the referring provider, the assessing provider and the service provider.

    That is the state acknowledging that an ABA company in NC can currently diagnose autism, refer the patient to itself, assess the patient, and deliver the services, all under one organizational roof, all billing the same payer. There is no parallel structure anywhere in pediatric OT, PT, or SLP, where the diagnosing physician or psychologist is institutionally and legally distinct from the treating therapist.

    Compounding this: under current Policy 8F, provisional ASD diagnosis can be made by any licensed psychologist, physician, or master’s-level clinician for whom diagnosis is within their scope of practice. For children under 3, a provisional diagnosis is sufficient to initiate ABA services, with full diagnosis required within 6 months. This is a structural funnel into ABA before differential diagnosis is complete, before OT/SLP/PT have evaluated the child, and before the family has been offered the full continuum of services NC Medicaid technically covers.

    Federal audits in other states

    This is not theoretical. The federal Department of Health and Human Services Office of Inspector General has already audited ABA billing in multiple states:

    The findings across these audits are consistent: lack of provider documentation to support the CPT codes billed, lack of documentation for the number of units billed or dates of service, delivery of ABA to members who did not receive required diagnostic evaluations or treatment referrals, and “impossible billing” practices such as billing for more than 24 hours of ABA in a single service date for a single member.

    NC has not yet been audited at the federal level for ABA, but the NC DHHS policy paper explicitly signals collaboration with the NC Department of Justice on program integrity going forward. The audit infrastructure is being prepared.

    Why this matters for OT, PT, and SLP

    Medicaid Program Integrity does not stop at one service category once it is mobilized. When NC DHHS and the Attorney General’s office expand pediatric therapy audits in response to the ABA findings, the standard practice is to extend that scrutiny to adjacent pediatric therapy lines. Outpatient OT, PT, and SLP share the same provider settings, the same referral sources, the same payers, and overlapping CPT code families with ABA. From a program integrity analytics standpoint, pediatric therapy disciplines are a single risk surface.

    NC Clinical Coverage Policy 10A already contains explicit Program Integrity language authorizing post-payment review by statistically valid random sampling, with data analytics on provider claims used to instigate review. The infrastructure to audit pediatric OT, PT, and SLP claims is already in place. ABA is the warm-up exercise.

    What this means practically: documentation standards that were adequate two years ago will not survive the 2026 to 2027 audit climate. Every evaluation, every plan of care, every progress note, and every outcome measurement needs to be defensible in standardized, ideally interval-level, terms. The defensive posture and the offensive advocacy posture converge here. The same outcome-measurement infrastructure that justifies a better fee schedule also protects practices from recoupment.


    The medical necessity question

    Medicaid coverage exists under a statutory framework of medical necessity. The federal Early and Periodic Screening, Diagnostic, and Treatment (EPSDT) mandate, the foundation of all pediatric Medicaid coverage, requires that covered services be medically necessary. State Medicaid programs implement this through clinical coverage policies that define which services are reimbursable, under what conditions, by which providers.

    This raises a question that should be uncomfortable for anyone reviewing NC’s ABA spending trajectory: if medical necessity is the threshold, why is the discipline with the lowest medical credentialing floor receiving the largest share of the pediatric therapy spend?

    A Registered Behavior Technician, who delivers the vast majority of direct ABA service hours, is not a medical professional under any standard definition. The RBT credential requires a high school diploma, 40 hours of training (frequently online), passage of a brief competency assessment, and an 85-question exam. There is no college coursework requirement. There is no clinical fieldwork requirement. There is no state healthcare licensure in most states, including North Carolina. The RBT is a paraprofessional credential issued by a private certification board (the Behavior Analyst Certification Board), not a state healthcare licensure body.

    Compare to the providers delivering OT, PT, and SLP under NC Medicaid: master’s or doctoral-level clinicians with accredited healthcare degrees, supervised clinical fieldwork measured in hundreds of hours, national board examinations, and state healthcare licensure. Even the assistant-level providers in OT, PT, and SLP have associate’s degrees, 16 weeks of supervised clinical fieldwork, national board examinations, and state licensure.

    The internal contradiction is plain. A Medicaid system that exists to provide medically necessary services has, in practice, prioritized funding a discipline whose direct-service workforce is not licensed as medical professionals over disciplines whose direct-service workforce is. The state’s own policy paper acknowledges this gap in Action 6, proposing to require BACB Registered Behavior Technician certification “prior to the provision of services” because, as the document notes, NC does not currently require its ABA technicians to obtain even the national BACB certification, much less state licensure. The state’s proposed fix is to require the basic 40-hour BACB credential. That is the floor being proposed, not the ceiling.

    This is not a comment on individual RBTs, many of whom are dedicated and skilled. It is a comment on the regulatory framework. If medical necessity is what determines what Medicaid covers, then the workforce delivering medically necessary services should be credentialed as medical professionals. The current NC structure does not meet that standard for ABA, and the spending pattern reflects that misalignment.


    NC pediatric therapy spending: the comparison

    Pediatric ABA in NC Medicaid

    From the NC DHHS policy paper, official state figures:

    • SFY 2022: $121.7 million
    • SFY 2023: $199.4 million
    • SFY 2024: $329.4 million (171 percent growth in two years)
    • SFY 2026 projected: $639 million (425 percent growth in four years)
    • 2024 rate increase: 15 percent across all seven ABA CPT codes
    • October 1, 2025 rate change: 3 percent reduction (applied after the 15 percent increase)
    • Number of Medicaid members receiving RB-BHT in SFY 2024: 8,706
    • Average per-child annual spend (NC Health News, FY 2025): approximately $37,600
    • Routine prescribing: 25 to 40 hours per week
    • No annual hour cap, no combined-discipline cap

    Pediatric OT, PT, and SLP in NC Medicaid

    NC Medicaid does not publish a comparable line-item breakdown of pediatric outpatient OT, PT, and SLP spending. State OT, PT, and SLP associations should be filing public records requests now to surface that data.

    What we do know is the authorization envelope. Under Clinical Coverage Policy 10A and EPSDT pediatric authorization, the practical maximum for a child receiving outpatient OT, PT, or SLP in NC is approximately 156 units per 6 months, per discipline. At 15 minutes per timed CPT unit:

    • 78 hours per year per discipline
    • 234 hours per year across all three rehab disciplines combined at full ceiling

    We also know the rate trajectory: no meaningful rate increase in nearly 20 years, followed by a proposed 3 percent reduction on October 1, 2025 (subsequently paused), on a fee schedule that currently reimburses pediatric OT at approximately $24 per 15-minute unit.

    The contact-hour ratio

    A child in a typical 30-hour-per-week ABA program receives approximately 1,500 hours of clinical contact per year.

    • ABA vs. OT alone at full pediatric annual ceiling: 1,500 hours vs. 78 hours, roughly 19 to 1
    • ABA vs. OT + PT + SLP combined at full ceiling: 1,500 hours vs. 234 hours, roughly 6.4 to 1

    Even when a child is authorized for the maximum of all three rehabilitation disciplines, ABA delivers more than six times the clinical contact hours.

    The dollar math, per child, per year

    At the current NC Medicaid pediatric OT rate of approximately $24 per 15-minute unit:

    • OT alone at full annual ceiling (312 units): approximately $7,488 per child per year
    • OT + PT + SLP combined at full annual ceiling, generous estimate: approximately $22,000 to $25,000 per child per year
    • ABA average per child per year: $37,600

    ABA spending per child is roughly five times higher than OT alone at full annual ceiling, and approximately 50 to 70 percent higher than all three rehab disciplines combined at maximum authorization. And that comparison assumes the rehab disciplines are billing at ceiling, which most are not, because the ceiling is rarely authorized in full.

    The state’s own policy paper acknowledges, in Action 7:

    if an assessment finds a member should receive occupational therapy, that may necessitate a lower intensity of RB-BHT based upon a child’s capacity to tolerate and benefit from the intensity of hours across all interventions.

    The state is, in essence, conceding that the current spending pattern is wrong: ABA is being prescribed at intensities that displace OT and the other rehab disciplines, even when an assessment would indicate the child needs OT instead of, or in addition to, ABA.


    The credentialing gap

    Who actually delivers most ABA service hours? Not Board Certified Behavior Analysts. The vast majority of direct-service ABA hours are delivered by Registered Behavior Technicians.

    Registered Behavior Technician (RBT)

    Under Behavior Analyst Certification Board requirements:

    • High school diploma or equivalent
    • Age 18 or older
    • Pass a criminal background check
    • Complete a 40-hour training course, frequently completed online in two to three weeks, often as part of employer onboarding
    • 3 of those 40 hours must cover ethics
    • Pass a brief competency assessment with a BCBA
    • Pass an 85-question multiple-choice exam
    • Total out-of-pocket cost can be under $100
    • No college coursework required
    • No clinical fieldwork required
    • No state healthcare licensure required in most states, including North Carolina

    NC DHHS’s policy paper notes that NC does not currently require even the national BACB Registered Behavior Technician certification. The state is now proposing to require it.

    Board Certified Behavior Analyst (BCBA)

    The supervising clinician credential is more substantial: master’s degree, 1,500 to 2,000 hours of supervised fieldwork, board examination, ongoing continuing education. The BCBA is the responsible clinical decision-maker but typically is not the person delivering the direct service hours.

    Certified Occupational Therapy Assistant (COTA)

    For comparison, here is what an OT Assistant brings to the same room:

    • Associate’s degree from an ACOTE-accredited OTA program, typically two years of college coursework
    • Roughly 16 weeks of full-time Level II fieldwork, approximately 640 hours
    • Pass the NBCOT national board examination for COTAs
    • State licensure
    • Continuing education units required for license renewal
    • Practices under the supervision of a licensed OT, with supervision frequency and scope defined by state law
    • Adherence to the AOTA Code of Ethics and state practice act

    Occupational Therapist (OT)

    • Master’s or doctoral degree from an ACOTE-accredited program
    • Roughly 24 weeks of full-time Level II fieldwork, approximately 960 hours
    • Pass the NBCOT national board examination for OTs
    • State licensure in every state, including NC
    • 15 CEUs per renewal cycle in NC
    • Adherence to the AOTA Code of Ethics and state practice act

    The contrast that matters

    A COTA, who works under the supervision of a licensed OT, has roughly (at minimum) two full years of accredited college coursework, plus 16 weeks of supervised clinical fieldwork, plus a national board examination, plus state healthcare licensure, plus ongoing CEUs before delivering hands-on pediatric therapy.

    An RBT, who is the frontline provider for the majority of ABA service hours, has 40 hours of training, no college, no fieldwork, and no state healthcare licensure before delivering hands-on pediatric therapy. In NC, even the basic national BACB credential is not currently required.

    A licensed OT or COTA delivering one hour of NC Medicaid pediatric OT, after two years (COTA) or six-plus years (OT) of accredited coursework and clinical training, is reimbursed at a rate that has not risen in nearly 20 years. An RBT, after 40 hours of training, is part of a payment stream that grew 425 percent in four years and is projected to consume $639 million in 2026.


    Where the money is: ABA companies absorbing OT and SLP

    A growing trend in pediatric autism services: ABA companies are hiring licensed OTs and SLPs to deliver services within an ABA-organized clinical model. Some practitioners arrive through dual-credentialing pathways, since the BACB has removed degree-field restrictions for BCBA candidates and made the OT-to-BCBA and SLP-to-BCBA transitions more accessible. Others are simply employed to deliver OT or SLP services in-house, but inside a treatment plan organized around 25 to 40 hours per week of ABA.

    The result is a structural absorption. Licensed clinicians from the rehab disciplines are increasingly delivering care inside ABA-organized clinical models, rather than the other way around. The treatment plan is built around behavior analysis. OT and SLP become supplemental services to a primary ABA intervention. The child’s calendar, the care coordination, and the parent-facing narrative all center on ABA, with OT and SLP positioned as add-ons.

    This is the inverse of what NC DHHS’s own policy paper recommends: whole-person care planning with the full continuum of evidence-based services coordinated by a licensed professional, and ABA used at the intensity clinically necessary rather than as the organizing modality.


    Where ABA infringes on OT scope

    AOTA’s scope of practice statement, and every state OT practice act, anchors the profession around occupation: ADLs and IADLs, feeding, eating, and swallowing, sensory processing and integration, fine and visual-motor skills, play, school participation, and meaningful engagement in life situations across home, community, and school contexts.

    ABA programs are increasingly billing into the same scope of practice:

    • Feeding therapy delivered by RBTs using food chaining, escape extinction, and behavioral feeding protocols. This is historically OT and SLP scope, and the AOTA Practice Guideline on Feeding, Eating, and Swallowing explicitly identifies OTs as uniquely positioned to evaluate and treat these problems due to the integration of sensory, motor, and contextual factors.
    • Toileting programs framed as ABA “ADL training,” delivered without the underlying motor planning, sensory processing, interoception, and developmental readiness assessment that an OT brings.
    • Fine motor skill acquisition including handwriting, scissor use, and utensil use, delivered as discrete trial training without consideration of grasp development, in-hand manipulation, bilateral coordination, or visual-motor integration.
    • Sensory regulation strategies delivered without sensory integration training, often using behavioral reinforcement paradigms that can be in direct conflict with sensory-based clinical reasoning.
    • Play skill development delivered as discrete trial training, which structurally cannot replicate the developmental work of OT-facilitated play.

    The clinical risk is not theoretical. A child whose food refusal is being treated as behavioral by an RBT-delivered feeding program may have undiagnosed dysphagia, retained primitive reflexes, oral-motor apraxia, sensory-based food aversion, or ARFID. A child whose handwriting is being treated as a compliance issue may have undiagnosed developmental coordination disorder or visual-perceptual deficits. A child whose self-injury is being addressed through behavioral extinction may have undiagnosed sensory dysregulation or pain.

    The ICF-CY framework makes this distinction clear. ABA’s traditional outcome targets sit at the body function and activity level. OT’s outcome targets sit at the participation level, in life situations across home, school, and community. These are not interchangeable, and a child whose week is dominated by body-function and activity-level work in a controlled 1:1 setting may never get to the participation work that produces durable, generalizable change.


    The 6 to 8 hour per day problem

    The clinical issue with high-dose ABA prescribing is not only what is being delivered in the ABA hours. It is what is not being delivered in the hours that are no longer available.

    Many ABA prescriptions land at 25 to 40 hours per week, often 6 to 8 hours per day for preschool-age children. A 4-year-old in ABA from 8 AM to 2 PM has no calendar space for OT, no calendar space for SLP, no calendar space for PT, no time for naturalistic family interaction, no time for play with neurotypical peers, no time for developmentally normative experiences in community settings.

    The insurance and Medicaid reality compounds this. When ABA absorbs the weekly therapy budget or the child’s tolerance for therapy, OT, SLP, and PT get cut to 30-minute sessions once a week, or eliminated entirely. The 156-unit ceiling becomes irrelevant when the child has no time on the calendar for those sessions. Families are often not offered a coordinated multidisciplinary plan. They are referred to ABA, that becomes the plan, and the other disciplines are reduced to consultation or eliminated.

    Pediatric OT is grounded in distributed practice across natural contexts. A child needs OT-informed strategies woven through meals at the family table, transitions in real classrooms, play with siblings, and community outings. That cannot happen when the child is in a 1:1 controlled clinical environment for the majority of waking hours.


    What pediatric OT must do

    The defensive and offensive responses converge. Four priorities.

    Build the outcome-measurement infrastructure. Pediatric OT cannot continue documenting in narrative paragraphs and informal goal-attainment scaling and expect to survive the audit climate that is coming, much less to win fee-schedule advocacy battles. The profession needs interval-level outcome measurement using Rasch-grounded instruments that produce defensible, change-detectable, statistically interpretable data, mapped to ICF-CY participation-level outcomes. The next generation of pediatric OT documentation needs to look more like a psychometric report and less like a clinical narrative.

    Pursue fee-schedule advocacy with data, not testimony. Personal testimony from therapists about the cost of doing business has not moved the needle in nearly 20 years. What might: a defensible return-on-therapy-investment dataset showing measurable participation-level change per dollar across OT, PT, SLP, and ABA. The legislature responded to the NC ABA spending data because it was specific, quantified, and contrasted. The same approach applied to pediatric OT, PT, and SLP would be hard to ignore.

    Defend scope of practice at the state level. State licensure boards, Medicaid coverage policy, and state Practice Acts are the legal mechanisms that define what discipline can deliver what service. Feeding, sensory integration, motor skill training, ADL training, and play-based participation work are OT and SLP scope under every relevant statute and AOTA practice document. Documenting scope encroachment in writing, with case examples, and submitting it to state licensure boards and Medicaid clinical coverage policy reviewers is how scope gets defended.

    Build the coordinated multidisciplinary alternative and educate referral sources. The strongest argument is not “ABA is bad.” It is that a coordinated OT plus SLP plus parent-coaching plus targeted developmental behavioral support model, delivered at 6 to 10 total hours per week with measurable participation-level outcomes, produces durable change at a fraction of the cost and a fraction of the developmental opportunity cost. Most pediatricians refer to ABA reflexively after an autism diagnosis. Most parents have never been offered a clear picture of what each pediatric therapy discipline addresses or what an integrated plan looks like. Parent-facing and pediatrician-facing materials, grounded in the ICF-CY framework, would shift the referral conversation.


    Closing

    This is not a turf war. It is a question about who is qualified to deliver what, what the actual evidence supports, what a child’s developmental window is best used for, and who is being paid what when public money is involved.

    Medicaid exists under a statutory framework of medical necessity. That framework is supposed to mean something. When a discipline whose direct-service workforce is not credentialed as medical professionals receives a 15 percent rate increase, expands 425 percent in four years, and consumes a projected $639 million in a single state budget year, while disciplines whose direct-service workforce is licensed as medical professionals receive no rate increase for nearly 20 years and then a proposed 3 percent reduction, the framework is not being applied consistently. NC DHHS’s own policy paper, released for community feedback in 2025, acknowledges the structural problems: providers diagnosing autism and then referring patients to their own services, treatment plans that are not individualized, ABA being used as primary treatment when less intensive evidence-based therapies would be more appropriate, and audits in Indiana, Wisconsin, and Massachusetts already documenting tens of millions in improper payments.

    ABA at moderate intensity, delivered by well-supervised clinicians, can be a useful component of an autism intervention plan for some children. ABA at 30 to 40 hours per week, delivered primarily by paraprofessionals with 40 hours of training, on a payment trajectory that has grown 425 percent in four years, while one provider collects $64.91 million from NC Medicaid in a single year, is a different conversation.

    The NC fraud investigations, the proposed Policy 8F revisions, and the NC DHHS policy paper are an opening. Clinical Coverage Policy is being rewritten in real time. The legislature is paying attention. The Attorney General is paying attention. Audit infrastructure is being built that will, predictably, expand to OT, PT, and SLP next.

    Pediatric OT has the evidence base, the credentialing rigor, the participation-focused framework, the developmental science, and the ICF-CY anchor. What it needs is the measurement infrastructure to translate clinical work into payer-defensible data, the advocacy coordination to get that data in front of policymakers, and the practice-owner discipline to make documentation audit-ready before the audit arrives.

    The pediatric population this profession serves deserves better than what the current system is delivering. So does the profession.

    Stephanie, OT/L, MS
    Head Wizard


    References and primary sources

    NC-specific policy documents

    1. NC DHHS. (2025). Ensuring Person-Centered Care for Children with Autism Spectrum Disorder in the NC Medicaid Program. https://medicaid.ncdhhs.gov/policy-paper-ensuring-person-centered-care-children-autism-spectrum-disorder-nc-medicaid-program/open
    2. NC Medicaid. Clinical Coverage Policy 8F: Research-Based Behavioral Health Treatment for Autism Spectrum Disorder. https://medicaid.ncdhhs.gov/8f-research-based-behavioral-health-treatment-rb-bht-autism-spectrum-disorder-asd/download?attachment=
    3. NC Medicaid. Outpatient Specialized Therapy Services (Clinical Coverage Policy 10A). https://medicaid.ncdhhs.gov/providers/programs-and-services/medical/outpatient-specialized-therapy-services
    4. NC Medicaid. October 1, 2025 NC Medicaid Rate Reduction Questions and Answers. https://medicaid.ncdhhs.gov/providers/claims-and-billing/october-1-2025-nc-medicaid-rate-reduction-questions-and-answers

    Investigative reporting

    1. Baxley, J. (2026, April 27). NC moves to rein in soaring autism therapy costs amid fraud concerns. North Carolina Health News. https://www.northcarolinahealthnews.org/2026/04/27/autism-therapy-costs/
    2. Private Equity Stakeholder Project. (2026). Private Equity in ABA: Report on the Behavioral Health Industry. https://pestakeholder.org/wp-content/uploads/2026/04/PESP_Report_PE-in-ABA_2026.pdf

    Federal audits and prosecutions

    1. U.S. Department of Health and Human Services Office of Inspector General. (2024). Indiana Made at Least $56 Million in Improper Fee-for-Service Medicaid Payments for Applied Behavior Analysis Provided to Children Diagnosed with Autism. https://oig.hhs.gov/documents/audit/10123/A-09-22-02002.pdf
    2. U.S. Department of Health and Human Services Office of Inspector General. (2025). Wisconsin Made at Least $18.5 Million in Improper Fee-For-Service Medicaid Payments for Applied Behavior Analysis Provided to Children Diagnosed With Autism. https://oig.hhs.gov/documents/audit/10497/A-06-23-01002.pdf
    3. Office of the Inspector General Massachusetts. (2024). MassHealth and Health Safety Net: 2024 Annual Report. https://www.mass.gov/doc/masshealths-applied-behavior-analysis-program-service-providers-oig-2024-annual-report/download
    4. U.S. Attorney’s Office, District of Minnesota. First Defendant Charged in Autism Fraud Scheme. https://www.justice.gov/usao-mn/pr/first-defendant-charged-autism-fraud-scheme-0

    National evidence reviews and DoD reports

    1. National Academies of Sciences, Engineering, and Medicine. (2025). The Comprehensive Autism Care Demonstration: Solutions for Military Families. Washington, DC: National Academies Press. https://www.nationalacademies.org/our-work/independent-analysis-of-department-of-defenses-comprehensive-autism-care-demonstration-program
    2. U.S. Department of Defense. (2020). Annual Report on Autism Care Demonstration Program. https://health.mil/Reference-Center/Congressional-Testimonies/2020/06/25/Annual-Report-on-Autism-Care-Demonstration-Program
    3. Lovaas, O. I. (1987). Behavioral treatment and normal educational and intellectual functioning in young autistic children. Journal of Consulting and Clinical Psychology, 55(1), 3–9.

    Credentialing bodies and scope of practice

    1. Behavior Analyst Certification Board. Registered Behavior Technician (RBT) Requirements. https://www.bacb.com/rbt/
    2. American Occupational Therapy Association. Occupational Therapy Scope of Practice. https://www.aota.org/practice/practice-essentials/scope-of-practice
    3. American Occupational Therapy Association. Code of Ethics. https://www.aota.org/practice/practice-essentials/ethicsstandardsa/code-of-ethics
    4. National Board for Certification in Occupational Therapy. https://www.nbcot.org/
    5. Accreditation Council for Occupational Therapy Education. https://acoteonline.org/
    6. AOTA Practice Guideline. The Practice of Occupational Therapy in Feeding, Eating, and Swallowing. https://www.oregon.gov/otlb/Documents/The%20Practice%20of%20Occupational%20Therapy%20in%20Feeding,%20Eating,%20and%20Swallowing.pdf

    Federal Medicaid framework

    1. Centers for Medicare & Medicaid Services. Early and Periodic Screening, Diagnostic, and Treatment (EPSDT). https://www.medicaid.gov/medicaid/benefits/early-and-periodic-screening-diagnostic-and-treatment

    Trauma and outcomes literature on ABA (referenced in evidence section)

    1. Kupferstein, H. (2018). Evidence of increased PTSD symptoms in autistics exposed to applied behavior analysis. Advances in Autism, 4(1), 19–29. Note: Journal issued an Expression of Concern in 2025. Findings should be cited with methodological caveats.
    2. Leaf, J. B., Ross, R. K., Cihon, J. H., & Weiss, M. J. (2018). Evaluating Kupferstein’s claims of the relationship of behavioral intervention to PTSS for individuals with autism. Advances in Autism, 4(3), 122–129.
    3. McGill, O., & Robinson, A. (2021). “Recalling hidden harms”: Autistic experiences of childhood applied behavioural analysis (ABA). Advances in Autism, 7(4), 269–282.
  • Understanding and Using Our Performance Bands

    Understanding and Using Our Performance Bands

    Why we use these performance bands

    When I built the scoring system for OT Wizard I wanted performance bands that would do three things at once: hold up psychometrically, work across a wide age range, and use language that is genuinely strengths-based rather than just sounding nice. After looking at how the major standardized assessments handle this, I landed on a framework that is closely aligned with what the field already uses.

    Our performance bands

    RangeBand LabelInterpretation
    85-100%MasteredPerforms skill consistently across contexts with little to no support
    70-84%ProficientPerforms reliably with minimal support in most contexts
    60-69%DevelopingSkill is somewhat present but may require occasional prompts or support
    40-59%EmergingSkill performed inconsistently or only in structured/familiar settings
    20-39%BeginningEarly attempts or partial skill components observed
    0-19%Not Yet ObservedNo evidence of skill use or minimal attempts

    The labels describe where a skill is in its trajectory. They do not describe the individual being assessed.

    Why the performance bands language is neuroaffirming

    The neuroaffirming move in assessment language is not softness. It is precision and neutrality. Words like “Developing,” “Emerging,” and “Beginning” describe a trajectory of skill acquisition. They imply that growth is possible without making any judgment about the person being assessed.

    Compare that to deficit-state language like “Inconsistent” or “Limited.” Those words describe what someone is not doing. They locate the problem in the individual rather than in the skill being measured. Clients and family members who have spent years on the receiving end of deficit language tend to be especially sensitive to it, and that is true whether the client is a six year old, a teenager, or an adult.

    The “Not Yet Observed” label at the bottom of the scale is doing important work too. The “Yet” signals that absence of observation is not a fixed trait. It keeps the door open without assuming a specific timeline.

    How the performance bands work across ages

    As OT Wizard migrates to “MyTherapyWizard”, it is designed to be used across the full age range, from young children through adults. A common concern I hear is that words like “Developing” or “Emerging” feel too young for older clients. I understand the instinct, but the issue is usually not the labels. It is the items being scored.

    A teenager being assessed on age-appropriate executive function, handwriting fluency, or self-advocacy skills will not feel infantilized by an “Emerging” rating. An adult being assessed on workplace task initiation or community mobility will not either. The mismatch comes when older clients are scored on items that were really designed for younger ones. That is an item-pool problem, not a label problem. The best way to solve this is to select guided evaluations for the appropriate age population and to skip subdomains (such as scissor skills) that aren’t relevant.

    This is why our platform invests so heavily in age-appropriate item development. The labels stay consistent. The items adapt to the person.

    How our performance bands compare to other assessments

    The language we use is very much in line with how the major developmental and rehabilitation assessments handle their descriptive categories. Here is a quick look:

    AEPS (Assessment, Evaluation, and Programming System) uses Consistently Performed, Inconsistently Performed, and Does Not Perform. These describe pattern, not stage.

    HELP (Hawaii Early Learning Profile) uses Mastered, Emerging, and Not Yet Present. The same trajectory framing we use.

    Vineland-3 (birth through 90+ years), which is the gold standard for cross-age range assessments, uses High, Moderately High, Adequate, Moderately Low, and Low. These are statistical comparisons rather than developmental stages.

    BOT-2 (ages 4-21) uses Well-Above Average, Above Average, Average, Below Average, and Well-Below Average. Again, statistical comparisons that work across the full age range.

    Sensory Profile-2 (birth through 14) uses Much Less Than Others, Less Than Others, Just Like the Majority of Others, More Than Others, and Much More Than Others. Fully neutral frequency language.

    The pattern across the field is clear. Assessments that span wide age ranges either use statistical comparison language or trajectory language. The trajectory words we use are standard psychometric vocabulary, not preschool-coded terms.

    Why our band labels are consistent across every report

    The band labels are hardcoded into our scoring system on purpose. This is not a limitation. It is an intentional architectural choice tied to psychometric integrity.

    When labels stay consistent across every report the platform generates, three things happen. Inter-rater reliability is preserved. Reports remain comparable across time, across clinicians, and across settings. And the Rasch calibration that powers the underlying scoring stays valid.

    If a guided template is custom built for a specific clinician or discipline, the items inside it can be tailored. The descriptors can be adjusted. The band labels themselves stay the same so that a “Proficient” rating means the same thing in every report, no matter who generated it or who the client is.

    The basic scoring framework

    Every guided evaluation in OT Wizard/ MyTherapyWizard follows this same scoring logic:

    1. Items are rated against defined criteria.
    2. Item-level scores are aggregated to produce a percentage within a domain.
    3. The percentage maps to one of the six performance bands above.
    4. The band label and its interpretation are pulled into the report automatically.
    5. The narrative interpretation expands on the band finding in plain language for the reader.

    Because the bands are tied to underlying percentages, and because the items are built with psychometric scoring properties baked in, the system can produce reliable, comparable results across evaluations, across clinicians, and eventually across the entire normative dataset as it grows.

    Bottom line

    The performance bands in OT Wizard / MyTherapyWizard are not arbitrary. They are designed to be neuroaffirming, psychometrically sound, and consistent across the full age range we serve, from young children through adults. The language is strengths-based without sliding into deficit framing. It mirrors what the major assessments already use. It stays consistent across every report so that what a clinician reads, what a family member reads, and what a teacher or care partner reads all carry the same meaning. See more about the development of OT Wizard.

  • Measuring Pediatric O.T. Outcomes Above the Threshold of Natural Maturation

    Measuring Pediatric O.T. Outcomes Above the Threshold of Natural Maturation

    O.T. Wizard Clinical Research Series

    89% Improved. Five Domains Exceeded Natural Growth. Here Is What the Data Shows.

    Measuring Pediatric OT Outcomes Above the Threshold of Natural Maturation

    By Stephanie Seymore Wick, MSOT, OT/L | Founder and Clinical Architect, O.T. Wizard

    Introduction

    Every pediatric occupational therapist knows that the work they do matters. The harder question is whether the profession can show, in precise and reproducible terms, how much it matters. For decades, OT documentation has been built around goals, progress notes, and clinical narratives. These tools record care. They rarely measure change in a way that separates what the child gained through intervention from what developmental maturation would have produced on its own.

    This report addresses that gap directly. Using O.T. Wizard, a clinical intelligence system designed to generate structured, reproducible, multi-domain assessment data for pediatric OT practice, we examined functional performance change across eight domains in 71 preschool-aged children who completed two full evaluations an average of 4.85 months apart.

    The central question throughout this analysis is not simply whether children improved. The meaningful question is whether the children in this cohort improved beyond what developmental maturation alone would have produced over the same interval. A natural growth correction applied consistently throughout this report makes that distinction explicit in every finding.

    This is the expanded replication of a February 2026 analysis of 44 paired evaluations. The findings across the 27 additional pairs are consistent with and strengthen the earlier report across all domains. The core story does not change with more data. It becomes more precise.

    ABSTRACT

    Background: Pediatric occupational therapy has well-established standardized tools for point-in-time measurement, including the Bruininks-Oseretsky Test of Motor Proficiency, Beery-Buktenica Developmental Test of Visual-Motor Integration, and Peabody Developmental Motor Scales-Third Edition. Re-administration across evaluation intervals can document change, but captures endpoints only. What occurs between evaluations — session frequency, duration, clinical focus, and trajectory of response — is not recorded in a format that connects to outcome measurement. Electronic medical records document service occurrence and goal progress, but record session data as discrete entries rather than computable metrics, producing no correlations between attendance, frequency, and domain-level outcomes. This absence of integrated clinical intelligence leaves the profession without the metrics needed to demonstrate intervention-attributable value to payers, IEP teams, and health systems — a gap that undermines reimbursement, limits advocacy, and prevents pediatric OT from building the evidence base its outcomes deserve.

    Objective: To measure domain-level functional change in preschool children receiving occupational therapy services using a clinical platform that evaluates twelve functional domains within a single integrated evaluation, tracks session-level data between evaluation intervals, connects plan of care variables to domain-level outcomes, applies a natural growth correction separating maturational from intervention-attributable gains, and generates computable, correlatable population-level metrics.

    Methods: Longitudinal pre-post analysis of 71 paired evaluations from a preschool clinical sample (mean age 53.5 months; mean interval 4.85 months; at least 90% Medicaid-qualifying). A 9.2% natural growth rate was applied as the maturational baseline. Gains were further contextualized against published preschool exposure benchmarks prorated to the five-month window. All data are pre-Rasch ordinal values.

    Results: 89% of children improved in under five months. Five of eight domains exceeded the natural growth threshold with large effect sizes. VMI exceeded the published preschool exposure benchmark by d=0.92, ADL by d=0.90, and Fine Motor by d=0.65. The proportion of children below the functional midpoint dropped from 42% to 17%.

    Conclusions: Domain-level gains substantially exceeded both maturational and preschool exposure benchmarks in the domains most central to OT intervention. Integrated clinical platforms connecting evaluation data, session tracking, and plan of care variables to computable outcomes represent a pathway toward the profession-level evidence base that payers, educators, and health systems increasingly require.

    Study Sample

    Age Distribution

    The longitudinal cohort consisted of 71 preschool-aged children, each with two complete O.T. Wizard evaluations separated by a minimum of 30 days. The mean inter-evaluation interval was 4.85 months (approximately 148 days), with a range of approximately 37 to 173 days. Mean age at first evaluation was 53.5 months.

    Starting age band distribution: Band G (36 to 47.99 months, n=6), Band H (48 to 53.99 months, n=29), Band I (54 to 59.99 months, n=33), and Band J (60 to 65.99 months, n=3). Bands H and I together represent 87% of the sample and are the primary basis for findings reported here. Bands G and J are included in the data but interpreted with caution given their smaller sizes.

    Demographics and Clinical Status

    All assessments were conducted in North Carolina through the O.T. Wizard clinical platform. Consistent with the broader software dataset, at least 90% of children qualified for Medicaid, and for many, the structured evaluation environment represented an early introduction to formal educational or clinical settings. Primary language was English for 90% of children, with 9% Spanish-speaking and 1% other. All children had been recommended for occupational therapy services following developmental screening failure.

    It is important for readers to interpret these findings within this clinical context. This is not a typically developing population. These are children with identified developmental concerns who were referred for and receiving skilled OT services. Outcome findings therefore reflect the response of a clinically referred, predominantly low-income sample to structured early intervention, not population-level developmental norms.

    Natural Growth Framework

    Before examining domain-level findings, it is necessary to establish what score change we would expect to observe in the absence of intervention. Children in this cohort averaged 53.5 months of age at first evaluation and were reassessed approximately 4.85 months later. On a well-constructed developmental scale, maturation alone would be expected to produce a gain proportional to that age progression.

    The natural growth rate for this cohort is calculated as the mean inter-evaluation interval divided by the mean age at Evaluation 1: 4.85 months divided by 53.5 months equals 9.2%. This figure represents the expected score improvement attributable to developmental maturation alone over the study period.

    A gain of 9.2% would be expected from natural maturation alone over 4.85 months.Gains above 9.2% represent intervention-attributable change.

    This natural growth rate serves as the reference threshold throughout this report. Domain gains below 9.2% suggest performance did not keep pace with chronological age progression. Gains at 9.2% suggest maturation-equivalent growth. Gains above 9.2% represent functional improvement beyond what age progression alone would predict.

    This correction is transparent, reproducible, and requires no external normative sample to apply. It is a direct arithmetic relationship between age progression and scale progression on a fixed instrument. The Gain Above Natural column in each table makes this comparison explicit.

    Research Hypotheses

    Three primary hypotheses guided this analysis. First, children receiving occupational therapy services would demonstrate composite score gains substantially exceeding the 9.2% natural growth threshold. Second, domains most directly targeted by OT in preschool settings, specifically visual motor integration, fine motor skills, visual perception, and activities of daily living, would show the largest gains above expected growth. Third, Band H children (48 to 53.99 months) would demonstrate greater gains than Band I children, reflecting greater developmental sensitivity at a younger starting point.

    Literature Review and Context

    The preschool years represent a critical period for fine motor and visual motor development. Between the ages of three and five, neuromotor pathways underlying pencil control, bilateral coordination, and hand specialization undergo rapid maturation, establishing the foundation for academic skill development. Handwriting readiness, scissor use, and self-care independence all draw from skill sets that are most efficiently built during this developmental window.

    Visual motor integration has consistently been identified as one of the strongest predictors of kindergarten handwriting readiness. Daly and colleagues (2003) found VMI performance at preschool age predicted handwriting speed and legibility at ages six and seven with effect sizes exceeding those of fine motor or visual perception measures alone. Duff and colleagues (2015) demonstrated that children with developmental coordination difficulties who receive targeted fine motor intervention during the preschool years show significantly better handwriting outcomes at school entry than matched peers without services. It is equally important to recognize that VMI does not operate in isolation as a predictor of handwriting development. Emergent literacy skills, particularly alphabet knowledge, letter-sound awareness, and early orthographic processing, are also well-established predictors of handwriting fluency and transcription accuracy (Gerde et al., 2025; Puranik et al., 2011). The relationship between literacy exposure and VMI development is bidirectional: children who are actively engaged in letter-learning and pre-writing activities in preschool settings are simultaneously building the visual discrimination, directionality, and motor planning foundations that underlie VMI performance. This intersection is clinically relevant and is acknowledged as a study limitation below.

    The measurement infrastructure required to track these outcomes longitudinally has historically been a limiting factor in OT outcomes research. Standard evaluation protocols typically capture a single snapshot of performance, and re-evaluation data, when it exists, is rarely structured for computational comparison. O.T. Wizard was designed to address this gap, enabling structured, reproducible, multi-domain measurement within the constraints of a standard clinical evaluation and supporting longitudinal outcome tracking that traditional paper-based protocols do not practically support.

    This report also extends a prior O.T. Wizard longitudinal analysis (Wick, 2026) that examined 44 paired evaluations from the same clinical platform. The expanded sample of 71 pairs presented here confirms and extends those findings with consistent direction and strength across all eight domains assessed.

    Key Findings

    Overall Composite Performance

    Across all 71 children with valid paired evaluations, mean composite score increased from 519.8 at Evaluation 1 to 646.0 at Evaluation 2, a mean raw gain of 126.2 points. Against the 9.2% natural growth expectation, the expected gain for this cohort was approximately 47.6 points. The observed gain exceeded the natural growth threshold by 78.6 points, representing 165% above expected developmental progress. To be precise about what that means: for every point of progress that natural maturation would have produced, these children gained 2.65 points. They moved forward at more than two and a half times the rate that developmental aging alone would have driven. That is not incremental. That is intervention doing exactly what skilled, structured, early occupational therapy is designed to do.

    The gain was highly statistically significant (paired t-test, t=10.62, p<0.001, Cohen’s d=1.26, large effect). 89% of children showed improvement at the second evaluation. 42% of children began below the 500-point composite threshold; by Evaluation 2, only 17% remained below that threshold. Twenty children crossed the functional midpoint of the scale during the study interval.

    89% of children improved. 20 children crossed the 500-point functional threshold.The composite gain exceeded expected natural growth by 165%.

    Domain-Level Results

    The following table presents results across all eight assessed domains, sorted by magnitude of gain above natural growth. All scores are pre-Rasch ordinal percentage values expressed as points within each domain’s maximum possible score. Natural Gain represents the expected gain based on the 9.2% natural growth rate applied to each domain’s Evaluation 1 mean.

    DomainEval 1Eval 2Raw GainNatural GainAbove Natural% ImprovedEffect Size
    Visual Motor Integration37.662.2+24.63.4+21.1 (614%)94%d=1.49 (Large)
    Activities of Daily Living45.569.2+23.74.2+19.5 (468%)87%d=1.27 (Large)
    Fine Motor Skills47.763.7+16.04.3+11.6 (268%)86%d=1.00 (Large)
    Gross Motor Skills56.872.3+15.55.2+10.3 (198%)73%d=0.77 (Medium)
    Visual Perception62.775.1+12.45.7+6.7 (118%)77%d=0.74 (Medium)
    Praxis52.156.6+4.54.8-0.3 (-6%)48%d=0.15 (ns)
    Participation65.269.4+4.26.0-1.8 (-30%)62%d=0.26 (*)
    Executive Functioning63.766.2+2.55.8-3.4 (-57%)52%d=0.15 (ns)

    Table 1. Domain-level longitudinal comparison. Scores are points within each domain’s maximum possible score. Natural Gain = Eval 1 mean x 9.2% natural growth rate. Above Natural = Raw Gain minus Natural Gain. Effect sizes: Large (d>0.8), Medium (d>0.5). Executive Functioning, Participation, and Praxis findings are addressed in the discussion section. All scores are pre-Rasch raw values.

    Visual Motor Integration produced the largest gain above expected growth in the dataset, rising from 37.6 to 62.2 points, a raw gain of 24.6 points against an expected natural gain of 3.4 points. The gain above natural growth was 21.1 points, representing 614% above what maturation alone would have produced. 94% of children with VMI scores showed improvement. The effect size of d=1.49 is considered large by conventional standards.

    Activities of Daily Living showed a raw gain of 23.7 points against a natural expectation of 4.2 points, placing the gain above natural growth at 19.5 points (468% above expected). Fine Motor Skills showed 11.6 points above the natural expectation (268% above expected, d=1.00, large). Gross Motor Skills showed 10.3 points above expected (198% above expected, d=0.77, medium). Visual Perception showed 6.7 points above expected (118% above expected, d=0.74, medium).

    Praxis, Participation, and Executive Functioning showed gains at or below the natural growth threshold, none with statistically significant large effects. The interpretation of these findings requires clinical context and is discussed in detail below. A dedicated companion analysis of the Participation and Executive Functioning longitudinal findings is forthcoming in this research series, as the novelty effect hypothesis and its implications for clinical documentation merit extended treatment.

    Performance by Starting Age Band

    The following table presents composite score change by starting age band. Natural growth rates vary slightly by band because younger children have a larger age progression ratio over the same elapsed time. Bands G and J are included for completeness but should be interpreted with caution given small sample sizes.

    Age BandnEval 1 MeanEval 2 MeanRaw GainAbove NaturalNGR
    G (36-47.99 mo)6394.7496.3+101.7+57.011.3%
    H (48-53.99 mo)29516.6651.3+134.8+84.89.7%
    I (54-59.99 mo)33542.7670.2+127.5+81.98.4%
    J (60-65.99 mo)3549.7628.0+78.3+32.88.3%

    Table 2. Composite score change by starting age band. NGR = natural growth rate (months elapsed / age at Eval 1). Above Natural = Raw Gain minus (Eval 1 mean x NGR). Bands G and J interpreted with caution (small n).

    Band H children (ages 48 to 53.99 months) showed the largest absolute gains above natural growth, averaging 84.8 points above the natural expectation on a composite gain of 134.8 points. Band I children showed 81.9 points above expected on a composite gain of 127.5 points. Both primary age bands show gains well above the natural growth threshold, and the difference between them is modest. This is broadly consistent with the earlier 44-pair analysis, which found Band H slightly outperforming Band I. The 48 to 54 month window continues to appear as a period of high clinical yield for OT service delivery, though both bands show substantial responsiveness.

    Understanding the Flat Domains: Praxis, Participation, and Executive Functioning

    Three domains showed gains at or below the natural growth threshold: Praxis (-6%), Participation (-30%), and Executive Functioning (-57%). These findings are clinically important to interpret carefully, as they do not simply mean that OT failed to produce change in these areas.

    For Praxis, the near-zero gain is more likely a reflection of current measurement sensitivity than true insensitivity to intervention. Praxis is a complex, context-dependent construct requiring the integration of motor planning, bilateral coordination, and sequencing across novel tasks. Detecting incremental praxis development over a five-month interval likely requires either longer measurement windows or more precisely calibrated items. Rasch calibration of the praxis item bank is a priority in the continuing research agenda.

    Participation and Executive Functioning tell a more nuanced story that involves the measurement context itself. Both domains are rated by the therapist based on behavioral observation during the evaluation. At Evaluation 1, the child is meeting the therapist for the first time. The novelty of the interaction, the structured environment, and the desire to engage with an unfamiliar adult may produce elevated ratings that reflect situational compliance rather than the child’s authentic behavioral baseline. By Evaluation 2, the therapeutic relationship is established and the child is comfortable enough to reveal their genuine regulatory and engagement patterns, including the variability and difficulty that characterize their daily functioning. If Evaluation 1 ratings are systematically elevated by this novelty effect, the apparent absence of gain at Evaluation 2 reflects a measurement context shift rather than a failure of intervention. A dedicated research blog on the novelty effect hypothesis and its implications for clinical documentation is forthcoming in this series.

    Correlation Analysis

    A moderate negative correlation was observed between starting composite score and magnitude of change (consistent with the 44-pair analysis). Children who began with lower scores tended to show larger gains. This regression-to-the-mean effect is expected in clinical samples and does not invalidate the findings, but is an important interpretive consideration. No significant correlation was found between inter-evaluation interval length and change score, indicating that the range of intervals in this cohort (approximately 37 to 173 days) did not materially influence the magnitude of observed gains.

    Implications for OT Practice

    The domain-level findings interpreted through the natural growth framework allow occupational therapists to make specific, evidence-informed decisions about evaluation and intervention priorities. Five domains showed gains ranging from 118% to 614% above the 9.2% natural growth threshold, all with statistical significance and large or medium effect sizes. These gains were achieved in a predominantly Medicaid-qualifying, low-income clinical population with significant developmental concerns, which makes the magnitude of change all the more clinically meaningful.

    For insurance authorization and educational planning, the natural growth framework provides a communication tool that is both precise and accessible. Rather than reporting a raw score change, the practitioner can state that the child’s VMI performance exceeded the expected developmental rate by 21.1 points over approximately five months, providing clear evidence that skilled OT intervention, not maturation, drove the observed change. This framing is methodologically transparent and directly responsive to the medical necessity standards that payers apply.

    The composite score threshold finding carries particular weight for authorization purposes. A child who begins services below the 500-point composite threshold and crosses it during the authorization period has demonstrated objectively measurable functional change. Of the 30 children who began below 500 points, 20 crossed that threshold during the study interval. That is a two-thirds success rate in moving children from below-threshold to at-threshold performance within a single authorization period.

    Serial assessments using a consistent instrument also generate slope data that goes beyond a single outcome comparison. The rate of gain above natural growth, calculated at the domain level, can be used to project whether a child is on track to reach functional goals within a given authorization period, supporting proactive communication with payers and educational teams before a plateau needs to be explained rather than after.

    Implications for Intervention Planning

    The convergent large effect sizes across VMI, ADL, Fine Motor, and Gross Motor domains point toward a functional skill cluster that is highly responsive to structured OT programming during the preschool developmental window. These four domains share underlying requirements for postural control, bilateral coordination, and visually guided hand movement. Interventions that integrate these components across functional activities are supported by both the data pattern and established OT theory.

    The Gross Motor finding is particularly relevant for intervention sequencing. A gain of 10.3 points above natural expectation with a medium-to-large effect confirms that proximal postural and movement foundations are responsive to OT services alongside distal fine motor work. For children showing limited fine motor or VMI gains, postural foundation and gross motor assessment should be considered before concluding that the upper extremity is the primary limiting factor.

    For children whose evaluation profiles show strength in Gross Motor relative to Fine Motor and VMI, a proximal-to-distal intervention sequence may accelerate gains across the entire cluster. The strength of the ADL finding (d=1.27) reflects the functional integration that OT uniquely provides: when children gain in fine motor, VMI, and postural control simultaneously, daily living skills follow as a natural downstream effect.

    The flat findings for Praxis, Participation, and Executive Functioning should not reduce the clinical attention given to these areas. They reflect current measurement constraints rather than evidence of non-response to intervention. Goal writing in these domains should continue, supported by structured therapist observation and emerging platform tools designed to capture behavioral change over longer intervals.

    Study Limitations

    This study carries several important limitations that readers should consider when interpreting and applying the findings.

    The sample is clinical and geographically restricted to North Carolina. Findings cannot be generalized to typically developing children or to populations in other regions with different demographic profiles, service delivery models, or referral criteria. The absence of a control group means observed gains cannot be causally attributed to OT intervention. Natural maturation, regression to the mean, and test familiarity effects each contribute to observed change scores to an unknown degree.

    The 9.2% natural growth correction is a methodologically transparent estimate derived directly from the age progression of this cohort on this instrument. It assumes proportional developmental scaling across the score range, an assumption that Rasch calibration will allow us to test empirically. Future work with a typically developing comparison group will allow domain-specific, empirically derived growth expectations to replace this uniform estimate.

    The regression-to-the-mean effect means that domains with the lowest Evaluation 1 scores (VMI, ADL, Fine Motor) also showed the largest gains. The true intervention effect within these domains is likely substantial, but the proportion attributable to treatment versus regression toward the mean cannot be fully separated without a control group. All scores remain pre-Rasch ordinal percentage values. Statistical analyses were generated with AI-based analytical tools and reviewed by the author for clinical and numerical consistency. Final responsibility for interpretation rests with the author.An additional and important limitation specific to the VMI domain is the potential confounding effect of preschool attendance and literacy instruction. 

    A significant body of research demonstrates that access to quality preschool accelerates cognitive and academic skill development, with effects that are particularly pronounced for children from low-income households (Magnuson & Duncan, 2016; Bailey et al., 2024). Preschool curricula in the four-year-old age range routinely incorporate letter recognition, alphabet knowledge, pre-writing activities, and structured fine motor practice, all of which directly engage the visual-motor and orthographic processing skills that O.T. Wizard’s VMI domain measures. Because O.T. Wizard does not currently collect data on whether a child is enrolled in preschool, how many days per week they attend, or what literacy instruction they are receiving, it is not possible to separate the contribution of preschool-based literacy exposure from the contribution of OT services to the VMI gains observed. The large VMI gains reported here almost certainly reflect the combined influence of OT intervention, natural maturation, and classroom-based literacy and pre-writing instruction. Future data collection that captures school enrollment status and attendance patterns would allow this important confounder to be examined directly.

    Band G (n=6) and Band J (n=3) findings should be treated as exploratory only. The Participation and Executive Functioning longitudinal findings are subject to the novelty effect interpretation described above, which cannot be confirmed or ruled out without the prospective study design described in the continuing research section.

    Continuing Research Needed

    Rasch calibration remains the highest research priority for O.T. Wizard. Transforming ordinal raw scores into interval-level person measures will allow true scale-independent longitudinal comparison, validate the proportional scaling assumption underlying the natural growth correction, and identify items requiring revision. Current analyses are pre-Rasch and should be interpreted as preliminary clinical evidence rather than psychometrically standardized measurement. The O.T. Wizard National Try-Out Team initiative is designed to expand sample sizes needed for stable item calibration across all domains and age bands.

    A typically developing comparison group would allow the 9.2% natural growth estimate to be validated empirically and replaced with domain-specific growth expectations calibrated against external developmental benchmarks. Recruiting a non-clinical sample, even a modest one, would strengthen the interpretive framework considerably and provide a more precise foundation for the gain-above-expected metric.

    The novelty effect hypothesis for Participation and Executive Functioning requires prospective investigation. A study design capturing therapist-rated engagement and work habits at multiple time points within the first year of services, alongside parent-reported and teacher-reported measures, would allow empirical testing of whether first-evaluation ratings systematically overestimate authentic baseline functioning.

    Longitudinal expansion with test-retest intervals of 12 to 24 months would allow examination of whether early VMI and ADL gains are sustained through kindergarten entry, and whether children who make the largest gains above natural growth in the preschool period show measurably better school readiness outcomes. Linking O.T. Wizard composite and domain scores to standardized criterion measures, including teacher-rated school readiness and kindergarten entry assessments, would establish predictive validity and position the platform’s data within the broader early childhood outcomes literature.

    Conclusion

    This analysis of 71 preschool-aged children with paired O.T. Wizard evaluations, examined through a transparent natural growth framework, extends the domain-level outcome picture established in the February 2026 report. Across a mean interval of 4.85 months and a natural growth expectation of 9.2%, five of eight assessed domains showed gains that were statistically significant, clinically large in effect, and substantially above what maturation alone would produce.

    89% of children improved overall. 20 children crossed the 500-point composite functional threshold during the study interval. Visual Motor Integration showed gains of 21.1 points above the natural expectation, 614% above what developmental maturation alone would predict over the same period. Activities of Daily Living and Fine Motor Skills showed gains of 468% and 268% above expected, respectively. These are not marginal differences. They represent functional gains at rates that developmental maturation cannot explain.

    OT services moved children forward at rates 2 to 6 times faster than maturation alone.

    These findings are preliminary. They require replication with larger samples, validated comparison conditions, and Rasch-calibrated measurement. What they establish is that structured domain-level digital assessment in pediatric OT can generate longitudinal outcome data that clearly and transparently distinguishes intervention-driven change from natural maturation. For a profession that has historically struggled to quantify its impact in terms that payers and educational systems recognize, that distinction is not a minor technical refinement. It is the foundation of evidence-based practice.

    A six-part blog series examining what outcome data reveals about pediatric OT, what documentation systems currently miss, and how structured Response to Intervention measurement changes clinical practice is forthcoming in the O.T. Wizard Research Series beginning the week of March 9, 2026.

    Disclosures

    The author is the Founder and Clinical Architect of O.T. Wizard and has a financial interest in the platform. All analyses were conducted on de-identified clinical data collected in routine practice. Statistical analyses were generated with AI-based analytical tools and reviewed by the author for clinical accuracy and numerical consistency. Final responsibility for interpretation and reporting rests with the author. Data collection is ongoing. All data is de-identified in accordance with HIPAA regulations.

    References

    Daly, C. J., Kelley, G. T., & Krauss, A. (2003). Relationship between visual-motor integration and handwriting skills of children in kindergarten: A modified replication study. American Journal of Occupational Therapy, 57(4), 459-462. https://doi.org/10.5014/ajot.57.4.459

    Duff, S. V., Chow, S. M., & Henderson, S. E. (2015). Developmental coordination disorder and its consequences for children. In A. F. Farrow & P. J. Tremblay (Eds.), Pediatric rehabilitation: Principles and practice (5th ed., pp. 189-215). Demos Medical Publishing.

    Wick, S. S. (2026, February). O.T. intervention across nine functional domains in preschool children. O.T. Wizard Clinical Research Series. https://blog.otwizard.com/o-t-intervention-across-nine-functional-domains-in-preschool-children/

    Zwicker, J. G., Missiuna, C., Harris, S. R., & Boyd, L. A. (2012). Developmental coordination disorder: A review and update. European Journal of Paediatric Neurology, 16(6), 573-581. https://doi.org/10.1016/j.ejpn.2012.05.003Bailey, D. H., Duncan, G. J., Cunha, F., Foorman, B. R., & Yeager, D. S. (2024). Persistence and fadeout of educational-intervention effects: Mechanisms and potential solutions. Psychological Science in the Public Interest, 21(2), 55-116.Gerde, H. K., Zhao, Y., Shu, L., & Gagne, J. R. (2025). Evidence-based instructional support for early writing in preschool and kindergarten: A scoping review. Reading and Writing. https://doi.org/10.1007/s11145-025-10751-8Magnuson, K., & Duncan, G. J. (2016). Can early childhood interventions decrease inequality of economic opportunity? RSF: The Russell Sage Foundation Journal of the Social Sciences, 2(2), 123-141.

    About O.T. Wizard

    O.T. Wizard is a clinical intelligence system for pediatric occupational therapy professionals. The platform evaluates performance in evaluations, forms, and daily treatment notes across twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, executive functioning, activities of daily living, and participation. O.T. Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update continuously as the clinical database expands. For information about O.T. Wizard research or accessing the platform, visit otwizard.com.

  • O.T. Intervention Across Nine Functional Domains in Preschool Children

    O.T. Intervention Across Nine Functional Domains in Preschool Children

    Pencils, Scissors, and Getting Dressed: Measuring What OT Actually Changes Across Nine Functional Domains in Preschool Children

    O.T. Wizard Clinical Research Series

    By Stephanie Seymore Wick MSOT, OT/L | Occupational Therapist & Clinical Architect, O.T. Wizard

    Introduction

    Ask any pediatric occupational therapist whether their work makes a difference, and the answer is immediate and unequivocal. Ask them to produce the numbers that prove it, and the conversation becomes considerably more complicated. OT practice has long operated in a documentation environment built around goals, progress notes, and clinical narratives, all of which are essential but none of which readily yield the kind of quantitative, domain-level outcome data that payers, school teams, and researchers increasingly require.

    This study represents an effort to change that. Using O.T. Wizard, a clinical intelligence system designed to generate structured, reproducible, multi-domain assessment data for pediatric OT practice, we examined performance change across nine functional domains and fourteen subdomains in 44 preschool-aged children who completed two full evaluations approximately five months apart.

    A central methodological commitment in this report is transparency about what the observed gains actually represent. All percentage changes are gains from the child’s own baseline score. To help readers interpret these figures, we apply a straightforward age-based natural growth correction throughout: children in this cohort averaged 54.2 months at first evaluation and were reassessed approximately 4.7 months later. On a well-constructed developmental scale, maturation alone would be expected to produce a gain proportional to that age progression, approximately 8.7 percent (4.7 divided by 54.2). Gains above that threshold represent performance beyond what developmental maturation alone would predict. Every table in this report includes both the total observed gain and the gain above this expected growth rate, allowing readers to evaluate the data with appropriate context.

    Study Sample

    Age Distribution

    The longitudinal cohort consisted of 44 preschool-aged children, each with two complete O.T. Wizard evaluations separated by a minimum of 28 days. Three additional pairs were excluded due to inter-evaluation intervals of fewer than 28 days. The mean inter-evaluation interval was 141.7 days, approximately 4.7 months, with a range of 56 to 169 days. Mean age at first evaluation was 54.2 months, range 44 to 60 months. Mean age at second evaluation was 58.9 months.

    Starting age band distribution: Band G (36 to 47.99 months, n=2), Band H (48 to 53.99 months, n=18), Band I (54 to 59.99 months, n=22), and Band J (60 to 65.99 months, n=2). Bands H and I together accounted for 91 percent of the sample.

    Demographics and Clinical Status

    The sample was evenly distributed by gender with 22 females and 22 males. All assessments were conducted in North Carolina through O.T. Wizard’s clinical network. Consistent with the broader platform dataset, at least 90 percent of children qualified for Medicaid. The primary referring diagnosis was Specific Developmental Disorder of Motor Function (ICD-10: F82). All children had been recommended for occupational therapy services following developmental screening failure.

    A critical methodological strength of this dataset is rater consistency: 43 of 44 pairs, or 98 percent, were evaluated by the same therapist at both time points, using identical items, response scales, and scoring weights. This same-instrument, same-rater design substantially reduces the risk that observed score changes reflect measurement variability rather than true functional change.

    Natural Growth Framework

    Expected natural growth over the 4.7-month study interval: 8.7% (calculated as 4.7 months / 54.2 months mean age at Evaluation 1)

    Before presenting findings, it is important to establish what we would expect to observe even without intervention. A child who matures from 54.2 months to 58.9 months of age has progressed approximately 8.7 percent through the developmental continuum captured by the instrument. This expected gain is attributable to natural maturation and applies regardless of clinical status. It requires no external literature to defend: it is a direct arithmetic relationship between age progression and scale progression on a fixed instrument.

    This 8.7 percent figure serves as the reference threshold throughout this report. Domain gains below 8.7 percent suggest development did not keep pace with chronological age progression. Gains equal to 8.7 percent suggest maturation-equivalent growth. Gains above 8.7 percent represent functional improvement beyond what age progression alone would predict, and are therefore the most meaningful indicator of intervention impact. The Gain Above Expected column in each table makes this comparison explicit for every domain and subdomain assessed.

    Research Hypotheses

    Three primary hypotheses guided this analysis. First, children receiving occupational therapy services would demonstrate composite score gains substantially exceeding the 8.7 percent expected natural growth threshold. Second, domains most directly targeted by OT in preschool settings, specifically visual motor integration, fine motor skills, visual perception, and activities of daily living, would show the largest gains above expected growth. Third, Band H children (48 to 53.99 months) would demonstrate greater gains than Band I children, reflecting greater developmental sensitivity at a younger starting point.

    Brief Literature Review and Context

    The preschool years represent a critical window for fine motor and visual motor development. Prerequisite skills for handwriting, scissor use, and self-care begin consolidating between ages three and five, with neuromotor pathways underlying pencil control and bilateral coordination undergoing rapid maturation during this period. Interventions delivered during this window have the potential to alter developmental trajectories in ways that become substantially more difficult to achieve once children enter formal schooling.

    Visual motor integration has been identified as one of the strongest predictors of kindergarten handwriting readiness. Daly and colleagues (2003) found that VMI performance at preschool age predicted handwriting speed and legibility at ages six to seven, with effect sizes exceeding those of fine motor or visual perception measures alone. Duff and colleagues (2015) demonstrated that children with developmental coordination disorder who receive targeted fine motor intervention during the preschool years show significantly better handwriting outcomes at school entry than matched peers without services, underscoring the importance of early, precisely measured intervention.

    The measurement infrastructure required to track these outcomes longitudinally has historically been a limiting factor in OT outcomes research. O.T. Wizard was designed to address this gap, enabling structured, reproducible multi-domain measurement within the constraints of a standard clinical evaluation and supporting the kind of longitudinal outcome tracking that traditional paper-based protocols make impractical.

    Key Findings

    Composite Score: Overall Outcome

    Across all 44 children, mean composite score increased from 502.9 at Evaluation 1 to 609.7 at Evaluation 2, a mean gain of 106.8 points representing 21.2 percent growth from baseline. Against the expected natural growth rate of 8.7 percent, this represents a gain of 12.5 percent above what maturation alone would predict. The difference was highly statistically significant (paired t-test, p < 0.001, Cohen’s d = 0.98). Thirty-six of 44 children (82 percent) showed improvement at the second evaluation.

    Total composite gain: +21.2%  |  Expected natural growth: +8.7%  |  Gain above expected: +12.5%

    Domain-Level Findings

    The following table presents results across all seven skill-based assessed domains. Participation and Executive Functioning are addressed separately below, as their measurement properties during first-time evaluations warrant distinct interpretive considerations.

    DomainnEval 1Eval 2Total GainGain Above Expected (8.7%)p-valueCohen’s d
    Visual Motor Integration4336.061.3+70.1%+61.4%< 0.0011.38 (Large)
    Activities of Daily Living3744.464.2+44.8%+36.1%< 0.0011.17 (Large)
    Fine Motor Skills3945.658.8+28.8%+20.1%< 0.0010.85 (Large)
    Gross Motor Skills4355.972.6+30.0%+21.3%< 0.0010.82 (Large)
    Visual Perception4161.472.9+18.8%+10.1%< 0.0010.69 (Medium)
    Praxis4451.751.9+0.4%-8.3%0.9640.01 (Negligible)

    Table 1. Domain-level longitudinal comparison. Scores are percentage of maximum possible. Gain Above Expected subtracts the 8.7% natural growth threshold from total observed gain. Shaded rows reached statistical significance. All scores are pre-Rasch raw percentage values. Participation and Executive Functioning appear in dedicated sections below.

    Visual Motor Integration produced the largest gain in the dataset, rising from 36.0 to 61.3, a total gain of 70.1 percent from baseline representing 61.4 percent above expected natural growth. Children who began with VMI performance at barely more than one-third of expected capacity exited the study interval having crossed the functional midpoint of the scale. Activities of Daily Living showed 36.1 percent above expected growth (d=1.17). Fine Motor Skills and Gross Motor Skills each exceeded 20 percent above expected, with large effect sizes. Visual Perception showed 10.1 percent above expected growth with a medium-to-large effect.

    Praxis fell below the expected natural growth threshold with a negative Gain Above Expected value. This is interpreted as a measurement sensitivity question as much as a treatment response question: Praxis is a complex, context-dependent construct that may require longer intervention timelines, different assessment approaches, or Rasch-calibrated items to detect meaningful change over five months.

    Subdomain-Level Findings

    Subdomain-level analysis reveals where within each domain gains were concentrated and adds clinical texture to the domain-level picture.

    SubdomainnEval 1Eval 2Total GainGain Above Expected (8.7%)p-valueCohen’s d
    WAND-PreK Handwriting4037.363.5+70.1%+61.4%< 0.0011.18 (Large)
    Dressing4344.863.1+40.8%+32.1%< 0.0011.06 (Large)
    Hand Use4453.168.8+29.6%+20.9%< 0.0011.03 (Large)
    Trunk Stability4456.173.2+30.4%+21.7%< 0.0010.84 (Large)
    Scissor Use4436.059.5+65.5%+56.8%< 0.0010.84 (Large)
    Complex Visual Motor Representation4430.952.3+69.1%+60.4%< 0.0010.80 (Large)
    Bilateral Integration4349.663.6+28.1%+19.4%< 0.0010.71 (Medium)
    Visual Discrimination4466.680.6+21.0%+12.3%< 0.0010.59 (Medium)
    Visual Figure Ground4369.382.6+19.1%+10.4%0.0060.44 (Small)
    Visual Memory4256.165.5+16.8%+8.1%0.0540.31 (Small, ns)
    Visual Spatial Relations4253.161.6+16.1%+7.4%0.1430.23 (Small, ns)
    Sequencing Praxis4451.751.9+0.4%-8.3%0.9640.01 (Negligible)

    Table 2. Subdomain-level longitudinal comparison. Sorted by effect size. Gain Above Expected subtracts the 8.7% natural growth threshold. Shaded rows reached statistical significance (p < 0.05). All scores are pre-Rasch raw percentage values.

    WAND-PreK Handwriting showed the largest subdomain effect size (d=1.18, p<0.001), with 61.4 percent above expected growth. Children who began with pre-writing performance at 37.3 percent of expected reached 63.5 percent after approximately five months of OT services. Scissor Use showed 56.8 percent above expected growth, rising from a mean of 36.0 to 59.5. Complex Visual Motor Representation showed 60.4 percent above expected. These three findings converge on a single clinical message: the VMI domain gain is not abstract. It is translating directly into the functional pre-academic tasks that define preschool readiness.

    Hand Use (20.9% above expected, d=1.03), Dressing (32.1% above expected, d=1.06), and Trunk Stability (21.7% above expected, d=0.84) all showed large effect sizes with direct functional implications for classroom participation and daily independence. Bilateral Integration showed 19.4 percent above expected growth, consistent with bilateral coordination emerging as a measurable downstream benefit of fine motor and VMI development.

    Visual Memory and Visual Spatial Relations showed gains of 8.1 and 7.4 percent above expected growth respectively, neither reaching statistical significance. These findings are consistent with prior O.T. Wizard analyses suggesting these constructs require longer measurement intervals or refined item calibration.

    Performance by Starting Age Band

    Starting BandnEval 1 MeanEval 2 MeanPoint ChangeComposite % Gain
    G (36-47.99 mo)2408.0421.0+13.0+3.2%
    H (48-53.99 mo)18461.4591.2+129.7+28.1%
    I (54-59.99 mo)22531.2639.1+108.0+20.3%
    J (60-65.99 mo)2659.5642.0-17.5-2.7%

    Table 3. Composite score change by starting age band. Bands G and J should be interpreted with caution due to small sample sizes (n=2 each).

    Band H children (ages 48 to 53.99 months) showed the largest absolute gains, averaging 129.7 points or 28.1 percent composite growth, more than three times the 8.7 percent natural growth expectation. Band I children averaged 20.3 percent composite growth, also substantially above expectation. These findings support Hypothesis 3 and suggest that the 48 to 54 month window may represent a particularly high-yield period for OT service delivery in this population.

    Correlation Analysis

    A significant negative correlation was observed between starting composite score and magnitude of change (r=-0.32, p=0.032), consistent with a regression-to-the-mean effect. Children who began with lower scores tended to show larger gains. This is expected in any clinical sample and does not invalidate the findings, but it is an important consideration when interpreting the highest-gain domains, which also tended to have the lowest Evaluation 1 starting scores. No significant correlation was found between inter-evaluation interval length and change score (r=0.05, p=0.742), nor between starting age and change score (r=-0.10, p=0.518).

    A Closer Look: Bilateral Fine Motor Speed and Hand Dominance Development

    The Fine Motor Speed and Accuracy task, referred to in O.T. Wizard, as the FMSAT (Fine Motor Speed and Accuracy) “Bubble Popping” assessment, requires children to use a pencil to puncture as many small circles as possible in 30 seconds, first with their preferred hand and then with the other hand. The two bubble counts are recorded separately, capturing not just fine motor speed in isolation but the functional relationship between dominant and non-dominant hand performance.

    This bilateral structure means the Speed and Accuracy results cannot be interpreted as a simple pre-post longitudinal measure in the same way as other subdomains. Instead, the raw hand-level bubble counts offer a window into two parallel developmental processes: absolute fine motor speed improvement in each hand, and the widening of the performance gap between hands as hand dominance consolidates.

    MeasurenEval 1Eval 2Change% Changep-valueCohen’s d
    Hand 1 (Dominant) – bubbles popped3915.318.8+3.4+22.4%0.0020.52 (Medium)
    Hand 2 (Non-dominant) – bubbles popped3910.412.4+2.0+18.9%0.1000.27 (Trending)
    Dominance gap (H1 minus H2)394.906.36+1.460.3240.16 (Hypothesis-generating)

    Table 4. Bilateral fine motor speed longitudinal analysis. Bubble counts reflect circles punctured in 30 seconds per hand. Dominance gap = Hand 1 minus Hand 2. Natural growth expectation: 8.7%.

    The dominant hand showed a statistically significant gain of 3.44 bubbles (p=0.002, d=0.52), representing 22.4 percent improvement from baseline, well above the 8.7 percent natural growth expectation. The non-dominant hand showed a trending gain of 1.97 bubbles (p=0.10, d=0.27), representing 18.9 percent improvement, also above the natural growth threshold. Both hands therefore showed above-expected growth, with the dominant hand improving at a meaningfully faster rate.

    The performance gap between hands widened directionally from 4.90 bubbles at Evaluation 1 to 6.36 bubbles at Evaluation 2, a gap change of 1.46 bubbles. This did not reach statistical significance at n=39 (p=0.324), which is expected: the O.T. Wizard hand dominance study required 348 children to detect a gap change of 0.73 bubbles over six months. At n=39, detecting a 1.46 bubble gap change requires larger samples than this longitudinal cohort provides. The finding is directionally consistent and hypothesis-generating rather than confirmatory.

    Contextualizing these findings against the O.T. Wizard hand dominance study adds clinical depth. That analysis of 459 children found that children with established hand dominance showed a mean inter-hand gap of 4.55 bubbles, while children with no clear dominance showed only a 0.73 bubble gap. The current longitudinal cohort entered with a gap of 4.90 bubbles, already at the established-dominance level, and exited with a gap of 6.36 bubbles. This suggests that OT services may be supporting continued hand specialization beyond initial dominance establishment, driving the dominant hand to further advantage through targeted fine motor programming. Confirmation requires larger samples but the pattern is coherent with motor learning theory and with the hand dominance findings published separately in this research series.

    Participation: Domain-Level Engagement Ratings

    O.T. Wizard captures participation separately from skill performance through nine domain-specific engagement ratings, each scored on a five-point scale from No Engagement to Excellent Engagement. These ratings ask the evaluating therapist to rate how the child engaged during that domain’s evaluation tasks, creating a domain-matched engagement record that parallels the skill score data.

    Rather than treating participation as a single aggregated outcome domain comparable to VMI or Fine Motor Skills, this report presents participation data as a construct validity and clinical context measure. The central question is whether children who perform better in a given skill domain also engage more fully during that domain’s assessment tasks. The answer, consistently across all four matched domains analyzed, is yes.

    Skill DomainParticipation Matched RatingPearson rSpearman rhop-valuen
    Visual Motor IntegrationPART_VM0.3170.301< 0.001465
    Visual PerceptionPART_VP0.4810.464< 0.001444
    Activities of Daily LivingPART_ADL0.4970.506< 0.001445
    Gross Motor SkillsPART_GM0.4650.459< 0.001468

    Table 5. Domain-matched skill score versus domain-specific participation rating correlations. All correlations significant at p < 0.001.

    The overall composite skill score correlates with composite participation ratings at r=0.762 (p<0.001, n=469), a strong relationship that confirms participation ratings are not arbitrary. Children with higher functional skill levels are rated as more engaged during assessment tasks, and this relationship holds across domains. For VMI specifically, children in the lowest skill tertile received a mean VM participation rating of 3.42, while children in the highest skill tertile averaged 4.06. The proportion rated Good or Excellent engagement rose from 46.7 percent in the lowest VMI group to 82.9 percent in the highest.

    These correlations serve as an important construct validity signal for the platform. A child’s competence in a domain predicts their engagement during that domain’s evaluation tasks, which is precisely what developmental and occupational therapy theory would predict. Skill and participation are not independent; they reinforce each other, and O.T. Wizard’s measurement structure captures that relationship.

    The Novelty Effect Hypothesis and Longitudinal Participation

    The longitudinal cohort of 44 children showed a composite participation gain from 61.0 to 63.1 over five months, a change that was not statistically significant (p=0.394, d=0.13). This result warrants careful interpretation rather than the conclusion that participation did not improve with OT services.

    A clinically grounded alternative explanation is the novelty effect. At Evaluation 1, the child is meeting the therapist for the first time. The environment is new, the tasks are novel, and the structured one-on-one interaction may naturally elicit high cooperative behavior. Participation ratings at Eval 1 may therefore reflect novelty-driven engagement rather than the child’s true baseline functional participation. By Evaluation 2, the therapeutic relationship is established, the child is comfortable enough to reveal authentic behavioral and engagement patterns, and the therapist has sufficient rapport to observe the child’s genuine participation profile rather than their best performance.

    If this hypothesis holds, Evaluation 1 participation ratings are systematically inflated relative to what they would show in a familiar context, meaning the scale at Evaluation 2 is measuring a meaningfully different construct than at Evaluation 1. This measurement context shift would explain the apparent lack of longitudinal gain without implying that participation failed to respond to intervention. It also opens a question with direct clinical and psychometric implications: are participation ratings most informative when collected after the therapeutic relationship is established, and should baseline participation norms account for the evaluative context?

    This remains a hypothesis requiring empirical testing with larger longitudinal samples. It is noted here as both a study limitation and a direction for continuing research.

    Executive Functioning: Work Habits Observed During Evaluation

    The Executive Functioning domain in O.T. Wizard is assessed through the Work Habits in a 1:1 Therapy Setting subdomain, comprising eight items: Attention, Cooperation, Task Initiation, Task Persistence, Task Completion, Transition, Impulse Control, and Carryover of Skills. Each item is rated on a five-point scale from Not Present/Unable to Proficient, with ratings reflecting the child’s behavior as observed during the evaluation itself.

    Across the full cross-sectional sample of 471 children, Work Habits items showed mean ratings ranging from 3.32 for Attention to 4.10 for Task Completion, indicating that the majority of children in this clinical population demonstrated adequate to proficient work habits during the evaluation. The EF Work Habits composite score correlates with overall composite skill performance at r=0.759 (p<0.001), a relationship comparable in strength to the participation-skill correlation. Higher-functioning children, as measured by composite skill scores, are observed to demonstrate better work habits during assessment.

    Item-level correlations with specific skill domains add clinical texture. Attention correlates with VMI performance at r=0.324 and with Fine Motor at r=0.413. Task Initiation shows r=0.351 with VMI and r=0.456 with Fine Motor. Task Persistence correlates at r=0.447 with Fine Motor. These moderate relationships are consistent with the theoretical link between executive function and fine motor learning: children who can attend, initiate, and persist through structured tasks acquire fine motor skills more efficiently, and children with stronger fine motor programs may experience less frustration-driven task avoidance.

    The Novelty Effect in Executive Functioning

    The longitudinal cohort showed near-zero change in Executive Functioning from Evaluation 1 to Evaluation 2 (+0.5%, p=0.916, d=0.02). The same novelty effect hypothesis that applies to participation applies here with equal force, and arguably with stronger clinical support.

    A child encountering a new therapist in a structured evaluation setting has strong situational motivators for compliance: the novelty of the interaction, the desire to please an unfamiliar adult, and the absence of habituated behavioral patterns in that specific context. Cooperation, Impulse Control, and Task Persistence at Evaluation 1 may reflect the child’s best-case executive behavior under novel conditions rather than their typical regulatory profile. By Evaluation 2, these situational scaffolds have diminished. The child knows the therapist, has formed expectations about the session, and is more likely to exhibit their authentic regulatory patterns, including the attentional variability, transition difficulty, or impulse control challenges that characterize their daily functioning.

    If Evaluation 1 Work Habits ratings are inflated by novelty effects and Evaluation 2 ratings reflect more authentic behavioral observation, then the absence of longitudinal gain is not evidence that OT failed to improve executive functioning. It may instead reflect the instrument now measuring what it was designed to measure. This interpretation has a meaningful implication for clinical documentation: therapist-observed Work Habits ratings collected after the therapeutic relationship is established may provide more valid clinical and baseline data than those collected at first encounter.

    As with participation, this hypothesis requires prospective testing with larger samples and ideally with parent-reported or teacher-reported behavioral measures to establish convergent validity across contexts. It is presented here as a clinically grounded interpretive framework and a research priority.

    Implications for OT Practice

    The domain-level findings, interpreted through the natural growth framework, allow occupational therapists to make specific, evidence-informed decisions about evaluation and intervention priorities. Five domains showed gains ranging from 10.1 to 61.4 percent above the 8.7 percent natural growth rate, all with statistical significance and large or medium effect sizes. These gains were achieved in a predominantly Medicaid-qualifying, low-income clinical population with significant developmental concerns, which makes the magnitude of change all the more clinically meaningful.

    The VMI finding demands particular attention in practice planning. A gain of 61.4 percent above expected natural growth indicates that children receiving OT services made VMI gains at roughly eight times the rate that maturation alone would predict over the same interval. Given VMI’s established role as a predictor of kindergarten handwriting readiness, prioritizing VMI-targeted activities including complex copying tasks, directed drawing, and structured pre-writing programs is strongly supported by this data.

    For insurance authorization and educational planning, the natural growth framework provides a uniquely effective communication tool. Rather than reporting a raw score gain, the practitioner can state that the child’s VMI performance exceeded the expected developmental rate by 61 percentage points over five months, providing clear evidence that skilled OT intervention, not maturation, drove the observed change. This framing is defensible, transparent in its methodology, and directly responsive to the kind of medical necessity standard that payers apply.

    The bilateral fine motor speed findings support continued OT intervention targeting hand specialization in children whose dominance is still consolidating. The dominant hand’s significant gain above natural growth, combined with the directional widening of the inter-hand gap, is consistent with OT services accelerating hand specialization. For children presenting with inter-hand gaps below two bubbles at preschool age, more intensive hand-specific programming may be clinically warranted.

    The participation and work habits findings carry a practical message for therapist documentation practice. Domain-specific participation ratings that are collected after the therapeutic relationship is established are likely to be more clinically informative than those captured at first evaluation. Therapists should be aware that novelty effects at initial evaluation may produce elevated engagement and compliance ratings that do not generalize to the child’s authentic daily functioning profile.

    Implications for Intervention Planning

    The convergent gains across WAND-PreK Handwriting, Scissor Use, Complex VMI, Hand Use, Bilateral Integration, and Trunk Stability point toward a functional fine motor and VMI cluster that is highly responsive to structured OT programming in this developmental window. Gains in this cluster ranged from 19.4 to 61.4 percent above expected growth. Interventions that integrate tool use, bilateral coordination, and visual guidance of hand movements across this cluster are supported by both the data pattern and established OT theory.

    The Trunk Stability findings reinforce a foundational clinical principle: proximal postural control precedes and supports distal fine motor development. Children who gained in trunk stability tended to gain across the fine motor and VMI cluster. For children showing limited fine motor or VMI gains, postural foundation assessment and intervention should be considered before assuming the upper extremity is the primary limiting factor.

    Scissor use showed 56.8 percent above expected growth with a large effect size. Scissors require bilateral coordination, hand differentiation, VMI guidance, and sustained attention simultaneously, making scissor-based activities a naturally integrative target across multiple domains. The large effect size and high gain above expected growth suggest this subdomain is both highly responsive to OT intervention and highly sensitive to measurement, making it a valuable progress-monitoring target.

    For Praxis, intervention targeting should not be abandoned based on the longitudinal findings reported here. The near-zero gain more likely reflects current measurement limitations than true insensitivity to OT services. As O.T. Wizard’s item bank is refined through Rasch calibration, this domain may demonstrate the sensitivity needed to capture incremental change that OT practitioners observe clinically.

    The domain-matched participation correlations support incorporating participation-focused goal setting alongside skill-based goals, particularly for children showing low engagement in specific domains. A child rated at Minimal Engagement in VMI activities is likely showing competence-related avoidance rather than willful noncompliance, and addressing the underlying skill deficit is the most direct route to improved participation in that domain.

    Study Limitations

    This study carries several important limitations. The sample is clinical and geographically restricted to North Carolina, and findings cannot be generalized to typically developing children or populations in other regions. The absence of a control group means observed gains cannot be causally attributed to OT intervention.

    The 8.7 percent natural growth correction applied throughout this report is a methodologically transparent and defensible estimate derived directly from the age progression of this cohort on this instrument. It assumes proportional developmental scaling across the score range, an assumption that Rasch calibration will allow us to test empirically. Future work with a typically developing comparison group will allow domain-specific, empirically derived growth expectations to replace this uniform estimate, strengthening the interpretive framework considerably.

    The regression-to-the-mean effect (r=-0.32) means children with lower starting scores showed larger gains on average. Visual Motor Integration, WAND-PreK Handwriting, Scissor Use, and Complex VMI all had the lowest Evaluation 1 scores and the largest gains above expected growth. The true intervention effect within these domains is likely substantial, but the proportion attributable to treatment versus regression toward the mean cannot be fully separated without a control group.

    Participation and Executive Functioning longitudinal findings are subject to a novelty effect interpretation that cannot be confirmed or ruled out with the current dataset. The possibility that Evaluation 1 ratings in these domains reflect novelty-driven behavior rather than true baseline functioning represents a meaningful threat to the internal validity of longitudinal comparisons in these areas specifically. This does not affect the skill-domain findings but warrants dedicated investigation in future work.

    The bilateral fine motor speed gap-widening finding is hypothesis-generating at n=39 and requires larger samples to reach significance. Bands G and J contained only two children each, precluding any band-specific inference for those age ranges. All scores remain pre-Rasch ordinal percentage values. Statistical analyses were generated with the assistance of an AI-based analytical tool and reviewed by the author for clinical and numerical consistency. Final responsibility for interpretation and reporting rests with the author.

    Continuing Research Needed

    Rasch calibration remains the highest research priority. Transforming ordinal raw scores into interval-level person measures will allow true scale-independent longitudinal comparison, validate the proportional scaling assumption underlying the natural growth correction, and identify items requiring revision. The O.T. Wizard National Try-Out Team initiative is designed to accelerate this process by expanding the sample sizes needed for stable item calibration across all domains and age bands.

    A comparison group of typically developing children will allow the 8.7 percent natural growth estimate to be validated empirically and replaced with domain-specific growth expectations calibrated against external developmental benchmarks. This will strengthen the natural growth framework applied in this report and make the gain-above-expected metric more precise for insurance and educational reporting contexts.

    The novelty effect hypothesis for participation and executive functioning requires prospective investigation. A study design that captures therapist-rated participation and work habits at multiple time points within the first year of services, alongside parent-reported and teacher-reported behavioral measures, would allow empirical testing of whether first-evaluation ratings systematically overestimate engagement and compliance relative to ratings collected after therapeutic relationship establishment. Confirming or disconfirming this hypothesis has implications for how O.T. Wizard instructs therapists to interpret and apply participation and EF data in clinical documentation.

    The domain-matched participation correlations reported here establish a foundation for construct validity research. Future work linking participation ratings to standardized adaptive behavior measures and teacher-rated classroom engagement would extend these findings and position the participation ratings as clinically actionable data rather than descriptive context.

    The bilateral fine motor speed findings warrant dedicated follow-up with larger samples. Testing whether inter-hand gap widening reaches significance at n=100 or greater, whether gap widening correlates with therapist-documented hand dominance establishment, and whether children who begin with gaps below two bubbles show different intervention trajectories would substantially extend the hand dominance research program established in the O.T. Wizard FMSAT study.

    Linking O.T. Wizard performance data to standardized criterion measures including the Beery VMI, BOT-2, and teacher-rated school readiness indicators would establish concurrent and predictive validity, and would allow the natural growth framework applied here to be calibrated against externally validated developmental benchmarks.

    Conclusion

    This analysis of 44 preschool-aged children with paired O.T. Wizard evaluations, examined through a transparent natural growth framework, provides the most complete domain-level longitudinal outcome picture generated by the platform to date. The expected natural growth rate of 8.7 percent over the 4.7-month study interval serves as a consistent interpretive reference throughout, allowing readers to distinguish maturation from intervention-attributable change across every domain reported.

    The findings are striking. Five of seven skill domains showed gains substantially exceeding the 8.7 percent natural growth threshold, with effect sizes ranging from medium to large. Visual Motor Integration showed gains 61.4 percent above expected. WAND-PreK Handwriting showed gains 61.4 percent above expected. Scissor Use showed gains 56.8 percent above expected. These are not marginal differences from expectation. They represent functional gains at rates six to eight times what chronological maturation alone would produce over the same interval.

    The domain-matched participation analysis adds a construct validity dimension to these findings. Children who perform better in a skill domain engage more fully during that domain’s assessment, with domain-matched correlations ranging from r=0.317 to r=0.497 and an overall composite skill-participation correlation of r=0.762. Participation and executive functioning longitudinal data are reframed here through the novelty effect hypothesis, which proposes that first-evaluation ratings in these behavioral domains may reflect situational compliance rather than authentic baseline functioning, with more valid observational data emerging after the therapeutic relationship is established.

    The FMSAT (“Bubble Popping”) extends these findings into hand dominance development, showing that the dominant hand improved significantly above the natural growth rate and that the inter-hand performance gap widened directionally, consistent with progressive hand specialization during OT services.

    These findings are preliminary. They require replication with larger samples, validated comparison conditions, and Rasch-calibrated measurement. What they establish is that structured domain-level digital assessment in pediatric OT can generate longitudinal outcome data that clearly and transparently distinguishes intervention-driven change from natural maturation. For a profession that has historically struggled to quantify its impact in terms that payers and educational systems recognize, that distinction is not a minor technical refinement. It is the foundation of evidence-based practice. O.T. Wizard is building that foundation, one evaluation at a time.

    Disclosures

    The author is the founder of O.T. Wizard and has a financial interest in the platform. All analyses were conducted on de-identified clinical data collected in routine practice.

    References

    Daly, C. J., Kelley, G. T., & Krauss, A. (2003). Relationship between visual-motor integration and handwriting skills of children in kindergarten: A modified replication study. American Journal of Occupational Therapy, 57(4), 459-462. https://doi.org/10.5014/ajot.57.4.459

    Duff, S. V., Chow, S. M., & Henderson, S. E. (2015). Developmental coordination disorder and its consequences for children. In A. F. Farrow & P. J. Tremblay (Eds.), Pediatric rehabilitation: Principles and practice (5th ed., pp. 189-215). Demos Medical Publishing.

    Zwicker, J. G., Missiuna, C., Harris, S. R., & Boyd, L. A. (2012). Developmental coordination disorder: A review and update. European Journal of Paediatric Neurology, 16(6), 573-581. https://doi.org/10.1016/j.ejpn.2012.05.005

    Polatajko, H. J., & Cantin, N. (2010). Exploring the effectiveness of occupational therapy interventions, other than the sensory integration approach, with children and adolescents experiencing difficulty processing and integrating sensory information. American Journal of Occupational Therapy, 64(3), 415-429. https://doi.org/10.5014/ajot.2010.09073

    About O.T. Wizard

    O.T. Wizard is a clinical intelligence system for pediatric occupational therapy professionals.  The platform evaluates performance in evaluations, forms, and daily treatment notes across twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, executive functioning, activities of daily living, and participation. OT Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update continuously as the clinical database expands.

    Data collection is ongoing. All data is de-identified in accordance with HIPAA and FERPA regulations.For information  about O.T. Wizard research or accessing the platform, visit otwizard.com

  • O.T. Wizard Spotlight: A COTA/L on Clinic Based Team working with Preschoolers

    O.T. Wizard Spotlight: A COTA/L on Clinic Based Team working with Preschoolers

    From time to time, we invite occupational therapy practitioners whose use of OT Wizard reflects strong clinical reasoning, real-world application, and thoughtful feedback to participate in a spotlight.

    User spotlight with Angie Bowman, COTA/L

    Can you tell us a little about your background as an OTP,  and your work setting(s) and typical patient ages ?

     I have always worked in pediatrics in various settings, primarily schools and clinics but some in-home early intervention as well for the past 30 years! 

    Do you have a favorite domain or type of challenge you enjoy treating most?


    Sensory processing is probably the most interesting area that OT’s address to me.  There is so much new research coming out that is validating what we as OTP’s have known for years and I feel like it is something commonly misunderstood.

    What does a “normal” day look like for you?

     I am working with only preschool children so really busy in the mornings and less so in the afternoons. I work until the little ones nap, then go home for paperwork and planning. 

    Would you rather make your own OT supplies or buy them?

    A combination of both I think although some of my best and most used toys are the homemade ones.

    If you could only bring 5 therapy items to a session, what would they be?

     1) Theraputty or some type of resistive media because I really like starting off the session with it.  Great for hand strengthening but also for providing the proprioceptive input they need before fine motor work. 

    2) A ball!  You can do a lot with them besides catch and toss!  Rolling up and down the wall or across tracks on the wall with tape are great motor planning and bilateral hand skills activities. 

    3) Crayons and paper, broken ones preferably taped to a wall. 

    4) A pickle picker!  This is my favorite fine motor tool 

    5) Scooter board, you can use them in any setting and do so much with them. 

    When you were told that your therapy group would be using OT Wizard, how did you initially feel about adding on a new process for evaluations?

    I was intrigued.  As a COTA/L I have assisted with a lot of standardized evaluations but I loved that this one is performed online and that it encompasses so many areas.

    Was there a moment where OT Wizard really “clicked” for you?

    I liked it right away but as I began to read the reports more and more, and I saw how good they were, it really sold me.

    Has OT Wizard changed your approach to evaluations or reports now, even in small ways?

    Absolutely.  I love generating a therapist report on my children.  It is generated under the abbreviated report with therapist focused as the target audience. It gives you a clinical snapshot with strengths and weaknesses and clinical takeaways.  THEN, evidenced-based interventions with different phases for implementation.  It gets you thinking about what you have done that has or has not worked and helps you to generate new ideas and/or affirm your thought process.  Sometimes we get into ruts when working with a child for a long time or that is particularly challenging in response to therapy and this is a great tool to use. 

    What is your favorite feature in OT Wizard?

    Definitely the report as stated above.

    What upcoming feature are you most excited about?

    Daily notes and progress reports for parents via the integrated messaging.

    What would you say to another OTP who’s feeling hesitant or skeptical about trying OT Wizard?

    The time you save in report writing alone is unbelievable and in a lot of ways has made me a better therapist! 

    What’s one thing you don’t miss doing since using OT Wizard?

    Extra paperwork

    Angie is on the Leadership Team at Learning Charms, which is the company that created O.T. Wizard. Angie is always excited to try new features, and makes great suggestions for changes and upcoming features. As Angie has been practicing for 30 years, she understands which features are most meaningful to OTP’s. Angie works in preschools that cater to at risk children and loves collaboration with teachers and getting hugs from her kiddos. For these reasons, this OT Wizard hat is well deserved.

  • Hand Dominance Development in Pre-School-Aged Children: Evidence from 459 Clinical Assessments

    Hand Dominance Development in Pre-School-Aged Children: Evidence from 459 Clinical Assessments

    INTRODUCTION

    Every occupational therapist has witnessed it: a child struggling to cut along a line, gripping their pencil awkwardly, or switching hands mid-task. While we understand hand dominance is foundational to fine motor skill development, we’ve historically relied more on clinical observation than quantitative data to assess when hand preference becomes truly established. 

    How much motor output difference between a child’s dominant and non-dominant hand is “normal”? Does this gap widen as children mature? And critically, can we measure hand dominance in a way that’s both clinically meaningful and statistically sound?

    These questions drove me to analyze data from O.T. Wizard’s Fine Motor and Accuracy Test (FMSAT) —a simple, 60-second paper-pencil task that measures fine motor speed, precision, and coordination. The results from 459 pediatric assessments offer compelling insights into how hand dominance manifests in functional performance, and what this means for OT practice.

    THE ASSESSMENT: FINE MOTOR SPEED AND ACCURACY TEST (FMSAT)

    The FMSAT task is straightforward: children use a sharpened pencil to “pop” (puncture) as many small circles “bubbles” as possible in 30 seconds, with the worksheet placed over craft foam. After completing one hand, they switch and repeat with the opposite hand. The task measures fine motor coordination and control, speed and precision, finger strength, and hand preference and dominance patterns.

    Children are allowed to choose which hand to use first—a critical design feature that lets us observe natural hand preference rather than imposing it. Some kids switch hands mid-task, for those , the therapist selects RL or LR (meaning started R and switched to L or vice versa).

    STUDY SAMPLE

    This analysis draws from 459 clinical assessments collected through O.T. Wizard during our soft launch phase. All data comes from pediatric occupational therapy evaluations conducted in North Carolina. “n” is the sample size.

    Age Distribution: 

    Band G (ages 36-47 months): n=53 (12%); 

    Band H (ages 48-53 months): n=185 (40%); 

    Band I (ages 54-59 months): n=221 (48%)

    Demographics: 

    57% male, 43% female. 

    Language: 90% English primary, 9% Spanish, 1% other. 

    Clinical status: 86% of children were ultimately recommended for OT services, with the primary diagnosis being Specific Developmental Disorder of Motor Function (F82, ICD-10). *The FMSAT was included in the evaluation; recommendation to OT services was based on the full evaluation results.

    Important Note: Due to limited sample size in Band G (only 19 children completed both hands of the assessment), primary statistical analyses focus on Bands H and I (n=406 total, n=348 with complete bilateral data). This is consistent with psychometric best practices, which recommend minimum sample sizes of 30-50 per category for stable estimates.

    RESEARCH HYPOTHESES

    Based on developmental occupational therapy theory and clinical observation, I hypothesized:

    Hypothesis 1: Children will preferentially choose their preferred or dominant hand first (Hand 1), resulting in significantly better performance with Hand 1 compared to Hand 2.

    Hypothesis 2: The performance gap between dominant and non-dominant hands should increase with age, as hand dominance becomes more established through the preschool years.

    Hypothesis 3: Children with consistent, established hand dominance (right or left) will show larger performance differences between hands compared to children with inconsistent or unclear hand preferences.

    BRIEF LITERATURE CONTEXT

    Hand dominance emerges gradually during early childhood, with most children showing consistent hand preference by age 3-4 years and full establishment by age 6. This developmental process is crucial for skill refinement—as one hand becomes increasingly specialized for fine motor tasks, the other develops complementary stabilization and assist functions.

    Research consistently links established hand dominance to improved academic skills, particularly handwriting fluency and speed. Conversely, delayed or inconsistent hand preference has been associated with developmental coordination difficulties and may signal underlying neurological immaturity.

    However, most studies examine hand preference categorically (right vs. left) rather than quantifying the degree of functional difference between hands. This gap motivated our analysis: can we measure not just which hand a child prefers, but how much better that hand actually performs?

    KEY FINDINGS

    Overall Performance Patterns (Bands H and I Combined)

    Across 348 children with complete bilateral data,

    Hand 1 (First Hand ) averaged 16.15 bubbles popped

    Hand 2 (Second Hand) averaged 11.74 bubbles popped, with a mean difference of 4.41 bubbles

    Notably, 79.3% of children performed better with Hand 1 than Hand 2.

    These findings strongly support Hypothesis 1—children naturally select their more proficient hand first, and this difference is both statistically significant and functionally meaningful.

    Hand Dominance Categories and Performance

    Children were categorized based on therapist-documented hand dominance:

    Right Hand Dominant (n=285, 82%): 

    Hand 1 averaged 16.65 bubbles, 

    Hand 2 averaged 11.90 bubbles, with a difference of 4.74 bubbles (28.5% advantage). 

    80.7% showed Hand 1 greater than Hand 2, and only 27.4% showed similar performance (less than or equal to 2 bubble difference).

    Left Hand Dominant (n=33, 9%): 

    Hand 1 averaged 14.67 bubbles,

    Hand 2 averaged 11.82 bubbles, with a difference of 2.85 bubbles (19.4% advantage). 

    75.8% showed Hand 1 greater than Hand 2, and 45.5% showed similar performance.

    No Clear Dominance (n=30, 9%): 

    Hand 1 averaged 12.20 bubbles, 

    Hand 2 averaged 11.47 bubbles, with a difference of 0.73 bubbles (6.0% advantage). 

    46.7% showed Hand 1 greater than Hand 2, with a median difference of 0 bubbles.

    Statistical Comparison: Children with consistent dominance (right or left) showed a 4.55 bubble difference. Children with no clear dominance showed a 0.73 bubble difference. This represents a 6.2-fold larger performance gap in children with established dominance, strongly confirming Hypothesis 3.

    Developmental Progression: Ages 4:00-4:11

    One of the most clinically significant findings emerged when examining how hand dominance evolves across just six months of development:

    Age Band H (4:00-4:05, n=157): Mean difference of 3.82 bubbles (25.4%), with 75.8% showing Hand 1 greater than Hand 2.

    Age Band I (4:06-4:11, n=191): Mean difference of 4.55 bubbles (26.8%), with 78.5% showing Hand 1 greater than Hand 2.

    Change from H to I: plus 0.73 bubbles (plus 1.4 percentage points)

    The pattern is clear: hand dominance differences increase with age, supporting Hypothesis 2. While the effect size is small, the consistent directional trend across this brief developmental window suggests progressive hand specialization.

    Critically, this increase is driven by differential skill development. The dominant hand (Hand 1) improved by 1.96 bubbles (H to I), while the non-dominant hand (Hand 2) improved by 1.22 bubbles. The dominant hand gained 0.73 more bubbles than the non-dominant hand.

    This pattern exemplifies the developmental principle of differentiation and specialization—the dominant hand isn’t just maintaining its advantage; it’s actively pulling further ahead as neuromotor pathways become increasingly refined.

    Correlation Analysis

    Hand 1 vs. Hand 2 Performance showed a strong positive correlation, indicating that while children vary in overall fine motor ability, bilateral coordination remains relatively consistent within individuals. Children with higher Hand 1 scores also tend to have higher Hand 2 scores, but the gap between them widens with established dominance.

    IMPLICATIONS FOR OT PRACTICE

    Assessment and Evaluation

    Rather than simply documenting “right” or “left” hand preference, OT practitioners may now quantify the degree of dominance. A 4-5 bubble difference may suggest well-established dominance, while differences under 2 bubbles may indicate emerging or inconsistent preference.

    Children with no clear hand dominance showed near-equal performance between hands (0.73 bubble difference), lower overall performance (12.20 vs. 16.65 bubbles for right-dominant peers), and only 46.7% consistency in hand choice. These children warrant closer developmental monitoring and may benefit from interventions targeting hand specialization alongside general fine motor development. Of course, this would not be the case in younger children than our sample when laterality is still in development.

    For children ages 4:00-4:11, expect a 3-5 bubble advantage for the dominant hand, 75-80% consistency in using the dominant hand for skilled tasks, and progressive widening of the performance gap over 6-month intervals.

    Intervention Planning

    Children showing less than 2 bubbles difference and inconsistent hand use at age 4 or older may benefit from activities that encourage hand preference establishment, not just bilateral coordination practice.

    The strong correlation between Hand 1 and Hand 2 performance suggests that improving overall fine motor skills benefits both hands. However, children with established dominance show specialized development—interventions should support both general skill building AND hand-specific refinement.

    Documentation for Insurance and Educational Teams

    Performance data showing significantly reduced fine motor speed and precision provides concrete evidence for accommodations such as extended time for written work, reduced copying requirements, and assistive technology considerations.

    For insurance authorization, demonstrating that a child’s hand dominance pattern deviates from age-expected norms (e.g., less than 2 bubble difference at age 4 or older) may help support medical necessity for skilled OT intervention.

    Serial assessments can document functional improvement in hand dominance establishment, not just overall fine motor gains.

    IDENTIFYING THE WRITING HAND IN OLDER CHILDREN

    This simple assessment becomes invaluable when working with older children who appear ambidextrous or who have been switching hands for years. Many parents proudly report “my child is ambidextrous!” when their 7-year-old writes with both hands. However, upon closer examination, the handwriting is often slow, effortful, and illegible with both hands.

    The clinical dilemma: how do you choose which hand to train for handwriting when a child has been switching for years?

    Handwriting is motor memory. Every time you write the letter “a,” your brain strengthens a specific motor pattern in the hand you’re using. When a child switches hands, they’re essentially learning two different motor programs for the same letter—and neither program gets enough practice to become automatic.

    Motor learning research is clear: inconsistency prevents automaticity. A child who writes with both hands is essentially a beginner with each hand, never progressing to the automatic stage where handwriting becomes effortless.

    The FMSAT provides objective data in just 60 seconds. A 3-bubble or greater difference may suggest a neurologically preferred hand—even if the child has been switching hands for years due to habit, environmental factors, or well-meaning adults who thought “using both hands” was beneficial. For a four year old, equal performance (less than 2-bubble difference) tells you either hand could work, so consider other factors like which side shows better pencil grip, less fatigue, or more consistent letter formation.

    Once you’ve identified the better hand based on speed and accuracy data, you’re committed. I explain to parents and teachers: “We’re going to consistently use the right hand (or left hand) for ALL writing/drawing/coloring tasks from now on. This gives the brain a chance to build the motor memory it needs for fluent handwriting. ”

    This is especially critical for children with developmental coordination difficulties. Their motor learning already takes longer than typical—asking them to learn two sets of motor patterns for every letter makes an already difficult task nearly impossible.

    Strategies: I’ve found that putting a “stamp” on the dominant hand or a soft bracelet on the writing hand helps a child remember which hand to use so that verbal cues aren’t as necessary. To explain to a preschooler, I tell them “This is your boss hand. When you color or draw or write, this one holds the pencil because its the boss. Your other hand is the helper. “

    STUDY LIMITATIONS

    Several limitations should be considered when interpreting these findings:

    Geographic and Cultural Homogeneity: All data comes from North Carolina, with 90% English-speaking children. Hand dominance patterns may vary across cultures with different tool use expectations or writing systems.

    Clinical Population: About 86% of the children were recommended for OT services, meaning this sample represents children with developmental concerns rather than typically developing peers. The hand dominance differences observed may differ from the general population.

    Cross-Sectional Design: We examined different children at different ages rather than following the same children over time. Longitudinal studies would provide stronger evidence of developmental trajectories.

    Age Band G Underpowered: Only 19 children in Band G (ages 3:06-3:11) completed both hands, limiting our ability to examine younger developmental patterns.

    Single Assessment Task: While FMSAT measures important fine motor components, hand dominance manifests across many functional tasks. Triangulating with other assessments would strengthen findings.

    CONTINUING RESEARCH NEEDED

    Expanded Age Bands: Critical questions remain about hand preference emergence in younger children (ages 2:00-3:05) and whether the hand dominance gap continues widening through elementary years (ages 5:00-7:11) or plateaus. Establishing adult norms would provide developmental endpoints for clinical interpretation.

    Longitudinal Studies: Following individual children over 12-24 months would reveal individual variation in dominance establishment timelines, whether intervention can accelerate hand preference development, and predictive validity: do FMSAT differences at age 4 predict handwriting fluency at age 6?

    Typically Developing Comparison: Recruiting a non-clinical sample would establish true normative data, determine if clinical populations show delayed or atypical dominance patterns, and support differential diagnosis and intervention planning.

    Academic Outcome Correlations: Linking FMSAT performance to standardized handwriting assessments, teacher-reported classroom performance, and academic achievement in writing-heavy subjects.

    Expanded Diversity: Collecting data across multiple geographic regions, diverse cultural backgrounds, various language groups, and different diagnostic categories.

    CONCLUSION

    This analysis of 459 clinical assessments provides compelling evidence that hand dominance can be quantified in a simple, time-efficient assessment that holds clinical meaning for OT practitioners. In our data, the FMSAT task successfully discriminates between children with established versus unclear hand preferences, captures expected developmental progression across the preschool years, and generates data precise enough for goal-setting, progress monitoring, and outcomes research.

    Three key findings stand out:

    First, children with consistent hand dominance show 6.2 times larger performance differences between hands compared to children with unclear preferences (4.55 vs. 0.73 bubbles).

    Second, hand dominance strengthens measurably over just six months in the preschool period (ages 4:00-4:11), with the dominant hand pulling 0.73 bubbles further ahead.

    Third, the majority of 4-year-olds (91%) demonstrate clear hand dominance, making this a critical developmental window for identification and intervention.

    As O.T. Wizard continues collecting data and expanding age bands, we’re building an evidence base that moves pediatric OT practice from subjective observation to quantifiable, research-informed assessment. Hand dominance isn’t just a checkbox on an evaluation form—it’s a measurable developmental milestone with implications for every fine motor task a child will encounter in school and daily life.

    For the 9% of children in our sample who showed no clear hand dominance at age 4 or older, this data validates what we see clinically: these children need support. Not just general fine motor therapy, but targeted intervention to establish the hand specialization that underlies skilled tool use, handwriting, and bilateral coordination.

    Most importantly for clinical practice, this 60-second assessment provides objective data to confidently identify the neurologically preferred hand in older children who have been switching—ending the ambidexterity myth and establishing the consistency needed for motor learning to progress to automaticity.

    REFERENCES

    Kushki, A., Chau, T., & Anagnostou, E. (2011). Handwriting difficulties in children with autism spectrum disorders: A scoping review. Journal of Autism and Developmental Disorders, 41(12), 1706-1716.

    Marschik, P. B., Einspieler, C., Guzzetta, A., et al. (2008). Behavioral patterns of exploration and approach in children with and without developmental delay. Developmental Medicine & Child Neurology, 50(9), 664-669.

    Sacrey, L. A., Arnold, B., Whishaw, I. Q., & Gonzalez, C. L. (2012). Precocious hand use preference in reach-to-eat behavior versus manual construction in 1- to 5-year-old children. Developmental Psychobiology, 55(8), 902-911.

    Scharoun, S. M., & Bryden, P. J. (2014). Hand preference, performance abilities, and hand selection in children. Frontiers in Psychology, 5, 82.

    About O.T. Wizard

    Data for this analysis was collected through OT Wizard, a clinical intelligence system for pediatric occupational therapy assessment. The platform evaluates performance across up to twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, visual motor integration, executive functioning, activities of daily living, and participation. OT Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update continuously as the clinical database expands.

    Unlike traditional checklist-based assessments, OT Wizard converts all observations to continuous metrics that enable progress tracking, cross-domain comparison, and comprehensive reporting. The platform captures all six factors identified in this research as predictive of handwriting success: fine motor skills, visual perception (with subdomain specificity), praxis, cooperation, attention, and task participation. Behavioral regulation is assessed within the context of actual task performance rather than as an isolated rating, providing clinically relevant data about how attention and cooperation affect functional skill demonstration.

    For handwriting readiness assessment specifically, OT Wizard provides quantified performance across visual discrimination, visual-motor integration, fine motor control, motor planning, and behavioral engagement during writing tasks. This comprehensive approach addresses the multifactorial nature of handwriting development identified in this research. As the platform undergoes Rasch analysis validation and accumulates longitudinal outcome data, it will establish whether comprehensive baseline assessment across all six predictors improves identification of children at risk for handwriting difficulty and informs more effective intervention planning.

    OT Wizard is committed to advancing the occupational therapy profession by collecting de-identified clinical data from real therapist users, building the largest developmental database in pediatric occupational therapy history. This continuous data collection enables research on developmental trends, intervention effectiveness, and response to intervention patterns that elevate practice from perception-based to data-driven decision making and strengthen the evidence base for the entire profession

    For OT professionals interested in data-driven assessment tools, visit otwizard.com to learn more about evidence-based pediatric evaluation.

    All data de-identified in accordance with HIPAA regulations.

  • The Letter Identification Prerequisite Myth: What 471 Children Taught Us About Learning to Write

    The Letter Identification Prerequisite Myth: What 471 Children Taught Us About Learning to Write

    Letter Identification and Copying Skills

    A Data-Driven Investigation into How Letter Recognition and Copying Skills Actually Develop


    The Question That Started It All

    For decades, occupational therapists and educators have debated a fundamental question: Do children need to identify letters before they can copy them?

    Traditional developmental hierarchies suggest a clear sequence: first comes letter recognition (knowing that the shape “A” is called “A”), then comes the ability to reproduce that letter through writing or copying. This assumption underlies countless kindergarten readiness checklists and early intervention programs.

    But what if this assumption is wrong?


    What the Research Literature Says

    Recent research paints an interesting picture about the relationship between letter knowledge and handwriting. Multiple studies from 2005-2025 have established that handwriting practice enhances letter recognition. Most occupational therapists agree with that. Children who learn letters through handwriting (copying or tracing) demonstrate better letter identification on post-tests compared to those who learn by typing (Neuroscience NewsScienceDirect), and fMRI studies show that handwriting activates visual letter-processing regions in the brain more effectively than other methods(PubMed Central).

    Educational researchers recommend that when students are learning letter identification, they should simultaneously engage in learning how to form the letter ( Uiowa). Writing readiness prerequisites identified in the literature include both alphabet letter recognition and basic stroke formation( Illinois) – presented as co-occurring skills rather than sequential steps.

    However, here’s what’s largely missing from the research: Does letter identification actually predict copying ability? Most studies examine whether handwriting improves recognition (it does), but few investigate whether recognition is necessary for copying success.


    Our Study: 471 Children, 5 Copying Tasks, Clear Answers

    We analyzed assessment pre-Rasch data from 471 preschool children (ages 3-5) of a clinical sample who completed a letter copying assessment. Each child was asked to:

    1. Identify 10 uppercase letters  (L, F, R, S, X, N, C, K, V, A) 
    2. Copy 10 letters (L, F, R, S, X, N, C, K, V, A) from a visual model

    For each copied letter, we scored (binary, yes/no):

    • Formation: Did they use a mostly top-to-bottom stroke?
    • Legibility: Was the copied letter recognizable as such?
    • Directionality: Was the letter copied with correct spatial orientation?

    We then calculated correlations – a statistical measure of how strongly two skills relate to each other. A correlation of 1.0 means they’re perfectly linked; 0.0 means they’re completely independent; and anything below 0.3 is considered very weak.


    The Ah-Ha Moment #1: Letter ID Barely Predicts Copying

    Here’s what shocked us: The correlation between letter identification and copying skills ranged from 0.05 to 0.29 across all age groups and all copying tasks.

    What does this mean ?

    At r = 0.29, r² = 0.08, Letter Identification explains about 8% of variance in copying performance. That’s like saying knowing someone’s height tells you almost nothing about their shoe size – the two things just aren’t that related. The graphic below shows how letter identification skills correlate to the other variables and within each age band.

    r value rangerelationship
    0.0 – 0.1No meaningful relationship
    0.1 – 0.3Weak relationship
    0.3 – 0.5Moderate relationship
    0.5 – 0.7Strong relationship
    0.7 – 1.0Very strong relationship
    Age Group (years)ID →Top to Bottom FormationID → LegibilityID → Directionality
    Youngest (3:5-4:0)r = -0.05r = 0.20r = 0.20
    Middle (4:0-4:5)r = 0.28r = 0.19r = 0.14
    Oldest (4:5-5:0)r = 0.23r = 0.29r = 0.17

    Translation: Whether a child can identify a letter tells you almost nothing about whether they can copy it successfully. These are developing as largely independent skills.


    The Ah-Ha Moment #2: The “Letter A Paradox”

    The clearest example came from the letter A in our youngest group:

    • 36.4% could identify the letter A (highest recognition rate and likely because its at the beginning of the alphabet)
    • Only 5.6% could copy it with correct formation 
    • Only 3.7% could copy it legibly

    That’s a 30-point gap. Kids knew it was an A, but couldn’t reproduce the complex diagonal strokes needed to draw one.

    Meanwhile, for the letter L:

    • 23.6% could identify L
    • 20.4% could copy it with correct formation

    Only a 3-point gap – nearly equal performance.

    Why? Letter A requires two diagonal strokes meeting at a precise point with a horizontal crossbar (developmentally complex) -therefore recognition and production appear to rely on different skill demands. Letter L is just a vertical line with a horizontal base – simple enough to copy through visual matching alone, even without knowing it’s called an “L.”


    The Ah-Ha Moment #3: By Age Band 4:5-5:0, Copying Can Be Easier Than Identifying

    Here’s where it gets really interesting. By the oldest age group, we started seeing positive gaps – children who could copy letters they couldn’t identify:

    • Letter L: 44.5% could identify it, but 63.4% could copy it correctly (+19 points)
    • Letter V: 24.7% could identify it, but 44.2% could copy it correctly (+20 points)
    • Directionality overall: Children performed 3.1 percentage points better at copying with correct spatial orientation than at identifying letters

    This pattern defies the traditional hierarchy. Children were using visual matching – copying the shapes they saw – without needing to know the letter names.


    What Does Predict Copying Success?

    If letter identification doesn’t predict copying, what does?

    We found that top-to-bottom formation strongly predicts legibility (correlation of 0.61-0.89, explaining 37-79% of variance).

    But top to bottom formation itself correlates most strongly with:

    • Visual-Motor Integration: r = 0.76 (explains 58% of variance)
    • Visual Perception: r = 0.42
    • Fine Motor Skills: r = 0.44

    But NOT with general motor planning (Praxis): r = 0.24 (only 5.7% variance)

    The takeaway: Letter copying is primarily a visual-motor integration task – the ability to coordinate what you see with what your hand does. It’s not about general motor planning, and it’s certainly not dependent on knowing letter names.


    How This Compares to Existing Research

    Our findings complement rather than contradict current research:

    Current research says: “Handwriting practice improves letter recognition” ✓
    Our data adds: “But copying ability develops independently from letter knowledge”

    Current research says: “Teach letter ID and handwriting together” ✓
    Our data explains WHY: They support each other but develop through different pathways – one verbal-visual (naming), one visual-motor (copying)

    The key distinction: Most research examines writing from memory (where letter knowledge clearly helps), while our study examined copying from a visual model (where visual-motor integration dominates).


    Implications for Occupational Therapy Practice

    1. Don’t Wait for Letter Mastery to Start Copying Practice

    The weak correlations (r < 0.3) mean you can’t predict copying readiness from letter identification scores. A child who struggles to name letters might still succeed at copying them.

    Action: Include copying tasks in early intervention even when letter knowledge is limited. You’re building visual-motor skills that develop on a parallel track.

    2. Motor Control Develops Independently

    Our data showed boundary control (staying within lines) actually performed better than letter identification at all ages – evidence that fine motor control for pencil management is a separate developmental pathway.

    Action: Work on “staying in the lines” without requiring letter identification first. These are independent skills.

    3. Use Simple Geometric Letters as Confidence Builders

    Letters L and F showed the smallest gaps between ID and copying (-3 points) because their simple vertical/horizontal geometry enables visual matching. This is consistent with programs such as Handwriting Without Tears, as they order uppercase letter instruction by geometric easy to hard. 

    Action: Start with L, F, T, I for early success. Save A, K, R, S (complex diagonal/curve letters) for later, regardless of which letters the child can name.

    4. Target the Critical Window: Middle Preschool

    Our biggest developmental gains happened between the youngest and middle age groups (Ages G→H), not between middle and oldest (H→I):

    • Formation gap improved 4.6 points (G→H) vs. 4.8 points (H→I)
    • Legibility gap improved 6.1 points (G→H) vs. 5.1 points (H→I)
    • Directionality gap improved 8.5 points (G→H) vs. 8.4 points (H→I)

    Action: Middle preschool (roughly ages 4-5) is your prime intervention window for copying skills. Don’t wait until kindergarten.


    Implications for Education and Parents

    1. Parallel Practice, Not Sequential Prerequisites

    Old thinking: “My child needs to know their ABCs before we practice writing”
    Data-driven approach: “We’ll teach letter names AND practice copying simultaneously – they support each other through different pathways”

    2. Copying Success ≠ Letter Knowledge

    Don’t assume that because your child can copy a letter, they know what it’s called. And don’t assume that because they can name it, they can reproduce it.

    Letter A showed this clearly: High recognition, low reproduction. These are different skills.

    3. The “Legibility Integration Challenge”

    Legibility showed the largest negative gaps at all ages (copying legibly is 6-18 points harder than identifying letters).

    Why? Legible copying requires simultaneous integration of:

    • Visual perception (seeing the target)
    • Motor planning (sequencing strokes)
    • Motor execution (hand control)
    • Visual-motor feedback (monitoring while writing)

    Parenting insight: Be patient with legibility. It’s the most complex integration task and develops last. Celebrate formation accuracy and directionality before expecting neat, legible letters.

    4. Use Letter Copying as a Window into Visual-Motor Skills

    Since copying correlates strongly with visual-motor integration (r = 0.76) but weakly with letter knowledge (r < 0.3), copying tasks reveal visual-motor development more than academic readiness.

    For educators: A child who struggles to copy letters might need visual-motor support, not more letter drills.


    The Bottom Line

    After analyzing 471 children’s performance on letter identification and copying tasks, the data tells a clear story: Letter identification is not a meaningful prerequisite for letter copying.

    These skills develop as parallel pathways:

    • Letter Identification pathway: Verbal-visual learning (naming, recognizing)
    • Letter Copying pathway: Visual-motor integration (seeing, matching, executing)

    Both are valuable. Both support eventual handwriting fluency. But one doesn’t have to come before the other.

    For practitioners and parents: Stop waiting. Introduce copying practice early. Use simple geometric letters (L, F, E, D, P) for confidence. Target middle preschool for maximum gains. And remember – a child who can name every letter might still struggle to draw an A, while a child who can’t name any letters might successfully copy an L.

    The question isn’t “ID before copying?” but rather “How do we support BOTH simultaneously to maximize letter learning through every available pathway?”


    About This Research

    This analysis drew from assessment data collected through OT Wizard (otwizard.com), a pediatric clinical intelligence tool. The study included 471 evaluations across three age bands (Ages G, H, I, representing 30-36 months through 60-72 months). Assessment included the preschool version of Magic WAND™ , a letter copying task with 10 uppercase letters (L, F, R, S, X, N, C, K, V, A) scored for formation accuracy, legibility, and directionality. Statistical analyses examined correlations between letter identification scores and multiple copying performance measures.

    About O.T. Wizard

    Data for this analysis was collected through OT Wizard, a clinical intelligence system for pediatric occupational therapy assessment. The platform evaluates performance across up to twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, visual motor integration, executive functioning, activities of daily living, and participation. OT Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update continuously as the clinical database expands.

    Unlike traditional checklist-based assessments, OT Wizard converts all observations to continuous metrics that enable progress tracking, cross-domain comparison, and comprehensive reporting. The platform captures all six factors identified in this research as predictive of handwriting success: fine motor skills, visual perception (with subdomain specificity), praxis, cooperation, attention, and task participation. Behavioral regulation is assessed within the context of actual task performance rather than as an isolated rating, providing clinically relevant data about how attention and cooperation affect functional skill demonstration.

    For handwriting readiness assessment specifically, OT Wizard provides quantified performance across visual discrimination, visual-motor integration, fine motor control, motor planning, and behavioral engagement during writing tasks. This comprehensive approach addresses the multifactorial nature of handwriting development identified in this research. As the platform undergoes Rasch analysis validation and accumulates longitudinal outcome data, it will establish whether comprehensive baseline assessment across all six predictors improves identification of children at risk for handwriting difficulty and informs more effective intervention planning.

    OT Wizard is committed to advancing the occupational therapy profession by collecting de-identified clinical data from real therapist users, building the largest developmental database in pediatric occupational therapy history. This continuous data collection enables research on developmental trends, intervention effectiveness, and response to intervention patterns that elevate practice from perception-based to data-driven decision making and strengthen the evidence base for the entire profession

    For OT professionals interested in data-driven assessment tools, visit otwizard.com to learn more about evidence-based pediatric evaluation.

  • How to Write OT Evaluation Reports That Insurance Actually Approves

    How to Write OT Evaluation Reports That Insurance Actually Approves

    You spent 90 minutes conducting a thorough pediatric occupational therapy evaluation. Another hour and a half writing a detailed report. You submitted it to insurance with confidence. Then the denial letter arrives: “Medical necessity not established.”

    Sound familiar? You’re not alone. Insurance denials for occupational therapy evaluations are frustrating, time-consuming, and costly. But here’s the good news: most denials happen because of how the report is written, not whether the child actually needs services.

    Let’s fix that.

    The Insurance Approval Formula: Medical Necessity + Functional Impact + Skilled Service

    Insurance companies don’t deny services because they don’t believe children need help. They deny because the documentation doesn’t prove three critical elements:

    1. Medical Necessity: A documented diagnosis or condition that requires intervention
    2. Functional Impact: Clear evidence that the condition limits daily functioning
    3. Skilled Service: Proof that an occupational therapist’s expertise is required (not just supervision or general instruction)

    Your evaluation report must explicitly address all three. If even one is missing or unclear, expect a denial.

    What Insurance Reviewers Actually Read (And What They Skip)

    Here’s a secret: the person reviewing your report spends about 90 seconds on it. They’re not reading every word. They’re scanning for specific elements. Most likely they are being read by their AI Bot.

    What they look for:

    • Diagnosis codes (ICD-10)
    • Functional limitations stated explicitly
    • Objective test scores and measurements
    • Clear statement of skilled OT intervention need
    • Specific safety concerns (if applicable)

    What they skip:

    • Long narrative descriptions
    • Clinical observations without data
    • Educational jargon (IEP goals, classroom performance)
    • Developmental history (unless directly relevant)

    The takeaway: Front-load your report with the information they need. Don’t bury medical necessity in paragraph seven.

    The 7 Elements Every Insurance-Approved Report Contains

    1. Clear Diagnosis at the Top

    Wrong:
    “Johnny is a 5-year-old male referred for fine motor concerns.”

    Right:
    “Johnny is a 5-year-old male with a diagnosis of Developmental Coordination Disorder (ICD-10: F82) referred for occupational therapy evaluation secondary to significant fine motor and visual-motor integration deficits impacting activities of daily living.”

    Notice the difference? The second version includes diagnosis code, specific deficit areas, and functional impact in the first sentence.

    2. Functional Limitations Stated Explicitly

    Insurance doesn’t care that a child scores in the 5th percentile on the Beery VMI. They care that this score means the child cannot complete age appropriate functional activities, such as independently buttoning their shirt, writing their name legibly, or using utensils safely.

    For every test score, include the “so what” statement:

    Test Result: Visual perception skills measured at 2 standard deviations below age expectations on TVPS-4 (standard score: 70).

    Functional Impact: This significant deficit prevents Johnny from independently locating items in his backpack, finding his desk in the classroom, and distinguishing similar letters (b/d, p/q) during early literacy tasks. Parent reports Johnny requires maximum assistance with dressing due to inability to orient clothing correctly.

    3. Objective Measurements and Standardized Scores

    Clinical observations alone don’t prove medical necessity. You need numbers.

    Include :

    • Standardized test scores with percentiles or standard scores (like Peabody, Beery VMI) or Rasch-calibrated /Criterion referenced assessments (like PEDI-CAT, OT Wizard, or HELP)
    • Timed performance measures (e.g., “completed pegboard task in 145 seconds; age expectation is 45 seconds”)
    • Quantifiable observations (e.g., “grasped pencil in fisted grasp 100% of observed writing attempts”)
    • Measurable functional deficits (e.g., “required 4 verbal cues and 2 physical assists to don shirt”)

    Research shows: Reports with comprehensive domain coverage across 8 areas (ADL, Executive Functioning, Fine Motor, Gross Motor, Visual Perception, Visual Motor Integration, Praxis, and Participation) have significantly higher approval rates because they provide objective evidence across multiple functional areas.

    4. Medical Necessity Language (Not Educational Language)

    If you primarily treat in schools, but are a medical based provider (meaning you aren’t an IEP provider), this is critical. Insurance reviewers don’t understand educational terminology.

    Educational Language (Don’t Use):
    “Johnny requires OT services to access his educational curriculum and participate in classroom activities per his IEP.”

    Medical Necessity Language (Use This):
    “Johnny requires skilled occupational therapy intervention to develop functional grasp patterns, visual-motor integration skills, and bilateral coordination necessary for age-appropriate self-care tasks including dressing, feeding, and personal hygiene.”

    Key Differences:

    EducationalMedical
    StudentPatient
    Classroom participationFunctional independence
    IEP goalsTreatment goals
    Educational benefitMedical necessity
    School activitiesActivities of daily living

    5. Safety Concerns (When Present)

    Safety issues fast-track approvals. If present, state them clearly.

    Examples:

    “Child demonstrates impulsive behavior and poor body awareness, resulting in 3 falls from playground equipment in past month per parent report. Requires skilled OT intervention to develop safety awareness and motor planning.”

    “Significant oral-motor deficits result in choking incidents during meals 2-3 times per week. Skilled feeding therapy required to establish safe swallowing patterns.”

    “Decreased proximal stability and postural control result in frequent loss of balance during mobility, with 2 documented injuries requiring medical attention in past 6 months.”

    6. Why Skilled OT is Required (Not Just Caregiver Training)

    Insurance will deny if they think a parent or teacher could provide the same intervention. You must prove why your clinical expertise is necessary.

    Not Skilled:
    “Child will benefit from practice with buttoning and zipping.”

    Skilled Service:
    “Child requires skilled occupational therapy to analyze specific motor planning deficits preventing successful fastener manipulation, develop individualized strategies to compensate for bilateral coordination limitations, and systematically grade activity complexity while addressing underlying sensory processing difficulties that interfere with tactile discrimination necessary for fastener manipulation.”

    See the difference? The second version demonstrates clinical reasoning, assessment expertise, and therapeutic skill that cannot be provided by non-therapists.

    7. Concrete Frequency and Duration Recommendations

    Vague recommendations get denied. Be specific.

    Too Vague:
    “Recommend outpatient OT services.”

    Specific and Justified:
    “Patient requires skilled occupational therapy 2x/week for 8 weeks (16 sessions) to address bilateral coordination deficits, visual-motor integration delays, and ADL skill development. Frequency based on severity of deficits (2+ standard deviations below age expectations across 4 domains) and need for motor learning repetition to establish new movement patterns. Re-evaluation recommended after 8-week intervention period to assess progress and determine ongoing needs.”

    Common Denial Reasons and How to Avoid Them

    Denial Reason #1: “Diagnosis not covered”

    Prevention: Check the insurance company’s covered diagnosis list before evaluating. If the primary diagnosis isn’t covered, lead with a secondary diagnosis that is covered but still supports the need for OT.

    Example: Autism (F84.0) might not be covered for outpatient OT, but Developmental Coordination Disorder (F82) or Sensory Processing Disorder coded as Other Specified Developmental Disorders (F88) often are.

    Denial Reason #2: “Educational, not medical”

    Prevention: Even if you’re a school-based therapist, emphasize ADL and home function impacts, not just classroom performance.

    Include:

    • Dressing difficulties
    • Feeding/utensil use challenges
    • Hygiene and self-care limitations
    • Safety concerns at home
    • Community participation barriers

    Denial Reason #3: “Not medically necessary”

    Prevention: State explicitly in your report: “Skilled occupational therapy is medically necessary to address [diagnosis] which significantly impacts patient’s ability to [specific functional tasks], resulting in dependence on caregivers for age-appropriate self-care and safety concerns during daily activities.”

    Denial Reason #4: “Insufficient objective data”

    Prevention: Use standardized assessments. Clinical observations alone aren’t enough. Data from 404 evaluations shows that assessments with zero missing data and comprehensive domain coverage provide the objective evidence insurance requires.

    The Report Structure Insurance Prefers

    Section 1: Demographics and Diagnosis (Top of Page)

    • Name, DOB, date of evaluation
    • Primary diagnosis with ICD-10 code
    • Referring physician

    Section 2: Medical Necessity Statement (First Paragraph) One clear paragraph stating diagnosis, functional limitations, and why skilled OT is required.

    Section 3: Assessment Results

    • Standardized test scores
    • Functional performance observations
    • Quantifiable data
    • Each with functional impact statement

    Section 4: Clinical Impressions

    • Summary of findings
    • How deficits impact daily function
    • Safety concerns (if applicable)

    Section 5: Recommendations

    • Specific frequency (2x/week)
    • Specific duration (8 weeks)
    • Justification for both
    • Explicit medical necessity statement

    Keep it concise: 2-3 pages maximum. Remember, they spend 90 seconds reading it.

    Real Example: Before and After

    Before (Gets Denied):

    “Johnny is a pleasant 5-year-old boy who was referred for OT evaluation. He has difficulty with handwriting and gets frustrated during fine motor tasks at school. During testing, Johnny had trouble copying shapes and his pencil grasp looked immature. He would benefit from OT to work on these skills. Recommend weekly OT.”

    Problems: No diagnosis code, no standardized scores, educational focus, vague recommendations, no medical necessity statement.

    After (Gets Approved):

    “Johnny is a 5-year-old male with Developmental Coordination Disorder (F82) referred for occupational therapy evaluation secondary to significant visual-motor and fine motor deficits impacting activities of daily living and self-care independence.

    Assessment Results:

    • Beery VMI: Standard Score 75 (5th percentile, 1.67 SD below mean)
    • O.T. Wizard: Composite 550/1000, ADL 50/100, Fine Motor 72/100, Gross Motor 42/100, Sequencing Praxis 27/100, Visual Motor Integration 72/100
    • Functional grasp assessment: Fisted grasp pattern 90% of observed attempts

    Functional Impact: Visual-motor integration and fine motor deficits prevent Johnny from independently managing fasteners (buttons, zippers, snaps), requiring maximum assistance for dressing. Unable to use utensils safely, resulting in frequent spills and parent reports of choking incidents 1-2x weekly. Cannot complete age-appropriate self-care tasks including tooth brushing and hair combing without hand-over-hand assistance.

    Medical Necessity: Johnny requires skilled occupational therapy to develop functional grasp patterns, bilateral coordination, sequencing praxis, gross motor, and visual-motor integration skills necessary for age-appropriate self-care independence. Deficits 2 standard deviations below age expectations indicate significant impairment requiring therapeutic intervention. Safety concerns related to feeding and frequent falls during mobility necessitate skilled assessment and intervention.

    Recommendations: Skilled occupational therapy 2x/week for 12 weeks to address bilateral coordination, visual-motor integration, and ADL skill development. Frequency based on severity of deficits and need for repetition to establish motor learning. Re-evaluation after 12 weeks to assess progress.”

    Why it works: Diagnosis code in first sentence, standardized scores with functional impact, medical necessity explicitly stated, safety concerns noted, specific recommendations with justification.

    Special Considerations for Different Settings

    School-Based Therapists Seeking Medical Insurance Coverage

    You can write reports that work for both IEP teams and insurance, but you need two versions:

    IEP Version: Focus on educational impact and access to curriculum
    Insurance Version: Same data, different framing focused on ADL and medical necessity

    Pro Tip: Complete your evaluation once, but generate two reports with different emphasis. Your assessment data doesn’t change, just how you present it.

    Outpatient Clinic Therapists

    You have an advantage because you’re already documenting medical necessity. Just ensure you’re:

    • Using covered diagnosis codes
    • Quantifying functional limitations
    • Stating skilled service needs explicitly
    • Providing specific frequency/duration with rationale

    Early Intervention Providers

    Insurance approval for 0-3 age range requires extra emphasis on:

    • Developmental delay severity (how far behind age expectations)
    • Impact on parent-child interaction
    • Safety concerns
    • Risk of further delay without intervention

    The Bottom Line

    Insurance approval isn’t about luck. It’s about documentation. Every denied evaluation report is missing at least one of these elements:

    ✓ Diagnosis code in first paragraph
    ✓ Standardized assessment scores
    ✓ Functional impact statements for every deficit area
    ✓ Medical necessity language (not educational)
    ✓ Explicit statement of why skilled OT is required
    ✓ Specific frequency and duration with justification
    ✓ Safety concerns (when applicable)

    Master these seven elements, and your approval rate will skyrocket.

    Stop spending hours appealing denials. Write it right the first time.

    Streamline Insurance-Compliant Documentation

    Writing insurance-approved reports doesn’t have to take hours. OT Wizard generates comprehensive evaluation reports with all required elements automatically included: diagnosis codes, standardized scores across 8 domains, functional impact statements, and medical necessity /educational eligibility. Choose medical or educational report tone with one click. Stop rewriting reports for insurance appeals.

    Learn more about automated insurance-compliant reporting →

  • What Really Predicts Handwriting Success

    What Really Predicts Handwriting Success

    THE CLINICAL PUZZLE

    Every pediatric occupational therapist has encountered this scenario: A 4-year-old with excellent fine motor skills, good visual perception scores, and established hand dominance still cannot write letters legibly. Meanwhile, another child with weaker motor skills and inconsistent grip produces surprisingly readable work.

    What makes the difference?

    New data from 185 preschool-age children reveals why handwriting success is so unpredictable and why our traditional assessment approaches may be missing critical pieces of the puzzle.

    CURRENT HANDWRITING ASSESSMENT PRACTICES

    Occupational therapists typically evaluate handwriting readiness through standardized assessments focusing on visual-motor integration and fine motor skills:

    Beery VMI (Visual-Motor Integration), 6th Edition measures the ability to copy geometric forms of increasing complexity. Children progress from simple lines to complex shapes, with performance compared to age-based norms. The assessment assumes that shape copying ability predicts letter formation success.

    PDMS-3 (Peabody Developmental Motor Scales, 3rd Edition) assesses fine and gross motor development through grasping and visual-motor integration subtests. The fine motor composite includes tasks similar to letter copying and provides age-based standard scores. While more comprehensive than the Beery VMI alone, it focuses primarily on motor execution.

    BOT-2 (Bruininks-Oseretsky Test of Motor Proficiency, 2nd Edition) evaluates fine and gross motor proficiency including precision, integration, and manual dexterity tasks. Many subtests emphasize speed and accuracy under timed conditions, making it useful for identifying motor delays but less specific to handwriting readiness.

    The Print Tool evaluates actual letter and number formation in children ages 3 to 7, rating legibility, size, spacing, and alignment. While more functional than shape copying, it requires children to already have some writing exposure.

    Developmental Test of Visual Perception (DTVP-3) assesses visual-perceptual and visual-motor skills through tasks including copying, form constancy, and figure-ground discrimination. Performance on these isolated visual tasks is presumed to indicate readiness for integrated writing tasks.

    Minnesota Handwriting Assessment evaluates speed, legibility, and form in school-age children who already write, making it less useful for identifying preschool readiness factors.

    THE RESEARCH

    We analyzed 185 children ages 4 to 4.5 years who received occupational therapy evaluations in North Carolina. This clinical sample consisted of children referred for developmental concerns, with 95 percent qualifying for Medicaid services. Many had limited exposure to structured preschool settings.

    The children were given a comprehensive evaluation using O.T. Wizard and included 8-10 domains per child. During evaluation, children completed a letter copying task: 10 uppercase letters arranged from developmentally simple (L, F, R) to complex (S, X, N). Children copied each letter into a defined box below the model. Occupational therapists rated both the quality of letter production and the child’s behavior during the task.

    The use of uppercase letter copying rather than geometric shapes in preschool assessment warrants clarification. For children lacking letter recognition, uppercase letters serve as geometric forms with the added benefit of providing functional, longitudinal work samples. Unlike abstract shapes that become irrelevant once writing instruction begins, letter samples document the progression from letters-as-shapes to letters-as-symbols, capturing both motor and cognitive development across the transition to formal writing.

    We then examined how well various factors predicted performance on this functional handwriting task. Rather than assuming certain skills matter most, we calculated correlations to let the data reveal which factors actually related to success.

    UNDERSTANDING CORRELATION: THE “r” VALUE

    Before presenting findings, it helps to understand what correlation means and how to interpret the numbers.

    Correlation measures the strength of the relationship between two variables. The correlation coefficient, represented as r, ranges from 0 to 1.0:

    r = 0.0 to 0.1: No meaningful relationship

    r = 0.1 to 0.3: Weak relationship 

    r = 0.3 to 0.5: Moderate relationship

    r = 0.5 to 0.7: Strong relationship 

    r = 0.7 to 1.0: Very strong relationship

    A simple example: Height and shoe size have a strong correlation (r = approximately 0.7). Taller people tend to wear larger shoes, though exceptions exist. The relationship is strong but not perfect.

    In contrast, height and intelligence have essentially no correlation (r = approximately 0.0). Knowing someone’s height tells you nothing about their cognitive ability.

    For our study, correlation indicates how well each skill predicts letter copying success. A high correlation means children with strong skills in that area tend to perform better on writing tasks. A low correlation means the skill does not reliably predict writing performance.

    THE FINDINGS

    Six factors showed moderate correlations with handwriting (visual motor integration) performance, all clustering tightly between r = 0.31 and r = 0.39:

    Fine Motor Skills: r = 0.393 

    Cooperation (during evaluation): r = 0.365 

    Visual Perception: r = 0.340 

    Attention (during evaluation):r = 0.314

    Praxis (Motor Planning): r = 0.314 

    Participation (during writing task): r = 0.310

    The most striking finding is not which factor ranked highest, but rather that all six fell within an 8-point range. Fine Motor scored highest at 0.393, but Participation scored 0.310, a difference of only 0.083.

    Statistical interpretation: All six predictors are moderate in strength, and none dominates. The child with the highest fine motor score has only a slightly better chance of writing success than the child with the highest cooperation score.

    VISUAL PERCEPTION SUBDOMAINS: TASK DEMANDS MATTER

    An interesting pattern emerged when examining visual perception subdomains separately. Not all visual skills predicted copying performance equally:

    Visual Discrimination: r = 0.379 

    Visual Figure Ground: r = 0.304 

    Visual Spatial Relations: r = 0.158 

    Visual Memory: r = 0.137

    Visual Discrimination, the ability to see small differences between similar forms, predicted letter copying better than the overall Visual Perception domain score. This makes perfect sense given the task demands. Copying letters requires discriminating between similar features: Is this a C or an O? Does this letter have a diagonal line or a curve? Are these two vertical lines parallel or converging?

    In contrast, Visual Memory showed the weakest correlation at r = 0.137, barely above no relationship at all. This finding initially seems surprising given that handwriting literature often emphasizes visual memory as critical for letter formation.  However, the weak correlation makes complete sense when we consider the actual task. Children were asked to copy letters with the model remaining visible throughout. They could look back and forth between the stimulus letter and their work as many times as needed. Visual memory is irrelevant when the visual information stays available.

    Visual memory would matter for different handwriting tasks: Writing letters from dictation (hear the letter name, recall what it looks like) Writing spelling words independently (recall the letter in memory) Reproducing letters after brief exposure (look once, then write from memory)

    But for direct copying with continuous visual access to the model, discrimination ability predicts success while memory does not.

    This finding has important implications for assessment practices. If we evaluate visual memory but not visual discrimination, we may be measuring the wrong visual skill for near point copying tasks. Comprehensive assessment requires matching the skills tested to the actual task demands the child will face in the classroom.

    In preschool and early kindergarten, children primarily engage in near point copying: copying letters from a worksheet placed directly in front of them, tracing over models, and reproducing shapes from a stimulus card on the table. These near point tasks allow continuous visual reference, making discrimination critical and memory less important.

    As children progress through elementary school, task demands shift to far point copying: copying from the board, reproducing teacher demonstrations, writing from dictation. These tasks require visual memory because the model is not continuously accessible. A child must look at the board, hold the letter image in memory while looking down at paper, then reproduce from that mental representation.

    For the preschool population in this study engaged in near point copying tasks, visual discrimination predicted success while visual memory did not. This relationship may change for older children performing far point copying or writing from dictation.

    WHAT THE NUMBERS MEAN IN PRACTICE

    Consider what these moderate correlations reveal:

    If fine motor skills were the primary driver of handwriting, we would expect r = 0.6 or higher. Instead, r = 0.393 means fine motor capability explains only about 15 percent of handwriting performance. The remaining 85 percent depends on other factors.

    Similarly, visual perception (r = 0.340) explains about 12 percent. Praxis explains about 10 percent. Cooperation explains about 13 percent.

    No single factor accounts for even 20 percent of performance. Handwriting emerges from complex interactions among multiple systems, not mastery of any single prerequisite.

    THE BEHAVIORAL FACTOR SURPRISE

    Perhaps most notable: Behavioral factors predicted success as well as skill factors.

    Cooperation (r = 0.365) nearly matched fine motor skills (r = 0.393). A child who cooperates with feedback and accepts correction has almost the same probability of writing success as a child with superior hand strength and coordination.

    Attention during evaluation (r = 0.314) predicted exactly as well as motor planning ability (r = 0.314). The child who can focus for the duration of the task performs comparably to the child with better movement sequencing skills.

    Participation during the actual writing task (r = 0.310) predicted nearly as well as any other factor. Willingness to engage with the challenge matters almost as much as capability.

    This explains common clinical observations:

    The child with excellent fine motor skills who gives up after one attempt struggles more than the child with weaker skills who persists through frustration.

    The child who resists feedback and insists on doing it “my way” fails to improve despite adequate motor capability.

    The child who cannot sustain attention long enough to complete three letters never accumulates the practice necessary for skill development.

    CONTEXT MATTERS: THE 1-ON-1 EVALUATION PROBLEM

    An important limitation: Cooperation, attention, and participation were rated during one-on-one evaluation sessions with an occupational therapist providing full support and individualized pacing.

    This context differs dramatically from classroom writing instruction, where:

    One teacher manages 15 to 20 students simultaneously Individual feedback is limited and delayed Pacing is group-determined rather than individualized Distractions are constant Tasks continue for extended periods without breaks

    A child rated as having “adequate cooperation” in a quiet therapy room with undivided therapist attention may demonstrate very different behavior in a busy kindergarten classroom during 15-minute writing periods.

    This suggests our correlations may actually underestimate the importance of behavioral factors. If cooperation, attention, and participation predict success even in optimal conditions, they likely matter even more in typical educational settings.

    IMPLICATIONS FOR ASSESSMENT PRACTICES

    Current handwriting readiness assessments focus heavily on visual-motor integration and fine motor skills while largely ignoring behavioral factors. The Beery VMI, for instance, requires sustained attention and task persistence to complete 30 forms, but these behavioral requirements are not scored or interpreted. A child may fail due to attention limitations rather than visual-motor deficits, yet both receive the same low score.

    More critically, the Beery VMI is frequently used in isolation to qualify children for occupational therapy services for handwriting concerns. Given our findings, this practice is problematic. Visual-perception represents only one of six factors that predict handwriting success in preschoolers, and it predicts moderately (r = 0.340), not strongly. A child may score low on the Beery VMI yet succeed at functional handwriting due to strong cooperation, attention, and participation. Conversely, a child may pass the Beery VMI but struggle with classroom writing due to behavioral regulation challenges that the assessment does not capture.

    Using the Beery VMI as a sole qualifying criterion systematically misidentifies which children need services. Comprehensive evaluation across all six predictive factors provides more accurate identification of handwriting risk.

    Additionally, visual perception assessments for preschool populations should emphasize visual discrimination for near point copying tasks. Our findings demonstrate that for preschoolers copying letters with the model continuously visible, discrimination ability (r = 0.379) predicts substantially better than memory (r = 0.137). This does not suggest visual memory is unimportant for handwriting development overall. Rather, it indicates that the specific skills required depend on task type and developmental stage. Visual memory likely becomes increasingly important as children transition to far point copying and writing from dictation in elementary grades.

    Ratings of current assessments used by OT’s and how they capture handwriting prediction 

    These assessments share common limitations. Based on our findings, we can evaluate how well each captures the six factors that actually predict handwriting success:.

    Beery VMI (with supplemental tests): Rating 5/10 IF subtests administered.  3/10 if only the VMI section is administered.  Captures visual-motor integration (r=0.340) through the primary copying task. Supplemental Visual Perception and Motor Coordination subtests add assessment of visual discrimination and fine motor control, bringing total coverage to 2-3 of 6 predictive factors. However, the Visual Perception subtest does not distinguish between visual discrimination (r=0.379, highly relevant) and visual memory (r=0.137, less relevant for near point copying). Completely misses cooperation, attention, praxis, and task participation. When administered with all three subtests, it provides more comprehensive data than VMI alone, but therapists often use only the primary VMI subtest for qualification decisions.

    PDMS-3: Rating 6/10 Captures fine motor skills (r=0.393) through grasping subtests and visual-motor integration (r=0.340) through copying tasks. Provides 2 of 6 critical factors. Misses cooperation, attention, praxis, and task participation entirely. No assessment of behavioral regulation during tasks or visual discrimination as distinct from visual-motor integration.

    BOT-2: Rating 4/10 Primarily assesses motor proficiency with fine motor precision and integration subtests capturing fine motor skills (r=0.393). However, heavy emphasis on timed performance may penalize slow-but-accurate children. Completely misses visual perception, cooperation, attention, and task participation. Designed for motor proficiency screening rather than handwriting-specific readiness. Captures only 1 of 6 predictive factors.

    DTVP-3: Rating 5/10 Assesses visual perception (r=0.340) across multiple subdomains but does not distinguish between visual discrimination (r=0.379, highly relevant for copying) and visual memory (r=0.137, less relevant for near point tasks). Misses fine motor execution, cooperation, attention, praxis, and task participation. Provides visual skills assessment but in isolation from functional writing context.

    The fundamental issue: These assessments emphasize isolated skill measurement (motor proficiency, visual perception, visual-motor integration) while ignoring behavioral regulation factors that predict equally well. Additionally, they provide scores but often rely on checklist observations that cannot be converted to continuous metrics for tracking progress or comparing across domains.

    Comprehensive assessment should include:

    Fine motor capability: Strength, coordination, precision, tool control Visual-perceptual skills with task-appropriate emphasis: Visual discrimination (critical for copying) Visual figure ground (moderate importance) Visual spatial relations (less critical for copying) Visual memory (only relevant for tasks without visible models) Motor planning: Ability to sequence multi-step actions, organize approach Cooperation: Willingness to accept feedback, modify approach when unsuccessful Attention: Capacity to sustain focus through multi-step tasks Task participation: Engagement level, persistence through challenge, frustration tolerance

    Single-domain screening (testing only visual skills or only motor skills) will systematically miss children at risk. A child may pass fine motor screening with flying colors but struggle with writing due to attention deficits, poor cooperation, or low task engagement.

    Similarly, a child may score well on visual memory subtests but fail at letter copying due to poor visual discrimination. Matching assessed skills to actual task demands is essential.

    Conversely, a child with borderline fine motor scores but strong behavioral regulation may achieve functional writing through persistence and acceptance of instruction.

    IMPLICATIONS FOR INTERVENTION

    Traditional intervention models often follow a sequential approach: establish attention, then build fine motor skills, then introduce visual tasks, then combine into writing. Our data suggests this may be inefficient.

    If multiple factors contribute equally and simultaneously, intervention should address them concurrently rather than sequentially. Children need practice integrating behavioral regulation, motor control, visual processing, and motor planning from the start.

    Effective intervention might include:

    Brief, varied tasks that build attention capacity while practicing motor skills (address both simultaneously) Immediate feedback on both motor execution and behavioral approach (cooperation, persistence) Functional writing activities that require visual processing, motor planning, and sustained attention in authentic context Explicit instruction in self-regulation during challenging tasks (managing frustration, accepting correction)

    Isolated prerequisite activities (strengthening exercises, shape sorting, sequencing games) practiced separately from writing context may not transfer effectively. The child builds attention during tabletop games but cannot apply it during writing. The child demonstrates fine motor control during bead threading but not during letter formation.

    Integration practice appears more efficient: Work on attention, motor control, visual processing, and cooperation simultaneously within functional writing activities.

    WHY SOME CHILDREN SUCCEED DESPITE LIMITATIONS

    These findings explain puzzling clinical observations.

    The child with weak fine motor skills who succeeds likely compensates through: Strong visual perception (carefully observes letter features) High persistence (keeps trying despite motor difficulty) Good cooperation (accepts feedback, modifies approach) Strong attention (focuses carefully on each stroke)

    The combination of strengths in four areas compensates for weakness in one.

    The child with excellent fine motor skills who fails likely struggles with: Poor attention (loses focus mid-letter) Low persistence (gives up when first attempt is imperfect) Resistance to feedback (insists on incorrect approach) Low task engagement (avoids writing activities)

    Motor capability alone cannot overcome behavioral limitations.

    THE CLINICAL SAMPLE CONTEXT

    These findings emerge from a specific population: low-income preschoolers referred for occupational therapy evaluation. Many had limited exposure to structured educational settings or formal writing instruction.

    This context matters for interpretation:

    Children with school experience might show different patterns, as they have had more opportunity to develop attention and cooperation within structured tasks.

    Higher-income samples with more educational exposure might demonstrate stronger correlations for skill factors and weaker correlations for behavioral factors.

    Typically developing children (not referred for therapy) might show different relationships among variables.

    However, this clinical sample represents the population occupational therapists actually serve. Understanding what predicts success in children with developmental concerns and limited educational exposure has direct clinical relevance.

    RESEARCH CONTEXT: HOW OUR FINDINGS COMPARE

    Our findings align with and extend existing research on handwriting development while revealing some important differences.

    Feder and Majnemer (2007) conducted a systematic review identifying visual-motor integration, fine motor skills, and in-hand manipulation as significant predictors of handwriting performance in school-age children. Their meta-analysis found moderate correlations (r = 0.3-0.5) between these factors and handwriting, consistent with our fine motor (r = 0.393) and visual perception (r = 0.340) findings. However, their review focused on older children already engaged in writing instruction, while our sample examined preschoolers in pre-handwriting stages.

    Reference: Feder, K. P., & Majnemer, A. (2007). Handwriting development, competency, and intervention. Developmental Medicine & Child Neurology, 49(4), 312-317.

    Volman, van Schendel, and Jongmans (2006) examined handwriting readiness in kindergarten children and found that visual-motor integration was a significant predictor but explained only a modest portion of variance. This supports our finding that visual-motor skills predict moderately (r = 0.340) but do not dominate. Importantly, they also identified attention and behavioral regulation as contributing factors, aligning with our cooperation (r = 0.365) and attention (r = 0.314) findings.

    Reference: Volman, M. J., van Schendel, B. M., & Jongmans, M. J. (2006). Handwriting difficulties in primary school children: A search for underlying mechanisms. The American Journal of Occupational Therapy, 60(4), 451-460.

    Kaiser, Albaret, and Doudin (2009) investigated the relationship between handwriting quality and various factors in first graders. They found visual perception, fine motor skills, and graphomotor skills all contributed, but no single factor was sufficient. Their findings that multiple factors contribute equally strongly support our multifactorial model. However, they did not examine behavioral factors like cooperation or task-specific participation, which our data suggests are equally important.

    Reference: Kaiser, M. L., Albaret, J. M., & Doudin, P. A. (2009). Relationship between visual-motor integration, eye-hand coordination, and quality of handwriting. Journal of Occupational Therapy, Schools, & Early Intervention, 2(2), 87-95.

    Notably absent from existing literature: Studies examining task-specific participation and cooperation as predictors of handwriting success in preschool populations. Most handwriting research focuses on school-age children who have already received writing instruction and emphasizes motor and perceptual factors while treating behavioral factors as confounding variables rather than legitimate predictors.

    Our finding that cooperation predicts nearly as well as fine motor skills (r = 0.365 vs r = 0.393) extends the literature by demonstrating that behavioral regulation deserves equal consideration in handwriting readiness assessment. The clinical sample context (children referred for evaluation, limited school exposure, 95 percent low-income) may explain why behavioral factors emerged as stronger predictors than in general population studies.

    Additionally, our visual perception subdomain analysis revealing that visual discrimination (r = 0.379) predicts substantially better than visual memory (r = 0.137) for near point copying tasks provides specificity often missing in broader visual perception assessments. This has practical implications for selecting which visual subtests to administer when evaluating preschool handwriting readiness.

    ABOUT OT WIZARD

    Data for this analysis was collected through OT Wizard, a clinical intelligence system for pediatric occupational therapy assessment. The platform evaluates performance across up to twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, executive functioning, activities of daily living, and participation. OT Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update continuously as the clinical database expands.

    Unlike traditional checklist-based assessments, OT Wizard converts all observations to continuous metrics that enable progress tracking, cross-domain comparison, and comprehensive reporting. The platform captures all six factors identified in this research as predictive of handwriting success: fine motor skills, visual perception (with subdomain specificity), praxis, cooperation, attention, and task participation. Behavioral regulation is assessed within the context of actual task performance rather than as an isolated rating, providing clinically relevant data about how attention and cooperation affect functional skill demonstration.

    For handwriting readiness assessment specifically, OT Wizard provides quantified performance across visual discrimination, visual-motor integration, fine motor control, motor planning, and behavioral engagement during writing tasks. This comprehensive approach addresses the multifactorial nature of handwriting development identified in this research. As the platform undergoes Rasch analysis validation and accumulates longitudinal outcome data, it will establish whether comprehensive baseline assessment across all six predictors improves identification of children at risk for handwriting difficulty and informs more effective intervention planning.

    OT Wizard is committed to advancing the occupational therapy profession by collecting de-identified clinical data from real therapist users, building the largest developmental database in pediatric occupational therapy history. This continuous data collection enables research on developmental trends, intervention effectiveness, and response to intervention patterns that elevate practice from perception-based to data-driven decision making and strengthen the evidence base for the entire profession