Tag: RTI

  • Value-Based Care Is Coming for Therapy Practices

    Value-Based Care Is Coming for Therapy Practices

    Value-Based Care Is Coming for Therapy. Here’s What It Means — and Why Some Practices Are Already Ready.


    A significant document landed this week. AOTA, APTA, and ASHA — the three largest therapy professional associations in the country — released a joint publication: Value-Based Care for Therapy: A Provider’s Guide. When three associations that rarely agree on anything publish something together, it is worth paying attention to.

    This post is for occupational therapists, physical therapists, speech-language pathologists, mental health professionals, and the practice owners who employ them. Whether you work in a clinic, a school, a hospital, or private practice — this shift is coming for you.


    What Is Value-Based Care, and Why Does It Matter Now?

    For decades, therapy has been paid under a fee-for-service model. You provide a service, you bill a code, you get paid. Volume drives revenue. The system rewards doing more, not necessarily achieving more.

    Value-based care flips that model. Under VBC, payment is tied to the value of care delivered, meaning the outcomes achieved relative to the cost of achieving them. Payers want evidence that patients improved, that care was efficient, and that therapy dollars produced measurable functional change.

    CMS has stated formally that by 2030, all Medicare plans and most Medicaid plans will include accountability for quality and total cost of care.

    And for anyone thinking this is only a Medicare issue, it is not. Commercial payers historically follow Medicare and Medicaid, typically within two to five years of federal adoption.Plans such as United Healthcare, Blue Cross, Aetna, United, and Cigna watch the federal model and build similar requirements into their own contracts. This is an incoming reality for your entire payer mix, regardless of the population you serve.


    What Will Payers Actually Require — and How Will They Get the Data?

    This is where it gets specific, and where most therapy providers are underprepared.

    Under value-based care contracts, payers will not rely solely on claims data. They will aggregate multiple data streams to generate a risk-adjusted performance score for each provider. That score determines reimbursement rates, bonus eligibility, penalty exposure, and increasingly — prior authorization requirements.

    The data sources payers will draw from include diagnostic codes, functional status scores at evaluation and discharge, episode length, caregiver and patient-reported outcome measures, and social determinants of health. Together, these paint a picture of who your patients are, how complex their needs are, and whether your billed intervention produced meaningful change.

    The risk adjustment piece deserves particular attention. Payers use a system called Hierarchical Condition Categories (HCC) to score the complexity of a provider’s caseload. The formula looks at diagnosis codes to predict how difficult and costly a patient’s care should be. The intent is fairness: a practice treating children with complex medical histories and significant functional deficits should not be benchmarked against a practice treating mild developmental delays.

    But here is the problem. If your documentation does not accurately capture patient complexity (ex: if your evaluations are narrative rather than structured, if diagnoses are under coded, if functional deficits are described rather than measured by metrics) the HCC formula underestimates your caseload. Your practice looks like it is treating simpler patients than it is. Your performance scores suffer. Your reimbursement suffers. Thorough, structured clinical documentation is no longer just professional best practice. It is financial protection.

    There is also a significant upside for high performers. Practices that consistently demonstrate measurable functional outcomes under VBC contracts will be rewarded with reduced prior authorization burdens. For any therapist who has spent hours writing auth appeals, had sessions denied mid-episode, or watched a child lose momentum because a payer delayed approval, that outcome is worth working toward.


    What the VBC Guide Specifically Calls For

    The joint guide outlines several infrastructure requirements for therapy providers preparing for value-based care. Here are the most consequential:

    1. Objective, structured outcome measurement. The guide is explicit: quality measures require scoreable, comparable data — not narrative notes. Payers need to be able to extract, aggregate, and benchmark outcome data across providers. Narrative documentation cannot be benchmarked. Structured, scored data can.

    2. Patient and caregiver-reported outcome measures. The guide specifically highlights the growing importance of capturing outcomes from the patient and family perspective — not just the clinician’s clinical observation. These measures capture health status, function, and quality of life as experienced by the people receiving care. They are becoming a required component of VBC quality scoring.

    3. Longitudinal data across the full episode of care. A snapshot evaluation is not sufficient. Payers need a before and after — baseline functional status, mid-episode progress, and discharge outcome. Without that full-episode arc, there is no way to calculate the value of the intervention.

    4. ICF framework for data exchange. The guide references the International Classification of Functioning, Disability, and Health — ICF — as the standard framework for exchanging functional status data across providers, settings, and payers. Providers whose documentation is built on ICF structure are already speaking the language payers are building their systems around.

    5. Social determinants of health. VBC models are increasingly required to capture nonmedical factors that influence outcomes — housing stability, transportation access, food security, economic stability. These factors affect therapy outcomes and will be factored into risk adjustment models.


    Why Most EMRs and EHRs Will Leave Practices Exposed

    Here is the uncomfortable truth: most electronic medical records and electronic health records were built for fee-for-service. They are fundamentally accounting systems — designed to track what was billed, scheduled, and coded. They do an adequate job of supporting claims submission. They do almost nothing to capture clinical intelligence.

    When value-based care contracts begin requiring structured outcome data, longitudinal functional measures, and caregiver-reported scores, practices running on a standard EMR will have nothing meaningful to submit. The documentation exists, but it is locked in narrative notes that cannot be extracted, scored, or benchmarked. That is a serious vulnerability.

    The practices that will navigate this transition well are the ones that have been capturing structured, measurable clinical data all along — not because a payer required it, but because good clinical practice demanded it.


    If You Are Using OT Wizard, You Are Already Ahead

    OT Wizard was not built to react to value-based care. It was built on the clinical and psychometric principles that value-based care is now catching up to.

    Here is how OT Wizard already delivers what the VBC guide calls for:

    1. Objective functional outcome data. OT Wizard generates structured documentation grounded in the ICF framework — the exact standard the VBC guide identifies for measuring and exchanging outcome data. Every evaluation produces metrics on each domain, subdomain and a composite score, structured functional data, not just narrative description.

    2. Caregiver and patient-reported measures. OT Wizard captures scored questionnaires from caregivers and families at baseline, mid-therapy, and discharge. That full-episode caregiver perspective is built into the platform workflow — not an add-on, not a separate form. It is the data payers will specifically look for.

    3. Longitudinal data across the episode of care. Because OT Wizard tracks from intake through discharge, pre and post functional data is built into every case. That full-episode arc is what payers need to calculate a risk-adjusted performance score accurately and fairly.

    This is what separates a clinical intelligence system from an accounting system. OT Wizard is not just recording what happened. It is measuring what changed.


    What Is Coming Next

    We are not stopping here. We are currently building Billing Wizard feature — a dedicated insurance billing feature integrated directly into OT Wizard. Your clinical outcome data and your claims will live in the same system. As value-based care contracts begin requiring outcome data alongside claims submissions, that integration will matter enormously.

    And not too far away, we will develop Therapy Wizard — an expansion that will bring this same clinical intelligence foundation to speech-language therapy, physical therapy, and mental health. One platform, one outcome framework, built for the full therapy team. Because value-based care does not silo disciplines — and neither should your documentation infrastructure.


    The Bottom Line

    The payment landscape is changing and the timeline is real. The practices that will thrive are not the ones that scramble to retrofit their documentation after the contracts change but are the ones that already have the infrastructure in place.

    If you are an OT, PT, SLP, or mental health professional asking what you can do right now: document with precision, capture complexity, measure function at intake and discharge, and make sure your platform is built to produce structured data… not just notes.

    If you are a practice owner: this is an infrastructure conversation, not just a clinical one. The system you are documenting in today will determine whether you can compete for value-based contracts tomorrow.

    The field is changing. OT Wizard was already here.


  • What We Document vs. What We Actually Need to Know

    What We Document vs. What We Actually Need to Know

    Clinical Takeaways | O.T. Wizard Research Series, Part 1

    By Stephanie Seymore Wick, MSOT, OT/L | Founder and Clinical Architect, O.T. Wizard


    Most pediatric OTs are excellent documenters. We write thorough evaluations, set meaningful goals, and log every session. The paper trail is solid. So why is it so hard to answer the one question that matters most?

    Is this working and if so, how much?

    Not “are we doing the right things?” Not “is this child making progress in a general sense?” The specific question: is this child changing, how fast, and is that fast enough?

    That question turns out to be surprisingly hard to answer with the tools most of us are using.

    The Snapshot Problem

    A standardized evaluation gives you a score at one point in time. A re-evaluation gives you another score. You compare the two and write a narrative about what changed. That is snapshot documentation, and it is useful. It tells you where a child started and where they landed.

    What it does not tell you is anything about the line between those two points.

    When did the change happen? Was progress steady, or did the child plateau for two months and then accelerate? Did a specific intervention approach produce better outcomes than others? Did an attendance gap create a measurable dip? Did the child actually cross a key functional threshold six weeks before the re-evaluation was even scheduled?

    Without structured session-level data linked to domain scores, you simply cannot see any of that. You have two dots. You do not have a trajectory.

    Why This Shows Up Differently in Medical vs. School Settings

    In a medical outpatient practice, the moment this gap becomes visible is usually an authorization request. You are being asked to justify continued services and the strongest argument is quantitative: here is where the child started, here is the rate at which they are improving, and here is where they are projected to land by the end of this period. Without RTI data, you fall back on clinical narrative. Narrative is defensible. It is not the same as a slope.

    In school-based practice, the moment arrives at the IEP table. You are sitting with a team that includes a parent, a classroom teacher, a special educator, and an administrator. They are deciding whether OT services should continue, increase, or be exited. The OT who arrives with a performance slope and a comparison to natural developmental growth is a different professional presence than the one who arrives with quarterly progress notes. Both care about the child. Only one has data that can change a decision.

    Exit recommendations are especially difficult without RTI data. Recommending that a child be exited from OT services is a clinical and ethical judgment call. With measurement data showing that the child has reached functional independence or participation and is maintaining gains without direct support, it becomes a defensible milestone. Without that data, it is an opinion.

    The Plateau Conversation

    Every pediatric OT has been here. A child who was making visible progress has leveled off. The parent is worried. The payer is skeptical. The school team is questioning whether to continue services.

    The problem is that a plateau looks the same on paper whether it is stagnation or consolidation. A child consolidating a new skill at a lower level of support may show no numerical gain for several sessions. That is not failure. It is a normal part of skill acquisition. But without session-level performance data, you cannot show anyone the difference. You can only explain it.

    With structured data, you can show the team exactly when the plateau began, what changed in the child’s routine or support structure around that time, and whether similar plateaus have resolved in this child’s history. That changes the conversation from “we think this is temporary” to “here is what the data shows.”

    What Parents Are Actually Asking

    When a parent asks whether therapy is working, the most honest answer most therapists can give without RTI infrastructure is a clinical impression. That impression may be completely accurate. But it is not the same as showing a parent a graph of their child’s performance over twenty sessions and saying: here is where he started, here is the rate at which he is moving, and here is what we project by the end of this period.

    For school-based OTs, the parent question arrives at the IEP table, in front of an entire team. The quality of your data shapes what parents understand, what they advocate for, and what they accept when the team recommends a service change. That matters.

    The Practical Distinction

    Documentation and measurement are not the same system. They are not competing systems either. They serve different purposes.

    Documentation records what happened, establishes compliance, and communicates clinical reasoning. Measurement tracks rate of change, identifies what conditions produce better performance, and determines whether progress is sufficient.

    Most EHRs were built for the first column. Very few were built for the second. The gap between them is not a failure of clinical intent. It is a gap in infrastructure. EMR’s are typically built by people who are interesting in billing insurance and keeping accounting records. They are not clinical intelligence systems.

    One Finding Worth Noting

    When the O.T. Wizard re-evaluation data was examined, Participation and Executive Functioning showed flat longitudinal profiles compared to the large gains seen in VMI, ADL, and fine motor domains. The initial interpretation might be that OT did not improve those areas. But there is another possibility worth taking seriously: the first evaluation rating in those domains may not have captured authentic baseline behavior. Children often present their best behavior when meeting a new therapist in a structured evaluation setting. By re-evaluation, the novelty has worn off. The rating that looks flat may simply be more accurate.

    That is a question that would not have surfaced without measurement data. It has real implications for how we interpret initial evaluation scores in observation-dependent domains. It is the kind of question that data raises and documentation alone cannot.

    A Few Things to Reflect On

    At re-evaluation, how do you determine the rate at which a child progressed? Can you identify which sessions produced the most meaningful gains? Can you show a parent the slope of improvement over an authorization period? Can you distinguish a true plateau from a reduction in required support?

    If those questions are hard to answer with your current system, the infrastructure gap is real, and it is worth thinking about.

    Parts 2 through 4 of this series will move from problem framing to evidence to practice, including composite clinical vignettes, full dataset patterns, and what RTI infrastructure looks like in day-to-day clinical workflow.


    Disclosures: The author is the Founder and Clinical Architect of O.T. Wizard and has a financial interest in the platform. All data referenced is de-identified clinical data collected through the O.T. Wizard software platform in routine practice.

    About O.T. Wizard: O.T. Wizard is a clinical intelligence system for pediatric occupational therapy professionals. The platform evaluates performance across twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, executive functioning, activities of daily living, and participation. For more information, visit otwizard.com.


  • Built on Evidence. Proven in Practice. Shoutout to Learning Charms’ Team

    Built on Evidence. Proven in Practice. Shoutout to Learning Charms’ Team

    How 3.8 Years of Systematic Clinical Measurement Demonstrates That Occupational Therapy Works

    Stephanie Seymore Wick, MSOT, OT/L | Founder and Clinical Architect, O.T. Wizard | Learning Charms, Inc., Charlotte, North Carolina

    The Problem With Checklists

    For years, occupational therapists working in early childhood settings were collecting data that told them almost nothing. Checklist-style evaluations produced a snapshot: present or absent, yes or no. They could not tell you whether a child improved. They could not tell you which counties had greater concentrations of developmental need. They could not tell you whether your team’s intervention was moving the needle or whether children were simply getting older.

    That was the reality facing Learning Charms in 2022. We were screening and evaluating large numbers of children across Head Start programs, NC Pre-K classrooms, and community settings throughout North Carolina, and we had nothing meaningful to show for it in terms of trend data, geographic insight, or outcome evidence.

    So I built something.

    From Nothing to 8,509 Screenings and Evaluations

    The FUNdamental Foundations (FF) screener was designed and developed by a managing pediatric occupational therapist with 25+years of clinical experience. It was built as a structured, multi-domain developmental tool designed from the outset to generate analyzable data. It was not designed for publication. It was designed to answer clinical questions: What does this population look like? Where are the gaps? Is what we are doing making a difference?

    Thirty clinicians on the Learning Charms team tested and used each version in the field, providing the real-world feedback that drove every refinement. They made the transition from paper-and-pencil evaluations to digital data entry on a tablet or laptop, mid-session, with children in front of them. That is not a small ask. The early weeks required support with the technology. There were growing pains. The team did it anyway, and they did it without much, if any complaint.

    The FF tool went through two versions, each refined based on team feedback. Version 6 ran from June 2022 through July 2023. Version 7, with improvements including date of birth capture and an updated item structure, ran from August 2023 through May 2025. In late 2025, the practice transitioned to O.T. Wizard, a fully rebuilt clinical intelligence platform designed and built by the same therapist.  O.T. Wizard was built with Rasch psychometric architecture, 15 guided evaluations, and integrated outcome tracking across 12 domains.

    The table below summarizes what 3.8 years of that effort produced.

    Table 1. Clinical Data Collected Across the Full Evidence Ecosystem (2022-2026)

    PlatformPeriodRecordsEvaluationsScreeningsE1-E2 Pairs
    FUNdamental Foundations V6Jun 2022 – Jul 20232,9281,5111,417324
    FUNdamental Foundations V7Aug 2023 – May 20254,9622,3682,594477
    O.T. WizardSep 2025 – Mar 202661961997
    TOTAL3.8 years8,5094,4984,011898

    Note. E1-E2 pairs = children with two complete evaluations allowing pre-to-post comparison. FF V6 pairs are V6-only matches. FF V7 pairs include V7-only and cross-version (V6 E1 to V7 E2) matches. OTW pairs matched by Student_ID. pp = percentage points.

    In total: 8,509 individual assessment records. 4,498 full evaluations. 4,011 developmental screenings. 898 pre-to-post evaluation pairs. Over 245,000 item-level data points. Collected by a single clinical team, through routine practice, over less than four years.

    What the Data Shows: Gains That Exceed Maturation

    The central question in any clinical outcome dataset without a randomized control group is this: how do you know the gains are from intervention and not just from children getting older?

    We address this directly.

    Using cross-sectional developmental data from our own E1 (Initial Evaluation) dataset, we calculated the expected rate of developmental growth per month for each skill area based on age alone. This gives us a maturation baseline specific to this population. We then compared that expected gain to the gains actually observed in children who received OT services between E1 and E2 (Re-evaluation), over a mean interval of 5.4 months.

    The results are consistent across all three measured domains and across both independent datasets.

    Table 2. Observed Gains vs. Expected Maturation Over Mean 5.4-Month Interval (FF n=801 pairs, OTW n=94 pairs)

    ItemFF Observed GainExpected (Maturation)RatioOTW Observed Gain
    Draw a Person (0-4 scale)+1.24 pts+0.39 pts3.2x+1.06 pts
    Functional Pencil Grasp+25.3 pp+9.7 pp2.6x+27.7 pp
    Finger Touching (54-mo milestone)+20.3 pp+11.7 pp1.7xn/a
    Cohen’s d (DAP)0.940.84

    Note. Expected gain calculated from cross-sectional linear regression of E1 scores on age in months using the full FF evaluated dataset. pp = percentage points. Cohen’s d: 0.2 = small, 0.5 = medium, 0.8 = large effect. OTW finger touching item not directly comparable due to different item structure.

    Draw a Person improved at 3.2 times the expected developmental rate in FF and 2.6 times in OTW. Functional pencil grasp improved at 2.6 times expected in FF and 3.1 times in OTW. These are not marginal differences from what maturation alone would predict. They are two to three times larger. And they replicate across an entirely independent dataset collected with a different tool, by the same team, with different children.

    Why This Is Not Just Children Getting Older

    If the gains above were driven primarily by maturation, we would expect children at all starting points to show similar improvement. A child who enters at score 0 would gain roughly as much as a child who enters at score 3, because age-related development does not care where you start.

    That is not what we see. The table below shows Draw a Person gains stratified by E1(Initial Evaluation)  score, combining FF and O.T. Wizard data. The pattern is unambiguous.

    Table 3. Draw a Person Gain by E1 Score: FF (n=801) + OTW (n=93) Combined

    E1 ScorenE2 MeanMean Gain% Improved% Same% Declined
    0 (no parts)365+291.74+1.7474%26%0%
    1 (approximations)203+162.37+1.3781%13%6%
    2 (head, no body)165+312.66+0.6654%37%9%
    3 (recognizable)60+112.96-0.0429%46%25%
    4 (6+ body parts)8+63.36-0.360%57%43%

    Note. n column shows FF count + OTW count at each E1 score level. E1 score 0 = no recognizable approximations. Score 4 = recognizable person with 6 or more body parts. Gains decline systematically as E1 score increases, reflecting ceiling effects at higher starting points rather than absence of progress.

    Children who started at score 0 improved by an average of 1.74 points, with 74% showing measurable gains. Children who started at score 3 or 4 were near the ceiling of the scale and showed flat or slightly negative scores at E2, exactly as ceiling effects predict.

    This score-dependent gain gradient is the signature of a real treatment effect. Maturation produces relatively uniform gains regardless of starting point. Intervention produces the largest gains in children with the most room to grow. That is what we observe, and it replicates point-for-point across both the FF and OTW datasets independently.

    The grasp and finger touching data tell the same story from a different angle.

    Table 4. Skill Transition Rates: What Happened Between E1 and E2

    ItemStatus at E1nOutcome at E2
    Pencil GraspNon-functional377 (FF) + 65 (OTW)60% converted to functional
    Pencil GraspFunctional417 (FF) + 29 (OTW)94% maintained functional
    Finger TouchingFail407 (FF)58% passed at E2
    Finger TouchingPass394 (FF)82% maintained pass

    60% of children with non-functional pencil grasp at E1 had functional grasp by E2. 94% of children with functional grasp at E1 maintained it. Skills were not fluctuating randomly. They were moving in one direction and holding. That is not maturation. That is intervention.

    Two Tools, Three Years Apart, Same Answer

    The FF screener and O.T. Wizard are different instruments. FF was a clinician-developed Google Form with embedded scoring anchors and standardized stimulus materials. O.T. Wizard is a fully architected clinical platform undergoing Rasch psychometric validation, with 596 data variables per evaluation and item-level calibration. The two tools share several core items, including Draw a Person, pencil grasp classification, and finger touching. They do not share overlapping children. The DAP scale is directly comparable across both tools at the 0 to 4 range, with identical scoring anchors at each level. O.T. Wizard extended the ceiling by adding two higher-level descriptors, bringing the OTW scale to 6 points total. For this analysis, OTW DAP scores were capped at 4 to ensure a valid cross-tool comparison.

    Yet when we calculate the cross-sectional developmental growth rate for Draw a Person from FF E1 data, we get 0.072 points per month. From OTW E1 data, we get 0.082 points per month. Two tools, thousands of children, the same underlying developmental trajectory captured within 0.01 points of each other per month.

    When two independent measurement systems produce convergent developmental slopes and convergent gain ratios, that is not a coincidence. That is construct validity. Each dataset serves as an independent replication of the other’s findings, and both point to the same conclusion.

    Danielle, an OTR/L out of NC, administers an evaluation with a 3 year old using OT Wizard

    What This Means for OT Practice and Clinical Infrastructure

    Pediatric occupational therapists have long known that their interventions make a difference. The challenge has been demonstrating it systematically, at scale, in a form that partners, funders, schools, and insurance providers find credible.

    The Learning Charms team built that demonstration over 3.8 years, starting from scratch, with no research funding, no university partnership, and no IRB. They built it by replacing meaningless checklists with structured clinical measurement, by training a team of 30 clinicians to collect data consistently, and by iterating their tools until the data was worth analyzing.

    O.T. Wizard is the current iteration of that infrastructure. It is not a platform claiming efficacy. It is a platform whose evidence base already exists, built by the same team that built the platform, using the same children, in the same communities, over the same years. The data in this paper is not a promise of what O.T. Wizard will eventually show. It is a record of what systematic clinical measurement has already demonstrated.

    OT works. The data, replicated across tools and years and nearly 900 pairs of children, shows it.

    Disclosure

    Stephanie Seymore Wick is the founder and clinical architect of O.T. Wizard and owner of Learning Charms, Inc. All data was collected through routine clinical practice and contracted screening partnerships. No external funding was received. The FUNdamental Foundations screener was a clinician-developed field tool and has not undergone formal psychometric validation. O.T. Wizard is currently undergoing Rasch analysis validation. All findings should be interpreted as practice-based clinical evidence rather than results from a randomized controlled trial.

    About O.T. Wizard

    O.T. Wizard is a clinical intelligence system for pediatric occupational therapy professionals. The platform supports evaluation, documentation, goal planning, and scheduling across 12 domains including fine motor skills, visual-motor integration, praxis, visual perception, executive functioning, activities of daily living, and participation. O.T. Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update as the clinical database expands. Learn more at otwizard.com.

    About Learning Charms

    Learning Charms is a pediatric occupational therapy group that employs roughly 25 OTP’s in the Charlotte , NC and surrounding counties. Learning Charms is now focused mainly on preschool aged children in their school environment.

  • What a Brief Scissor Skills Assessment Reveals in Preschool-Aged Children

    What a Brief Scissor Skills Assessment Reveals in Preschool-Aged Children

    More Than a Milestone: What a Brief Scissor Skills Assessment Reveals About Tool Use, Hand Dominance, and Cutting Development in Preschool-Aged Children

    Stephanie Seymore Wick, MSOT, OT/L  |  Founder and Clinical Architect, O.T. Wizard

    March 2026

    Abstract

    Aims: To examine scissor cutting performance across preschool age bands (36 to 65 months) in a clinical sample, identify relationships between cutting accuracy, hand dominance, and hand positioning, assess cross-domain correlations, and evaluate longitudinal progress from initial to re-evaluation.

    Methods: Descriptive and correlational analysis of 541 pediatric occupational therapy evaluations (Age Bands G through J) using structured cutting tasks scored on a standardized 14-point rubric. Participants were children referred for OT services, predominantly ages 48 to 59 months, with at least 90% qualifying for Medicaid.

    Results: Cutting accuracy followed a clear developmental progression across age bands. Thumb-up dominant hand positioning was a large-effect predictor of cutting accuracy (Cohen d=1.17). Established hand dominance and writing-to-scissor hand consistency were strongly associated with performance. Scissor performance correlated significantly with fine motor, visual motor integration, ADL, visual perception, gross motor, bilateral integration, and praxis domains. Longitudinal gains of 5.21 points over 4.8 months exceeded the expected natural growth rate of 1.51 points.

    Conclusions: A structured scissor skills assessment captures clinically meaningful variation in cutting skill and supports response-to-intervention documentation, goal writing, and cross-domain clinical reasoning in pediatric OT practice.

    Keywords: scissor skills, hand dominance, fine motor development, pediatric occupational therapy, response to intervention, preschool

    Scissors are one of the most commonly targeted skills in pediatric occupational therapy, yet they are rarely assessed with the precision required to drive goal writing, track progress, or demonstrate response to intervention. In most clinical settings, a child either can cut or cannot cut. That binary framing misses the developmental story that unfolds across the preschool years and leaves practitioners without the data needed to communicate clinical value to families, educators, and payers.

    The preschool years represent the primary window for scissor skill acquisition. Cutting a straight line with 1-inch tolerance is typically expected by 41 months, precision cutting within a 1/4-inch boundary emerges around 48 months, and smooth curvy-line cutting at a 1/4-inch tolerance is an expectation by 60 months (O.T. Wizard Scissor Skills Assessment, v6.2). Despite this well-established developmental sequence, most standardized pediatric OT assessments address scissor skills with limited granularity, and published outcome data on scissor skill development in referred clinical populations remains sparse.

    This article presents findings from 541 pediatric occupational therapy evaluations collected through O.T. Wizard, a clinical intelligence platform for pediatric occupational therapy professionals. Using a standardized scissor skills assessment embedded within the evaluation process, this analysis examines cutting performance across age bands, its relationship to hand dominance and hand positioning, its correlation with multiple developmental domains, and longitudinal gains across re-evaluation. The evidence supports scissor skill assessment as a window into neuromotor organization, tool use learning, and cross-domain functional development.

    Methods

    Participants

    This analysis includes 541 pediatric occupational therapy evaluations representing children ages 36 through 65 months (Age Bands G through J): Band G (36 to 47 months, n=61), Band H (48 to 53 months, n=199), Band I (54 to 59 months, n=256), and Band J (60 to 65 months, n=69). An additional 82 children had paired evaluations (E1 and E2), with a mean interval of 4.8 months between assessments, enabling longitudinal analysis. The sample was 57% male and 43% female. Primary language was English for 90% of participants and Spanish for 9%. At least 90% of children qualified for Medicaid, and approximately 86% were recommended for OT services following evaluation. All evaluations were conducted in North Carolina.

    Data were collected through routine pediatric occupational therapy evaluations conducted in clinical practice using O.T. Wizard. Parents provided informed consent as part of standard clinical care. As this analysis represents clinical outcomes data rather than human subjects research, institutional review board approval was not required. All data were de-identified in accordance with HIPAA regulations.

    Measures

    The O.T. Wizard Scissor Skills Assessment (v6.2) is a structured, standardized tool embedded within the O.T. Wizard multi-domain pediatric OT evaluation platform. It measures cutting performance across a progression of task demands anchored to published developmental milestones: holding scissors with one hand (24 months), snipping paper (34 months), opening and closing scissors (36 months), cutting a 1-inch straight line (41 months), 3/4-inch and 1/2-inch straight lines (43 and 45 months), a 1/4-inch straight line (48 months), a 1/4-inch curvy line (60 months), and smooth cutting quality (72 months). Cutting accuracy was scored using a standardized rubric with 0, 0.5, and 1.0 point values per segment, yielding a maximum score of 14 points per cutting task. Therapists also documented scissor type used, dominant hand position (thumb up versus thumb down or absent), stabilizer hand description, and qualitative hand positioning ratings.

    This analysis focuses primarily on FM_SCIS_48 (1/4-inch straight line), selected as the primary analytic item due to near-complete data across all age bands (n=541) and its anchor at the 48-month developmental expectation. FM_SCIS_41 (1-inch straight line) and FM_SCIS_60C (1/4-inch curvy line) were included where sample sizes permitted. Scissor type data indicated that 90.6% of children were assessed with school or safety scissors; adaptive equipment (spring-assisted, loop handle) was used in fewer than 1% of cases and is not analyzed separately.

    Data Analysis

    Descriptive statistics were calculated for FM_SCIS_48 by age band, dominance level, and thumb positioning group. Pearson and Spearman correlations were computed between FM_SCIS_48 and all available domain and subdomain scores. Between-group comparisons used independent samples t-tests; effect sizes are reported as Cohen d. Longitudinal analysis used paired t-tests comparing E1 and E2 scores among children with paired evaluations. A cross-sectional growth rate was used to estimate expected natural maturation over the mean evaluation interval, following the methodology established in the response-to-intervention outcomes article in this series. Artificial intelligence writing assistance (Claude, Anthropic, version Sonnet 4.6) was used in the preparation of this manuscript for language editing and formatting; all analytical decisions, clinical interpretations, and conclusions are those of the author.

    All data is pre-normative. Rasch analysis validation is ongoing and will be reported in subsequent publications.

    Results

    Developmental Progression of Cutting Accuracy

    Across all age bands, mean scores on FM_SCIS_48 increased steadily, with the steepest growth occurring between Bands G and H, the window when this skill is developmentally expected to emerge (Table 1). At Band H (the anchor age for this task), 35.6% of children in this clinical sample scored zero, reflecting the referred nature of the population. The bimodal distribution within Band H is more clinically informative than a pass/fail classification. Of 165 children scoring 10 or above on FM_SCIS_48, 162 (98.2%) were also administered FM_SCIS_60C on the same evaluation day. Among children who scored 14 on FM_SCIS_48, scores on FM_SCIS_60C ranged across the full spectrum (mean 8.61), confirming that the curvy task captures a meaningfully more demanding level of motor control.

    Table 1. FM_SCIS_48 (1/4″ straight line) performance by age band. Data from a clinical sample of children referred for OT evaluation, prior to intervention.

    Age BandAge RangenMean /14% ScoreFloor (0)Ceiling (14)
    G36-47m451.7912.8%64.4%2.2%
    H48-53m1804.6433.1%35.6%10.6%
    I54-59m2486.5246.6%18.1%21.4%
    J60-65m688.0757.7%13.2%26.5%

    Hand Positioning and Hand Dominance

    Dominant hand thumb-up positioning was associated with substantially higher cutting performance. Children with thumb-up positioning averaged 8.16 on FM_SCIS_48 compared to 2.74 for children without thumb-up positioning (Cohen d=1.17, p<0.001; median scores 9.0 versus 1.0). Within Band H, 28% of children who scored zero had thumb-up positioning compared to 89% of children who scored 14. No children who scored 14 on FM_SCIS_48 were rated Never for overall hand positioning quality.

    Hand dominance status was consistently associated with cutting performance (Table 2). Children with established dominance scored more than five points higher on average than children with inconsistent dominance. The Spearman correlation between dominance level and FM_SCIS_48 was rho=0.360 (p<0.001). Writing-to-scissor hand consistency produced a large effect: children who used the same hand for writing and cutting averaged 7.06 on FM_SCIS_48 compared to 2.25 for children who used different hands (Cohen d=1.05, p<0.001).

    Table 2. FM_SCIS_48 mean score by documented hand dominance level (Bands G-J, n=534). Dominance level based on therapist rubric rating at time of evaluation.

    Dominance LevelnMean /14% ScoreFloor (scored 0)
    Emerging181.5010.7%56%
    Inconsistent702.3817.0%49%
    Strong preference2335.1736.9%28%
    Established2137.7655.4%16%

    Cross-Domain Correlations

    Scissor cutting performance correlated significantly with all major developmental domain scores (Table 3). Fine Motor domain showed the strongest correlation (r=0.772). Visual Motor Integration (r=0.511) and Visual Perception (r=0.461) correlations reflect the visual guidance demands of cutting within a narrow boundary. The ADL correlation (r=0.511) speaks to generalization of tool use skill across daily living contexts. The Gross Motor correlation (r=0.430) reflects the role of proximal stability as a foundation for distal precision. Praxis showed the weakest correlation (r=0.271), consistent with motor planning contributing primarily to early skill acquisition rather than refined execution. The FMSAT Speed and Accuracy subdomain showed no significant relationship (Spearman rho=-0.012, not significant), confirming that scissor performance and pencil speed capture distinct aspects of fine motor control and contribute independent clinical information.

    Table 3. Pearson and Spearman correlations between FM_SCIS_48 and domain/subdomain scores (Bands G-J). *** p<0.001. FMSAT Spearman correlation not significant.

    Domain / SubdomainnPearson rSpearman rhoClinical Interpretation
    Fine Motor domain4540.772***0.787Strong — cutting as fine motor expression
    Hand Use subdomain5410.541***0.549Bilateral tool use as hand use marker
    Visual Motor Integration5370.511***0.519Visual guidance of cutting path
    ADL domain5100.511***0.530Tool use generalizes to daily living
    Visual Perception domain5150.461***0.469Line discrimination supports accuracy
    Gross Motor domain5400.430***0.444Proximal stability drives distal precision
    Bilateral Integration5380.398***0.395Two-hand coordination demand
    Praxis domain5390.271***0.281Motor planning, weaker once program formed
    Speed & Accuracy (FMSAT)459r=-0.105*rho=-0.012 (ns)Distinct skill; independent clinical value

    Longitudinal Gains

    Among 82 children with paired evaluations, FM_SCIS_48 showed a mean gain of 5.21 points over an average interval of 4.8 months (paired t-test t=8.10, p<0.001; Cohen d=0.895). The curvy task (FM_SCIS_60C) showed a mean gain of 3.75 points over the same interval (t=7.20, p<0.001), with 76.4% of children improving. Using the cross-sectional growth rate as a natural maturation baseline, the expected natural growth on FM_SCIS_48 over 4.8 months was approximately 1.51 points. The observed mean gain of 5.21 points exceeded this expected rate by 3.71 points. Ceiling effects were noted among children who had scored at or near 14 at E1; therapists correctly applied FM_SCIS_60C as the clinically sensitive measure for near-ceiling children in 98.2% of applicable cases.

    Discussion

    This analysis of 541 pediatric OT evaluations demonstrates that a brief, standardized scissor skills assessment generates clinically meaningful data across multiple dimensions of preschool development. The developmental progression of FM_SCIS_48 scores across age bands aligns with published milestone expectations and provides clinical benchmarks for a referred population that are not currently available in the literature. Floor effects in younger bands and the bimodal distribution at the anchor age band reflect the nature of scissor skill acquisition in children referred for OT services, where skill emergence is delayed relative to normative expectations. These distributions support documentation of functional deficit and medical necessity in ways that pass/fail classifications cannot.

    The large effect of thumb-up dominant hand positioning (Cohen d=1.17) elevates hand positioning from a clinical observation to a primary, modifiable intervention target. Establishing correct scissor grip before focusing on path accuracy is supported by the data: children without functional thumb-up positioning have mean scores below three regardless of age band, suggesting that grip orientation is a near-prerequisite for achieving cutting accuracy at the mastery level.

    The relationship between hand dominance and scissor performance connects these findings to a broader literature on neuromotor specialization in the preschool years (Scharoun & Bryden, 2014). A child who has not yet organized a consistent preferred hand for tool use is reflecting an underlying developmental state that affects performance across all tool-based tasks. Writing-to-scissor hand consistency findings have direct implications for school-based practice: addressing hand consistency across classroom tool use activities, not only during designated scissor tasks, is consistent with both the data and motor learning principles and supports IEP goal development that reflects the child’s functional performance across settings.

    The cross-domain correlation profile challenges the framing of scissor skills as an isolated fine motor task. The significant Gross Motor correlation (r=0.430) reinforces that proximal stability, including trunk support and shoulder girdle control, is foundational to distal precision, consistent with developmental neuroscience frameworks (Stoodley, 2016). The VMI and Visual Perception correlations reflect the visual guidance demands of path-following. Intervention plans that address only the distal cutting task without considering foundational postural, visual, and neuromotor systems may produce slower or less durable gains. The absence of a significant FMSAT correlation confirms that scissor accuracy and pencil precision speed are complementary measures rather than redundant ones, and that both contribute independent information to a comprehensive fine motor profile.

    Longitudinal gains of 5.21 points over 4.8 months, exceeding the expected natural growth rate by 3.71 points in a population with limited scissor practice outside of OT sessions, provide meaningful support for OT as a driver of skill development. This above-expected gain pattern is consistent with the RTI methodology established in this article series, which separates natural developmental progress from intervention-attributable change using cross-sectional growth rates as a baseline correction. For medical model practitioners, this framing supports quantifiable medical necessity documentation. For school-based practitioners, it provides data to support RTI tier documentation and progress monitoring language consistent with IDEA requirements.

    Several methodological factors warrant consideration. This sample represents children referred for OT evaluation in North Carolina and should not be generalized to typically developing children or other geographic populations. The cross-sectional age-band comparisons reflect group differences rather than individual trajectories. Adaptive scissor data was insufficient for analysis. Without a controlled comparison group, causal claims about OT effectiveness cannot be made; above-expected gains in a referred population with limited community practice are suggestive but not definitive. Future research should include a typically developing comparison sample, session-level dosage data, and expanded age coverage through early elementary years to determine whether scissor precision continues to develop beyond 66 months and at what point the 1/4-inch straight line becomes a floor item for older children (Cameron et al., 2012; Zhang et al., 2025).

    Conclusions

    A structured scissor skills assessment generates clinically meaningful data for pediatric OT practice. Cutting accuracy in a referred preschool population follows a measurable developmental trajectory, is strongly predicted by thumb-up dominant hand positioning and established hand dominance, correlates significantly with fine motor, gross motor, visual motor, visual perception, ADL, and praxis domains, and shows gains that exceed expected natural growth rates over a 4.8-month evaluation interval in a population with limited community scissor access. These findings support the use of standardized, quantitative scissor skills assessment as a component of comprehensive pediatric OT evaluation and as a practical tool for RTI documentation, goal writing, and cross-domain clinical reasoning in both school-based and medical model practice settings.

    Disclosure of Interest

    Stephanie Seymore Wick is the Founder and Clinical Architect of O.T. Wizard, the platform from which all data in this article was collected. Data collection is ongoing under clinical quality improvement protocols. All data were de-identified in accordance with HIPAA regulations. The author reports no other competing interests.

    Data Availability Statement

    De-identified aggregate data supporting the findings of this study are available from the corresponding author upon reasonable request. Individual-level data cannot be shared due to HIPAA de-identification obligations.

    Biographical Note

    Stephanie Seymore Wick, MSOT, OT/L is the Founder and Clinical Architect of O.T. Wizard, a clinical intelligence platform for pediatric occupational therapy professionals, and the founder of Learning Charms, Inc. Her clinical and research focus is the development of psychometrically sound, computable measurement tools that quantify pediatric OT outcomes across multiple developmental domains. She practices and conducts research in North Carolina.

    References

    Cameron, C. E., Brock, L. L., Murrah, W. M., Bell, L. H., Worzalla, S. L., Grissmer, D., & Morrison, F. J. (2012). Fine motor skills and executive function both contribute to kindergarten achievement. Child Development, 83(4), 1229-1244. https://doi.org/10.1111/j.1467-8624.2012.01768.x

    Scharoun, S. M., & Bryden, P. J. (2014). Hand preference, performance abilities, and hand selection in children. Frontiers in Psychology, 5, 82. https://doi.org/10.3389/fpsyg.2014.00082

    Stoodley, C. J. (2016). The cerebellum and neurodevelopmental disorders. Cerebellum, 15(1), 34-37. https://doi.org/10.1007/s12311-015-0715-3

    Zhang, B.-F., Lin, Z.-C., & Li, C. (2025). Fine motor skills assessment instruments for preschool children with typical development: A scoping review. Frontiers in Psychology, 16, 1620235. https://doi.org/10.3389/fpsyg.2025.1620235

  • Scissor Skills Are Telling You More Than You Think

    Scissor Skills Are Telling You More Than You Think

    What 541 preschool OT evaluations revealed about scissor skill development, hand dominance, and the whole child

    Stephanie Seymore Wick, MSOT, OT/L  |  Founder and Clinical Architect, O.T. Wizard

    March 2026

    This post summarizes clinical findings from O.T. Wizard data. For full statistical detail, methodology, and references, read the companion research article: “What a Brief Scissor Skills Assessment Reveals About Tool Use, Hand Dominance, and Scissor Skill Development in Preschool-Aged Children.

    We Have Always Known Scissor Skills Matter. Now We Have the Numbers.

    You already know scissor skills are more than a craft milestone. You know a child who cannot hold scissors with thumb up, who switches hands mid-cut, or who can barely score a line with school scissors is telling you something important about their neuromotor development. What we have not always had is the data to say exactly what they are telling us, and to show that picture to the families, teachers, and payers who need to understand it.

    That is what this analysis is about.

    O.T. Wizard collected structured scissor skills assessment data across 541 pediatric OT evaluations from preschool-aged children (ages 3 to 5 and a half) in North Carolina. At least 90% of children qualified for Medicaid. All were referred for OT services after failing a developmental screening. For many, OT sessions were the primary or only setting where scissors were regularly available. Teachers limit scissor time in large classroom groups. Parents restrict use at home. That context matters a lot when we look at what the data shows.

    What We Measured

    The O.T. Wizard Scissor Skills Assessment is a structured, scored tool built into the evaluation process. It uses a developmental progression of cutting tasks, from snipping (around 34 months) through cutting a 1/4-inch straight line (around 48 months) to cutting a 1/4-inch curvy line (around 60 months). Each task is scored on a 0 to 14 point scale using a consistent segment-by-segment rubric. Therapists also document thumb orientation, stabilizer hand, and overall hand positioning quality.

    Unlike a standardized evaluation tool that captures a single snapshot in one domain, O.T. Wizard tracks performance across twelve domains simultaneously including fine motor, gross motor, visual motor integration, visual perception, activities of daily living, praxis, and executive functioning. This is what made the findings below possible. We were not just looking at how a child cuts. We were looking at what scissor skill tells us about the rest of the child.

    What We Found

    Scissor skill development follows a clear developmental path, and most 4-year-olds referred for OT are not yet where we would expect.

    All data in this table reflects first evaluations (E1) conducted before OT intervention began. These are the starting points, the picture of where children arrived for their first evaluation. On the 1/4-inch straight line task (developmentally anchored at 48 months), here is how children in this clinical sample performed at initial evaluation:

    Age GroupAvg Score /14% AccurateScored Zero (0/14)Scored Perfectly (14/14)
    3:0 to 3:11 (Band G)1.813%64%2%
    4:0 to 4:5 (Band H)4.633%36%11%
    4:6 to 4:11 (Band I)6.547%18%21%
    5:0 to 5:5 (Band J)8.158%13%27%

    Table 1. Scissor skills performance on 1/4-inch straight line task at first evaluation (E1), before OT intervention. Clinical sample referred for OT evaluation. Data from O.T. Wizard, pre-normative.

    At the anchor age for this task, 36% of children arrived at their first evaluation scoring zero on the 1/4-inch straight line, and another 21% scored between 1 and 3 out of 14. That means that at initial evaluation, before OT had even begun, the majority of 4-year-olds referred for services could not yet cut a 1/4-inch line with any consistency. That is not a failure of intervention. That is a description of who we serve and why they need us. And now we can describe it precisely, in numbers, instead of writing “emerging scissor skills” in a narrative.

    These are clinical benchmarks for a referred population, not norms for typically developing children. They show where our kids start. That starting point is worth documenting and measuring.

    The multi-task design works. Advancing to harder tasks captures the full range of scissor skill ability.

    A common concern with developmental assessments is ceiling effects: what happens when a child is too skilled for the task in front of them? The data confirmed that therapists in this sample handled this correctly. Of children who scored 10 or higher on the 1/4-inch straight line task, 98% were also given the curvy line task on the same evaluation day. Among children who scored a perfect 14 on the straight line, the curvy line scores ranged across the full spectrum, averaging 8.6 out of 14 (61% accuracy). The harder task captured real, meaningful variation that the straight line task could not.

    GroupnStraight Line Avg /14Straight Line % AccurateCurvy Line Avg /14Curvy Line % Accurate
    Scored 10+ on straight line16512.589%6.849%
    Scored 14 on straight line9014.0100%8.661%

    Table 2. Curvy line performance among children approaching or reaching ceiling on the straight line task. Confirms that advancing to the harder task captures meaningful clinical variation.

    Thumb up is not just a cue. It is the difference between scissor skills and not having them.

    This was the strongest predictor in the entire dataset. Hand positioning was not just associated with better cutting accuracy. It was associated with a more than 3-fold difference in performance.

    Thumb PositioningnAvg Score /14% AccurateMedian Score /14
    Thumb-up (correct)~2708.259%9.0
    Not thumb-up~2702.719%1.0

    Table 3. Scissor skills performance by dominant hand thumb positioning. Bands G through J. Cohen d=1.17 (large effect).

    Looking at Band H children specifically: among those who scored a perfect 14 on the 1/4-inch line, 89% had correct thumb-up positioning. Among those who scored zero, only 28% did. Not a single child at zero had positioning rated as “always” correct. The positioning is not the finishing touch on scissor skill development. It is the prerequisite.

    For intervention planning, this data strongly supports prioritizing grip and hand orientation before focusing on line-following accuracy. Therapists who have structured treatment this way have been right all along. Now there are numbers to back it up, and to share with families and IEP teams.

    Hand dominance predicts scissor skill accuracy, and the connection runs deeper than which hand holds the scissors.

    Children with established hand dominance performed dramatically better on the cutting task than children with inconsistent or emerging dominance. The data below reflects all children in Bands G through J at initial evaluation.

    Dominance LevelnAvg Score /14% Accurate% Scored Zero
    Emerging181.511%56%
    Inconsistent702.417%49%
    Strong preference2335.237%28%
    Established2137.856%16%

    Table 4. Scissor skills performance by documented hand dominance level at initial evaluation. Spearman rho=0.360, p<0.001.

    The hand consistency finding was equally striking. Children who used the same hand for both writing and cutting averaged 7.1 out of 14 (51% accuracy). Children who used different hands for writing versus cutting averaged only 2.3 out of 14 (16% accuracy). That is a more than 3-fold performance gap, and it connects directly to the Hand Dominance article in this series.

    A child who switches hands between writing and cutting is not just making an inconsistent tool choice. They are reflecting an unresolved neuromotor organization question that affects all tool use tasks. For school-based OTs, this is IEP-relevant data. Documenting hand consistency across writing and cutting tasks as part of the evaluation supports accommodation planning for the full classroom day, not just during scissor activities.

    Scissor skills are a whole-body skill. The domain correlation data proves it.

    This is where O.T. Wizard’s multi-domain approach made findings possible that are simply not achievable with a single standardized assessment tool that only looks at one domain or captures a single point in time.

    Cutting performance was correlated against every other domain scored in the same evaluation. Here is the picture that emerged:

    DomainCorrelation with Scissor Skills ScoreWhat This Means Clinically
    Fine MotorStrong (r=0.77)Expected: distal precision underlies cutting
    Visual Motor IntegrationModerate (r=0.51)Visual guidance is needed to follow a line
    ADL (Daily Living)Moderate (r=0.51)Tool use skill generalizes across daily tasks
    Visual PerceptionModerate (r=0.46)Seeing and interpreting the line drives cutting accuracy
    Gross MotorModerate (r=0.43)Trunk and shoulder stability support distal hand control
    Bilateral IntegrationModerate (r=0.40)Confirmed: cutting is a two-handed coordinated task
    PraxisWeak-moderate (r=0.27)Motor planning matters early; less so once the program is established
    FMSAT (Fine Motor Speed and Accuracy Test)Near zero (r=0.00)Pencil speed and cutting accuracy are distinct skills

    Table 5. Correlations between scissor skills performance (1/4-inch straight line) and O.T. Wizard domain scores. All correlations p<0.001 except FMSAT which was not significant. Bands G through J, n=454 to 541 depending on domain.

    The gross motor correlation deserves a specific callout. Cutting is a distal fine motor task, but proximal stability drives distal precision. Trunk support and shoulder girdle control provide the foundation from which hand precision is expressed. A child with poor postural stability will have reduced arm control, which directly limits how precisely they can guide scissors along a line. This is the body-supports-hand principle that experienced OTs understand clinically. This dataset quantifies it.

    The FMSAT finding is equally important in a different way. FMSAT “Bubble-popping” measures open-field pencil speed and precision. Scissor skills assessment measures controlled path-following with a bilateral tool. These two tasks draw on related but distinct aspects of fine motor function, and both contribute independent clinical information. Having both in the same evaluation is not redundancy. It is clinical depth.

    This is the kind of cross-domain picture you cannot build from a BOT-2 or a Beery VMI alone. Those tools give you a score in an isolated domain. O.T. Wizard gives you a developmental profile across twelve domains, from the same child, on the same day, every time you complete an evaluation.

    Children Made Real Progress, and OT Likely Drove Most of It.

    Among the 97 children with two evaluations on file, those with scissor skills data at both time points showed an average gain of 5.2 points on the 1/4-inch straight line over an average of 4.8 months between evaluations. Nearly 70% showed improvement. However, not all of the sample had scissor related goals on their plan of care. 

    MeasureE1 (Before OT)E2 (After ~5 months OT)Gain% Who Improved
    Avg score /14 (straight line)4.4 / 149.6 / 14+5.2 pts69.5%
    % Accuracy31%69%+38 pts
    Avg score /14 (curvy line)2.5 / 146.2 / 14+3.8 pts76.4%
    % Accuracy (curvy)18%44%+27 pts

    Table 6. Scissor skills gains from initial evaluation (E1) to re-evaluation (E2). n=82 for straight line, n=72 for curvy line. Mean interval 4.8 months. Clinical sample referred for OT services.

    Based on cross-sectional growth data, we would expect natural developmental growth of about 1.5 points over a 4.8-month window. These children gained 5.2 points. That is 3.7 points above the expected natural rate, representing more than triple the growth that maturation alone would predict.

    And remember the context: these children were largely practicing scissor skills only during OT sessions. Teachers avoid large-group scissor time. Parents restrict home use. If nearly all scissor practice was happening in OT, and children gained more than three times the expected natural growth rate, that is meaningful evidence for OT-driven outcomes, even without a randomized controlled trial.

    For school-based and medical model OTs alike, this is the kind of data that supports medical necessity, justifies continuation of services, and answers the parent question: Is this working?

    What Makes This Different From a Standard Evaluation

    Standardized tools like the BOT-2, PDMS-3, or Beery VMI are valuable. They are not being replaced. But they have a structural limitation: they capture performance in isolated domains, on a single day, at a single point in time. They cannot show how a child’s scissor skills score connects to their gross motor stability or ADL function. They cannot track how a child changes from evaluation to re-evaluation. And they cannot build a growing evidence base across hundreds of children that gets more precise over time.

    O.T. Wizard was designed to do all of those things. Every evaluation adds to the clinical intelligence base. Scissor skills scores sit alongside fine motor, gross motor, visual perception, ADL, VMI, praxis, executive functioning, and participation data from the same child on the same day. The cross-domain correlations in this article were only possible because the platform was built to be computable, not just documentable.

    This is the difference between documenting therapy and understanding it.

    What This Means for Your Practice

    If you work in a school setting: Scissor skills accuracy scores contextualized within a developmental progression give you IEP-ready language. A score of 4.6 out of 14 at Band H is not just “emerging” and is a measurable starting point with room to define a meaningful, achievable goal. The hand consistency finding connects directly to classroom accommodations: if a child switches hands between writing and cutting tasks, that is relevant accommodation data for the IEP team, not just a therapy note.

    If you work in a medical model setting: The 5.2-point average gain over 4.8 months, in a population with almost no between-session scissor exposure, supports medical necessity documentation with concrete numbers. The cross-domain correlations support a whole-child framing in your evaluation report: scissor skills difficulty is not just a fine motor problem. It reflects neuromotor organization, visual guidance, postural stability, and bilateral coordination working together.

    For both settings: Thumb-up hand positioning is the most actionable clinical target in the dataset. If thumb orientation is not being documented and targeted as a prerequisite for cutting accuracy, this data makes the case for starting there.

    Scissor Skills Are the Child’s Story

    A pair of scissors in a preschooler’s hand is a window. It shows you how well the brain has organized a preferred side, how the trunk is supporting the arms, how the eyes are guiding the hands, and how much the child has internalized the motor program for this specific tool. It is one of the richest clinical observations we make, and one of the least quantified.

    That is changing. Data from over 500 evaluations now gives us a picture of what scissor skill development looks like in the children we actually serve, what predicts success, what domains are implicated, and what progress looks like over time. This is the beginning of an evidence base that the profession has needed.

    Click to read the full research article with statistics, tables, and references.

    Disclosure of Interest

    Stephanie Seymore Wick is the Founder and Clinical Architect of O.T. Wizard, the platform from which all data in this article was collected. Data collection is ongoing under clinical quality improvement protocols. All data has been de-identified in accordance with HIPAA regulations.

    About O.T. Wizard

    O.T. Wizard is a clinical intelligence system for pediatric occupational therapy professionals. The platform evaluates performance in evaluations, forms, and daily treatment notes across twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, executive functioning, activities of daily living, and participation. O.T. Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update continuously as the clinical database expands. Learn more at otwizard.com.

  • Measuring Pediatric O.T. Outcomes Above the Threshold of Natural Maturation

    Measuring Pediatric O.T. Outcomes Above the Threshold of Natural Maturation

    O.T. Wizard Clinical Research Series

    89% Improved. Five Domains Exceeded Natural Growth. Here Is What the Data Shows.

    Measuring Pediatric OT Outcomes Above the Threshold of Natural Maturation

    By Stephanie Seymore Wick, MSOT, OT/L | Founder and Clinical Architect, O.T. Wizard

    Introduction

    Every pediatric occupational therapist knows that the work they do matters. The harder question is whether the profession can show, in precise and reproducible terms, how much it matters. For decades, OT documentation has been built around goals, progress notes, and clinical narratives. These tools record care. They rarely measure change in a way that separates what the child gained through intervention from what developmental maturation would have produced on its own.

    This report addresses that gap directly. Using O.T. Wizard, a clinical intelligence system designed to generate structured, reproducible, multi-domain assessment data for pediatric OT practice, we examined functional performance change across eight domains in 71 preschool-aged children who completed two full evaluations an average of 4.85 months apart.

    The central question throughout this analysis is not simply whether children improved. The meaningful question is whether the children in this cohort improved beyond what developmental maturation alone would have produced over the same interval. A natural growth correction applied consistently throughout this report makes that distinction explicit in every finding.

    This is the expanded replication of a February 2026 analysis of 44 paired evaluations. The findings across the 27 additional pairs are consistent with and strengthen the earlier report across all domains. The core story does not change with more data. It becomes more precise.

    ABSTRACT

    Background: Pediatric occupational therapy has well-established standardized tools for point-in-time measurement, including the Bruininks-Oseretsky Test of Motor Proficiency, Beery-Buktenica Developmental Test of Visual-Motor Integration, and Peabody Developmental Motor Scales-Third Edition. Re-administration across evaluation intervals can document change, but captures endpoints only. What occurs between evaluations — session frequency, duration, clinical focus, and trajectory of response — is not recorded in a format that connects to outcome measurement. Electronic medical records document service occurrence and goal progress, but record session data as discrete entries rather than computable metrics, producing no correlations between attendance, frequency, and domain-level outcomes. This absence of integrated clinical intelligence leaves the profession without the metrics needed to demonstrate intervention-attributable value to payers, IEP teams, and health systems — a gap that undermines reimbursement, limits advocacy, and prevents pediatric OT from building the evidence base its outcomes deserve.

    Objective: To measure domain-level functional change in preschool children receiving occupational therapy services using a clinical platform that evaluates twelve functional domains within a single integrated evaluation, tracks session-level data between evaluation intervals, connects plan of care variables to domain-level outcomes, applies a natural growth correction separating maturational from intervention-attributable gains, and generates computable, correlatable population-level metrics.

    Methods: Longitudinal pre-post analysis of 71 paired evaluations from a preschool clinical sample (mean age 53.5 months; mean interval 4.85 months; at least 90% Medicaid-qualifying). A 9.2% natural growth rate was applied as the maturational baseline. Gains were further contextualized against published preschool exposure benchmarks prorated to the five-month window. All data are pre-Rasch ordinal values.

    Results: 89% of children improved in under five months. Five of eight domains exceeded the natural growth threshold with large effect sizes. VMI exceeded the published preschool exposure benchmark by d=0.92, ADL by d=0.90, and Fine Motor by d=0.65. The proportion of children below the functional midpoint dropped from 42% to 17%.

    Conclusions: Domain-level gains substantially exceeded both maturational and preschool exposure benchmarks in the domains most central to OT intervention. Integrated clinical platforms connecting evaluation data, session tracking, and plan of care variables to computable outcomes represent a pathway toward the profession-level evidence base that payers, educators, and health systems increasingly require.

    Study Sample

    Age Distribution

    The longitudinal cohort consisted of 71 preschool-aged children, each with two complete O.T. Wizard evaluations separated by a minimum of 30 days. The mean inter-evaluation interval was 4.85 months (approximately 148 days), with a range of approximately 37 to 173 days. Mean age at first evaluation was 53.5 months.

    Starting age band distribution: Band G (36 to 47.99 months, n=6), Band H (48 to 53.99 months, n=29), Band I (54 to 59.99 months, n=33), and Band J (60 to 65.99 months, n=3). Bands H and I together represent 87% of the sample and are the primary basis for findings reported here. Bands G and J are included in the data but interpreted with caution given their smaller sizes.

    Demographics and Clinical Status

    All assessments were conducted in North Carolina through the O.T. Wizard clinical platform. Consistent with the broader software dataset, at least 90% of children qualified for Medicaid, and for many, the structured evaluation environment represented an early introduction to formal educational or clinical settings. Primary language was English for 90% of children, with 9% Spanish-speaking and 1% other. All children had been recommended for occupational therapy services following developmental screening failure.

    It is important for readers to interpret these findings within this clinical context. This is not a typically developing population. These are children with identified developmental concerns who were referred for and receiving skilled OT services. Outcome findings therefore reflect the response of a clinically referred, predominantly low-income sample to structured early intervention, not population-level developmental norms.

    Natural Growth Framework

    Before examining domain-level findings, it is necessary to establish what score change we would expect to observe in the absence of intervention. Children in this cohort averaged 53.5 months of age at first evaluation and were reassessed approximately 4.85 months later. On a well-constructed developmental scale, maturation alone would be expected to produce a gain proportional to that age progression.

    The natural growth rate for this cohort is calculated as the mean inter-evaluation interval divided by the mean age at Evaluation 1: 4.85 months divided by 53.5 months equals 9.2%. This figure represents the expected score improvement attributable to developmental maturation alone over the study period.

    A gain of 9.2% would be expected from natural maturation alone over 4.85 months.Gains above 9.2% represent intervention-attributable change.

    This natural growth rate serves as the reference threshold throughout this report. Domain gains below 9.2% suggest performance did not keep pace with chronological age progression. Gains at 9.2% suggest maturation-equivalent growth. Gains above 9.2% represent functional improvement beyond what age progression alone would predict.

    This correction is transparent, reproducible, and requires no external normative sample to apply. It is a direct arithmetic relationship between age progression and scale progression on a fixed instrument. The Gain Above Natural column in each table makes this comparison explicit.

    Research Hypotheses

    Three primary hypotheses guided this analysis. First, children receiving occupational therapy services would demonstrate composite score gains substantially exceeding the 9.2% natural growth threshold. Second, domains most directly targeted by OT in preschool settings, specifically visual motor integration, fine motor skills, visual perception, and activities of daily living, would show the largest gains above expected growth. Third, Band H children (48 to 53.99 months) would demonstrate greater gains than Band I children, reflecting greater developmental sensitivity at a younger starting point.

    Literature Review and Context

    The preschool years represent a critical period for fine motor and visual motor development. Between the ages of three and five, neuromotor pathways underlying pencil control, bilateral coordination, and hand specialization undergo rapid maturation, establishing the foundation for academic skill development. Handwriting readiness, scissor use, and self-care independence all draw from skill sets that are most efficiently built during this developmental window.

    Visual motor integration has consistently been identified as one of the strongest predictors of kindergarten handwriting readiness. Daly and colleagues (2003) found VMI performance at preschool age predicted handwriting speed and legibility at ages six and seven with effect sizes exceeding those of fine motor or visual perception measures alone. Duff and colleagues (2015) demonstrated that children with developmental coordination difficulties who receive targeted fine motor intervention during the preschool years show significantly better handwriting outcomes at school entry than matched peers without services. It is equally important to recognize that VMI does not operate in isolation as a predictor of handwriting development. Emergent literacy skills, particularly alphabet knowledge, letter-sound awareness, and early orthographic processing, are also well-established predictors of handwriting fluency and transcription accuracy (Gerde et al., 2025; Puranik et al., 2011). The relationship between literacy exposure and VMI development is bidirectional: children who are actively engaged in letter-learning and pre-writing activities in preschool settings are simultaneously building the visual discrimination, directionality, and motor planning foundations that underlie VMI performance. This intersection is clinically relevant and is acknowledged as a study limitation below.

    The measurement infrastructure required to track these outcomes longitudinally has historically been a limiting factor in OT outcomes research. Standard evaluation protocols typically capture a single snapshot of performance, and re-evaluation data, when it exists, is rarely structured for computational comparison. O.T. Wizard was designed to address this gap, enabling structured, reproducible, multi-domain measurement within the constraints of a standard clinical evaluation and supporting longitudinal outcome tracking that traditional paper-based protocols do not practically support.

    This report also extends a prior O.T. Wizard longitudinal analysis (Wick, 2026) that examined 44 paired evaluations from the same clinical platform. The expanded sample of 71 pairs presented here confirms and extends those findings with consistent direction and strength across all eight domains assessed.

    Key Findings

    Overall Composite Performance

    Across all 71 children with valid paired evaluations, mean composite score increased from 519.8 at Evaluation 1 to 646.0 at Evaluation 2, a mean raw gain of 126.2 points. Against the 9.2% natural growth expectation, the expected gain for this cohort was approximately 47.6 points. The observed gain exceeded the natural growth threshold by 78.6 points, representing 165% above expected developmental progress. To be precise about what that means: for every point of progress that natural maturation would have produced, these children gained 2.65 points. They moved forward at more than two and a half times the rate that developmental aging alone would have driven. That is not incremental. That is intervention doing exactly what skilled, structured, early occupational therapy is designed to do.

    The gain was highly statistically significant (paired t-test, t=10.62, p<0.001, Cohen’s d=1.26, large effect). 89% of children showed improvement at the second evaluation. 42% of children began below the 500-point composite threshold; by Evaluation 2, only 17% remained below that threshold. Twenty children crossed the functional midpoint of the scale during the study interval.

    89% of children improved. 20 children crossed the 500-point functional threshold.The composite gain exceeded expected natural growth by 165%.

    Domain-Level Results

    The following table presents results across all eight assessed domains, sorted by magnitude of gain above natural growth. All scores are pre-Rasch ordinal percentage values expressed as points within each domain’s maximum possible score. Natural Gain represents the expected gain based on the 9.2% natural growth rate applied to each domain’s Evaluation 1 mean.

    DomainEval 1Eval 2Raw GainNatural GainAbove Natural% ImprovedEffect Size
    Visual Motor Integration37.662.2+24.63.4+21.1 (614%)94%d=1.49 (Large)
    Activities of Daily Living45.569.2+23.74.2+19.5 (468%)87%d=1.27 (Large)
    Fine Motor Skills47.763.7+16.04.3+11.6 (268%)86%d=1.00 (Large)
    Gross Motor Skills56.872.3+15.55.2+10.3 (198%)73%d=0.77 (Medium)
    Visual Perception62.775.1+12.45.7+6.7 (118%)77%d=0.74 (Medium)
    Praxis52.156.6+4.54.8-0.3 (-6%)48%d=0.15 (ns)
    Participation65.269.4+4.26.0-1.8 (-30%)62%d=0.26 (*)
    Executive Functioning63.766.2+2.55.8-3.4 (-57%)52%d=0.15 (ns)

    Table 1. Domain-level longitudinal comparison. Scores are points within each domain’s maximum possible score. Natural Gain = Eval 1 mean x 9.2% natural growth rate. Above Natural = Raw Gain minus Natural Gain. Effect sizes: Large (d>0.8), Medium (d>0.5). Executive Functioning, Participation, and Praxis findings are addressed in the discussion section. All scores are pre-Rasch raw values.

    Visual Motor Integration produced the largest gain above expected growth in the dataset, rising from 37.6 to 62.2 points, a raw gain of 24.6 points against an expected natural gain of 3.4 points. The gain above natural growth was 21.1 points, representing 614% above what maturation alone would have produced. 94% of children with VMI scores showed improvement. The effect size of d=1.49 is considered large by conventional standards.

    Activities of Daily Living showed a raw gain of 23.7 points against a natural expectation of 4.2 points, placing the gain above natural growth at 19.5 points (468% above expected). Fine Motor Skills showed 11.6 points above the natural expectation (268% above expected, d=1.00, large). Gross Motor Skills showed 10.3 points above expected (198% above expected, d=0.77, medium). Visual Perception showed 6.7 points above expected (118% above expected, d=0.74, medium).

    Praxis, Participation, and Executive Functioning showed gains at or below the natural growth threshold, none with statistically significant large effects. The interpretation of these findings requires clinical context and is discussed in detail below. A dedicated companion analysis of the Participation and Executive Functioning longitudinal findings is forthcoming in this research series, as the novelty effect hypothesis and its implications for clinical documentation merit extended treatment.

    Performance by Starting Age Band

    The following table presents composite score change by starting age band. Natural growth rates vary slightly by band because younger children have a larger age progression ratio over the same elapsed time. Bands G and J are included for completeness but should be interpreted with caution given small sample sizes.

    Age BandnEval 1 MeanEval 2 MeanRaw GainAbove NaturalNGR
    G (36-47.99 mo)6394.7496.3+101.7+57.011.3%
    H (48-53.99 mo)29516.6651.3+134.8+84.89.7%
    I (54-59.99 mo)33542.7670.2+127.5+81.98.4%
    J (60-65.99 mo)3549.7628.0+78.3+32.88.3%

    Table 2. Composite score change by starting age band. NGR = natural growth rate (months elapsed / age at Eval 1). Above Natural = Raw Gain minus (Eval 1 mean x NGR). Bands G and J interpreted with caution (small n).

    Band H children (ages 48 to 53.99 months) showed the largest absolute gains above natural growth, averaging 84.8 points above the natural expectation on a composite gain of 134.8 points. Band I children showed 81.9 points above expected on a composite gain of 127.5 points. Both primary age bands show gains well above the natural growth threshold, and the difference between them is modest. This is broadly consistent with the earlier 44-pair analysis, which found Band H slightly outperforming Band I. The 48 to 54 month window continues to appear as a period of high clinical yield for OT service delivery, though both bands show substantial responsiveness.

    Understanding the Flat Domains: Praxis, Participation, and Executive Functioning

    Three domains showed gains at or below the natural growth threshold: Praxis (-6%), Participation (-30%), and Executive Functioning (-57%). These findings are clinically important to interpret carefully, as they do not simply mean that OT failed to produce change in these areas.

    For Praxis, the near-zero gain is more likely a reflection of current measurement sensitivity than true insensitivity to intervention. Praxis is a complex, context-dependent construct requiring the integration of motor planning, bilateral coordination, and sequencing across novel tasks. Detecting incremental praxis development over a five-month interval likely requires either longer measurement windows or more precisely calibrated items. Rasch calibration of the praxis item bank is a priority in the continuing research agenda.

    Participation and Executive Functioning tell a more nuanced story that involves the measurement context itself. Both domains are rated by the therapist based on behavioral observation during the evaluation. At Evaluation 1, the child is meeting the therapist for the first time. The novelty of the interaction, the structured environment, and the desire to engage with an unfamiliar adult may produce elevated ratings that reflect situational compliance rather than the child’s authentic behavioral baseline. By Evaluation 2, the therapeutic relationship is established and the child is comfortable enough to reveal their genuine regulatory and engagement patterns, including the variability and difficulty that characterize their daily functioning. If Evaluation 1 ratings are systematically elevated by this novelty effect, the apparent absence of gain at Evaluation 2 reflects a measurement context shift rather than a failure of intervention. A dedicated research blog on the novelty effect hypothesis and its implications for clinical documentation is forthcoming in this series.

    Correlation Analysis

    A moderate negative correlation was observed between starting composite score and magnitude of change (consistent with the 44-pair analysis). Children who began with lower scores tended to show larger gains. This regression-to-the-mean effect is expected in clinical samples and does not invalidate the findings, but is an important interpretive consideration. No significant correlation was found between inter-evaluation interval length and change score, indicating that the range of intervals in this cohort (approximately 37 to 173 days) did not materially influence the magnitude of observed gains.

    Implications for OT Practice

    The domain-level findings interpreted through the natural growth framework allow occupational therapists to make specific, evidence-informed decisions about evaluation and intervention priorities. Five domains showed gains ranging from 118% to 614% above the 9.2% natural growth threshold, all with statistical significance and large or medium effect sizes. These gains were achieved in a predominantly Medicaid-qualifying, low-income clinical population with significant developmental concerns, which makes the magnitude of change all the more clinically meaningful.

    For insurance authorization and educational planning, the natural growth framework provides a communication tool that is both precise and accessible. Rather than reporting a raw score change, the practitioner can state that the child’s VMI performance exceeded the expected developmental rate by 21.1 points over approximately five months, providing clear evidence that skilled OT intervention, not maturation, drove the observed change. This framing is methodologically transparent and directly responsive to the medical necessity standards that payers apply.

    The composite score threshold finding carries particular weight for authorization purposes. A child who begins services below the 500-point composite threshold and crosses it during the authorization period has demonstrated objectively measurable functional change. Of the 30 children who began below 500 points, 20 crossed that threshold during the study interval. That is a two-thirds success rate in moving children from below-threshold to at-threshold performance within a single authorization period.

    Serial assessments using a consistent instrument also generate slope data that goes beyond a single outcome comparison. The rate of gain above natural growth, calculated at the domain level, can be used to project whether a child is on track to reach functional goals within a given authorization period, supporting proactive communication with payers and educational teams before a plateau needs to be explained rather than after.

    Implications for Intervention Planning

    The convergent large effect sizes across VMI, ADL, Fine Motor, and Gross Motor domains point toward a functional skill cluster that is highly responsive to structured OT programming during the preschool developmental window. These four domains share underlying requirements for postural control, bilateral coordination, and visually guided hand movement. Interventions that integrate these components across functional activities are supported by both the data pattern and established OT theory.

    The Gross Motor finding is particularly relevant for intervention sequencing. A gain of 10.3 points above natural expectation with a medium-to-large effect confirms that proximal postural and movement foundations are responsive to OT services alongside distal fine motor work. For children showing limited fine motor or VMI gains, postural foundation and gross motor assessment should be considered before concluding that the upper extremity is the primary limiting factor.

    For children whose evaluation profiles show strength in Gross Motor relative to Fine Motor and VMI, a proximal-to-distal intervention sequence may accelerate gains across the entire cluster. The strength of the ADL finding (d=1.27) reflects the functional integration that OT uniquely provides: when children gain in fine motor, VMI, and postural control simultaneously, daily living skills follow as a natural downstream effect.

    The flat findings for Praxis, Participation, and Executive Functioning should not reduce the clinical attention given to these areas. They reflect current measurement constraints rather than evidence of non-response to intervention. Goal writing in these domains should continue, supported by structured therapist observation and emerging platform tools designed to capture behavioral change over longer intervals.

    Study Limitations

    This study carries several important limitations that readers should consider when interpreting and applying the findings.

    The sample is clinical and geographically restricted to North Carolina. Findings cannot be generalized to typically developing children or to populations in other regions with different demographic profiles, service delivery models, or referral criteria. The absence of a control group means observed gains cannot be causally attributed to OT intervention. Natural maturation, regression to the mean, and test familiarity effects each contribute to observed change scores to an unknown degree.

    The 9.2% natural growth correction is a methodologically transparent estimate derived directly from the age progression of this cohort on this instrument. It assumes proportional developmental scaling across the score range, an assumption that Rasch calibration will allow us to test empirically. Future work with a typically developing comparison group will allow domain-specific, empirically derived growth expectations to replace this uniform estimate.

    The regression-to-the-mean effect means that domains with the lowest Evaluation 1 scores (VMI, ADL, Fine Motor) also showed the largest gains. The true intervention effect within these domains is likely substantial, but the proportion attributable to treatment versus regression toward the mean cannot be fully separated without a control group. All scores remain pre-Rasch ordinal percentage values. Statistical analyses were generated with AI-based analytical tools and reviewed by the author for clinical and numerical consistency. Final responsibility for interpretation rests with the author.An additional and important limitation specific to the VMI domain is the potential confounding effect of preschool attendance and literacy instruction. 

    A significant body of research demonstrates that access to quality preschool accelerates cognitive and academic skill development, with effects that are particularly pronounced for children from low-income households (Magnuson & Duncan, 2016; Bailey et al., 2024). Preschool curricula in the four-year-old age range routinely incorporate letter recognition, alphabet knowledge, pre-writing activities, and structured fine motor practice, all of which directly engage the visual-motor and orthographic processing skills that O.T. Wizard’s VMI domain measures. Because O.T. Wizard does not currently collect data on whether a child is enrolled in preschool, how many days per week they attend, or what literacy instruction they are receiving, it is not possible to separate the contribution of preschool-based literacy exposure from the contribution of OT services to the VMI gains observed. The large VMI gains reported here almost certainly reflect the combined influence of OT intervention, natural maturation, and classroom-based literacy and pre-writing instruction. Future data collection that captures school enrollment status and attendance patterns would allow this important confounder to be examined directly.

    Band G (n=6) and Band J (n=3) findings should be treated as exploratory only. The Participation and Executive Functioning longitudinal findings are subject to the novelty effect interpretation described above, which cannot be confirmed or ruled out without the prospective study design described in the continuing research section.

    Continuing Research Needed

    Rasch calibration remains the highest research priority for O.T. Wizard. Transforming ordinal raw scores into interval-level person measures will allow true scale-independent longitudinal comparison, validate the proportional scaling assumption underlying the natural growth correction, and identify items requiring revision. Current analyses are pre-Rasch and should be interpreted as preliminary clinical evidence rather than psychometrically standardized measurement. The O.T. Wizard National Try-Out Team initiative is designed to expand sample sizes needed for stable item calibration across all domains and age bands.

    A typically developing comparison group would allow the 9.2% natural growth estimate to be validated empirically and replaced with domain-specific growth expectations calibrated against external developmental benchmarks. Recruiting a non-clinical sample, even a modest one, would strengthen the interpretive framework considerably and provide a more precise foundation for the gain-above-expected metric.

    The novelty effect hypothesis for Participation and Executive Functioning requires prospective investigation. A study design capturing therapist-rated engagement and work habits at multiple time points within the first year of services, alongside parent-reported and teacher-reported measures, would allow empirical testing of whether first-evaluation ratings systematically overestimate authentic baseline functioning.

    Longitudinal expansion with test-retest intervals of 12 to 24 months would allow examination of whether early VMI and ADL gains are sustained through kindergarten entry, and whether children who make the largest gains above natural growth in the preschool period show measurably better school readiness outcomes. Linking O.T. Wizard composite and domain scores to standardized criterion measures, including teacher-rated school readiness and kindergarten entry assessments, would establish predictive validity and position the platform’s data within the broader early childhood outcomes literature.

    Conclusion

    This analysis of 71 preschool-aged children with paired O.T. Wizard evaluations, examined through a transparent natural growth framework, extends the domain-level outcome picture established in the February 2026 report. Across a mean interval of 4.85 months and a natural growth expectation of 9.2%, five of eight assessed domains showed gains that were statistically significant, clinically large in effect, and substantially above what maturation alone would produce.

    89% of children improved overall. 20 children crossed the 500-point composite functional threshold during the study interval. Visual Motor Integration showed gains of 21.1 points above the natural expectation, 614% above what developmental maturation alone would predict over the same period. Activities of Daily Living and Fine Motor Skills showed gains of 468% and 268% above expected, respectively. These are not marginal differences. They represent functional gains at rates that developmental maturation cannot explain.

    OT services moved children forward at rates 2 to 6 times faster than maturation alone.

    These findings are preliminary. They require replication with larger samples, validated comparison conditions, and Rasch-calibrated measurement. What they establish is that structured domain-level digital assessment in pediatric OT can generate longitudinal outcome data that clearly and transparently distinguishes intervention-driven change from natural maturation. For a profession that has historically struggled to quantify its impact in terms that payers and educational systems recognize, that distinction is not a minor technical refinement. It is the foundation of evidence-based practice.

    A six-part blog series examining what outcome data reveals about pediatric OT, what documentation systems currently miss, and how structured Response to Intervention measurement changes clinical practice is forthcoming in the O.T. Wizard Research Series beginning the week of March 9, 2026.

    Disclosures

    The author is the Founder and Clinical Architect of O.T. Wizard and has a financial interest in the platform. All analyses were conducted on de-identified clinical data collected in routine practice. Statistical analyses were generated with AI-based analytical tools and reviewed by the author for clinical accuracy and numerical consistency. Final responsibility for interpretation and reporting rests with the author. Data collection is ongoing. All data is de-identified in accordance with HIPAA regulations.

    References

    Daly, C. J., Kelley, G. T., & Krauss, A. (2003). Relationship between visual-motor integration and handwriting skills of children in kindergarten: A modified replication study. American Journal of Occupational Therapy, 57(4), 459-462. https://doi.org/10.5014/ajot.57.4.459

    Duff, S. V., Chow, S. M., & Henderson, S. E. (2015). Developmental coordination disorder and its consequences for children. In A. F. Farrow & P. J. Tremblay (Eds.), Pediatric rehabilitation: Principles and practice (5th ed., pp. 189-215). Demos Medical Publishing.

    Wick, S. S. (2026, February). O.T. intervention across nine functional domains in preschool children. O.T. Wizard Clinical Research Series. https://blog.otwizard.com/o-t-intervention-across-nine-functional-domains-in-preschool-children/

    Zwicker, J. G., Missiuna, C., Harris, S. R., & Boyd, L. A. (2012). Developmental coordination disorder: A review and update. European Journal of Paediatric Neurology, 16(6), 573-581. https://doi.org/10.1016/j.ejpn.2012.05.003Bailey, D. H., Duncan, G. J., Cunha, F., Foorman, B. R., & Yeager, D. S. (2024). Persistence and fadeout of educational-intervention effects: Mechanisms and potential solutions. Psychological Science in the Public Interest, 21(2), 55-116.Gerde, H. K., Zhao, Y., Shu, L., & Gagne, J. R. (2025). Evidence-based instructional support for early writing in preschool and kindergarten: A scoping review. Reading and Writing. https://doi.org/10.1007/s11145-025-10751-8Magnuson, K., & Duncan, G. J. (2016). Can early childhood interventions decrease inequality of economic opportunity? RSF: The Russell Sage Foundation Journal of the Social Sciences, 2(2), 123-141.

    About O.T. Wizard

    O.T. Wizard is a clinical intelligence system for pediatric occupational therapy professionals. The platform evaluates performance in evaluations, forms, and daily treatment notes across twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, executive functioning, activities of daily living, and participation. O.T. Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update continuously as the clinical database expands. For information about O.T. Wizard research or accessing the platform, visit otwizard.com.

  • O.T. Intervention Across Nine Functional Domains in Preschool Children

    O.T. Intervention Across Nine Functional Domains in Preschool Children

    Pencils, Scissors, and Getting Dressed: Measuring What OT Actually Changes Across Nine Functional Domains in Preschool Children

    O.T. Wizard Clinical Research Series

    By Stephanie Seymore Wick MSOT, OT/L | Occupational Therapist & Clinical Architect, O.T. Wizard

    Introduction

    Ask any pediatric occupational therapist whether their work makes a difference, and the answer is immediate and unequivocal. Ask them to produce the numbers that prove it, and the conversation becomes considerably more complicated. OT practice has long operated in a documentation environment built around goals, progress notes, and clinical narratives, all of which are essential but none of which readily yield the kind of quantitative, domain-level outcome data that payers, school teams, and researchers increasingly require.

    This study represents an effort to change that. Using O.T. Wizard, a clinical intelligence system designed to generate structured, reproducible, multi-domain assessment data for pediatric OT practice, we examined performance change across nine functional domains and fourteen subdomains in 44 preschool-aged children who completed two full evaluations approximately five months apart.

    A central methodological commitment in this report is transparency about what the observed gains actually represent. All percentage changes are gains from the child’s own baseline score. To help readers interpret these figures, we apply a straightforward age-based natural growth correction throughout: children in this cohort averaged 54.2 months at first evaluation and were reassessed approximately 4.7 months later. On a well-constructed developmental scale, maturation alone would be expected to produce a gain proportional to that age progression, approximately 8.7 percent (4.7 divided by 54.2). Gains above that threshold represent performance beyond what developmental maturation alone would predict. Every table in this report includes both the total observed gain and the gain above this expected growth rate, allowing readers to evaluate the data with appropriate context.

    Study Sample

    Age Distribution

    The longitudinal cohort consisted of 44 preschool-aged children, each with two complete O.T. Wizard evaluations separated by a minimum of 28 days. Three additional pairs were excluded due to inter-evaluation intervals of fewer than 28 days. The mean inter-evaluation interval was 141.7 days, approximately 4.7 months, with a range of 56 to 169 days. Mean age at first evaluation was 54.2 months, range 44 to 60 months. Mean age at second evaluation was 58.9 months.

    Starting age band distribution: Band G (36 to 47.99 months, n=2), Band H (48 to 53.99 months, n=18), Band I (54 to 59.99 months, n=22), and Band J (60 to 65.99 months, n=2). Bands H and I together accounted for 91 percent of the sample.

    Demographics and Clinical Status

    The sample was evenly distributed by gender with 22 females and 22 males. All assessments were conducted in North Carolina through O.T. Wizard’s clinical network. Consistent with the broader platform dataset, at least 90 percent of children qualified for Medicaid. The primary referring diagnosis was Specific Developmental Disorder of Motor Function (ICD-10: F82). All children had been recommended for occupational therapy services following developmental screening failure.

    A critical methodological strength of this dataset is rater consistency: 43 of 44 pairs, or 98 percent, were evaluated by the same therapist at both time points, using identical items, response scales, and scoring weights. This same-instrument, same-rater design substantially reduces the risk that observed score changes reflect measurement variability rather than true functional change.

    Natural Growth Framework

    Expected natural growth over the 4.7-month study interval: 8.7% (calculated as 4.7 months / 54.2 months mean age at Evaluation 1)

    Before presenting findings, it is important to establish what we would expect to observe even without intervention. A child who matures from 54.2 months to 58.9 months of age has progressed approximately 8.7 percent through the developmental continuum captured by the instrument. This expected gain is attributable to natural maturation and applies regardless of clinical status. It requires no external literature to defend: it is a direct arithmetic relationship between age progression and scale progression on a fixed instrument.

    This 8.7 percent figure serves as the reference threshold throughout this report. Domain gains below 8.7 percent suggest development did not keep pace with chronological age progression. Gains equal to 8.7 percent suggest maturation-equivalent growth. Gains above 8.7 percent represent functional improvement beyond what age progression alone would predict, and are therefore the most meaningful indicator of intervention impact. The Gain Above Expected column in each table makes this comparison explicit for every domain and subdomain assessed.

    Research Hypotheses

    Three primary hypotheses guided this analysis. First, children receiving occupational therapy services would demonstrate composite score gains substantially exceeding the 8.7 percent expected natural growth threshold. Second, domains most directly targeted by OT in preschool settings, specifically visual motor integration, fine motor skills, visual perception, and activities of daily living, would show the largest gains above expected growth. Third, Band H children (48 to 53.99 months) would demonstrate greater gains than Band I children, reflecting greater developmental sensitivity at a younger starting point.

    Brief Literature Review and Context

    The preschool years represent a critical window for fine motor and visual motor development. Prerequisite skills for handwriting, scissor use, and self-care begin consolidating between ages three and five, with neuromotor pathways underlying pencil control and bilateral coordination undergoing rapid maturation during this period. Interventions delivered during this window have the potential to alter developmental trajectories in ways that become substantially more difficult to achieve once children enter formal schooling.

    Visual motor integration has been identified as one of the strongest predictors of kindergarten handwriting readiness. Daly and colleagues (2003) found that VMI performance at preschool age predicted handwriting speed and legibility at ages six to seven, with effect sizes exceeding those of fine motor or visual perception measures alone. Duff and colleagues (2015) demonstrated that children with developmental coordination disorder who receive targeted fine motor intervention during the preschool years show significantly better handwriting outcomes at school entry than matched peers without services, underscoring the importance of early, precisely measured intervention.

    The measurement infrastructure required to track these outcomes longitudinally has historically been a limiting factor in OT outcomes research. O.T. Wizard was designed to address this gap, enabling structured, reproducible multi-domain measurement within the constraints of a standard clinical evaluation and supporting the kind of longitudinal outcome tracking that traditional paper-based protocols make impractical.

    Key Findings

    Composite Score: Overall Outcome

    Across all 44 children, mean composite score increased from 502.9 at Evaluation 1 to 609.7 at Evaluation 2, a mean gain of 106.8 points representing 21.2 percent growth from baseline. Against the expected natural growth rate of 8.7 percent, this represents a gain of 12.5 percent above what maturation alone would predict. The difference was highly statistically significant (paired t-test, p < 0.001, Cohen’s d = 0.98). Thirty-six of 44 children (82 percent) showed improvement at the second evaluation.

    Total composite gain: +21.2%  |  Expected natural growth: +8.7%  |  Gain above expected: +12.5%

    Domain-Level Findings

    The following table presents results across all seven skill-based assessed domains. Participation and Executive Functioning are addressed separately below, as their measurement properties during first-time evaluations warrant distinct interpretive considerations.

    DomainnEval 1Eval 2Total GainGain Above Expected (8.7%)p-valueCohen’s d
    Visual Motor Integration4336.061.3+70.1%+61.4%< 0.0011.38 (Large)
    Activities of Daily Living3744.464.2+44.8%+36.1%< 0.0011.17 (Large)
    Fine Motor Skills3945.658.8+28.8%+20.1%< 0.0010.85 (Large)
    Gross Motor Skills4355.972.6+30.0%+21.3%< 0.0010.82 (Large)
    Visual Perception4161.472.9+18.8%+10.1%< 0.0010.69 (Medium)
    Praxis4451.751.9+0.4%-8.3%0.9640.01 (Negligible)

    Table 1. Domain-level longitudinal comparison. Scores are percentage of maximum possible. Gain Above Expected subtracts the 8.7% natural growth threshold from total observed gain. Shaded rows reached statistical significance. All scores are pre-Rasch raw percentage values. Participation and Executive Functioning appear in dedicated sections below.

    Visual Motor Integration produced the largest gain in the dataset, rising from 36.0 to 61.3, a total gain of 70.1 percent from baseline representing 61.4 percent above expected natural growth. Children who began with VMI performance at barely more than one-third of expected capacity exited the study interval having crossed the functional midpoint of the scale. Activities of Daily Living showed 36.1 percent above expected growth (d=1.17). Fine Motor Skills and Gross Motor Skills each exceeded 20 percent above expected, with large effect sizes. Visual Perception showed 10.1 percent above expected growth with a medium-to-large effect.

    Praxis fell below the expected natural growth threshold with a negative Gain Above Expected value. This is interpreted as a measurement sensitivity question as much as a treatment response question: Praxis is a complex, context-dependent construct that may require longer intervention timelines, different assessment approaches, or Rasch-calibrated items to detect meaningful change over five months.

    Subdomain-Level Findings

    Subdomain-level analysis reveals where within each domain gains were concentrated and adds clinical texture to the domain-level picture.

    SubdomainnEval 1Eval 2Total GainGain Above Expected (8.7%)p-valueCohen’s d
    WAND-PreK Handwriting4037.363.5+70.1%+61.4%< 0.0011.18 (Large)
    Dressing4344.863.1+40.8%+32.1%< 0.0011.06 (Large)
    Hand Use4453.168.8+29.6%+20.9%< 0.0011.03 (Large)
    Trunk Stability4456.173.2+30.4%+21.7%< 0.0010.84 (Large)
    Scissor Use4436.059.5+65.5%+56.8%< 0.0010.84 (Large)
    Complex Visual Motor Representation4430.952.3+69.1%+60.4%< 0.0010.80 (Large)
    Bilateral Integration4349.663.6+28.1%+19.4%< 0.0010.71 (Medium)
    Visual Discrimination4466.680.6+21.0%+12.3%< 0.0010.59 (Medium)
    Visual Figure Ground4369.382.6+19.1%+10.4%0.0060.44 (Small)
    Visual Memory4256.165.5+16.8%+8.1%0.0540.31 (Small, ns)
    Visual Spatial Relations4253.161.6+16.1%+7.4%0.1430.23 (Small, ns)
    Sequencing Praxis4451.751.9+0.4%-8.3%0.9640.01 (Negligible)

    Table 2. Subdomain-level longitudinal comparison. Sorted by effect size. Gain Above Expected subtracts the 8.7% natural growth threshold. Shaded rows reached statistical significance (p < 0.05). All scores are pre-Rasch raw percentage values.

    WAND-PreK Handwriting showed the largest subdomain effect size (d=1.18, p<0.001), with 61.4 percent above expected growth. Children who began with pre-writing performance at 37.3 percent of expected reached 63.5 percent after approximately five months of OT services. Scissor Use showed 56.8 percent above expected growth, rising from a mean of 36.0 to 59.5. Complex Visual Motor Representation showed 60.4 percent above expected. These three findings converge on a single clinical message: the VMI domain gain is not abstract. It is translating directly into the functional pre-academic tasks that define preschool readiness.

    Hand Use (20.9% above expected, d=1.03), Dressing (32.1% above expected, d=1.06), and Trunk Stability (21.7% above expected, d=0.84) all showed large effect sizes with direct functional implications for classroom participation and daily independence. Bilateral Integration showed 19.4 percent above expected growth, consistent with bilateral coordination emerging as a measurable downstream benefit of fine motor and VMI development.

    Visual Memory and Visual Spatial Relations showed gains of 8.1 and 7.4 percent above expected growth respectively, neither reaching statistical significance. These findings are consistent with prior O.T. Wizard analyses suggesting these constructs require longer measurement intervals or refined item calibration.

    Performance by Starting Age Band

    Starting BandnEval 1 MeanEval 2 MeanPoint ChangeComposite % Gain
    G (36-47.99 mo)2408.0421.0+13.0+3.2%
    H (48-53.99 mo)18461.4591.2+129.7+28.1%
    I (54-59.99 mo)22531.2639.1+108.0+20.3%
    J (60-65.99 mo)2659.5642.0-17.5-2.7%

    Table 3. Composite score change by starting age band. Bands G and J should be interpreted with caution due to small sample sizes (n=2 each).

    Band H children (ages 48 to 53.99 months) showed the largest absolute gains, averaging 129.7 points or 28.1 percent composite growth, more than three times the 8.7 percent natural growth expectation. Band I children averaged 20.3 percent composite growth, also substantially above expectation. These findings support Hypothesis 3 and suggest that the 48 to 54 month window may represent a particularly high-yield period for OT service delivery in this population.

    Correlation Analysis

    A significant negative correlation was observed between starting composite score and magnitude of change (r=-0.32, p=0.032), consistent with a regression-to-the-mean effect. Children who began with lower scores tended to show larger gains. This is expected in any clinical sample and does not invalidate the findings, but it is an important consideration when interpreting the highest-gain domains, which also tended to have the lowest Evaluation 1 starting scores. No significant correlation was found between inter-evaluation interval length and change score (r=0.05, p=0.742), nor between starting age and change score (r=-0.10, p=0.518).

    A Closer Look: Bilateral Fine Motor Speed and Hand Dominance Development

    The Fine Motor Speed and Accuracy task, referred to in O.T. Wizard, as the FMSAT (Fine Motor Speed and Accuracy) “Bubble Popping” assessment, requires children to use a pencil to puncture as many small circles as possible in 30 seconds, first with their preferred hand and then with the other hand. The two bubble counts are recorded separately, capturing not just fine motor speed in isolation but the functional relationship between dominant and non-dominant hand performance.

    This bilateral structure means the Speed and Accuracy results cannot be interpreted as a simple pre-post longitudinal measure in the same way as other subdomains. Instead, the raw hand-level bubble counts offer a window into two parallel developmental processes: absolute fine motor speed improvement in each hand, and the widening of the performance gap between hands as hand dominance consolidates.

    MeasurenEval 1Eval 2Change% Changep-valueCohen’s d
    Hand 1 (Dominant) – bubbles popped3915.318.8+3.4+22.4%0.0020.52 (Medium)
    Hand 2 (Non-dominant) – bubbles popped3910.412.4+2.0+18.9%0.1000.27 (Trending)
    Dominance gap (H1 minus H2)394.906.36+1.460.3240.16 (Hypothesis-generating)

    Table 4. Bilateral fine motor speed longitudinal analysis. Bubble counts reflect circles punctured in 30 seconds per hand. Dominance gap = Hand 1 minus Hand 2. Natural growth expectation: 8.7%.

    The dominant hand showed a statistically significant gain of 3.44 bubbles (p=0.002, d=0.52), representing 22.4 percent improvement from baseline, well above the 8.7 percent natural growth expectation. The non-dominant hand showed a trending gain of 1.97 bubbles (p=0.10, d=0.27), representing 18.9 percent improvement, also above the natural growth threshold. Both hands therefore showed above-expected growth, with the dominant hand improving at a meaningfully faster rate.

    The performance gap between hands widened directionally from 4.90 bubbles at Evaluation 1 to 6.36 bubbles at Evaluation 2, a gap change of 1.46 bubbles. This did not reach statistical significance at n=39 (p=0.324), which is expected: the O.T. Wizard hand dominance study required 348 children to detect a gap change of 0.73 bubbles over six months. At n=39, detecting a 1.46 bubble gap change requires larger samples than this longitudinal cohort provides. The finding is directionally consistent and hypothesis-generating rather than confirmatory.

    Contextualizing these findings against the O.T. Wizard hand dominance study adds clinical depth. That analysis of 459 children found that children with established hand dominance showed a mean inter-hand gap of 4.55 bubbles, while children with no clear dominance showed only a 0.73 bubble gap. The current longitudinal cohort entered with a gap of 4.90 bubbles, already at the established-dominance level, and exited with a gap of 6.36 bubbles. This suggests that OT services may be supporting continued hand specialization beyond initial dominance establishment, driving the dominant hand to further advantage through targeted fine motor programming. Confirmation requires larger samples but the pattern is coherent with motor learning theory and with the hand dominance findings published separately in this research series.

    Participation: Domain-Level Engagement Ratings

    O.T. Wizard captures participation separately from skill performance through nine domain-specific engagement ratings, each scored on a five-point scale from No Engagement to Excellent Engagement. These ratings ask the evaluating therapist to rate how the child engaged during that domain’s evaluation tasks, creating a domain-matched engagement record that parallels the skill score data.

    Rather than treating participation as a single aggregated outcome domain comparable to VMI or Fine Motor Skills, this report presents participation data as a construct validity and clinical context measure. The central question is whether children who perform better in a given skill domain also engage more fully during that domain’s assessment tasks. The answer, consistently across all four matched domains analyzed, is yes.

    Skill DomainParticipation Matched RatingPearson rSpearman rhop-valuen
    Visual Motor IntegrationPART_VM0.3170.301< 0.001465
    Visual PerceptionPART_VP0.4810.464< 0.001444
    Activities of Daily LivingPART_ADL0.4970.506< 0.001445
    Gross Motor SkillsPART_GM0.4650.459< 0.001468

    Table 5. Domain-matched skill score versus domain-specific participation rating correlations. All correlations significant at p < 0.001.

    The overall composite skill score correlates with composite participation ratings at r=0.762 (p<0.001, n=469), a strong relationship that confirms participation ratings are not arbitrary. Children with higher functional skill levels are rated as more engaged during assessment tasks, and this relationship holds across domains. For VMI specifically, children in the lowest skill tertile received a mean VM participation rating of 3.42, while children in the highest skill tertile averaged 4.06. The proportion rated Good or Excellent engagement rose from 46.7 percent in the lowest VMI group to 82.9 percent in the highest.

    These correlations serve as an important construct validity signal for the platform. A child’s competence in a domain predicts their engagement during that domain’s evaluation tasks, which is precisely what developmental and occupational therapy theory would predict. Skill and participation are not independent; they reinforce each other, and O.T. Wizard’s measurement structure captures that relationship.

    The Novelty Effect Hypothesis and Longitudinal Participation

    The longitudinal cohort of 44 children showed a composite participation gain from 61.0 to 63.1 over five months, a change that was not statistically significant (p=0.394, d=0.13). This result warrants careful interpretation rather than the conclusion that participation did not improve with OT services.

    A clinically grounded alternative explanation is the novelty effect. At Evaluation 1, the child is meeting the therapist for the first time. The environment is new, the tasks are novel, and the structured one-on-one interaction may naturally elicit high cooperative behavior. Participation ratings at Eval 1 may therefore reflect novelty-driven engagement rather than the child’s true baseline functional participation. By Evaluation 2, the therapeutic relationship is established, the child is comfortable enough to reveal authentic behavioral and engagement patterns, and the therapist has sufficient rapport to observe the child’s genuine participation profile rather than their best performance.

    If this hypothesis holds, Evaluation 1 participation ratings are systematically inflated relative to what they would show in a familiar context, meaning the scale at Evaluation 2 is measuring a meaningfully different construct than at Evaluation 1. This measurement context shift would explain the apparent lack of longitudinal gain without implying that participation failed to respond to intervention. It also opens a question with direct clinical and psychometric implications: are participation ratings most informative when collected after the therapeutic relationship is established, and should baseline participation norms account for the evaluative context?

    This remains a hypothesis requiring empirical testing with larger longitudinal samples. It is noted here as both a study limitation and a direction for continuing research.

    Executive Functioning: Work Habits Observed During Evaluation

    The Executive Functioning domain in O.T. Wizard is assessed through the Work Habits in a 1:1 Therapy Setting subdomain, comprising eight items: Attention, Cooperation, Task Initiation, Task Persistence, Task Completion, Transition, Impulse Control, and Carryover of Skills. Each item is rated on a five-point scale from Not Present/Unable to Proficient, with ratings reflecting the child’s behavior as observed during the evaluation itself.

    Across the full cross-sectional sample of 471 children, Work Habits items showed mean ratings ranging from 3.32 for Attention to 4.10 for Task Completion, indicating that the majority of children in this clinical population demonstrated adequate to proficient work habits during the evaluation. The EF Work Habits composite score correlates with overall composite skill performance at r=0.759 (p<0.001), a relationship comparable in strength to the participation-skill correlation. Higher-functioning children, as measured by composite skill scores, are observed to demonstrate better work habits during assessment.

    Item-level correlations with specific skill domains add clinical texture. Attention correlates with VMI performance at r=0.324 and with Fine Motor at r=0.413. Task Initiation shows r=0.351 with VMI and r=0.456 with Fine Motor. Task Persistence correlates at r=0.447 with Fine Motor. These moderate relationships are consistent with the theoretical link between executive function and fine motor learning: children who can attend, initiate, and persist through structured tasks acquire fine motor skills more efficiently, and children with stronger fine motor programs may experience less frustration-driven task avoidance.

    The Novelty Effect in Executive Functioning

    The longitudinal cohort showed near-zero change in Executive Functioning from Evaluation 1 to Evaluation 2 (+0.5%, p=0.916, d=0.02). The same novelty effect hypothesis that applies to participation applies here with equal force, and arguably with stronger clinical support.

    A child encountering a new therapist in a structured evaluation setting has strong situational motivators for compliance: the novelty of the interaction, the desire to please an unfamiliar adult, and the absence of habituated behavioral patterns in that specific context. Cooperation, Impulse Control, and Task Persistence at Evaluation 1 may reflect the child’s best-case executive behavior under novel conditions rather than their typical regulatory profile. By Evaluation 2, these situational scaffolds have diminished. The child knows the therapist, has formed expectations about the session, and is more likely to exhibit their authentic regulatory patterns, including the attentional variability, transition difficulty, or impulse control challenges that characterize their daily functioning.

    If Evaluation 1 Work Habits ratings are inflated by novelty effects and Evaluation 2 ratings reflect more authentic behavioral observation, then the absence of longitudinal gain is not evidence that OT failed to improve executive functioning. It may instead reflect the instrument now measuring what it was designed to measure. This interpretation has a meaningful implication for clinical documentation: therapist-observed Work Habits ratings collected after the therapeutic relationship is established may provide more valid clinical and baseline data than those collected at first encounter.

    As with participation, this hypothesis requires prospective testing with larger samples and ideally with parent-reported or teacher-reported behavioral measures to establish convergent validity across contexts. It is presented here as a clinically grounded interpretive framework and a research priority.

    Implications for OT Practice

    The domain-level findings, interpreted through the natural growth framework, allow occupational therapists to make specific, evidence-informed decisions about evaluation and intervention priorities. Five domains showed gains ranging from 10.1 to 61.4 percent above the 8.7 percent natural growth rate, all with statistical significance and large or medium effect sizes. These gains were achieved in a predominantly Medicaid-qualifying, low-income clinical population with significant developmental concerns, which makes the magnitude of change all the more clinically meaningful.

    The VMI finding demands particular attention in practice planning. A gain of 61.4 percent above expected natural growth indicates that children receiving OT services made VMI gains at roughly eight times the rate that maturation alone would predict over the same interval. Given VMI’s established role as a predictor of kindergarten handwriting readiness, prioritizing VMI-targeted activities including complex copying tasks, directed drawing, and structured pre-writing programs is strongly supported by this data.

    For insurance authorization and educational planning, the natural growth framework provides a uniquely effective communication tool. Rather than reporting a raw score gain, the practitioner can state that the child’s VMI performance exceeded the expected developmental rate by 61 percentage points over five months, providing clear evidence that skilled OT intervention, not maturation, drove the observed change. This framing is defensible, transparent in its methodology, and directly responsive to the kind of medical necessity standard that payers apply.

    The bilateral fine motor speed findings support continued OT intervention targeting hand specialization in children whose dominance is still consolidating. The dominant hand’s significant gain above natural growth, combined with the directional widening of the inter-hand gap, is consistent with OT services accelerating hand specialization. For children presenting with inter-hand gaps below two bubbles at preschool age, more intensive hand-specific programming may be clinically warranted.

    The participation and work habits findings carry a practical message for therapist documentation practice. Domain-specific participation ratings that are collected after the therapeutic relationship is established are likely to be more clinically informative than those captured at first evaluation. Therapists should be aware that novelty effects at initial evaluation may produce elevated engagement and compliance ratings that do not generalize to the child’s authentic daily functioning profile.

    Implications for Intervention Planning

    The convergent gains across WAND-PreK Handwriting, Scissor Use, Complex VMI, Hand Use, Bilateral Integration, and Trunk Stability point toward a functional fine motor and VMI cluster that is highly responsive to structured OT programming in this developmental window. Gains in this cluster ranged from 19.4 to 61.4 percent above expected growth. Interventions that integrate tool use, bilateral coordination, and visual guidance of hand movements across this cluster are supported by both the data pattern and established OT theory.

    The Trunk Stability findings reinforce a foundational clinical principle: proximal postural control precedes and supports distal fine motor development. Children who gained in trunk stability tended to gain across the fine motor and VMI cluster. For children showing limited fine motor or VMI gains, postural foundation assessment and intervention should be considered before assuming the upper extremity is the primary limiting factor.

    Scissor use showed 56.8 percent above expected growth with a large effect size. Scissors require bilateral coordination, hand differentiation, VMI guidance, and sustained attention simultaneously, making scissor-based activities a naturally integrative target across multiple domains. The large effect size and high gain above expected growth suggest this subdomain is both highly responsive to OT intervention and highly sensitive to measurement, making it a valuable progress-monitoring target.

    For Praxis, intervention targeting should not be abandoned based on the longitudinal findings reported here. The near-zero gain more likely reflects current measurement limitations than true insensitivity to OT services. As O.T. Wizard’s item bank is refined through Rasch calibration, this domain may demonstrate the sensitivity needed to capture incremental change that OT practitioners observe clinically.

    The domain-matched participation correlations support incorporating participation-focused goal setting alongside skill-based goals, particularly for children showing low engagement in specific domains. A child rated at Minimal Engagement in VMI activities is likely showing competence-related avoidance rather than willful noncompliance, and addressing the underlying skill deficit is the most direct route to improved participation in that domain.

    Study Limitations

    This study carries several important limitations. The sample is clinical and geographically restricted to North Carolina, and findings cannot be generalized to typically developing children or populations in other regions. The absence of a control group means observed gains cannot be causally attributed to OT intervention.

    The 8.7 percent natural growth correction applied throughout this report is a methodologically transparent and defensible estimate derived directly from the age progression of this cohort on this instrument. It assumes proportional developmental scaling across the score range, an assumption that Rasch calibration will allow us to test empirically. Future work with a typically developing comparison group will allow domain-specific, empirically derived growth expectations to replace this uniform estimate, strengthening the interpretive framework considerably.

    The regression-to-the-mean effect (r=-0.32) means children with lower starting scores showed larger gains on average. Visual Motor Integration, WAND-PreK Handwriting, Scissor Use, and Complex VMI all had the lowest Evaluation 1 scores and the largest gains above expected growth. The true intervention effect within these domains is likely substantial, but the proportion attributable to treatment versus regression toward the mean cannot be fully separated without a control group.

    Participation and Executive Functioning longitudinal findings are subject to a novelty effect interpretation that cannot be confirmed or ruled out with the current dataset. The possibility that Evaluation 1 ratings in these domains reflect novelty-driven behavior rather than true baseline functioning represents a meaningful threat to the internal validity of longitudinal comparisons in these areas specifically. This does not affect the skill-domain findings but warrants dedicated investigation in future work.

    The bilateral fine motor speed gap-widening finding is hypothesis-generating at n=39 and requires larger samples to reach significance. Bands G and J contained only two children each, precluding any band-specific inference for those age ranges. All scores remain pre-Rasch ordinal percentage values. Statistical analyses were generated with the assistance of an AI-based analytical tool and reviewed by the author for clinical and numerical consistency. Final responsibility for interpretation and reporting rests with the author.

    Continuing Research Needed

    Rasch calibration remains the highest research priority. Transforming ordinal raw scores into interval-level person measures will allow true scale-independent longitudinal comparison, validate the proportional scaling assumption underlying the natural growth correction, and identify items requiring revision. The O.T. Wizard National Try-Out Team initiative is designed to accelerate this process by expanding the sample sizes needed for stable item calibration across all domains and age bands.

    A comparison group of typically developing children will allow the 8.7 percent natural growth estimate to be validated empirically and replaced with domain-specific growth expectations calibrated against external developmental benchmarks. This will strengthen the natural growth framework applied in this report and make the gain-above-expected metric more precise for insurance and educational reporting contexts.

    The novelty effect hypothesis for participation and executive functioning requires prospective investigation. A study design that captures therapist-rated participation and work habits at multiple time points within the first year of services, alongside parent-reported and teacher-reported behavioral measures, would allow empirical testing of whether first-evaluation ratings systematically overestimate engagement and compliance relative to ratings collected after therapeutic relationship establishment. Confirming or disconfirming this hypothesis has implications for how O.T. Wizard instructs therapists to interpret and apply participation and EF data in clinical documentation.

    The domain-matched participation correlations reported here establish a foundation for construct validity research. Future work linking participation ratings to standardized adaptive behavior measures and teacher-rated classroom engagement would extend these findings and position the participation ratings as clinically actionable data rather than descriptive context.

    The bilateral fine motor speed findings warrant dedicated follow-up with larger samples. Testing whether inter-hand gap widening reaches significance at n=100 or greater, whether gap widening correlates with therapist-documented hand dominance establishment, and whether children who begin with gaps below two bubbles show different intervention trajectories would substantially extend the hand dominance research program established in the O.T. Wizard FMSAT study.

    Linking O.T. Wizard performance data to standardized criterion measures including the Beery VMI, BOT-2, and teacher-rated school readiness indicators would establish concurrent and predictive validity, and would allow the natural growth framework applied here to be calibrated against externally validated developmental benchmarks.

    Conclusion

    This analysis of 44 preschool-aged children with paired O.T. Wizard evaluations, examined through a transparent natural growth framework, provides the most complete domain-level longitudinal outcome picture generated by the platform to date. The expected natural growth rate of 8.7 percent over the 4.7-month study interval serves as a consistent interpretive reference throughout, allowing readers to distinguish maturation from intervention-attributable change across every domain reported.

    The findings are striking. Five of seven skill domains showed gains substantially exceeding the 8.7 percent natural growth threshold, with effect sizes ranging from medium to large. Visual Motor Integration showed gains 61.4 percent above expected. WAND-PreK Handwriting showed gains 61.4 percent above expected. Scissor Use showed gains 56.8 percent above expected. These are not marginal differences from expectation. They represent functional gains at rates six to eight times what chronological maturation alone would produce over the same interval.

    The domain-matched participation analysis adds a construct validity dimension to these findings. Children who perform better in a skill domain engage more fully during that domain’s assessment, with domain-matched correlations ranging from r=0.317 to r=0.497 and an overall composite skill-participation correlation of r=0.762. Participation and executive functioning longitudinal data are reframed here through the novelty effect hypothesis, which proposes that first-evaluation ratings in these behavioral domains may reflect situational compliance rather than authentic baseline functioning, with more valid observational data emerging after the therapeutic relationship is established.

    The FMSAT (“Bubble Popping”) extends these findings into hand dominance development, showing that the dominant hand improved significantly above the natural growth rate and that the inter-hand performance gap widened directionally, consistent with progressive hand specialization during OT services.

    These findings are preliminary. They require replication with larger samples, validated comparison conditions, and Rasch-calibrated measurement. What they establish is that structured domain-level digital assessment in pediatric OT can generate longitudinal outcome data that clearly and transparently distinguishes intervention-driven change from natural maturation. For a profession that has historically struggled to quantify its impact in terms that payers and educational systems recognize, that distinction is not a minor technical refinement. It is the foundation of evidence-based practice. O.T. Wizard is building that foundation, one evaluation at a time.

    Disclosures

    The author is the founder of O.T. Wizard and has a financial interest in the platform. All analyses were conducted on de-identified clinical data collected in routine practice.

    References

    Daly, C. J., Kelley, G. T., & Krauss, A. (2003). Relationship between visual-motor integration and handwriting skills of children in kindergarten: A modified replication study. American Journal of Occupational Therapy, 57(4), 459-462. https://doi.org/10.5014/ajot.57.4.459

    Duff, S. V., Chow, S. M., & Henderson, S. E. (2015). Developmental coordination disorder and its consequences for children. In A. F. Farrow & P. J. Tremblay (Eds.), Pediatric rehabilitation: Principles and practice (5th ed., pp. 189-215). Demos Medical Publishing.

    Zwicker, J. G., Missiuna, C., Harris, S. R., & Boyd, L. A. (2012). Developmental coordination disorder: A review and update. European Journal of Paediatric Neurology, 16(6), 573-581. https://doi.org/10.1016/j.ejpn.2012.05.005

    Polatajko, H. J., & Cantin, N. (2010). Exploring the effectiveness of occupational therapy interventions, other than the sensory integration approach, with children and adolescents experiencing difficulty processing and integrating sensory information. American Journal of Occupational Therapy, 64(3), 415-429. https://doi.org/10.5014/ajot.2010.09073

    About O.T. Wizard

    O.T. Wizard is a clinical intelligence system for pediatric occupational therapy professionals.  The platform evaluates performance in evaluations, forms, and daily treatment notes across twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, executive functioning, activities of daily living, and participation. OT Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update continuously as the clinical database expands.

    Data collection is ongoing. All data is de-identified in accordance with HIPAA and FERPA regulations.For information  about O.T. Wizard research or accessing the platform, visit otwizard.com

  • What Really Predicts Handwriting Success

    What Really Predicts Handwriting Success

    THE CLINICAL PUZZLE

    Every pediatric occupational therapist has encountered this scenario: A 4-year-old with excellent fine motor skills, good visual perception scores, and established hand dominance still cannot write letters legibly. Meanwhile, another child with weaker motor skills and inconsistent grip produces surprisingly readable work.

    What makes the difference?

    New data from 185 preschool-age children reveals why handwriting success is so unpredictable and why our traditional assessment approaches may be missing critical pieces of the puzzle.

    CURRENT HANDWRITING ASSESSMENT PRACTICES

    Occupational therapists typically evaluate handwriting readiness through standardized assessments focusing on visual-motor integration and fine motor skills:

    Beery VMI (Visual-Motor Integration), 6th Edition measures the ability to copy geometric forms of increasing complexity. Children progress from simple lines to complex shapes, with performance compared to age-based norms. The assessment assumes that shape copying ability predicts letter formation success.

    PDMS-3 (Peabody Developmental Motor Scales, 3rd Edition) assesses fine and gross motor development through grasping and visual-motor integration subtests. The fine motor composite includes tasks similar to letter copying and provides age-based standard scores. While more comprehensive than the Beery VMI alone, it focuses primarily on motor execution.

    BOT-2 (Bruininks-Oseretsky Test of Motor Proficiency, 2nd Edition) evaluates fine and gross motor proficiency including precision, integration, and manual dexterity tasks. Many subtests emphasize speed and accuracy under timed conditions, making it useful for identifying motor delays but less specific to handwriting readiness.

    The Print Tool evaluates actual letter and number formation in children ages 3 to 7, rating legibility, size, spacing, and alignment. While more functional than shape copying, it requires children to already have some writing exposure.

    Developmental Test of Visual Perception (DTVP-3) assesses visual-perceptual and visual-motor skills through tasks including copying, form constancy, and figure-ground discrimination. Performance on these isolated visual tasks is presumed to indicate readiness for integrated writing tasks.

    Minnesota Handwriting Assessment evaluates speed, legibility, and form in school-age children who already write, making it less useful for identifying preschool readiness factors.

    THE RESEARCH

    We analyzed 185 children ages 4 to 4.5 years who received occupational therapy evaluations in North Carolina. This clinical sample consisted of children referred for developmental concerns, with 95 percent qualifying for Medicaid services. Many had limited exposure to structured preschool settings.

    The children were given a comprehensive evaluation using O.T. Wizard and included 8-10 domains per child. During evaluation, children completed a letter copying task: 10 uppercase letters arranged from developmentally simple (L, F, R) to complex (S, X, N). Children copied each letter into a defined box below the model. Occupational therapists rated both the quality of letter production and the child’s behavior during the task.

    The use of uppercase letter copying rather than geometric shapes in preschool assessment warrants clarification. For children lacking letter recognition, uppercase letters serve as geometric forms with the added benefit of providing functional, longitudinal work samples. Unlike abstract shapes that become irrelevant once writing instruction begins, letter samples document the progression from letters-as-shapes to letters-as-symbols, capturing both motor and cognitive development across the transition to formal writing.

    We then examined how well various factors predicted performance on this functional handwriting task. Rather than assuming certain skills matter most, we calculated correlations to let the data reveal which factors actually related to success.

    UNDERSTANDING CORRELATION: THE “r” VALUE

    Before presenting findings, it helps to understand what correlation means and how to interpret the numbers.

    Correlation measures the strength of the relationship between two variables. The correlation coefficient, represented as r, ranges from 0 to 1.0:

    r = 0.0 to 0.1: No meaningful relationship

    r = 0.1 to 0.3: Weak relationship 

    r = 0.3 to 0.5: Moderate relationship

    r = 0.5 to 0.7: Strong relationship 

    r = 0.7 to 1.0: Very strong relationship

    A simple example: Height and shoe size have a strong correlation (r = approximately 0.7). Taller people tend to wear larger shoes, though exceptions exist. The relationship is strong but not perfect.

    In contrast, height and intelligence have essentially no correlation (r = approximately 0.0). Knowing someone’s height tells you nothing about their cognitive ability.

    For our study, correlation indicates how well each skill predicts letter copying success. A high correlation means children with strong skills in that area tend to perform better on writing tasks. A low correlation means the skill does not reliably predict writing performance.

    THE FINDINGS

    Six factors showed moderate correlations with handwriting (visual motor integration) performance, all clustering tightly between r = 0.31 and r = 0.39:

    Fine Motor Skills: r = 0.393 

    Cooperation (during evaluation): r = 0.365 

    Visual Perception: r = 0.340 

    Attention (during evaluation):r = 0.314

    Praxis (Motor Planning): r = 0.314 

    Participation (during writing task): r = 0.310

    The most striking finding is not which factor ranked highest, but rather that all six fell within an 8-point range. Fine Motor scored highest at 0.393, but Participation scored 0.310, a difference of only 0.083.

    Statistical interpretation: All six predictors are moderate in strength, and none dominates. The child with the highest fine motor score has only a slightly better chance of writing success than the child with the highest cooperation score.

    VISUAL PERCEPTION SUBDOMAINS: TASK DEMANDS MATTER

    An interesting pattern emerged when examining visual perception subdomains separately. Not all visual skills predicted copying performance equally:

    Visual Discrimination: r = 0.379 

    Visual Figure Ground: r = 0.304 

    Visual Spatial Relations: r = 0.158 

    Visual Memory: r = 0.137

    Visual Discrimination, the ability to see small differences between similar forms, predicted letter copying better than the overall Visual Perception domain score. This makes perfect sense given the task demands. Copying letters requires discriminating between similar features: Is this a C or an O? Does this letter have a diagonal line or a curve? Are these two vertical lines parallel or converging?

    In contrast, Visual Memory showed the weakest correlation at r = 0.137, barely above no relationship at all. This finding initially seems surprising given that handwriting literature often emphasizes visual memory as critical for letter formation.  However, the weak correlation makes complete sense when we consider the actual task. Children were asked to copy letters with the model remaining visible throughout. They could look back and forth between the stimulus letter and their work as many times as needed. Visual memory is irrelevant when the visual information stays available.

    Visual memory would matter for different handwriting tasks: Writing letters from dictation (hear the letter name, recall what it looks like) Writing spelling words independently (recall the letter in memory) Reproducing letters after brief exposure (look once, then write from memory)

    But for direct copying with continuous visual access to the model, discrimination ability predicts success while memory does not.

    This finding has important implications for assessment practices. If we evaluate visual memory but not visual discrimination, we may be measuring the wrong visual skill for near point copying tasks. Comprehensive assessment requires matching the skills tested to the actual task demands the child will face in the classroom.

    In preschool and early kindergarten, children primarily engage in near point copying: copying letters from a worksheet placed directly in front of them, tracing over models, and reproducing shapes from a stimulus card on the table. These near point tasks allow continuous visual reference, making discrimination critical and memory less important.

    As children progress through elementary school, task demands shift to far point copying: copying from the board, reproducing teacher demonstrations, writing from dictation. These tasks require visual memory because the model is not continuously accessible. A child must look at the board, hold the letter image in memory while looking down at paper, then reproduce from that mental representation.

    For the preschool population in this study engaged in near point copying tasks, visual discrimination predicted success while visual memory did not. This relationship may change for older children performing far point copying or writing from dictation.

    WHAT THE NUMBERS MEAN IN PRACTICE

    Consider what these moderate correlations reveal:

    If fine motor skills were the primary driver of handwriting, we would expect r = 0.6 or higher. Instead, r = 0.393 means fine motor capability explains only about 15 percent of handwriting performance. The remaining 85 percent depends on other factors.

    Similarly, visual perception (r = 0.340) explains about 12 percent. Praxis explains about 10 percent. Cooperation explains about 13 percent.

    No single factor accounts for even 20 percent of performance. Handwriting emerges from complex interactions among multiple systems, not mastery of any single prerequisite.

    THE BEHAVIORAL FACTOR SURPRISE

    Perhaps most notable: Behavioral factors predicted success as well as skill factors.

    Cooperation (r = 0.365) nearly matched fine motor skills (r = 0.393). A child who cooperates with feedback and accepts correction has almost the same probability of writing success as a child with superior hand strength and coordination.

    Attention during evaluation (r = 0.314) predicted exactly as well as motor planning ability (r = 0.314). The child who can focus for the duration of the task performs comparably to the child with better movement sequencing skills.

    Participation during the actual writing task (r = 0.310) predicted nearly as well as any other factor. Willingness to engage with the challenge matters almost as much as capability.

    This explains common clinical observations:

    The child with excellent fine motor skills who gives up after one attempt struggles more than the child with weaker skills who persists through frustration.

    The child who resists feedback and insists on doing it “my way” fails to improve despite adequate motor capability.

    The child who cannot sustain attention long enough to complete three letters never accumulates the practice necessary for skill development.

    CONTEXT MATTERS: THE 1-ON-1 EVALUATION PROBLEM

    An important limitation: Cooperation, attention, and participation were rated during one-on-one evaluation sessions with an occupational therapist providing full support and individualized pacing.

    This context differs dramatically from classroom writing instruction, where:

    One teacher manages 15 to 20 students simultaneously Individual feedback is limited and delayed Pacing is group-determined rather than individualized Distractions are constant Tasks continue for extended periods without breaks

    A child rated as having “adequate cooperation” in a quiet therapy room with undivided therapist attention may demonstrate very different behavior in a busy kindergarten classroom during 15-minute writing periods.

    This suggests our correlations may actually underestimate the importance of behavioral factors. If cooperation, attention, and participation predict success even in optimal conditions, they likely matter even more in typical educational settings.

    IMPLICATIONS FOR ASSESSMENT PRACTICES

    Current handwriting readiness assessments focus heavily on visual-motor integration and fine motor skills while largely ignoring behavioral factors. The Beery VMI, for instance, requires sustained attention and task persistence to complete 30 forms, but these behavioral requirements are not scored or interpreted. A child may fail due to attention limitations rather than visual-motor deficits, yet both receive the same low score.

    More critically, the Beery VMI is frequently used in isolation to qualify children for occupational therapy services for handwriting concerns. Given our findings, this practice is problematic. Visual-perception represents only one of six factors that predict handwriting success in preschoolers, and it predicts moderately (r = 0.340), not strongly. A child may score low on the Beery VMI yet succeed at functional handwriting due to strong cooperation, attention, and participation. Conversely, a child may pass the Beery VMI but struggle with classroom writing due to behavioral regulation challenges that the assessment does not capture.

    Using the Beery VMI as a sole qualifying criterion systematically misidentifies which children need services. Comprehensive evaluation across all six predictive factors provides more accurate identification of handwriting risk.

    Additionally, visual perception assessments for preschool populations should emphasize visual discrimination for near point copying tasks. Our findings demonstrate that for preschoolers copying letters with the model continuously visible, discrimination ability (r = 0.379) predicts substantially better than memory (r = 0.137). This does not suggest visual memory is unimportant for handwriting development overall. Rather, it indicates that the specific skills required depend on task type and developmental stage. Visual memory likely becomes increasingly important as children transition to far point copying and writing from dictation in elementary grades.

    Ratings of current assessments used by OT’s and how they capture handwriting prediction 

    These assessments share common limitations. Based on our findings, we can evaluate how well each captures the six factors that actually predict handwriting success:.

    Beery VMI (with supplemental tests): Rating 5/10 IF subtests administered.  3/10 if only the VMI section is administered.  Captures visual-motor integration (r=0.340) through the primary copying task. Supplemental Visual Perception and Motor Coordination subtests add assessment of visual discrimination and fine motor control, bringing total coverage to 2-3 of 6 predictive factors. However, the Visual Perception subtest does not distinguish between visual discrimination (r=0.379, highly relevant) and visual memory (r=0.137, less relevant for near point copying). Completely misses cooperation, attention, praxis, and task participation. When administered with all three subtests, it provides more comprehensive data than VMI alone, but therapists often use only the primary VMI subtest for qualification decisions.

    PDMS-3: Rating 6/10 Captures fine motor skills (r=0.393) through grasping subtests and visual-motor integration (r=0.340) through copying tasks. Provides 2 of 6 critical factors. Misses cooperation, attention, praxis, and task participation entirely. No assessment of behavioral regulation during tasks or visual discrimination as distinct from visual-motor integration.

    BOT-2: Rating 4/10 Primarily assesses motor proficiency with fine motor precision and integration subtests capturing fine motor skills (r=0.393). However, heavy emphasis on timed performance may penalize slow-but-accurate children. Completely misses visual perception, cooperation, attention, and task participation. Designed for motor proficiency screening rather than handwriting-specific readiness. Captures only 1 of 6 predictive factors.

    DTVP-3: Rating 5/10 Assesses visual perception (r=0.340) across multiple subdomains but does not distinguish between visual discrimination (r=0.379, highly relevant for copying) and visual memory (r=0.137, less relevant for near point tasks). Misses fine motor execution, cooperation, attention, praxis, and task participation. Provides visual skills assessment but in isolation from functional writing context.

    The fundamental issue: These assessments emphasize isolated skill measurement (motor proficiency, visual perception, visual-motor integration) while ignoring behavioral regulation factors that predict equally well. Additionally, they provide scores but often rely on checklist observations that cannot be converted to continuous metrics for tracking progress or comparing across domains.

    Comprehensive assessment should include:

    Fine motor capability: Strength, coordination, precision, tool control Visual-perceptual skills with task-appropriate emphasis: Visual discrimination (critical for copying) Visual figure ground (moderate importance) Visual spatial relations (less critical for copying) Visual memory (only relevant for tasks without visible models) Motor planning: Ability to sequence multi-step actions, organize approach Cooperation: Willingness to accept feedback, modify approach when unsuccessful Attention: Capacity to sustain focus through multi-step tasks Task participation: Engagement level, persistence through challenge, frustration tolerance

    Single-domain screening (testing only visual skills or only motor skills) will systematically miss children at risk. A child may pass fine motor screening with flying colors but struggle with writing due to attention deficits, poor cooperation, or low task engagement.

    Similarly, a child may score well on visual memory subtests but fail at letter copying due to poor visual discrimination. Matching assessed skills to actual task demands is essential.

    Conversely, a child with borderline fine motor scores but strong behavioral regulation may achieve functional writing through persistence and acceptance of instruction.

    IMPLICATIONS FOR INTERVENTION

    Traditional intervention models often follow a sequential approach: establish attention, then build fine motor skills, then introduce visual tasks, then combine into writing. Our data suggests this may be inefficient.

    If multiple factors contribute equally and simultaneously, intervention should address them concurrently rather than sequentially. Children need practice integrating behavioral regulation, motor control, visual processing, and motor planning from the start.

    Effective intervention might include:

    Brief, varied tasks that build attention capacity while practicing motor skills (address both simultaneously) Immediate feedback on both motor execution and behavioral approach (cooperation, persistence) Functional writing activities that require visual processing, motor planning, and sustained attention in authentic context Explicit instruction in self-regulation during challenging tasks (managing frustration, accepting correction)

    Isolated prerequisite activities (strengthening exercises, shape sorting, sequencing games) practiced separately from writing context may not transfer effectively. The child builds attention during tabletop games but cannot apply it during writing. The child demonstrates fine motor control during bead threading but not during letter formation.

    Integration practice appears more efficient: Work on attention, motor control, visual processing, and cooperation simultaneously within functional writing activities.

    WHY SOME CHILDREN SUCCEED DESPITE LIMITATIONS

    These findings explain puzzling clinical observations.

    The child with weak fine motor skills who succeeds likely compensates through: Strong visual perception (carefully observes letter features) High persistence (keeps trying despite motor difficulty) Good cooperation (accepts feedback, modifies approach) Strong attention (focuses carefully on each stroke)

    The combination of strengths in four areas compensates for weakness in one.

    The child with excellent fine motor skills who fails likely struggles with: Poor attention (loses focus mid-letter) Low persistence (gives up when first attempt is imperfect) Resistance to feedback (insists on incorrect approach) Low task engagement (avoids writing activities)

    Motor capability alone cannot overcome behavioral limitations.

    THE CLINICAL SAMPLE CONTEXT

    These findings emerge from a specific population: low-income preschoolers referred for occupational therapy evaluation. Many had limited exposure to structured educational settings or formal writing instruction.

    This context matters for interpretation:

    Children with school experience might show different patterns, as they have had more opportunity to develop attention and cooperation within structured tasks.

    Higher-income samples with more educational exposure might demonstrate stronger correlations for skill factors and weaker correlations for behavioral factors.

    Typically developing children (not referred for therapy) might show different relationships among variables.

    However, this clinical sample represents the population occupational therapists actually serve. Understanding what predicts success in children with developmental concerns and limited educational exposure has direct clinical relevance.

    RESEARCH CONTEXT: HOW OUR FINDINGS COMPARE

    Our findings align with and extend existing research on handwriting development while revealing some important differences.

    Feder and Majnemer (2007) conducted a systematic review identifying visual-motor integration, fine motor skills, and in-hand manipulation as significant predictors of handwriting performance in school-age children. Their meta-analysis found moderate correlations (r = 0.3-0.5) between these factors and handwriting, consistent with our fine motor (r = 0.393) and visual perception (r = 0.340) findings. However, their review focused on older children already engaged in writing instruction, while our sample examined preschoolers in pre-handwriting stages.

    Reference: Feder, K. P., & Majnemer, A. (2007). Handwriting development, competency, and intervention. Developmental Medicine & Child Neurology, 49(4), 312-317.

    Volman, van Schendel, and Jongmans (2006) examined handwriting readiness in kindergarten children and found that visual-motor integration was a significant predictor but explained only a modest portion of variance. This supports our finding that visual-motor skills predict moderately (r = 0.340) but do not dominate. Importantly, they also identified attention and behavioral regulation as contributing factors, aligning with our cooperation (r = 0.365) and attention (r = 0.314) findings.

    Reference: Volman, M. J., van Schendel, B. M., & Jongmans, M. J. (2006). Handwriting difficulties in primary school children: A search for underlying mechanisms. The American Journal of Occupational Therapy, 60(4), 451-460.

    Kaiser, Albaret, and Doudin (2009) investigated the relationship between handwriting quality and various factors in first graders. They found visual perception, fine motor skills, and graphomotor skills all contributed, but no single factor was sufficient. Their findings that multiple factors contribute equally strongly support our multifactorial model. However, they did not examine behavioral factors like cooperation or task-specific participation, which our data suggests are equally important.

    Reference: Kaiser, M. L., Albaret, J. M., & Doudin, P. A. (2009). Relationship between visual-motor integration, eye-hand coordination, and quality of handwriting. Journal of Occupational Therapy, Schools, & Early Intervention, 2(2), 87-95.

    Notably absent from existing literature: Studies examining task-specific participation and cooperation as predictors of handwriting success in preschool populations. Most handwriting research focuses on school-age children who have already received writing instruction and emphasizes motor and perceptual factors while treating behavioral factors as confounding variables rather than legitimate predictors.

    Our finding that cooperation predicts nearly as well as fine motor skills (r = 0.365 vs r = 0.393) extends the literature by demonstrating that behavioral regulation deserves equal consideration in handwriting readiness assessment. The clinical sample context (children referred for evaluation, limited school exposure, 95 percent low-income) may explain why behavioral factors emerged as stronger predictors than in general population studies.

    Additionally, our visual perception subdomain analysis revealing that visual discrimination (r = 0.379) predicts substantially better than visual memory (r = 0.137) for near point copying tasks provides specificity often missing in broader visual perception assessments. This has practical implications for selecting which visual subtests to administer when evaluating preschool handwriting readiness.

    ABOUT OT WIZARD

    Data for this analysis was collected through OT Wizard, a clinical intelligence system for pediatric occupational therapy assessment. The platform evaluates performance across up to twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, executive functioning, activities of daily living, and participation. OT Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update continuously as the clinical database expands.

    Unlike traditional checklist-based assessments, OT Wizard converts all observations to continuous metrics that enable progress tracking, cross-domain comparison, and comprehensive reporting. The platform captures all six factors identified in this research as predictive of handwriting success: fine motor skills, visual perception (with subdomain specificity), praxis, cooperation, attention, and task participation. Behavioral regulation is assessed within the context of actual task performance rather than as an isolated rating, providing clinically relevant data about how attention and cooperation affect functional skill demonstration.

    For handwriting readiness assessment specifically, OT Wizard provides quantified performance across visual discrimination, visual-motor integration, fine motor control, motor planning, and behavioral engagement during writing tasks. This comprehensive approach addresses the multifactorial nature of handwriting development identified in this research. As the platform undergoes Rasch analysis validation and accumulates longitudinal outcome data, it will establish whether comprehensive baseline assessment across all six predictors improves identification of children at risk for handwriting difficulty and informs more effective intervention planning.

    OT Wizard is committed to advancing the occupational therapy profession by collecting de-identified clinical data from real therapist users, building the largest developmental database in pediatric occupational therapy history. This continuous data collection enables research on developmental trends, intervention effectiveness, and response to intervention patterns that elevate practice from perception-based to data-driven decision making and strengthen the evidence base for the entire profession