Tag: assessment

  • Why OTP’s need better assessments-update on the FMSAT

    Why OTP’s need better assessments-update on the FMSAT

    Measuring Neuromotor Lateralization Across the Lifespan: Progress on the FMSAT Norming Project (& Why OTP’s need better assessments)

    How a one-minute screener for hand dominance, modern Rasch psychometrics, and the shift to value-based care are converging to change how occupational therapists, educators, and clinicians measure what we do.

    Published May 24, 2026 by Stephanie Seymore Wick, MSOT, OT/L · Founder, Learning Charms and O.T. Wizard

    A short progress note before we go deeper

    Every week I send out an update to the therapists, educators, and clinicians contributing to the FMSAT norming project. The updates focus on data, leaderboards, and the age bands we still need to fill. This week the data deserves a deeper look than a weekly email can carry. The findings touch on questions that go well beyond a single screener, including how our profession measures what we do, how those measurements connect to insurance reimbursement, and why occupational therapy salaries have not kept pace with the cost of becoming an OT.

    If you are a pediatric occupational therapist who has ever had to defend a Beery score at an IEP meeting that did not match the child sitting in front of you, this post is for you. If you are an educator looking for a fast, fair screening tool to identify students who may benefit from earlier support, this is for you. If you are an adult-focused OT, a neurology specialist, or a clinician working with progressive motor conditions, this is for you too. And if you are an OT, OTA, or therapy leader thinking about the coming shift to value-based care and what it means for your practice, this is also for you.

    Here is what is in this post:

    • What the FMSAT is, what it actually measures, and why it works across the lifespan
    • How the FMSAT can be used as a one-time screener, a progress monitoring tool, and a lifelong tracking instrument
    • What Rasch analysis is, in plain English, and why it is considered the gold standard in modern assessment
    • The first Rasch results on the FMPRS, our companion rating scale
    • What ecological validity means and why your favorite assessments may be missing it
    • Why the FMSAT itself uses Classical Test Theory and the FMPRS uses Rasch
    • Why value-based care is about to make all of this matter more than it ever has before
    • How the O.T. Wizard platform, soon to be rebranded as MyTherapyWizard, fits into the bigger picture

    A note on terminology before we go further

    Throughout this post you will see three terms that are sometimes used interchangeably but mean different things in occupational therapy practice and in psychometrics. Getting the distinction right matters because the FMSAT is one of those terms and not the others.

    A screener is a brief, low-burden tool designed to flag people who may benefit from further evaluation. Screeners take a minute or two, can often be administered by people without specialized clinical training, and produce a simple signal that says proceed to next step or no further action needed. Screeners do not diagnose, do not establish eligibility, and do not produce a comprehensive clinical picture. Examples in healthcare include the M-CHAT for autism, the PHQ-9 for depression, and vision and hearing screens in schools.

    An assessment is a more comprehensive structured evaluation that produces detailed information sufficient to support diagnosis, eligibility decisions, and treatment planning. Assessments typically take 30 to 90 minutes, require trained administrators, and produce multiple subscores or domain scores. Examples in pediatric OT include the Beery-Buktenica Developmental Test of Visual-Motor Integration, the Peabody Developmental Motor Scales, the Bruininks-Oseretsky Test of Motor Proficiency, the Sensory Processing Measure, and the Pediatric Evaluation of Disability Inventory Computer Adaptive Test (PediCAT).

    An evaluation is the broader clinical process that uses one or more screeners and assessments together with observation, interview, and chart review to produce a full clinical picture and recommendations.

    The FMSAT (Fine Motor Speed and Accuracy Test) is a screener. It is not designed to replace any assessment in your toolkit. It is designed to do the screener job well, which means flagging test takers who may benefit from a comprehensive evaluation when their lateralization or fine motor speed scores fall outside expected ranges. The assessments that follow a positive FMSAT screen will be whatever your clinical reasoning and your setting indicate. The point of a good screener is to make sure the right people get to those assessments faster than they would have otherwise.

    The FMPRS (Fine Motor Participation Rating Scale), on the other hand, is functioning more like a brief assessment instrument during the validation phase of this project, because its job is to produce the multi-domain rating data needed to establish ecological validity for the FMSAT. After validation is complete, the FMPRS will not typically be used alongside the FMSAT in routine clinical screening.

    What the FMSAT measures and why it works across the lifespan

    The FMSAT, which stands for Fine Motor Speed and Accuracy Test, is a one-minute screener. The test taker is given a bubble-popping worksheet and asked to pop as many bubbles as they can in 30 seconds with one hand, then 30 seconds with the other hand. The score for each hand is the number of bubbles popped. That is it. No rater training, no expensive test kits, no proprietary materials beyond a printed worksheet.

    The name says fine motor speed and accuracy, and that is what each hand score reflects. But the construct the instrument was designed to measure, and the reason the two hands are tested separately, is neuromotor lateralization. Lateralization is the degree to which a person has committed one side of the brain, and therefore one hand, to specialized motor work. Strong lateralization means the dominant hand performs precision tasks fluently while the non-dominant hand serves as a stabilizer. Weak or absent lateralization means the two hands perform more similarly, hand preference is inconsistent, or the person switches hands mid-task. Fine motor speed is the metric we use to detect that pattern, because timed performance under demand reveals the lateralization signal more clearly than untimed observation does.

    This matters because lateralization is not just a developmental milestone of early childhood. It is a lifelong neuromotor property that emerges during the preschool years, consolidates through school age, holds through adulthood, and can shift or erode in response to neurological changes later in life. The FMSAT was designed to capture that lateralization signal at any age, which is why our normative dataset spans from preschool through older adulthood rather than stopping at age 12 or 17.

    What makes the FMSAT useful is not the bubble popping itself. It is what the two hand scores together tell us about how well a person has consolidated hand dominance, and how that pattern compares to typical lateralization for their age. The dominant hand score reflects fine motor capability. The difference between the two hands reflects lateralization. Both pieces of information matter for clinical and educational practice, and neither one is captured well by the standardized assessments most occupational therapists currently use.

    How the FMSAT can be used: one-time screener, progress monitoring, and lifelong tracking

    Because the FMSAT is fast, standardized, and produces a numeric score, it can support several clinical and educational use cases that most current assessments cannot. Three are worth naming explicitly.

    Use case 1: One-time screening for hand dominance and fine motor concerns

    This is the most familiar use case. A pediatric occupational therapist, school OT, or educator administers the FMSAT to a student who has been flagged for fine motor concerns, handwriting difficulty, or unclear hand dominance. The score, compared to age-appropriate normative bands, gives a quick objective marker of whether the student’s performance and lateralization fall within expected ranges. This is the use case the validation manuscript will focus on first.

    Use case 2: Response to Intervention and progress monitoring

    Because the FMSAT takes one minute and produces a numeric score, it can be re-administered at intervals to track whether a person is making measurable progress over time. This is exactly the kind of brief repeatable measurement that Response to Intervention frameworks and Multi-Tiered Systems of Support models require. A school OT could administer the FMSAT at the start of an intervention block, midway through, and at the end, and have objective data showing whether the lateralization or fine motor speed measure is changing in response to the intervention. A clinic-based OT could do the same across a course of therapy. The same minimum-detectable-change thresholds that value-based payment models are increasingly requiring would apply directly. Once the normative dataset is finalized, the FMSAT becomes one of the few fine motor measures fast enough to use for routine progress monitoring.

    Use case 3: Educator-administered screening in school settings

    The FMSAT requires no rater training and no clinical interpretation to administer. A teacher, paraprofessional, school nurse, or interventionist can give the test and record the scores. With enough normative data in place, the score itself does the screening work and a teacher does not need to be an OT to identify a student who falls below the expected band. This opens the door to universal screening at the classroom or grade level, the same way schools currently screen vision and hearing. Students flagged by the screener can then be referred to occupational therapy for full evaluation. This is a long-term vision rather than an immediate use case, but it is exactly the kind of MTSS Tier 1 screening application that schools have been asking for and that occupational therapy has not yet had the tools to support.

    Use case 4: Lifespan monitoring, including progressive neuromotor conditions

    This is the use case that emerges directly from norming the FMSAT across all ages rather than only in pediatrics. Lateralization can shift across the lifespan in response to neurological events and conditions. A person recovering from stroke may show changed dominant-hand performance. A person living with multiple sclerosis, Parkinson’s disease, or another progressive neuromotor condition may show gradual erosion of the lateralization signal as the condition advances. A person experiencing age-related changes in motor control may show compression of the dominant-hand advantage. The FMSAT, administered periodically, could detect those shifts earlier and more objectively than self-report or general clinical observation. With enough normative data across all ages, the same one-minute test that screens a five-year-old for emerging hand dominance could monitor a 55-year-old neurologist for early signs of motor change in the years after a diagnosis. That is the broader clinical reach the lifespan normative dataset makes possible.

    None of these use cases require a separate test. They all use the same one-minute bubble-popping task. What changes is who is administering it, how often, and what the score is being compared against. That flexibility is one of the reasons we are investing in the validation work the way we are.

    What is Rasch analysis and why is it considered the gold standard?

    Many occupational therapy assessments you have used in your career, including those most commonly used to demonstrate progress in pediatric settings, were developed using Classical Test Theory, often shortened to CTT. The Beery-Buktenica Developmental Test of Visual-Motor Integration, the Bruininks-Oseretsky Test of Motor Proficiency, the Peabody Developmental Motor Scales, and the Sensory Processing Measure are CTT-based instruments. CTT produces a total score that is then compared to a normative sample. The score tells you where a test taker falls relative to peers, but it has some real limitations.

    The biggest limitation is that CTT treats every item on a test as if it were equally difficult, and every point on the score scale as if it represented the same amount of skill. A test taker who scores 84 versus one who scores 89 may differ by a meaningful amount of skill, or may differ by almost nothing, depending on where on the scale those scores fall and which specific items they got right. CTT cannot tell you which.

    This is true for both screeners and assessments built under CTT, though the limitation is more consequential for assessments because assessments are doing more of the clinical decision-making work. A screener producing a CTT score is still useful as a yes/maybe/no signal. An assessment producing CTT scores is being asked to support diagnosis, eligibility, and treatment planning decisions on the same imprecise measurement scale.

    Rasch analysis, developed by Danish mathematician Georg Rasch in the 1960s, takes a fundamentally different approach. Rasch places each test item and each person on the same interval scale, called the logit scale. This means the distance between a score of 5 logits and 10 logits represents the same amount of skill change as the distance between 20 and 25. Rasch also calibrates each item individually, telling you which items are easy, which are hard, and whether each item is actually pulling its weight in measuring the construct.

    Rasch is to assessment what a ruler is to measurement. CTT scores tell you a child is somewhere in the middle. Rasch scores tell you exactly where, on a scale where the units are equal.

    Rasch analysis is considered the gold standard for modern instrument development because it is the framework used by some of the most respected and most defensible pediatric assessments in the field. The PEDI-CAT, the AMPS or Assessment of Motor and Process Skills, the School Function Assessment, and the WeeFIM are all Rasch-based or Item Response Theory based instruments. These are the tools that produce data insurance companies and researchers trust. If you have ever wondered why some assessments seem to have stronger research backing than others, the framework behind them is usually a big part of the answer.

    Behind the scenes: the first Rasch checkup on the FMPRS

    The FMPRS, or Fine Motor Participation Rating Scale, is the 12-item observer-rated companion to the FMSAT. It captures four dimensions of fine motor function across real-world tasks: fine motor speed, fine motor precision, laterality and bilateral differentiation, and participation in daily roles. Our Rasch consultant, Angie, ran the first calibration on 139 paired FMSAT and FMPRS records earlier this month. The results were strong for a first calibration on a relatively small sample.

    Finding one: the item difficulty order matches the developmental theory

    Rasch ordered the 12 items from easiest to hardest based on how raters actually responded to them. The Laterality items, which ask about hand dominance and bilateral hand use, came out as the easiest. The Participation items, which ask about sustained engagement in fine motor tasks throughout the day, came out as the hardest. This is exactly what we would predict developmentally. Hand dominance consolidates earlier in childhood than sustained occupational engagement, so on a rating scale measuring fine motor function across the developmental arc, laterality items should be easier to endorse than participation items. The data confirmed the theory.

    Finding two: the instrument separates clinical and typical test takers cleanly

    On the Rasch-derived person measure, typical preschoolers scored more than two logits higher than clinical preschoolers. In plain terms, the typical preschooler scored higher than approximately 98 percent of the clinical sample. This is what is called a known-groups validity effect, and a Cohen’s d effect size above 2.0 is unusually strong. Most pediatric assessments are pleased to show known-groups effects in the 0.5 to 0.8 range. The FMPRS is producing separation that is roughly three times stronger. As the normative dataset grows beyond preschool ages, we expect the same separation pattern to extend across older age bands and into adult populations where clinical and typical comparison groups can be defined.

    Finding three: the FMPRS and the FMSAT are picking up the same underlying construct

    The Rasch-derived person measure on the FMPRS correlated with FMSAT dominant hand scores at r = 0.575. That correlation tells us that two completely different methods of measurement, a one-minute performance task and a 12-item observer rating scale, are picking up the same underlying construct. This is exactly the kind of cross-method convergence a strong validation manuscript needs.

    What still needs work

    Five items in the FMPRS came back as overfitting in the Rasch model, which means they are too internally redundant. They are not bad items. They are just not adding as much new information as they could. Those items will be revised for the next version of the FMPRS based on what the calibration showed us. This kind of iterative refinement is how serious instrument development works, and it is exactly why we are doing this calibration now, before the manuscript is finalized.

    Ever wondered why Beery, PDMS-3, or BOT-2 scores do not match real-world function?

    Here is a question every occupational therapist has wrestled with at some point in their career. Why do scores from the Beery-Buktenica VMI, the Peabody Developmental Motor Scales, or the Bruininks-Oseretsky Test sometimes fail to line up with what you actually see in the classroom, at home, on the playground, or in adult daily life?

    You are not imagining it. A student can score below average on a tabletop visual motor task and still write legibly, manage their lunchbox, and participate fully in PE. Another can score in the average range and still struggle every single day to keep up with handwriting demands or self-care routines. The same pattern shows up in adult assessments. The mismatch is real, and it has a name in the psychometric literature. It is called an ecological validity gap.

    What ecological validity actually means

    Ecological validity is a psychometric term that answers a simple question: does this assessment measure something that actually matters in real life? An instrument with strong ecological validity produces scores that connect to how a person functions in their daily environment. An instrument with weak ecological validity produces scores that connect mainly to how a person performs on the test itself, with limited evidence that the score predicts daily function.

    Many of the assessments OTs use most often were designed to measure isolated motor performance, not real-world participation. They are good at what they measure. They were just never built to answer the participation question. When a school-based therapist is asked to defend a Beery standard score at an IEP meeting, or when a clinic-based therapist is asked to justify medical necessity to an insurer using a PDMS-3 score, the underlying problem is often that the score is being asked to do something the assessment was not designed to do.

    Why the FMPRS is essential to FMSAT validation, even though clinicians will not use both in practice

    Here is a question that comes up almost every time I explain this project. If the FMSAT is a screener, why pair it with the FMPRS at all? Won’t clinicians be expected to do both?

    The answer is no. The FMPRS is doing critical work right now, during validation, so that the FMSAT will not need to be paired with it in clinical practice later. The whole value of the FMSAT as a screener depends on it being a one-minute, paper-and-pencil, standalone score. Asking clinicians to also complete a 12-item rating scale every time they administered the screener would defeat the entire point of having a screener in the first place.

    What the validation work establishes, and what every paired FMSAT and FMPRS submission helps establish, is that when the FMSAT score is elevated or compressed in a particular way, it is reflecting something that shows up in the test taker’s real-world functional life. Once that link is established and published in a peer-reviewed manuscript, the FMSAT score on its own carries that ecological meaning forward. Clinicians using the FMSAT in practice will be able to point to the published validation evidence rather than having to demonstrate the connection every time.

    This is the same approach used to validate other widely accepted screeners. The Modified Checklist for Autism in Toddlers, known as the M-CHAT, was validated against full ADOS and ADI-R diagnostic batteries. Pediatricians using the M-CHAT today do not run an ADOS alongside it. The validation work was done once and the screener now stands on its own. The PHQ-9 depression screener was validated against structured psychiatric interviews. Primary care providers use it on its own today. The pattern is consistent across well-validated screeners. Validate against a richer companion measure once, publish the validity evidence, then use the screener on its own.

    Why the FMSAT itself uses Classical Test Theory and the FMPRS uses Rasch

    This is a question that any sharp reader will be asking by now. If Rasch is the gold standard, why is the FMSAT being calibrated using Classical Test Theory rather than Rasch?

    The answer comes down to the measurement structure of each instrument. The FMSAT produces raw bubble counts on a zero to 80 scale for each hand. That kind of continuous count data is well-suited to CTT-style descriptive statistics, percentile norms, and known-groups validity comparisons. These are the analyses the FMSAT validation manuscript will lean on, and they are the analyses that produce the percentile bands and severity cutoffs clinicians actually use at the point of care.

    Rasch is the right framework for the FMPRS because the FMPRS uses ordered category responses on a four-point scale, and each item can be at a different difficulty level on the same underlying trait. That is exactly the kind of measurement structure Rasch was designed to handle. The two instruments are built differently on purpose, and each one is being analyzed using the framework that fits its measurement structure.

    The Rasch work on the FMPRS gives the manuscript the modern psychometric backbone peer reviewers expect. The CTT work on the FMSAT keeps the screener simple and interpretable for the clinicians who will actually use it. Both pieces matter, and together they create a defensible validation argument.

    Why value-based care is about to make all of this matter much more than it has before

    If you have been practicing for more than a few years, you already know that occupational therapy reimbursement has been under pressure for a long time. The 2026 Medicare Physician Fee Schedule final rule from the Centers for Medicare and Medicaid Services, released October 31, 2025, continued a trend of flat or declining payment rates for outpatient OT. According to OT Potential’s 2026 reimbursement analysis, the proposed 1% decrease to OT and PT relative value units for 2026 came after a 0% increase in 2025 and a 3% decrease in 2024. The trajectory is real and most OTs feel it directly in their paychecks.

    What is changing right now, and what most clinicians have not fully internalized, is that the structure of the entire payment system is shifting underneath us. The shift is called value-based care.

    What value-based care actually means

    Value-based care is a payment model where providers and health systems are reimbursed based on the outcomes their patients achieve, not the volume of services delivered. Under traditional fee-for-service payment, an OT bills for each visit and gets paid for each visit. Under value-based care, payment is increasingly tied to whether the patient demonstrated measurable functional improvement against established benchmarks.

    CMS has been driving this shift through several specific programs. The Quality Payment Program, the Merit-based Incentive Payment System known as MIPS, and an expanding suite of Alternative Payment Models are all moving rehabilitation services toward outcomes-based reimbursement. The 2026 payment updates included a 0.75% increase for qualified APM participants and a 0.25% increase for everyone else, an early but clear signal that participating in alternative payment models will increasingly be where the financial upside is.

    The era of writing patient made progress toward goals in a discharge note and being reimbursed for it is ending. The era of demonstrating measurable functional change against defensible benchmarks is beginning.

    Why occupational therapy is structurally underprepared for this shift

    Here is the connection most clinicians have not drawn explicitly. The shift to value-based care requires outcomes data, and outcomes data is only as good as the assessments and progress-monitoring instruments producing it. Most of the assessments occupational therapists use to demonstrate progress, and most of the screeners they use to identify who needs services in the first place, were developed under CTT frameworks that produce raw scores, percentile bands, and standard scores. None of those formats give insurers what value-based payment models actually require.

    Insurers under value-based care want interval-level evidence of measurable functional change against established minimum-detectable-change thresholds. When a third-party reviewer asks whether a patient made meaningful progress, the answer they want is not, the patient’s standard score improved from 84 to 89. The answer they want is, the patient’s interval-level fine motor measure shifted by 0.45 logits, which exceeds the minimum detectable change threshold of 0.30 logits established in the calibration sample. One of those answers is opinion-vulnerable. The other is data.

    Closing this evidence gap requires investment at every level of the measurement pipeline: better screeners that identify who needs services earlier and more accurately, better assessments that produce the diagnostic and eligibility data on a defensible measurement scale, and better outcomes instruments that document functional change over an episode of care. The FMSAT and the FMPRS sit at the screener and ecological validity ends of that pipeline. They are one contribution among many that the profession needs.

    This evidence gap is one of the underrecognized reasons our profession has struggled to make the reimbursement case at the level of physical therapy or speech-language pathology. Both adjacent professions have invested more heavily in Rasch-calibrated, IRT-based assessment development over the last 20 years. The PEDI-CAT, the AM-PAC, and similar tools represent what that investment looks like. Occupational therapy has far fewer Rasch-calibrated tools across the screener, assessment, and outcomes layers, and that thinness in our measurement infrastructure shows up downstream as flatter reimbursement, narrower coverage policies, and ultimately compensation that has not kept pace with the cost of the training required to enter the field.

    How the O.T. Wizard platform fits into this picture

    The FMSAT and the FMPRS are not standalone projects. They are pieces of a larger evidence infrastructure being built into the O.T. Wizard platform, which is being rebranded as MyTherapyWizard.

    O.T. Wizard is a digital evaluation and outcomes platform built from the ground up on modern psychometric standards. The FMSAT lives inside the platform as a fast, defensible screener with clear research foundations. The FMPRS lives alongside it as the ecological validity companion during validation. The broader platform houses structured evaluation templates designed to produce the kind of data that holds up under value-based payment scrutiny. The architecture is PHI-free and operates under a 1EdTech-approved legal framework, which means the platform itself functions as a passive-accrual research engine. Every paired evaluation contributes to the dataset that makes the next generation of assessments stronger.

    The rebrand to MyTherapyWizard reflects the platform’s expanding scope beyond occupational therapy into a multi-discipline space for pediatric therapy professionals. The underlying mission stays the same. Build the measurement infrastructure our profession needs to move forward, in step with where reimbursement is going rather than chasing it after the fact.

    What you can do

    Our profession needs more Rasch-calibrated assessments. It needs more validated rating scales with strong ecological validity. It needs more normative datasets large enough to defend in peer review. And it needs more clinicians and educators willing to contribute the data that makes all of that possible.

    If you are an occupational therapist, an educator, a clinician working with adult or geriatric populations, or anyone interested in supporting the development of evidence-based assessment tools, here is how to get involved:

    • Request to be on the Norming Tryout Team and Contribute FMSAT data. The screener takes one minute per test taker. If you administer it after a session or screening you would have run anyway, the marginal time cost is essentially zero.
    • Complete the FMPRS when you can. The paired data is what makes the validation manuscript possible. Every paired submission directly strengthens the published evidence base our profession will use.
    • Look for the bands we need most. As of this week, the most urgent recruitment gaps are adolescents ages 12 to 17, both clinical and typical, three-year-olds in both groups, adults age 50 and older, and left-dominant test takers at every age.
    • Share this work with colleagues. The bigger and more representative the dataset, the stronger the eventual screener will be for the people you serve.

    Every paired submission you contribute is a small but real piece of building the measurement infrastructure our profession needs. Building a Rasch-calibrated rating scale and a CTT-validated performance screener together, on a normative dataset large enough to defend in peer review and broad enough to span the lifespan, is exactly the foundational psychometric work the field has needed for years. If we want occupational therapy to be reimbursed at the level our training and clinical expertise warrant under the new value-based payment models, we have to produce the kind of evidence other professions have already produced.

    That work does not happen in conference panels or position papers. It happens in datasets, calibrations, and validation manuscripts. It happens in projects like this one. And the people producing it are not academics in distant labs. They are clinicians and educators like you who choose to spend a few minutes on a Tuesday afternoon contributing to something larger than a single evaluation.

    About this project

    The FMSAT, Fine Motor Speed and Accuracy Test, is a one-minute screener for neuromotor lateralization and fine motor speed, currently in active normative data collection across the lifespan toward a peer-reviewed validation manuscript. The FMPRS, Fine Motor Participation Rating Scale, is the 12-item observer-rated companion used to establish ecological validity during the validation phase. Both instruments are part of the O.T. Wizard platform, rebranding to MyTherapyWizard. Pearl IRB Not Human Subjects Research determination on file (ID 2026-0154).

    Related topics Neuromotor lateralization assessment, hand dominance evaluation across the lifespan, Rasch analysis in rehabilitation, fine motor screening for educators, Response to Intervention RTI fine motor measures, MTSS Tier 1 and Tier 2 screening, ecological validity in occupational therapy, value-based care for outpatient therapy, Medicare Physician Fee Schedule 2026, evidence-based occupational therapy practice, school-based occupational therapy, OT reimbursement, alternative payment models for rehabilitation, progressive neuromotor condition monitoring, multiple sclerosis fine motor tracking, stroke rehabilitation outcomes measurement, lifespan motor assessment

  • What a Brief Scissor Skills Assessment Reveals in Preschool-Aged Children

    What a Brief Scissor Skills Assessment Reveals in Preschool-Aged Children

    More Than a Milestone: What a Brief Scissor Skills Assessment Reveals About Tool Use, Hand Dominance, and Cutting Development in Preschool-Aged Children

    Stephanie Seymore Wick, MSOT, OT/L  |  Founder and Clinical Architect, O.T. Wizard

    March 2026

    Abstract

    Aims: To examine scissor cutting performance across preschool age bands (36 to 65 months) in a clinical sample, identify relationships between cutting accuracy, hand dominance, and hand positioning, assess cross-domain correlations, and evaluate longitudinal progress from initial to re-evaluation.

    Methods: Descriptive and correlational analysis of 541 pediatric occupational therapy evaluations (Age Bands G through J) using structured cutting tasks scored on a standardized 14-point rubric. Participants were children referred for OT services, predominantly ages 48 to 59 months, with at least 90% qualifying for Medicaid.

    Results: Cutting accuracy followed a clear developmental progression across age bands. Thumb-up dominant hand positioning was a large-effect predictor of cutting accuracy (Cohen d=1.17). Established hand dominance and writing-to-scissor hand consistency were strongly associated with performance. Scissor performance correlated significantly with fine motor, visual motor integration, ADL, visual perception, gross motor, bilateral integration, and praxis domains. Longitudinal gains of 5.21 points over 4.8 months exceeded the expected natural growth rate of 1.51 points.

    Conclusions: A structured scissor skills assessment captures clinically meaningful variation in cutting skill and supports response-to-intervention documentation, goal writing, and cross-domain clinical reasoning in pediatric OT practice.

    Keywords: scissor skills, hand dominance, fine motor development, pediatric occupational therapy, response to intervention, preschool

    Scissors are one of the most commonly targeted skills in pediatric occupational therapy, yet they are rarely assessed with the precision required to drive goal writing, track progress, or demonstrate response to intervention. In most clinical settings, a child either can cut or cannot cut. That binary framing misses the developmental story that unfolds across the preschool years and leaves practitioners without the data needed to communicate clinical value to families, educators, and payers.

    The preschool years represent the primary window for scissor skill acquisition. Cutting a straight line with 1-inch tolerance is typically expected by 41 months, precision cutting within a 1/4-inch boundary emerges around 48 months, and smooth curvy-line cutting at a 1/4-inch tolerance is an expectation by 60 months (O.T. Wizard Scissor Skills Assessment, v6.2). Despite this well-established developmental sequence, most standardized pediatric OT assessments address scissor skills with limited granularity, and published outcome data on scissor skill development in referred clinical populations remains sparse.

    This article presents findings from 541 pediatric occupational therapy evaluations collected through O.T. Wizard, a clinical intelligence platform for pediatric occupational therapy professionals. Using a standardized scissor skills assessment embedded within the evaluation process, this analysis examines cutting performance across age bands, its relationship to hand dominance and hand positioning, its correlation with multiple developmental domains, and longitudinal gains across re-evaluation. The evidence supports scissor skill assessment as a window into neuromotor organization, tool use learning, and cross-domain functional development.

    Methods

    Participants

    This analysis includes 541 pediatric occupational therapy evaluations representing children ages 36 through 65 months (Age Bands G through J): Band G (36 to 47 months, n=61), Band H (48 to 53 months, n=199), Band I (54 to 59 months, n=256), and Band J (60 to 65 months, n=69). An additional 82 children had paired evaluations (E1 and E2), with a mean interval of 4.8 months between assessments, enabling longitudinal analysis. The sample was 57% male and 43% female. Primary language was English for 90% of participants and Spanish for 9%. At least 90% of children qualified for Medicaid, and approximately 86% were recommended for OT services following evaluation. All evaluations were conducted in North Carolina.

    Data were collected through routine pediatric occupational therapy evaluations conducted in clinical practice using O.T. Wizard. Parents provided informed consent as part of standard clinical care. As this analysis represents clinical outcomes data rather than human subjects research, institutional review board approval was not required. All data were de-identified in accordance with HIPAA regulations.

    Measures

    The O.T. Wizard Scissor Skills Assessment (v6.2) is a structured, standardized tool embedded within the O.T. Wizard multi-domain pediatric OT evaluation platform. It measures cutting performance across a progression of task demands anchored to published developmental milestones: holding scissors with one hand (24 months), snipping paper (34 months), opening and closing scissors (36 months), cutting a 1-inch straight line (41 months), 3/4-inch and 1/2-inch straight lines (43 and 45 months), a 1/4-inch straight line (48 months), a 1/4-inch curvy line (60 months), and smooth cutting quality (72 months). Cutting accuracy was scored using a standardized rubric with 0, 0.5, and 1.0 point values per segment, yielding a maximum score of 14 points per cutting task. Therapists also documented scissor type used, dominant hand position (thumb up versus thumb down or absent), stabilizer hand description, and qualitative hand positioning ratings.

    This analysis focuses primarily on FM_SCIS_48 (1/4-inch straight line), selected as the primary analytic item due to near-complete data across all age bands (n=541) and its anchor at the 48-month developmental expectation. FM_SCIS_41 (1-inch straight line) and FM_SCIS_60C (1/4-inch curvy line) were included where sample sizes permitted. Scissor type data indicated that 90.6% of children were assessed with school or safety scissors; adaptive equipment (spring-assisted, loop handle) was used in fewer than 1% of cases and is not analyzed separately.

    Data Analysis

    Descriptive statistics were calculated for FM_SCIS_48 by age band, dominance level, and thumb positioning group. Pearson and Spearman correlations were computed between FM_SCIS_48 and all available domain and subdomain scores. Between-group comparisons used independent samples t-tests; effect sizes are reported as Cohen d. Longitudinal analysis used paired t-tests comparing E1 and E2 scores among children with paired evaluations. A cross-sectional growth rate was used to estimate expected natural maturation over the mean evaluation interval, following the methodology established in the response-to-intervention outcomes article in this series. Artificial intelligence writing assistance (Claude, Anthropic, version Sonnet 4.6) was used in the preparation of this manuscript for language editing and formatting; all analytical decisions, clinical interpretations, and conclusions are those of the author.

    All data is pre-normative. Rasch analysis validation is ongoing and will be reported in subsequent publications.

    Results

    Developmental Progression of Cutting Accuracy

    Across all age bands, mean scores on FM_SCIS_48 increased steadily, with the steepest growth occurring between Bands G and H, the window when this skill is developmentally expected to emerge (Table 1). At Band H (the anchor age for this task), 35.6% of children in this clinical sample scored zero, reflecting the referred nature of the population. The bimodal distribution within Band H is more clinically informative than a pass/fail classification. Of 165 children scoring 10 or above on FM_SCIS_48, 162 (98.2%) were also administered FM_SCIS_60C on the same evaluation day. Among children who scored 14 on FM_SCIS_48, scores on FM_SCIS_60C ranged across the full spectrum (mean 8.61), confirming that the curvy task captures a meaningfully more demanding level of motor control.

    Table 1. FM_SCIS_48 (1/4″ straight line) performance by age band. Data from a clinical sample of children referred for OT evaluation, prior to intervention.

    Age BandAge RangenMean /14% ScoreFloor (0)Ceiling (14)
    G36-47m451.7912.8%64.4%2.2%
    H48-53m1804.6433.1%35.6%10.6%
    I54-59m2486.5246.6%18.1%21.4%
    J60-65m688.0757.7%13.2%26.5%

    Hand Positioning and Hand Dominance

    Dominant hand thumb-up positioning was associated with substantially higher cutting performance. Children with thumb-up positioning averaged 8.16 on FM_SCIS_48 compared to 2.74 for children without thumb-up positioning (Cohen d=1.17, p<0.001; median scores 9.0 versus 1.0). Within Band H, 28% of children who scored zero had thumb-up positioning compared to 89% of children who scored 14. No children who scored 14 on FM_SCIS_48 were rated Never for overall hand positioning quality.

    Hand dominance status was consistently associated with cutting performance (Table 2). Children with established dominance scored more than five points higher on average than children with inconsistent dominance. The Spearman correlation between dominance level and FM_SCIS_48 was rho=0.360 (p<0.001). Writing-to-scissor hand consistency produced a large effect: children who used the same hand for writing and cutting averaged 7.06 on FM_SCIS_48 compared to 2.25 for children who used different hands (Cohen d=1.05, p<0.001).

    Table 2. FM_SCIS_48 mean score by documented hand dominance level (Bands G-J, n=534). Dominance level based on therapist rubric rating at time of evaluation.

    Dominance LevelnMean /14% ScoreFloor (scored 0)
    Emerging181.5010.7%56%
    Inconsistent702.3817.0%49%
    Strong preference2335.1736.9%28%
    Established2137.7655.4%16%

    Cross-Domain Correlations

    Scissor cutting performance correlated significantly with all major developmental domain scores (Table 3). Fine Motor domain showed the strongest correlation (r=0.772). Visual Motor Integration (r=0.511) and Visual Perception (r=0.461) correlations reflect the visual guidance demands of cutting within a narrow boundary. The ADL correlation (r=0.511) speaks to generalization of tool use skill across daily living contexts. The Gross Motor correlation (r=0.430) reflects the role of proximal stability as a foundation for distal precision. Praxis showed the weakest correlation (r=0.271), consistent with motor planning contributing primarily to early skill acquisition rather than refined execution. The FMSAT Speed and Accuracy subdomain showed no significant relationship (Spearman rho=-0.012, not significant), confirming that scissor performance and pencil speed capture distinct aspects of fine motor control and contribute independent clinical information.

    Table 3. Pearson and Spearman correlations between FM_SCIS_48 and domain/subdomain scores (Bands G-J). *** p<0.001. FMSAT Spearman correlation not significant.

    Domain / SubdomainnPearson rSpearman rhoClinical Interpretation
    Fine Motor domain4540.772***0.787Strong — cutting as fine motor expression
    Hand Use subdomain5410.541***0.549Bilateral tool use as hand use marker
    Visual Motor Integration5370.511***0.519Visual guidance of cutting path
    ADL domain5100.511***0.530Tool use generalizes to daily living
    Visual Perception domain5150.461***0.469Line discrimination supports accuracy
    Gross Motor domain5400.430***0.444Proximal stability drives distal precision
    Bilateral Integration5380.398***0.395Two-hand coordination demand
    Praxis domain5390.271***0.281Motor planning, weaker once program formed
    Speed & Accuracy (FMSAT)459r=-0.105*rho=-0.012 (ns)Distinct skill; independent clinical value

    Longitudinal Gains

    Among 82 children with paired evaluations, FM_SCIS_48 showed a mean gain of 5.21 points over an average interval of 4.8 months (paired t-test t=8.10, p<0.001; Cohen d=0.895). The curvy task (FM_SCIS_60C) showed a mean gain of 3.75 points over the same interval (t=7.20, p<0.001), with 76.4% of children improving. Using the cross-sectional growth rate as a natural maturation baseline, the expected natural growth on FM_SCIS_48 over 4.8 months was approximately 1.51 points. The observed mean gain of 5.21 points exceeded this expected rate by 3.71 points. Ceiling effects were noted among children who had scored at or near 14 at E1; therapists correctly applied FM_SCIS_60C as the clinically sensitive measure for near-ceiling children in 98.2% of applicable cases.

    Discussion

    This analysis of 541 pediatric OT evaluations demonstrates that a brief, standardized scissor skills assessment generates clinically meaningful data across multiple dimensions of preschool development. The developmental progression of FM_SCIS_48 scores across age bands aligns with published milestone expectations and provides clinical benchmarks for a referred population that are not currently available in the literature. Floor effects in younger bands and the bimodal distribution at the anchor age band reflect the nature of scissor skill acquisition in children referred for OT services, where skill emergence is delayed relative to normative expectations. These distributions support documentation of functional deficit and medical necessity in ways that pass/fail classifications cannot.

    The large effect of thumb-up dominant hand positioning (Cohen d=1.17) elevates hand positioning from a clinical observation to a primary, modifiable intervention target. Establishing correct scissor grip before focusing on path accuracy is supported by the data: children without functional thumb-up positioning have mean scores below three regardless of age band, suggesting that grip orientation is a near-prerequisite for achieving cutting accuracy at the mastery level.

    The relationship between hand dominance and scissor performance connects these findings to a broader literature on neuromotor specialization in the preschool years (Scharoun & Bryden, 2014). A child who has not yet organized a consistent preferred hand for tool use is reflecting an underlying developmental state that affects performance across all tool-based tasks. Writing-to-scissor hand consistency findings have direct implications for school-based practice: addressing hand consistency across classroom tool use activities, not only during designated scissor tasks, is consistent with both the data and motor learning principles and supports IEP goal development that reflects the child’s functional performance across settings.

    The cross-domain correlation profile challenges the framing of scissor skills as an isolated fine motor task. The significant Gross Motor correlation (r=0.430) reinforces that proximal stability, including trunk support and shoulder girdle control, is foundational to distal precision, consistent with developmental neuroscience frameworks (Stoodley, 2016). The VMI and Visual Perception correlations reflect the visual guidance demands of path-following. Intervention plans that address only the distal cutting task without considering foundational postural, visual, and neuromotor systems may produce slower or less durable gains. The absence of a significant FMSAT correlation confirms that scissor accuracy and pencil precision speed are complementary measures rather than redundant ones, and that both contribute independent information to a comprehensive fine motor profile.

    Longitudinal gains of 5.21 points over 4.8 months, exceeding the expected natural growth rate by 3.71 points in a population with limited scissor practice outside of OT sessions, provide meaningful support for OT as a driver of skill development. This above-expected gain pattern is consistent with the RTI methodology established in this article series, which separates natural developmental progress from intervention-attributable change using cross-sectional growth rates as a baseline correction. For medical model practitioners, this framing supports quantifiable medical necessity documentation. For school-based practitioners, it provides data to support RTI tier documentation and progress monitoring language consistent with IDEA requirements.

    Several methodological factors warrant consideration. This sample represents children referred for OT evaluation in North Carolina and should not be generalized to typically developing children or other geographic populations. The cross-sectional age-band comparisons reflect group differences rather than individual trajectories. Adaptive scissor data was insufficient for analysis. Without a controlled comparison group, causal claims about OT effectiveness cannot be made; above-expected gains in a referred population with limited community practice are suggestive but not definitive. Future research should include a typically developing comparison sample, session-level dosage data, and expanded age coverage through early elementary years to determine whether scissor precision continues to develop beyond 66 months and at what point the 1/4-inch straight line becomes a floor item for older children (Cameron et al., 2012; Zhang et al., 2025).

    Conclusions

    A structured scissor skills assessment generates clinically meaningful data for pediatric OT practice. Cutting accuracy in a referred preschool population follows a measurable developmental trajectory, is strongly predicted by thumb-up dominant hand positioning and established hand dominance, correlates significantly with fine motor, gross motor, visual motor, visual perception, ADL, and praxis domains, and shows gains that exceed expected natural growth rates over a 4.8-month evaluation interval in a population with limited community scissor access. These findings support the use of standardized, quantitative scissor skills assessment as a component of comprehensive pediatric OT evaluation and as a practical tool for RTI documentation, goal writing, and cross-domain clinical reasoning in both school-based and medical model practice settings.

    Disclosure of Interest

    Stephanie Seymore Wick is the Founder and Clinical Architect of O.T. Wizard, the platform from which all data in this article was collected. Data collection is ongoing under clinical quality improvement protocols. All data were de-identified in accordance with HIPAA regulations. The author reports no other competing interests.

    Data Availability Statement

    De-identified aggregate data supporting the findings of this study are available from the corresponding author upon reasonable request. Individual-level data cannot be shared due to HIPAA de-identification obligations.

    Biographical Note

    Stephanie Seymore Wick, MSOT, OT/L is the Founder and Clinical Architect of O.T. Wizard, a clinical intelligence platform for pediatric occupational therapy professionals, and the founder of Learning Charms, Inc. Her clinical and research focus is the development of psychometrically sound, computable measurement tools that quantify pediatric OT outcomes across multiple developmental domains. She practices and conducts research in North Carolina.

    References

    Cameron, C. E., Brock, L. L., Murrah, W. M., Bell, L. H., Worzalla, S. L., Grissmer, D., & Morrison, F. J. (2012). Fine motor skills and executive function both contribute to kindergarten achievement. Child Development, 83(4), 1229-1244. https://doi.org/10.1111/j.1467-8624.2012.01768.x

    Scharoun, S. M., & Bryden, P. J. (2014). Hand preference, performance abilities, and hand selection in children. Frontiers in Psychology, 5, 82. https://doi.org/10.3389/fpsyg.2014.00082

    Stoodley, C. J. (2016). The cerebellum and neurodevelopmental disorders. Cerebellum, 15(1), 34-37. https://doi.org/10.1007/s12311-015-0715-3

    Zhang, B.-F., Lin, Z.-C., & Li, C. (2025). Fine motor skills assessment instruments for preschool children with typical development: A scoping review. Frontiers in Psychology, 16, 1620235. https://doi.org/10.3389/fpsyg.2025.1620235