Tag: Rasch

  • Why OTP’s need better assessments-update on the FMSAT

    Why OTP’s need better assessments-update on the FMSAT

    Measuring Neuromotor Lateralization Across the Lifespan: Progress on the FMSAT Norming Project (& Why OTP’s need better assessments)

    How a one-minute screener for hand dominance, modern Rasch psychometrics, and the shift to value-based care are converging to change how occupational therapists, educators, and clinicians measure what we do.

    Published May 24, 2026 by Stephanie Seymore Wick, MSOT, OT/L · Founder, Learning Charms and O.T. Wizard

    A short progress note before we go deeper

    Every week I send out an update to the therapists, educators, and clinicians contributing to the FMSAT norming project. The updates focus on data, leaderboards, and the age bands we still need to fill. This week the data deserves a deeper look than a weekly email can carry. The findings touch on questions that go well beyond a single screener, including how our profession measures what we do, how those measurements connect to insurance reimbursement, and why occupational therapy salaries have not kept pace with the cost of becoming an OT.

    If you are a pediatric occupational therapist who has ever had to defend a Beery score at an IEP meeting that did not match the child sitting in front of you, this post is for you. If you are an educator looking for a fast, fair screening tool to identify students who may benefit from earlier support, this is for you. If you are an adult-focused OT, a neurology specialist, or a clinician working with progressive motor conditions, this is for you too. And if you are an OT, OTA, or therapy leader thinking about the coming shift to value-based care and what it means for your practice, this is also for you.

    Here is what is in this post:

    • What the FMSAT is, what it actually measures, and why it works across the lifespan
    • How the FMSAT can be used as a one-time screener, a progress monitoring tool, and a lifelong tracking instrument
    • What Rasch analysis is, in plain English, and why it is considered the gold standard in modern assessment
    • The first Rasch results on the FMPRS, our companion rating scale
    • What ecological validity means and why your favorite assessments may be missing it
    • Why the FMSAT itself uses Classical Test Theory and the FMPRS uses Rasch
    • Why value-based care is about to make all of this matter more than it ever has before
    • How the O.T. Wizard platform, soon to be rebranded as MyTherapyWizard, fits into the bigger picture

    A note on terminology before we go further

    Throughout this post you will see three terms that are sometimes used interchangeably but mean different things in occupational therapy practice and in psychometrics. Getting the distinction right matters because the FMSAT is one of those terms and not the others.

    A screener is a brief, low-burden tool designed to flag people who may benefit from further evaluation. Screeners take a minute or two, can often be administered by people without specialized clinical training, and produce a simple signal that says proceed to next step or no further action needed. Screeners do not diagnose, do not establish eligibility, and do not produce a comprehensive clinical picture. Examples in healthcare include the M-CHAT for autism, the PHQ-9 for depression, and vision and hearing screens in schools.

    An assessment is a more comprehensive structured evaluation that produces detailed information sufficient to support diagnosis, eligibility decisions, and treatment planning. Assessments typically take 30 to 90 minutes, require trained administrators, and produce multiple subscores or domain scores. Examples in pediatric OT include the Beery-Buktenica Developmental Test of Visual-Motor Integration, the Peabody Developmental Motor Scales, the Bruininks-Oseretsky Test of Motor Proficiency, the Sensory Processing Measure, and the Pediatric Evaluation of Disability Inventory Computer Adaptive Test (PediCAT).

    An evaluation is the broader clinical process that uses one or more screeners and assessments together with observation, interview, and chart review to produce a full clinical picture and recommendations.

    The FMSAT (Fine Motor Speed and Accuracy Test) is a screener. It is not designed to replace any assessment in your toolkit. It is designed to do the screener job well, which means flagging test takers who may benefit from a comprehensive evaluation when their lateralization or fine motor speed scores fall outside expected ranges. The assessments that follow a positive FMSAT screen will be whatever your clinical reasoning and your setting indicate. The point of a good screener is to make sure the right people get to those assessments faster than they would have otherwise.

    The FMPRS (Fine Motor Participation Rating Scale), on the other hand, is functioning more like a brief assessment instrument during the validation phase of this project, because its job is to produce the multi-domain rating data needed to establish ecological validity for the FMSAT. After validation is complete, the FMPRS will not typically be used alongside the FMSAT in routine clinical screening.

    What the FMSAT measures and why it works across the lifespan

    The FMSAT, which stands for Fine Motor Speed and Accuracy Test, is a one-minute screener. The test taker is given a bubble-popping worksheet and asked to pop as many bubbles as they can in 30 seconds with one hand, then 30 seconds with the other hand. The score for each hand is the number of bubbles popped. That is it. No rater training, no expensive test kits, no proprietary materials beyond a printed worksheet.

    The name says fine motor speed and accuracy, and that is what each hand score reflects. But the construct the instrument was designed to measure, and the reason the two hands are tested separately, is neuromotor lateralization. Lateralization is the degree to which a person has committed one side of the brain, and therefore one hand, to specialized motor work. Strong lateralization means the dominant hand performs precision tasks fluently while the non-dominant hand serves as a stabilizer. Weak or absent lateralization means the two hands perform more similarly, hand preference is inconsistent, or the person switches hands mid-task. Fine motor speed is the metric we use to detect that pattern, because timed performance under demand reveals the lateralization signal more clearly than untimed observation does.

    This matters because lateralization is not just a developmental milestone of early childhood. It is a lifelong neuromotor property that emerges during the preschool years, consolidates through school age, holds through adulthood, and can shift or erode in response to neurological changes later in life. The FMSAT was designed to capture that lateralization signal at any age, which is why our normative dataset spans from preschool through older adulthood rather than stopping at age 12 or 17.

    What makes the FMSAT useful is not the bubble popping itself. It is what the two hand scores together tell us about how well a person has consolidated hand dominance, and how that pattern compares to typical lateralization for their age. The dominant hand score reflects fine motor capability. The difference between the two hands reflects lateralization. Both pieces of information matter for clinical and educational practice, and neither one is captured well by the standardized assessments most occupational therapists currently use.

    How the FMSAT can be used: one-time screener, progress monitoring, and lifelong tracking

    Because the FMSAT is fast, standardized, and produces a numeric score, it can support several clinical and educational use cases that most current assessments cannot. Three are worth naming explicitly.

    Use case 1: One-time screening for hand dominance and fine motor concerns

    This is the most familiar use case. A pediatric occupational therapist, school OT, or educator administers the FMSAT to a student who has been flagged for fine motor concerns, handwriting difficulty, or unclear hand dominance. The score, compared to age-appropriate normative bands, gives a quick objective marker of whether the student’s performance and lateralization fall within expected ranges. This is the use case the validation manuscript will focus on first.

    Use case 2: Response to Intervention and progress monitoring

    Because the FMSAT takes one minute and produces a numeric score, it can be re-administered at intervals to track whether a person is making measurable progress over time. This is exactly the kind of brief repeatable measurement that Response to Intervention frameworks and Multi-Tiered Systems of Support models require. A school OT could administer the FMSAT at the start of an intervention block, midway through, and at the end, and have objective data showing whether the lateralization or fine motor speed measure is changing in response to the intervention. A clinic-based OT could do the same across a course of therapy. The same minimum-detectable-change thresholds that value-based payment models are increasingly requiring would apply directly. Once the normative dataset is finalized, the FMSAT becomes one of the few fine motor measures fast enough to use for routine progress monitoring.

    Use case 3: Educator-administered screening in school settings

    The FMSAT requires no rater training and no clinical interpretation to administer. A teacher, paraprofessional, school nurse, or interventionist can give the test and record the scores. With enough normative data in place, the score itself does the screening work and a teacher does not need to be an OT to identify a student who falls below the expected band. This opens the door to universal screening at the classroom or grade level, the same way schools currently screen vision and hearing. Students flagged by the screener can then be referred to occupational therapy for full evaluation. This is a long-term vision rather than an immediate use case, but it is exactly the kind of MTSS Tier 1 screening application that schools have been asking for and that occupational therapy has not yet had the tools to support.

    Use case 4: Lifespan monitoring, including progressive neuromotor conditions

    This is the use case that emerges directly from norming the FMSAT across all ages rather than only in pediatrics. Lateralization can shift across the lifespan in response to neurological events and conditions. A person recovering from stroke may show changed dominant-hand performance. A person living with multiple sclerosis, Parkinson’s disease, or another progressive neuromotor condition may show gradual erosion of the lateralization signal as the condition advances. A person experiencing age-related changes in motor control may show compression of the dominant-hand advantage. The FMSAT, administered periodically, could detect those shifts earlier and more objectively than self-report or general clinical observation. With enough normative data across all ages, the same one-minute test that screens a five-year-old for emerging hand dominance could monitor a 55-year-old neurologist for early signs of motor change in the years after a diagnosis. That is the broader clinical reach the lifespan normative dataset makes possible.

    None of these use cases require a separate test. They all use the same one-minute bubble-popping task. What changes is who is administering it, how often, and what the score is being compared against. That flexibility is one of the reasons we are investing in the validation work the way we are.

    What is Rasch analysis and why is it considered the gold standard?

    Many occupational therapy assessments you have used in your career, including those most commonly used to demonstrate progress in pediatric settings, were developed using Classical Test Theory, often shortened to CTT. The Beery-Buktenica Developmental Test of Visual-Motor Integration, the Bruininks-Oseretsky Test of Motor Proficiency, the Peabody Developmental Motor Scales, and the Sensory Processing Measure are CTT-based instruments. CTT produces a total score that is then compared to a normative sample. The score tells you where a test taker falls relative to peers, but it has some real limitations.

    The biggest limitation is that CTT treats every item on a test as if it were equally difficult, and every point on the score scale as if it represented the same amount of skill. A test taker who scores 84 versus one who scores 89 may differ by a meaningful amount of skill, or may differ by almost nothing, depending on where on the scale those scores fall and which specific items they got right. CTT cannot tell you which.

    This is true for both screeners and assessments built under CTT, though the limitation is more consequential for assessments because assessments are doing more of the clinical decision-making work. A screener producing a CTT score is still useful as a yes/maybe/no signal. An assessment producing CTT scores is being asked to support diagnosis, eligibility, and treatment planning decisions on the same imprecise measurement scale.

    Rasch analysis, developed by Danish mathematician Georg Rasch in the 1960s, takes a fundamentally different approach. Rasch places each test item and each person on the same interval scale, called the logit scale. This means the distance between a score of 5 logits and 10 logits represents the same amount of skill change as the distance between 20 and 25. Rasch also calibrates each item individually, telling you which items are easy, which are hard, and whether each item is actually pulling its weight in measuring the construct.

    Rasch is to assessment what a ruler is to measurement. CTT scores tell you a child is somewhere in the middle. Rasch scores tell you exactly where, on a scale where the units are equal.

    Rasch analysis is considered the gold standard for modern instrument development because it is the framework used by some of the most respected and most defensible pediatric assessments in the field. The PEDI-CAT, the AMPS or Assessment of Motor and Process Skills, the School Function Assessment, and the WeeFIM are all Rasch-based or Item Response Theory based instruments. These are the tools that produce data insurance companies and researchers trust. If you have ever wondered why some assessments seem to have stronger research backing than others, the framework behind them is usually a big part of the answer.

    Behind the scenes: the first Rasch checkup on the FMPRS

    The FMPRS, or Fine Motor Participation Rating Scale, is the 12-item observer-rated companion to the FMSAT. It captures four dimensions of fine motor function across real-world tasks: fine motor speed, fine motor precision, laterality and bilateral differentiation, and participation in daily roles. Our Rasch consultant, Angie, ran the first calibration on 139 paired FMSAT and FMPRS records earlier this month. The results were strong for a first calibration on a relatively small sample.

    Finding one: the item difficulty order matches the developmental theory

    Rasch ordered the 12 items from easiest to hardest based on how raters actually responded to them. The Laterality items, which ask about hand dominance and bilateral hand use, came out as the easiest. The Participation items, which ask about sustained engagement in fine motor tasks throughout the day, came out as the hardest. This is exactly what we would predict developmentally. Hand dominance consolidates earlier in childhood than sustained occupational engagement, so on a rating scale measuring fine motor function across the developmental arc, laterality items should be easier to endorse than participation items. The data confirmed the theory.

    Finding two: the instrument separates clinical and typical test takers cleanly

    On the Rasch-derived person measure, typical preschoolers scored more than two logits higher than clinical preschoolers. In plain terms, the typical preschooler scored higher than approximately 98 percent of the clinical sample. This is what is called a known-groups validity effect, and a Cohen’s d effect size above 2.0 is unusually strong. Most pediatric assessments are pleased to show known-groups effects in the 0.5 to 0.8 range. The FMPRS is producing separation that is roughly three times stronger. As the normative dataset grows beyond preschool ages, we expect the same separation pattern to extend across older age bands and into adult populations where clinical and typical comparison groups can be defined.

    Finding three: the FMPRS and the FMSAT are picking up the same underlying construct

    The Rasch-derived person measure on the FMPRS correlated with FMSAT dominant hand scores at r = 0.575. That correlation tells us that two completely different methods of measurement, a one-minute performance task and a 12-item observer rating scale, are picking up the same underlying construct. This is exactly the kind of cross-method convergence a strong validation manuscript needs.

    What still needs work

    Five items in the FMPRS came back as overfitting in the Rasch model, which means they are too internally redundant. They are not bad items. They are just not adding as much new information as they could. Those items will be revised for the next version of the FMPRS based on what the calibration showed us. This kind of iterative refinement is how serious instrument development works, and it is exactly why we are doing this calibration now, before the manuscript is finalized.

    Ever wondered why Beery, PDMS-3, or BOT-2 scores do not match real-world function?

    Here is a question every occupational therapist has wrestled with at some point in their career. Why do scores from the Beery-Buktenica VMI, the Peabody Developmental Motor Scales, or the Bruininks-Oseretsky Test sometimes fail to line up with what you actually see in the classroom, at home, on the playground, or in adult daily life?

    You are not imagining it. A student can score below average on a tabletop visual motor task and still write legibly, manage their lunchbox, and participate fully in PE. Another can score in the average range and still struggle every single day to keep up with handwriting demands or self-care routines. The same pattern shows up in adult assessments. The mismatch is real, and it has a name in the psychometric literature. It is called an ecological validity gap.

    What ecological validity actually means

    Ecological validity is a psychometric term that answers a simple question: does this assessment measure something that actually matters in real life? An instrument with strong ecological validity produces scores that connect to how a person functions in their daily environment. An instrument with weak ecological validity produces scores that connect mainly to how a person performs on the test itself, with limited evidence that the score predicts daily function.

    Many of the assessments OTs use most often were designed to measure isolated motor performance, not real-world participation. They are good at what they measure. They were just never built to answer the participation question. When a school-based therapist is asked to defend a Beery standard score at an IEP meeting, or when a clinic-based therapist is asked to justify medical necessity to an insurer using a PDMS-3 score, the underlying problem is often that the score is being asked to do something the assessment was not designed to do.

    Why the FMPRS is essential to FMSAT validation, even though clinicians will not use both in practice

    Here is a question that comes up almost every time I explain this project. If the FMSAT is a screener, why pair it with the FMPRS at all? Won’t clinicians be expected to do both?

    The answer is no. The FMPRS is doing critical work right now, during validation, so that the FMSAT will not need to be paired with it in clinical practice later. The whole value of the FMSAT as a screener depends on it being a one-minute, paper-and-pencil, standalone score. Asking clinicians to also complete a 12-item rating scale every time they administered the screener would defeat the entire point of having a screener in the first place.

    What the validation work establishes, and what every paired FMSAT and FMPRS submission helps establish, is that when the FMSAT score is elevated or compressed in a particular way, it is reflecting something that shows up in the test taker’s real-world functional life. Once that link is established and published in a peer-reviewed manuscript, the FMSAT score on its own carries that ecological meaning forward. Clinicians using the FMSAT in practice will be able to point to the published validation evidence rather than having to demonstrate the connection every time.

    This is the same approach used to validate other widely accepted screeners. The Modified Checklist for Autism in Toddlers, known as the M-CHAT, was validated against full ADOS and ADI-R diagnostic batteries. Pediatricians using the M-CHAT today do not run an ADOS alongside it. The validation work was done once and the screener now stands on its own. The PHQ-9 depression screener was validated against structured psychiatric interviews. Primary care providers use it on its own today. The pattern is consistent across well-validated screeners. Validate against a richer companion measure once, publish the validity evidence, then use the screener on its own.

    Why the FMSAT itself uses Classical Test Theory and the FMPRS uses Rasch

    This is a question that any sharp reader will be asking by now. If Rasch is the gold standard, why is the FMSAT being calibrated using Classical Test Theory rather than Rasch?

    The answer comes down to the measurement structure of each instrument. The FMSAT produces raw bubble counts on a zero to 80 scale for each hand. That kind of continuous count data is well-suited to CTT-style descriptive statistics, percentile norms, and known-groups validity comparisons. These are the analyses the FMSAT validation manuscript will lean on, and they are the analyses that produce the percentile bands and severity cutoffs clinicians actually use at the point of care.

    Rasch is the right framework for the FMPRS because the FMPRS uses ordered category responses on a four-point scale, and each item can be at a different difficulty level on the same underlying trait. That is exactly the kind of measurement structure Rasch was designed to handle. The two instruments are built differently on purpose, and each one is being analyzed using the framework that fits its measurement structure.

    The Rasch work on the FMPRS gives the manuscript the modern psychometric backbone peer reviewers expect. The CTT work on the FMSAT keeps the screener simple and interpretable for the clinicians who will actually use it. Both pieces matter, and together they create a defensible validation argument.

    Why value-based care is about to make all of this matter much more than it has before

    If you have been practicing for more than a few years, you already know that occupational therapy reimbursement has been under pressure for a long time. The 2026 Medicare Physician Fee Schedule final rule from the Centers for Medicare and Medicaid Services, released October 31, 2025, continued a trend of flat or declining payment rates for outpatient OT. According to OT Potential’s 2026 reimbursement analysis, the proposed 1% decrease to OT and PT relative value units for 2026 came after a 0% increase in 2025 and a 3% decrease in 2024. The trajectory is real and most OTs feel it directly in their paychecks.

    What is changing right now, and what most clinicians have not fully internalized, is that the structure of the entire payment system is shifting underneath us. The shift is called value-based care.

    What value-based care actually means

    Value-based care is a payment model where providers and health systems are reimbursed based on the outcomes their patients achieve, not the volume of services delivered. Under traditional fee-for-service payment, an OT bills for each visit and gets paid for each visit. Under value-based care, payment is increasingly tied to whether the patient demonstrated measurable functional improvement against established benchmarks.

    CMS has been driving this shift through several specific programs. The Quality Payment Program, the Merit-based Incentive Payment System known as MIPS, and an expanding suite of Alternative Payment Models are all moving rehabilitation services toward outcomes-based reimbursement. The 2026 payment updates included a 0.75% increase for qualified APM participants and a 0.25% increase for everyone else, an early but clear signal that participating in alternative payment models will increasingly be where the financial upside is.

    The era of writing patient made progress toward goals in a discharge note and being reimbursed for it is ending. The era of demonstrating measurable functional change against defensible benchmarks is beginning.

    Why occupational therapy is structurally underprepared for this shift

    Here is the connection most clinicians have not drawn explicitly. The shift to value-based care requires outcomes data, and outcomes data is only as good as the assessments and progress-monitoring instruments producing it. Most of the assessments occupational therapists use to demonstrate progress, and most of the screeners they use to identify who needs services in the first place, were developed under CTT frameworks that produce raw scores, percentile bands, and standard scores. None of those formats give insurers what value-based payment models actually require.

    Insurers under value-based care want interval-level evidence of measurable functional change against established minimum-detectable-change thresholds. When a third-party reviewer asks whether a patient made meaningful progress, the answer they want is not, the patient’s standard score improved from 84 to 89. The answer they want is, the patient’s interval-level fine motor measure shifted by 0.45 logits, which exceeds the minimum detectable change threshold of 0.30 logits established in the calibration sample. One of those answers is opinion-vulnerable. The other is data.

    Closing this evidence gap requires investment at every level of the measurement pipeline: better screeners that identify who needs services earlier and more accurately, better assessments that produce the diagnostic and eligibility data on a defensible measurement scale, and better outcomes instruments that document functional change over an episode of care. The FMSAT and the FMPRS sit at the screener and ecological validity ends of that pipeline. They are one contribution among many that the profession needs.

    This evidence gap is one of the underrecognized reasons our profession has struggled to make the reimbursement case at the level of physical therapy or speech-language pathology. Both adjacent professions have invested more heavily in Rasch-calibrated, IRT-based assessment development over the last 20 years. The PEDI-CAT, the AM-PAC, and similar tools represent what that investment looks like. Occupational therapy has far fewer Rasch-calibrated tools across the screener, assessment, and outcomes layers, and that thinness in our measurement infrastructure shows up downstream as flatter reimbursement, narrower coverage policies, and ultimately compensation that has not kept pace with the cost of the training required to enter the field.

    How the O.T. Wizard platform fits into this picture

    The FMSAT and the FMPRS are not standalone projects. They are pieces of a larger evidence infrastructure being built into the O.T. Wizard platform, which is being rebranded as MyTherapyWizard.

    O.T. Wizard is a digital evaluation and outcomes platform built from the ground up on modern psychometric standards. The FMSAT lives inside the platform as a fast, defensible screener with clear research foundations. The FMPRS lives alongside it as the ecological validity companion during validation. The broader platform houses structured evaluation templates designed to produce the kind of data that holds up under value-based payment scrutiny. The architecture is PHI-free and operates under a 1EdTech-approved legal framework, which means the platform itself functions as a passive-accrual research engine. Every paired evaluation contributes to the dataset that makes the next generation of assessments stronger.

    The rebrand to MyTherapyWizard reflects the platform’s expanding scope beyond occupational therapy into a multi-discipline space for pediatric therapy professionals. The underlying mission stays the same. Build the measurement infrastructure our profession needs to move forward, in step with where reimbursement is going rather than chasing it after the fact.

    What you can do

    Our profession needs more Rasch-calibrated assessments. It needs more validated rating scales with strong ecological validity. It needs more normative datasets large enough to defend in peer review. And it needs more clinicians and educators willing to contribute the data that makes all of that possible.

    If you are an occupational therapist, an educator, a clinician working with adult or geriatric populations, or anyone interested in supporting the development of evidence-based assessment tools, here is how to get involved:

    • Request to be on the Norming Tryout Team and Contribute FMSAT data. The screener takes one minute per test taker. If you administer it after a session or screening you would have run anyway, the marginal time cost is essentially zero.
    • Complete the FMPRS when you can. The paired data is what makes the validation manuscript possible. Every paired submission directly strengthens the published evidence base our profession will use.
    • Look for the bands we need most. As of this week, the most urgent recruitment gaps are adolescents ages 12 to 17, both clinical and typical, three-year-olds in both groups, adults age 50 and older, and left-dominant test takers at every age.
    • Share this work with colleagues. The bigger and more representative the dataset, the stronger the eventual screener will be for the people you serve.

    Every paired submission you contribute is a small but real piece of building the measurement infrastructure our profession needs. Building a Rasch-calibrated rating scale and a CTT-validated performance screener together, on a normative dataset large enough to defend in peer review and broad enough to span the lifespan, is exactly the foundational psychometric work the field has needed for years. If we want occupational therapy to be reimbursed at the level our training and clinical expertise warrant under the new value-based payment models, we have to produce the kind of evidence other professions have already produced.

    That work does not happen in conference panels or position papers. It happens in datasets, calibrations, and validation manuscripts. It happens in projects like this one. And the people producing it are not academics in distant labs. They are clinicians and educators like you who choose to spend a few minutes on a Tuesday afternoon contributing to something larger than a single evaluation.

    About this project

    The FMSAT, Fine Motor Speed and Accuracy Test, is a one-minute screener for neuromotor lateralization and fine motor speed, currently in active normative data collection across the lifespan toward a peer-reviewed validation manuscript. The FMPRS, Fine Motor Participation Rating Scale, is the 12-item observer-rated companion used to establish ecological validity during the validation phase. Both instruments are part of the O.T. Wizard platform, rebranding to MyTherapyWizard. Pearl IRB Not Human Subjects Research determination on file (ID 2026-0154).

    Related topics Neuromotor lateralization assessment, hand dominance evaluation across the lifespan, Rasch analysis in rehabilitation, fine motor screening for educators, Response to Intervention RTI fine motor measures, MTSS Tier 1 and Tier 2 screening, ecological validity in occupational therapy, value-based care for outpatient therapy, Medicare Physician Fee Schedule 2026, evidence-based occupational therapy practice, school-based occupational therapy, OT reimbursement, alternative payment models for rehabilitation, progressive neuromotor condition monitoring, multiple sclerosis fine motor tracking, stroke rehabilitation outcomes measurement, lifespan motor assessment

  • OT Wizard Psychometric Validation: What It Means for Evidence-Based Practice

    OT Wizard Psychometric Validation: What It Means for Evidence-Based Practice

    I’m thrilled to announce that OT Wizard has officially begun psychometric validation through Rasch analysis – a major milestone in our journey to elevate evidence-based practice in pediatric occupational therapy!

    What’s Happening Now

    We’ve sent evaluation data for ages 4-5 years (Age Bands H & I) to an independent psychometrician for comprehensive analysis. With over 400 evaluations from 18 therapists across North Carolina, we have robust data to validate that OT Wizard measures what we say it measures & accurately, reliably, and fairly.

    What Is Rasch Analysis?

    Rasch analysis is a sophisticated psychometric approach that goes beyond traditional test validation. Unlike norm-referenced assessments that simply compare students to each other, Rasch analysis creates an interval-level measurement scale, similar to measuring temperature or weight.

    Think of assessments you may know that use Rasch methodology:

      • PEDI-CAT (Pediatric Evaluation of Disability Inventory – Computer Adaptive Test) – Rasch-calibrated functional assessment

      • AMPS (Assessment of Motor and Process Skills) – fully Rasch-calibrated for ADL performance

      • COPM (Canadian Occupational Performance Measure) – uses Rasch principles for measuring occupational performance

      • HELP (Hawaii Early Learning Profile) – Rasch-validated developmental assessment

      • Original PEDI – Rasch-based functional assessment

    These assessments are considered gold standards because Rasch analysis ensures:

      • Equal intervals: A 10-point gain at any level represents the same amount of growth

      • Sample-independent measurement: Item difficulty doesn’t depend on who takes the test

      • Missing data handling: Scores are valid even when not all items are administered (adaptive testing)

      • Precise error estimation: Know exactly how confident you can be in each score

      • Item hierarchy validation: Confirms items are developmentally sequenced correctly

    Why Rasch Instead of Traditional Norming?

    Traditional norm-referenced tests (like BOT-3, PDMS-3) require testing typically-developing children to create percentile ranks. That’s valuable, but has limitations:

    Norms become outdated (tests re-normed every 15-20 years)
    Percentiles are ordinal, not interval (85th→95th ≠ 15th→25th in actual ability)
    Can’t track growth accurately across different ability levels
    Require complete test administration

    Rasch analysis provides:

    Continuous measurement scale – Track growth precisely over time
    Adaptive testing – Administer only relevant items, still get accurate scores
    Sample-independent – Item difficulty stays stable regardless of who’s tested
    Living calibration – Can update and refine continuously with new data
    Clinical utility – Scores directly interpretable for intervention planning

    We’re building OT Wizard to work like the PEDI-CAT and AMPS.   These are tools that OTs trust because they’re built on rigorous Rasch foundations.

    What We Already Know About Our Data

    Before even sending data to our psychometrician, we conducted preliminary analysis on our 404 evaluations to ensure data quality. Here’s what we’ve learned:

    Strong Sample Characteristics

      • Well-balanced age distribution: 184 evaluations (ages 4.0-4.4) and 220 evaluations (ages 4.5-4.9)

      • 18 therapists contributing: Average of 22 evaluations each, with range of 7-45 per therapist (good for inter-rater reliability)

      • Gender representation: 56% male, 44% female (reflects typical OT referral patterns)

      • Diverse language backgrounds: 90% English, 8% Spanish, 2% other languages

      • Clinical population validity: 96% recommended for OT services

    Excellent Data Completeness

      • Zero missing responses – every administered item was answered

      • 75.5% average completion rate – our adaptive basal/ceiling rules are working perfectly

      • 19,062 total data points across 74 items and 8 domains

    Strong Domain Coverage

      • Visual Perception: 15 items

      • Activities of Daily Living: 14 items

      • Gross Motor & Fine Motor: 10 items each

      • Participation: 9 items

      • Executive Functioning: 8 items

      • Visual Motor Integration: 5 items

      • Praxis: 3 items

    Areas for Improvement Identified

    Our preliminary analysis flagged several items for the psychometrician to examine closely:

    Ceiling Effects in Visual Perception (Preschool evaluation): 56% of responses scored at ceiling (mastered), with 33% at floor (not yet observed). This bimodal pattern suggests we may need more mid-difficulty items to better differentiate students in the middle range.

    Rating Scale Consistency: A small number of responses (3.3%) showed raw scores instead of normalized scores, indicating a formula issue we’ve already corrected.

    Developmental Anchor Gaps: About 54% of items (primarily Executive Functioning and Participation domains) lack developmental anchors. The Rasch analysis will empirically determine difficulty levels so we can assign appropriate anchors.

    Item-Age Alignment: Many items administered are anchored above student age ranges (51-55% above range). This is actually expected and appropriate.  Students with developmental delays are working on skills typically seen at older ages. However, Rasch will help us recalibrate anchors based on clinical population performance vs. typical development.

    ✅ Best-Performing Domains

      • ADL: Excellent distribution with only 8% floor and 8% ceiling – items are well-targeted

      • Executive Functioning: Minimal floor effect (0.7%), good spread across ability levels

      • Participation: Near-zero floor (0.2%), strong measurement potential

    This preliminary work means we’re sending clean, robust data to our psychometrician.  This is maximizing the value of the Rasch analysis and ensuring reliable results.

    What’s Being Analyzed

    Our psychometrician is conducting comprehensive Rasch analysis across multiple dimensions:

    1. Construct Validity (Unidimensionality)

    Do items within each domain (Gross Motor, Fine Motor, Visual Perception, etc.) measure a single, coherent construct? This is critical for Rasch – if items don’t “hang together,” they can’t be on the same measurement scale.

    2. Item Fit

    Which items contribute to reliable measurement? Rasch provides specific fit statistics (infit/outfit MNSQ) showing whether each item:

      • Is too predictable (doesn’t add information)

      • Is too unpredictable (confuses the measurement)

      • Functions optimally (contributes to precise measurement)

    Items outside acceptable ranges get flagged for revision or removal.

    3. Rating Scale Functioning

    Do our 5-point performance bands (Beginning → Mastered) function as intended? Rasch examines:

      • Are all categories used appropriately?

      • Do response thresholds advance in the right order?

      • Should categories be collapsed (e.g., 5-point → 3-point)?

    This is similar to how AMPS validates its 4-point scoring scale.

    4. Item Hierarchy

    Rasch places all items on a single difficulty scale (measured in logits). We’ll see if:

      • Items anchored at 48 months are empirically easier than 54-month items

      • Our developmental sequencing matches actual difficulty

      • Gaps exist where we need additional items

    This will be especially important for items currently lacking developmental anchors – the Rasch analysis will tell us where they belong.

    5. Measurement Precision

    Unlike traditional reliability (one number for whole test), Rasch shows precision at every ability level:

      • Where is measurement most accurate?

      • What’s the standard error at our 70% clinical threshold?

      • Can we distinguish between students with small ability differences?

    6. Differential Item Functioning (DIF)

    Do items work the same way for:

      • Boys vs girls?

      • 4-year-olds vs 5-year-olds?

      • English vs Spanish speakers?

      • Different diagnoses?

    Items showing bias get flagged or removed – ensuring fairness.

    7. Person Separation

    Can we reliably distinguish between students at different ability levels? Rasch provides a separation index showing how many distinct ability levels we can measure. Higher separation = more precise clinical distinctions.

    8. Addressing Known Issues

    The psychometrician will specifically examine:

      • Visual Perception’s ceiling effects : do we need additional mid-difficulty items?

      • Praxis domain with only 3 items : is this sufficient or should it combine with another domain?

      • Executive Functioning and Participation rating scales : do they function as separate constructs from performance-based items?

    Why This Matters for You

    Rasch validation transforms how you can use OT Wizard scores:

    Meaningful Progress Monitoring

    Because Rasch creates interval-level measurement, you can confidently say:

      • “Student gained 0.8 logits in 6 months”

      • “This represents clinically significant progress”

      • “Growth rate exceeds typical intervention response”

    Traditional percentage scores can’t make these claims and a jump from 40% to 50% isn’t necessarily the same growth as 70% to 80%.

    Adaptive Testing Validation

    Like the PEDI-CAT and AMPS, OT Wizard uses basal/ceiling rules so students aren’t frustrated with too-hard items or bored with too-easy ones. Rasch analysis confirms:

      • Scores are comparable even when different items are administered

      • Our 75% completion rate is optimal

      • Missing items are appropriately “not administered,” not missing data

    Credible Clinical Decisions

    When you document that a child’s gross motor ability is at -1.2 logits:

      • Insurance companies recognize Rasch-based measurement

      • School districts understand the methodology (same as PEDI-CAT/HELP)

      • You can defend your clinical reasoning with published psychometric evidence

    Item-Level Interpretation

    Rasch analysis creates item hierarchy maps showing exactly which skills a child has mastered, which are emerging, and which aren’t yet present. This directly informs intervention planning  just like how AMPS users identify specific ADL breakdowns or PEDI-CAT shows functional skill patterns.

    Why Start with Ages 4-5?

    We strategically chose this age range because:

      1. Sufficient sample size: 400+ evaluations provide robust statistical power for Rasch analysis

      1. Diverse representation: Students with various diagnoses, languages, and ability levels

      1. Multiple raters: 18 different therapists ensure inter-rater reliability analysis

      1. Item overlap: Many items in this age range also appear in adjacent ages, so findings inform the entire platform

      1. Strong data quality: Our preliminary analysis confirmed excellent completion rates and coverage

    The psychometrician will identify any problematic items, validate our developmental anchors, assign anchors to items missing them, and ensure rating scales function optimally. We’ll implement improvements before these issues cascade into other age bands.

    What Happens Next

    Based on the psychometrician’s findings (expected in 4-5 weeks), we’ll:

      1. Remove or revise misfitting items that don’t meet Rasch fit criteria

      1. Add mid-difficulty items to Visual Perception domain to address ceiling effects

      1. Optimize rating scales if analysis shows categories aren’t functioning as intended

      1. Recalibrate existing anchors if clinical population performance differs from typical development

      1. Establish measurement precision estimates at different ability levels

      1. Publish validation statistics you can cite in reports and presentations

    This refined version becomes the foundation for validating additional age bands.

    Expanding Validation: Ages 3 Months to 12 Years

    Over the next 12 months, as we reach 200+ evaluations per age band, we’ll validate each additional age group. This will create a comprehensive, linked measurement system similar to how PEDI-CAT links across age ranges  where we can:

      • Track individual students across multiple years on the same logit scale

      • Provide age-equivalent scores based on item difficulty calibration

      • Create clinical reference data comparing students receiving OT services

      • Document growth trajectories with true interval-level measurement

      • Demonstrate outcomes with unprecedented precision

    By Month 12, we’ll conduct a comprehensive linking study that places all age bands (3 months through 12 years) on a single, continuous measurement scale. Items that appear in multiple age bands will “anchor” the scales together, ensuring continuity.

    This approach mirrors how major Rasch-based assessments (PEDI-CAT, AMPS, HELP) maintain measurement continuity across ages and versions.

    How You Can Support This Work

    1. Keep Using OT Wizard

    Every evaluation you complete contributes to our growing database. Rasch analysis becomes more robust with larger samples and the more data we collect, the more confident we can be in item calibrations.

    administering OT evaluation with child.

    Brittany B., administers an eval with OT Wizard

    2. Share Your Clinical Insights

    If you notice items that seem:

      • Confusing or ambiguous to score

      • Too easy or too hard for the age range

      • Misaligned with what you observe clinically

      • Culturally biased or inappropriate

    Please let us know! Your real-world feedback is invaluable. Rasch analysis will identify statistical misfits, but your clinical judgment helps us understand why items aren’t working.

    What This Means for Our Profession

    Most therapy documentation tools rely on subjective clinical observation. While tools like PEDI-CAT, AMPS, COPM, and HELP exist and use Rasch methodology, they’re limited in scope and focus on narrow or specific functional domains, requiring specialized training, or covering narrow age ranges.

    OT Wizard is different: We’re creating a comprehensive, Rasch-validated clinical intelligence platform that covers:

      • Multiple domains (gross motor, fine motor, visual perception, praxis, executive functioning, ADL, participation)

      • Birth through early adulthood (3 months – 18 years)

      • Performance-based, participational based, and functional assessment

      • Integrated into everyday clinical workflow

    By pursuing rigorous Rasch validation, we’re increasing psychometric rigor to comprehensive pediatric OT assessment.

    When this validation is complete, you’ll be able to say:

    “I use OT Wizard, a comprehensive Rasch-validated pediatric OT assessment platform with published psychometric evidence across 2,000+ evaluations providing the same measurement quality as tools like PEDI-CAT and AMPS, but covering all developmental domains.”

    The Vision: Living Calibration and Real-Time Data

    Here’s what excites me most about Rasch methodology: Unlike traditional norm-referenced tests that become frozen in time, Rasch-calibrated assessments can be continuously refined.

    The PEDI-CAT has demonstrated this – as more data is collected, item calibrations can be updated, new items added, and measurement precision improved all while maintaining the same measurement scale.

    OT Wizard will have “living calibration”:

      • Continuous item refinement as we collect more data

      • New items added to fill gaps in difficulty coverage

      • Real-time quality monitoring

      • Annual recalibration studies

      • Regional and demographic analyses

    Imagine:

      • Item difficulties that reflect current populations

      • Outcome analytics showing which interventions are most effective

      • Predictive data identifying which early skills best predict later success

      • The world’s largest real-time Rasch-calibrated pediatric development database

    Every evaluation you complete contributes to this unprecedented resource.

    💕Thank you for building the future of evidence-based OT assessment with O.T. Wizard.