Tag: evidence-based

  • Built on Evidence. Proven in Practice. Shoutout to Learning Charms’ Team

    Built on Evidence. Proven in Practice. Shoutout to Learning Charms’ Team

    How 3.8 Years of Systematic Clinical Measurement Demonstrates That Occupational Therapy Works

    Stephanie Seymore Wick, MSOT, OT/L | Founder and Clinical Architect, O.T. Wizard | Learning Charms, Inc., Charlotte, North Carolina

    The Problem With Checklists

    For years, occupational therapists working in early childhood settings were collecting data that told them almost nothing. Checklist-style evaluations produced a snapshot: present or absent, yes or no. They could not tell you whether a child improved. They could not tell you which counties had greater concentrations of developmental need. They could not tell you whether your team’s intervention was moving the needle or whether children were simply getting older.

    That was the reality facing Learning Charms in 2022. We were screening and evaluating large numbers of children across Head Start programs, NC Pre-K classrooms, and community settings throughout North Carolina, and we had nothing meaningful to show for it in terms of trend data, geographic insight, or outcome evidence.

    So I built something.

    From Nothing to 8,509 Screenings and Evaluations

    The FUNdamental Foundations (FF) screener was designed and developed by a managing pediatric occupational therapist with 25+years of clinical experience. It was built as a structured, multi-domain developmental tool designed from the outset to generate analyzable data. It was not designed for publication. It was designed to answer clinical questions: What does this population look like? Where are the gaps? Is what we are doing making a difference?

    Thirty clinicians on the Learning Charms team tested and used each version in the field, providing the real-world feedback that drove every refinement. They made the transition from paper-and-pencil evaluations to digital data entry on a tablet or laptop, mid-session, with children in front of them. That is not a small ask. The early weeks required support with the technology. There were growing pains. The team did it anyway, and they did it without much, if any complaint.

    The FF tool went through two versions, each refined based on team feedback. Version 6 ran from June 2022 through July 2023. Version 7, with improvements including date of birth capture and an updated item structure, ran from August 2023 through May 2025. In late 2025, the practice transitioned to O.T. Wizard, a fully rebuilt clinical intelligence platform designed and built by the same therapist.  O.T. Wizard was built with Rasch psychometric architecture, 15 guided evaluations, and integrated outcome tracking across 12 domains.

    The table below summarizes what 3.8 years of that effort produced.

    Table 1. Clinical Data Collected Across the Full Evidence Ecosystem (2022-2026)

    PlatformPeriodRecordsEvaluationsScreeningsE1-E2 Pairs
    FUNdamental Foundations V6Jun 2022 – Jul 20232,9281,5111,417324
    FUNdamental Foundations V7Aug 2023 – May 20254,9622,3682,594477
    O.T. WizardSep 2025 – Mar 202661961997
    TOTAL3.8 years8,5094,4984,011898

    Note. E1-E2 pairs = children with two complete evaluations allowing pre-to-post comparison. FF V6 pairs are V6-only matches. FF V7 pairs include V7-only and cross-version (V6 E1 to V7 E2) matches. OTW pairs matched by Student_ID. pp = percentage points.

    In total: 8,509 individual assessment records. 4,498 full evaluations. 4,011 developmental screenings. 898 pre-to-post evaluation pairs. Over 245,000 item-level data points. Collected by a single clinical team, through routine practice, over less than four years.

    What the Data Shows: Gains That Exceed Maturation

    The central question in any clinical outcome dataset without a randomized control group is this: how do you know the gains are from intervention and not just from children getting older?

    We address this directly.

    Using cross-sectional developmental data from our own E1 (Initial Evaluation) dataset, we calculated the expected rate of developmental growth per month for each skill area based on age alone. This gives us a maturation baseline specific to this population. We then compared that expected gain to the gains actually observed in children who received OT services between E1 and E2 (Re-evaluation), over a mean interval of 5.4 months.

    The results are consistent across all three measured domains and across both independent datasets.

    Table 2. Observed Gains vs. Expected Maturation Over Mean 5.4-Month Interval (FF n=801 pairs, OTW n=94 pairs)

    ItemFF Observed GainExpected (Maturation)RatioOTW Observed Gain
    Draw a Person (0-4 scale)+1.24 pts+0.39 pts3.2x+1.06 pts
    Functional Pencil Grasp+25.3 pp+9.7 pp2.6x+27.7 pp
    Finger Touching (54-mo milestone)+20.3 pp+11.7 pp1.7xn/a
    Cohen’s d (DAP)0.940.84

    Note. Expected gain calculated from cross-sectional linear regression of E1 scores on age in months using the full FF evaluated dataset. pp = percentage points. Cohen’s d: 0.2 = small, 0.5 = medium, 0.8 = large effect. OTW finger touching item not directly comparable due to different item structure.

    Draw a Person improved at 3.2 times the expected developmental rate in FF and 2.6 times in OTW. Functional pencil grasp improved at 2.6 times expected in FF and 3.1 times in OTW. These are not marginal differences from what maturation alone would predict. They are two to three times larger. And they replicate across an entirely independent dataset collected with a different tool, by the same team, with different children.

    Why This Is Not Just Children Getting Older

    If the gains above were driven primarily by maturation, we would expect children at all starting points to show similar improvement. A child who enters at score 0 would gain roughly as much as a child who enters at score 3, because age-related development does not care where you start.

    That is not what we see. The table below shows Draw a Person gains stratified by E1(Initial Evaluation)  score, combining FF and O.T. Wizard data. The pattern is unambiguous.

    Table 3. Draw a Person Gain by E1 Score: FF (n=801) + OTW (n=93) Combined

    E1 ScorenE2 MeanMean Gain% Improved% Same% Declined
    0 (no parts)365+291.74+1.7474%26%0%
    1 (approximations)203+162.37+1.3781%13%6%
    2 (head, no body)165+312.66+0.6654%37%9%
    3 (recognizable)60+112.96-0.0429%46%25%
    4 (6+ body parts)8+63.36-0.360%57%43%

    Note. n column shows FF count + OTW count at each E1 score level. E1 score 0 = no recognizable approximations. Score 4 = recognizable person with 6 or more body parts. Gains decline systematically as E1 score increases, reflecting ceiling effects at higher starting points rather than absence of progress.

    Children who started at score 0 improved by an average of 1.74 points, with 74% showing measurable gains. Children who started at score 3 or 4 were near the ceiling of the scale and showed flat or slightly negative scores at E2, exactly as ceiling effects predict.

    This score-dependent gain gradient is the signature of a real treatment effect. Maturation produces relatively uniform gains regardless of starting point. Intervention produces the largest gains in children with the most room to grow. That is what we observe, and it replicates point-for-point across both the FF and OTW datasets independently.

    The grasp and finger touching data tell the same story from a different angle.

    Table 4. Skill Transition Rates: What Happened Between E1 and E2

    ItemStatus at E1nOutcome at E2
    Pencil GraspNon-functional377 (FF) + 65 (OTW)60% converted to functional
    Pencil GraspFunctional417 (FF) + 29 (OTW)94% maintained functional
    Finger TouchingFail407 (FF)58% passed at E2
    Finger TouchingPass394 (FF)82% maintained pass

    60% of children with non-functional pencil grasp at E1 had functional grasp by E2. 94% of children with functional grasp at E1 maintained it. Skills were not fluctuating randomly. They were moving in one direction and holding. That is not maturation. That is intervention.

    Two Tools, Three Years Apart, Same Answer

    The FF screener and O.T. Wizard are different instruments. FF was a clinician-developed Google Form with embedded scoring anchors and standardized stimulus materials. O.T. Wizard is a fully architected clinical platform undergoing Rasch psychometric validation, with 596 data variables per evaluation and item-level calibration. The two tools share several core items, including Draw a Person, pencil grasp classification, and finger touching. They do not share overlapping children. The DAP scale is directly comparable across both tools at the 0 to 4 range, with identical scoring anchors at each level. O.T. Wizard extended the ceiling by adding two higher-level descriptors, bringing the OTW scale to 6 points total. For this analysis, OTW DAP scores were capped at 4 to ensure a valid cross-tool comparison.

    Yet when we calculate the cross-sectional developmental growth rate for Draw a Person from FF E1 data, we get 0.072 points per month. From OTW E1 data, we get 0.082 points per month. Two tools, thousands of children, the same underlying developmental trajectory captured within 0.01 points of each other per month.

    When two independent measurement systems produce convergent developmental slopes and convergent gain ratios, that is not a coincidence. That is construct validity. Each dataset serves as an independent replication of the other’s findings, and both point to the same conclusion.

    Danielle, an OTR/L out of NC, administers an evaluation with a 3 year old using OT Wizard

    What This Means for OT Practice and Clinical Infrastructure

    Pediatric occupational therapists have long known that their interventions make a difference. The challenge has been demonstrating it systematically, at scale, in a form that partners, funders, schools, and insurance providers find credible.

    The Learning Charms team built that demonstration over 3.8 years, starting from scratch, with no research funding, no university partnership, and no IRB. They built it by replacing meaningless checklists with structured clinical measurement, by training a team of 30 clinicians to collect data consistently, and by iterating their tools until the data was worth analyzing.

    O.T. Wizard is the current iteration of that infrastructure. It is not a platform claiming efficacy. It is a platform whose evidence base already exists, built by the same team that built the platform, using the same children, in the same communities, over the same years. The data in this paper is not a promise of what O.T. Wizard will eventually show. It is a record of what systematic clinical measurement has already demonstrated.

    OT works. The data, replicated across tools and years and nearly 900 pairs of children, shows it.

    Disclosure

    Stephanie Seymore Wick is the founder and clinical architect of O.T. Wizard and owner of Learning Charms, Inc. All data was collected through routine clinical practice and contracted screening partnerships. No external funding was received. The FUNdamental Foundations screener was a clinician-developed field tool and has not undergone formal psychometric validation. O.T. Wizard is currently undergoing Rasch analysis validation. All findings should be interpreted as practice-based clinical evidence rather than results from a randomized controlled trial.

    About O.T. Wizard

    O.T. Wizard is a clinical intelligence system for pediatric occupational therapy professionals. The platform supports evaluation, documentation, goal planning, and scheduling across 12 domains including fine motor skills, visual-motor integration, praxis, visual perception, executive functioning, activities of daily living, and participation. O.T. Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update as the clinical database expands. Learn more at otwizard.com.

    About Learning Charms

    Learning Charms is a pediatric occupational therapy group that employs roughly 25 OTP’s in the Charlotte , NC and surrounding counties. Learning Charms is now focused mainly on preschool aged children in their school environment.

  • OT Wizard Psychometric Validation: What It Means for Evidence-Based Practice

    OT Wizard Psychometric Validation: What It Means for Evidence-Based Practice

    I’m thrilled to announce that OT Wizard has officially begun psychometric validation through Rasch analysis – a major milestone in our journey to elevate evidence-based practice in pediatric occupational therapy!

    What’s Happening Now

    We’ve sent evaluation data for ages 4-5 years (Age Bands H & I) to an independent psychometrician for comprehensive analysis. With over 400 evaluations from 18 therapists across North Carolina, we have robust data to validate that OT Wizard measures what we say it measures & accurately, reliably, and fairly.

    What Is Rasch Analysis?

    Rasch analysis is a sophisticated psychometric approach that goes beyond traditional test validation. Unlike norm-referenced assessments that simply compare students to each other, Rasch analysis creates an interval-level measurement scale, similar to measuring temperature or weight.

    Think of assessments you may know that use Rasch methodology:

      • PEDI-CAT (Pediatric Evaluation of Disability Inventory – Computer Adaptive Test) – Rasch-calibrated functional assessment

      • AMPS (Assessment of Motor and Process Skills) – fully Rasch-calibrated for ADL performance

      • COPM (Canadian Occupational Performance Measure) – uses Rasch principles for measuring occupational performance

      • HELP (Hawaii Early Learning Profile) – Rasch-validated developmental assessment

      • Original PEDI – Rasch-based functional assessment

    These assessments are considered gold standards because Rasch analysis ensures:

      • Equal intervals: A 10-point gain at any level represents the same amount of growth

      • Sample-independent measurement: Item difficulty doesn’t depend on who takes the test

      • Missing data handling: Scores are valid even when not all items are administered (adaptive testing)

      • Precise error estimation: Know exactly how confident you can be in each score

      • Item hierarchy validation: Confirms items are developmentally sequenced correctly

    Why Rasch Instead of Traditional Norming?

    Traditional norm-referenced tests (like BOT-3, PDMS-3) require testing typically-developing children to create percentile ranks. That’s valuable, but has limitations:

    Norms become outdated (tests re-normed every 15-20 years)
    Percentiles are ordinal, not interval (85th→95th ≠ 15th→25th in actual ability)
    Can’t track growth accurately across different ability levels
    Require complete test administration

    Rasch analysis provides:

    Continuous measurement scale – Track growth precisely over time
    Adaptive testing – Administer only relevant items, still get accurate scores
    Sample-independent – Item difficulty stays stable regardless of who’s tested
    Living calibration – Can update and refine continuously with new data
    Clinical utility – Scores directly interpretable for intervention planning

    We’re building OT Wizard to work like the PEDI-CAT and AMPS.   These are tools that OTs trust because they’re built on rigorous Rasch foundations.

    What We Already Know About Our Data

    Before even sending data to our psychometrician, we conducted preliminary analysis on our 404 evaluations to ensure data quality. Here’s what we’ve learned:

    Strong Sample Characteristics

      • Well-balanced age distribution: 184 evaluations (ages 4.0-4.4) and 220 evaluations (ages 4.5-4.9)

      • 18 therapists contributing: Average of 22 evaluations each, with range of 7-45 per therapist (good for inter-rater reliability)

      • Gender representation: 56% male, 44% female (reflects typical OT referral patterns)

      • Diverse language backgrounds: 90% English, 8% Spanish, 2% other languages

      • Clinical population validity: 96% recommended for OT services

    Excellent Data Completeness

      • Zero missing responses – every administered item was answered

      • 75.5% average completion rate – our adaptive basal/ceiling rules are working perfectly

      • 19,062 total data points across 74 items and 8 domains

    Strong Domain Coverage

      • Visual Perception: 15 items

      • Activities of Daily Living: 14 items

      • Gross Motor & Fine Motor: 10 items each

      • Participation: 9 items

      • Executive Functioning: 8 items

      • Visual Motor Integration: 5 items

      • Praxis: 3 items

    Areas for Improvement Identified

    Our preliminary analysis flagged several items for the psychometrician to examine closely:

    Ceiling Effects in Visual Perception (Preschool evaluation): 56% of responses scored at ceiling (mastered), with 33% at floor (not yet observed). This bimodal pattern suggests we may need more mid-difficulty items to better differentiate students in the middle range.

    Rating Scale Consistency: A small number of responses (3.3%) showed raw scores instead of normalized scores, indicating a formula issue we’ve already corrected.

    Developmental Anchor Gaps: About 54% of items (primarily Executive Functioning and Participation domains) lack developmental anchors. The Rasch analysis will empirically determine difficulty levels so we can assign appropriate anchors.

    Item-Age Alignment: Many items administered are anchored above student age ranges (51-55% above range). This is actually expected and appropriate.  Students with developmental delays are working on skills typically seen at older ages. However, Rasch will help us recalibrate anchors based on clinical population performance vs. typical development.

    ✅ Best-Performing Domains

      • ADL: Excellent distribution with only 8% floor and 8% ceiling – items are well-targeted

      • Executive Functioning: Minimal floor effect (0.7%), good spread across ability levels

      • Participation: Near-zero floor (0.2%), strong measurement potential

    This preliminary work means we’re sending clean, robust data to our psychometrician.  This is maximizing the value of the Rasch analysis and ensuring reliable results.

    What’s Being Analyzed

    Our psychometrician is conducting comprehensive Rasch analysis across multiple dimensions:

    1. Construct Validity (Unidimensionality)

    Do items within each domain (Gross Motor, Fine Motor, Visual Perception, etc.) measure a single, coherent construct? This is critical for Rasch – if items don’t “hang together,” they can’t be on the same measurement scale.

    2. Item Fit

    Which items contribute to reliable measurement? Rasch provides specific fit statistics (infit/outfit MNSQ) showing whether each item:

      • Is too predictable (doesn’t add information)

      • Is too unpredictable (confuses the measurement)

      • Functions optimally (contributes to precise measurement)

    Items outside acceptable ranges get flagged for revision or removal.

    3. Rating Scale Functioning

    Do our 5-point performance bands (Beginning → Mastered) function as intended? Rasch examines:

      • Are all categories used appropriately?

      • Do response thresholds advance in the right order?

      • Should categories be collapsed (e.g., 5-point → 3-point)?

    This is similar to how AMPS validates its 4-point scoring scale.

    4. Item Hierarchy

    Rasch places all items on a single difficulty scale (measured in logits). We’ll see if:

      • Items anchored at 48 months are empirically easier than 54-month items

      • Our developmental sequencing matches actual difficulty

      • Gaps exist where we need additional items

    This will be especially important for items currently lacking developmental anchors – the Rasch analysis will tell us where they belong.

    5. Measurement Precision

    Unlike traditional reliability (one number for whole test), Rasch shows precision at every ability level:

      • Where is measurement most accurate?

      • What’s the standard error at our 70% clinical threshold?

      • Can we distinguish between students with small ability differences?

    6. Differential Item Functioning (DIF)

    Do items work the same way for:

      • Boys vs girls?

      • 4-year-olds vs 5-year-olds?

      • English vs Spanish speakers?

      • Different diagnoses?

    Items showing bias get flagged or removed – ensuring fairness.

    7. Person Separation

    Can we reliably distinguish between students at different ability levels? Rasch provides a separation index showing how many distinct ability levels we can measure. Higher separation = more precise clinical distinctions.

    8. Addressing Known Issues

    The psychometrician will specifically examine:

      • Visual Perception’s ceiling effects : do we need additional mid-difficulty items?

      • Praxis domain with only 3 items : is this sufficient or should it combine with another domain?

      • Executive Functioning and Participation rating scales : do they function as separate constructs from performance-based items?

    Why This Matters for You

    Rasch validation transforms how you can use OT Wizard scores:

    Meaningful Progress Monitoring

    Because Rasch creates interval-level measurement, you can confidently say:

      • “Student gained 0.8 logits in 6 months”

      • “This represents clinically significant progress”

      • “Growth rate exceeds typical intervention response”

    Traditional percentage scores can’t make these claims and a jump from 40% to 50% isn’t necessarily the same growth as 70% to 80%.

    Adaptive Testing Validation

    Like the PEDI-CAT and AMPS, OT Wizard uses basal/ceiling rules so students aren’t frustrated with too-hard items or bored with too-easy ones. Rasch analysis confirms:

      • Scores are comparable even when different items are administered

      • Our 75% completion rate is optimal

      • Missing items are appropriately “not administered,” not missing data

    Credible Clinical Decisions

    When you document that a child’s gross motor ability is at -1.2 logits:

      • Insurance companies recognize Rasch-based measurement

      • School districts understand the methodology (same as PEDI-CAT/HELP)

      • You can defend your clinical reasoning with published psychometric evidence

    Item-Level Interpretation

    Rasch analysis creates item hierarchy maps showing exactly which skills a child has mastered, which are emerging, and which aren’t yet present. This directly informs intervention planning  just like how AMPS users identify specific ADL breakdowns or PEDI-CAT shows functional skill patterns.

    Why Start with Ages 4-5?

    We strategically chose this age range because:

      1. Sufficient sample size: 400+ evaluations provide robust statistical power for Rasch analysis

      1. Diverse representation: Students with various diagnoses, languages, and ability levels

      1. Multiple raters: 18 different therapists ensure inter-rater reliability analysis

      1. Item overlap: Many items in this age range also appear in adjacent ages, so findings inform the entire platform

      1. Strong data quality: Our preliminary analysis confirmed excellent completion rates and coverage

    The psychometrician will identify any problematic items, validate our developmental anchors, assign anchors to items missing them, and ensure rating scales function optimally. We’ll implement improvements before these issues cascade into other age bands.

    What Happens Next

    Based on the psychometrician’s findings (expected in 4-5 weeks), we’ll:

      1. Remove or revise misfitting items that don’t meet Rasch fit criteria

      1. Add mid-difficulty items to Visual Perception domain to address ceiling effects

      1. Optimize rating scales if analysis shows categories aren’t functioning as intended

      1. Recalibrate existing anchors if clinical population performance differs from typical development

      1. Establish measurement precision estimates at different ability levels

      1. Publish validation statistics you can cite in reports and presentations

    This refined version becomes the foundation for validating additional age bands.

    Expanding Validation: Ages 3 Months to 12 Years

    Over the next 12 months, as we reach 200+ evaluations per age band, we’ll validate each additional age group. This will create a comprehensive, linked measurement system similar to how PEDI-CAT links across age ranges  where we can:

      • Track individual students across multiple years on the same logit scale

      • Provide age-equivalent scores based on item difficulty calibration

      • Create clinical reference data comparing students receiving OT services

      • Document growth trajectories with true interval-level measurement

      • Demonstrate outcomes with unprecedented precision

    By Month 12, we’ll conduct a comprehensive linking study that places all age bands (3 months through 12 years) on a single, continuous measurement scale. Items that appear in multiple age bands will “anchor” the scales together, ensuring continuity.

    This approach mirrors how major Rasch-based assessments (PEDI-CAT, AMPS, HELP) maintain measurement continuity across ages and versions.

    How You Can Support This Work

    1. Keep Using OT Wizard

    Every evaluation you complete contributes to our growing database. Rasch analysis becomes more robust with larger samples and the more data we collect, the more confident we can be in item calibrations.

    administering OT evaluation with child.

    Brittany B., administers an eval with OT Wizard

    2. Share Your Clinical Insights

    If you notice items that seem:

      • Confusing or ambiguous to score

      • Too easy or too hard for the age range

      • Misaligned with what you observe clinically

      • Culturally biased or inappropriate

    Please let us know! Your real-world feedback is invaluable. Rasch analysis will identify statistical misfits, but your clinical judgment helps us understand why items aren’t working.

    What This Means for Our Profession

    Most therapy documentation tools rely on subjective clinical observation. While tools like PEDI-CAT, AMPS, COPM, and HELP exist and use Rasch methodology, they’re limited in scope and focus on narrow or specific functional domains, requiring specialized training, or covering narrow age ranges.

    OT Wizard is different: We’re creating a comprehensive, Rasch-validated clinical intelligence platform that covers:

      • Multiple domains (gross motor, fine motor, visual perception, praxis, executive functioning, ADL, participation)

      • Birth through early adulthood (3 months – 18 years)

      • Performance-based, participational based, and functional assessment

      • Integrated into everyday clinical workflow

    By pursuing rigorous Rasch validation, we’re increasing psychometric rigor to comprehensive pediatric OT assessment.

    When this validation is complete, you’ll be able to say:

    “I use OT Wizard, a comprehensive Rasch-validated pediatric OT assessment platform with published psychometric evidence across 2,000+ evaluations providing the same measurement quality as tools like PEDI-CAT and AMPS, but covering all developmental domains.”

    The Vision: Living Calibration and Real-Time Data

    Here’s what excites me most about Rasch methodology: Unlike traditional norm-referenced tests that become frozen in time, Rasch-calibrated assessments can be continuously refined.

    The PEDI-CAT has demonstrated this – as more data is collected, item calibrations can be updated, new items added, and measurement precision improved all while maintaining the same measurement scale.

    OT Wizard will have “living calibration”:

      • Continuous item refinement as we collect more data

      • New items added to fill gaps in difficulty coverage

      • Real-time quality monitoring

      • Annual recalibration studies

      • Regional and demographic analyses

    Imagine:

      • Item difficulties that reflect current populations

      • Outcome analytics showing which interventions are most effective

      • Predictive data identifying which early skills best predict later success

      • The world’s largest real-time Rasch-calibrated pediatric development database

    Every evaluation you complete contributes to this unprecedented resource.

    💕Thank you for building the future of evidence-based OT assessment with O.T. Wizard.