Tag: validation

  • OT Wizard Psychometric Validation: What It Means for Evidence-Based Practice

    OT Wizard Psychometric Validation: What It Means for Evidence-Based Practice

    I’m thrilled to announce that OT Wizard has officially begun psychometric validation through Rasch analysis – a major milestone in our journey to elevate evidence-based practice in pediatric occupational therapy!

    What’s Happening Now

    We’ve sent evaluation data for ages 4-5 years (Age Bands H & I) to an independent psychometrician for comprehensive analysis. With over 400 evaluations from 18 therapists across North Carolina, we have robust data to validate that OT Wizard measures what we say it measures & accurately, reliably, and fairly.

    What Is Rasch Analysis?

    Rasch analysis is a sophisticated psychometric approach that goes beyond traditional test validation. Unlike norm-referenced assessments that simply compare students to each other, Rasch analysis creates an interval-level measurement scale, similar to measuring temperature or weight.

    Think of assessments you may know that use Rasch methodology:

      • PEDI-CAT (Pediatric Evaluation of Disability Inventory – Computer Adaptive Test) – Rasch-calibrated functional assessment

      • AMPS (Assessment of Motor and Process Skills) – fully Rasch-calibrated for ADL performance

      • COPM (Canadian Occupational Performance Measure) – uses Rasch principles for measuring occupational performance

      • HELP (Hawaii Early Learning Profile) – Rasch-validated developmental assessment

      • Original PEDI – Rasch-based functional assessment

    These assessments are considered gold standards because Rasch analysis ensures:

      • Equal intervals: A 10-point gain at any level represents the same amount of growth

      • Sample-independent measurement: Item difficulty doesn’t depend on who takes the test

      • Missing data handling: Scores are valid even when not all items are administered (adaptive testing)

      • Precise error estimation: Know exactly how confident you can be in each score

      • Item hierarchy validation: Confirms items are developmentally sequenced correctly

    Why Rasch Instead of Traditional Norming?

    Traditional norm-referenced tests (like BOT-3, PDMS-3) require testing typically-developing children to create percentile ranks. That’s valuable, but has limitations:

    Norms become outdated (tests re-normed every 15-20 years)
    Percentiles are ordinal, not interval (85th→95th ≠ 15th→25th in actual ability)
    Can’t track growth accurately across different ability levels
    Require complete test administration

    Rasch analysis provides:

    Continuous measurement scale – Track growth precisely over time
    Adaptive testing – Administer only relevant items, still get accurate scores
    Sample-independent – Item difficulty stays stable regardless of who’s tested
    Living calibration – Can update and refine continuously with new data
    Clinical utility – Scores directly interpretable for intervention planning

    We’re building OT Wizard to work like the PEDI-CAT and AMPS.   These are tools that OTs trust because they’re built on rigorous Rasch foundations.

    What We Already Know About Our Data

    Before even sending data to our psychometrician, we conducted preliminary analysis on our 404 evaluations to ensure data quality. Here’s what we’ve learned:

    Strong Sample Characteristics

      • Well-balanced age distribution: 184 evaluations (ages 4.0-4.4) and 220 evaluations (ages 4.5-4.9)

      • 18 therapists contributing: Average of 22 evaluations each, with range of 7-45 per therapist (good for inter-rater reliability)

      • Gender representation: 56% male, 44% female (reflects typical OT referral patterns)

      • Diverse language backgrounds: 90% English, 8% Spanish, 2% other languages

      • Clinical population validity: 96% recommended for OT services

    Excellent Data Completeness

      • Zero missing responses – every administered item was answered

      • 75.5% average completion rate – our adaptive basal/ceiling rules are working perfectly

      • 19,062 total data points across 74 items and 8 domains

    Strong Domain Coverage

      • Visual Perception: 15 items

      • Activities of Daily Living: 14 items

      • Gross Motor & Fine Motor: 10 items each

      • Participation: 9 items

      • Executive Functioning: 8 items

      • Visual Motor Integration: 5 items

      • Praxis: 3 items

    Areas for Improvement Identified

    Our preliminary analysis flagged several items for the psychometrician to examine closely:

    Ceiling Effects in Visual Perception (Preschool evaluation): 56% of responses scored at ceiling (mastered), with 33% at floor (not yet observed). This bimodal pattern suggests we may need more mid-difficulty items to better differentiate students in the middle range.

    Rating Scale Consistency: A small number of responses (3.3%) showed raw scores instead of normalized scores, indicating a formula issue we’ve already corrected.

    Developmental Anchor Gaps: About 54% of items (primarily Executive Functioning and Participation domains) lack developmental anchors. The Rasch analysis will empirically determine difficulty levels so we can assign appropriate anchors.

    Item-Age Alignment: Many items administered are anchored above student age ranges (51-55% above range). This is actually expected and appropriate.  Students with developmental delays are working on skills typically seen at older ages. However, Rasch will help us recalibrate anchors based on clinical population performance vs. typical development.

    ✅ Best-Performing Domains

      • ADL: Excellent distribution with only 8% floor and 8% ceiling – items are well-targeted

      • Executive Functioning: Minimal floor effect (0.7%), good spread across ability levels

      • Participation: Near-zero floor (0.2%), strong measurement potential

    This preliminary work means we’re sending clean, robust data to our psychometrician.  This is maximizing the value of the Rasch analysis and ensuring reliable results.

    What’s Being Analyzed

    Our psychometrician is conducting comprehensive Rasch analysis across multiple dimensions:

    1. Construct Validity (Unidimensionality)

    Do items within each domain (Gross Motor, Fine Motor, Visual Perception, etc.) measure a single, coherent construct? This is critical for Rasch – if items don’t “hang together,” they can’t be on the same measurement scale.

    2. Item Fit

    Which items contribute to reliable measurement? Rasch provides specific fit statistics (infit/outfit MNSQ) showing whether each item:

      • Is too predictable (doesn’t add information)

      • Is too unpredictable (confuses the measurement)

      • Functions optimally (contributes to precise measurement)

    Items outside acceptable ranges get flagged for revision or removal.

    3. Rating Scale Functioning

    Do our 5-point performance bands (Beginning → Mastered) function as intended? Rasch examines:

      • Are all categories used appropriately?

      • Do response thresholds advance in the right order?

      • Should categories be collapsed (e.g., 5-point → 3-point)?

    This is similar to how AMPS validates its 4-point scoring scale.

    4. Item Hierarchy

    Rasch places all items on a single difficulty scale (measured in logits). We’ll see if:

      • Items anchored at 48 months are empirically easier than 54-month items

      • Our developmental sequencing matches actual difficulty

      • Gaps exist where we need additional items

    This will be especially important for items currently lacking developmental anchors – the Rasch analysis will tell us where they belong.

    5. Measurement Precision

    Unlike traditional reliability (one number for whole test), Rasch shows precision at every ability level:

      • Where is measurement most accurate?

      • What’s the standard error at our 70% clinical threshold?

      • Can we distinguish between students with small ability differences?

    6. Differential Item Functioning (DIF)

    Do items work the same way for:

      • Boys vs girls?

      • 4-year-olds vs 5-year-olds?

      • English vs Spanish speakers?

      • Different diagnoses?

    Items showing bias get flagged or removed – ensuring fairness.

    7. Person Separation

    Can we reliably distinguish between students at different ability levels? Rasch provides a separation index showing how many distinct ability levels we can measure. Higher separation = more precise clinical distinctions.

    8. Addressing Known Issues

    The psychometrician will specifically examine:

      • Visual Perception’s ceiling effects : do we need additional mid-difficulty items?

      • Praxis domain with only 3 items : is this sufficient or should it combine with another domain?

      • Executive Functioning and Participation rating scales : do they function as separate constructs from performance-based items?

    Why This Matters for You

    Rasch validation transforms how you can use OT Wizard scores:

    Meaningful Progress Monitoring

    Because Rasch creates interval-level measurement, you can confidently say:

      • “Student gained 0.8 logits in 6 months”

      • “This represents clinically significant progress”

      • “Growth rate exceeds typical intervention response”

    Traditional percentage scores can’t make these claims and a jump from 40% to 50% isn’t necessarily the same growth as 70% to 80%.

    Adaptive Testing Validation

    Like the PEDI-CAT and AMPS, OT Wizard uses basal/ceiling rules so students aren’t frustrated with too-hard items or bored with too-easy ones. Rasch analysis confirms:

      • Scores are comparable even when different items are administered

      • Our 75% completion rate is optimal

      • Missing items are appropriately “not administered,” not missing data

    Credible Clinical Decisions

    When you document that a child’s gross motor ability is at -1.2 logits:

      • Insurance companies recognize Rasch-based measurement

      • School districts understand the methodology (same as PEDI-CAT/HELP)

      • You can defend your clinical reasoning with published psychometric evidence

    Item-Level Interpretation

    Rasch analysis creates item hierarchy maps showing exactly which skills a child has mastered, which are emerging, and which aren’t yet present. This directly informs intervention planning  just like how AMPS users identify specific ADL breakdowns or PEDI-CAT shows functional skill patterns.

    Why Start with Ages 4-5?

    We strategically chose this age range because:

      1. Sufficient sample size: 400+ evaluations provide robust statistical power for Rasch analysis

      1. Diverse representation: Students with various diagnoses, languages, and ability levels

      1. Multiple raters: 18 different therapists ensure inter-rater reliability analysis

      1. Item overlap: Many items in this age range also appear in adjacent ages, so findings inform the entire platform

      1. Strong data quality: Our preliminary analysis confirmed excellent completion rates and coverage

    The psychometrician will identify any problematic items, validate our developmental anchors, assign anchors to items missing them, and ensure rating scales function optimally. We’ll implement improvements before these issues cascade into other age bands.

    What Happens Next

    Based on the psychometrician’s findings (expected in 4-5 weeks), we’ll:

      1. Remove or revise misfitting items that don’t meet Rasch fit criteria

      1. Add mid-difficulty items to Visual Perception domain to address ceiling effects

      1. Optimize rating scales if analysis shows categories aren’t functioning as intended

      1. Recalibrate existing anchors if clinical population performance differs from typical development

      1. Establish measurement precision estimates at different ability levels

      1. Publish validation statistics you can cite in reports and presentations

    This refined version becomes the foundation for validating additional age bands.

    Expanding Validation: Ages 3 Months to 12 Years

    Over the next 12 months, as we reach 200+ evaluations per age band, we’ll validate each additional age group. This will create a comprehensive, linked measurement system similar to how PEDI-CAT links across age ranges  where we can:

      • Track individual students across multiple years on the same logit scale

      • Provide age-equivalent scores based on item difficulty calibration

      • Create clinical reference data comparing students receiving OT services

      • Document growth trajectories with true interval-level measurement

      • Demonstrate outcomes with unprecedented precision

    By Month 12, we’ll conduct a comprehensive linking study that places all age bands (3 months through 12 years) on a single, continuous measurement scale. Items that appear in multiple age bands will “anchor” the scales together, ensuring continuity.

    This approach mirrors how major Rasch-based assessments (PEDI-CAT, AMPS, HELP) maintain measurement continuity across ages and versions.

    How You Can Support This Work

    1. Keep Using OT Wizard

    Every evaluation you complete contributes to our growing database. Rasch analysis becomes more robust with larger samples and the more data we collect, the more confident we can be in item calibrations.

    administering OT evaluation with child.

    Brittany B., administers an eval with OT Wizard

    2. Share Your Clinical Insights

    If you notice items that seem:

      • Confusing or ambiguous to score

      • Too easy or too hard for the age range

      • Misaligned with what you observe clinically

      • Culturally biased or inappropriate

    Please let us know! Your real-world feedback is invaluable. Rasch analysis will identify statistical misfits, but your clinical judgment helps us understand why items aren’t working.

    What This Means for Our Profession

    Most therapy documentation tools rely on subjective clinical observation. While tools like PEDI-CAT, AMPS, COPM, and HELP exist and use Rasch methodology, they’re limited in scope and focus on narrow or specific functional domains, requiring specialized training, or covering narrow age ranges.

    OT Wizard is different: We’re creating a comprehensive, Rasch-validated clinical intelligence platform that covers:

      • Multiple domains (gross motor, fine motor, visual perception, praxis, executive functioning, ADL, participation)

      • Birth through early adulthood (3 months – 18 years)

      • Performance-based, participational based, and functional assessment

      • Integrated into everyday clinical workflow

    By pursuing rigorous Rasch validation, we’re increasing psychometric rigor to comprehensive pediatric OT assessment.

    When this validation is complete, you’ll be able to say:

    “I use OT Wizard, a comprehensive Rasch-validated pediatric OT assessment platform with published psychometric evidence across 2,000+ evaluations providing the same measurement quality as tools like PEDI-CAT and AMPS, but covering all developmental domains.”

    The Vision: Living Calibration and Real-Time Data

    Here’s what excites me most about Rasch methodology: Unlike traditional norm-referenced tests that become frozen in time, Rasch-calibrated assessments can be continuously refined.

    The PEDI-CAT has demonstrated this – as more data is collected, item calibrations can be updated, new items added, and measurement precision improved all while maintaining the same measurement scale.

    OT Wizard will have “living calibration”:

      • Continuous item refinement as we collect more data

      • New items added to fill gaps in difficulty coverage

      • Real-time quality monitoring

      • Annual recalibration studies

      • Regional and demographic analyses

    Imagine:

      • Item difficulties that reflect current populations

      • Outcome analytics showing which interventions are most effective

      • Predictive data identifying which early skills best predict later success

      • The world’s largest real-time Rasch-calibrated pediatric development database

    Every evaluation you complete contributes to this unprecedented resource.

    💕Thank you for building the future of evidence-based OT assessment with O.T. Wizard.