Clinical Takeaways | O.T. Wizard Research Series, Part 1
By Stephanie Seymore Wick, MSOT, OT/L | Founder and Clinical Architect, O.T. Wizard
Most pediatric OTs are excellent documenters. We write thorough evaluations, set meaningful goals, and log every session. The paper trail is solid. So why is it so hard to answer the one question that matters most?
Is this workingand if so, how much?
Not “are we doing the right things?” Not “is this child making progress in a general sense?” The specific question: is this child changing, how fast, and is that fast enough?
That question turns out to be surprisingly hard to answer with the tools most of us are using.
The Snapshot Problem
A standardized evaluation gives you a score at one point in time. A re-evaluation gives you another score. You compare the two and write a narrative about what changed. That is snapshot documentation, and it is useful. It tells you where a child started and where they landed.
What it does not tell you is anything about the line between those two points.
When did the change happen? Was progress steady, or did the child plateau for two months and then accelerate? Did a specific intervention approach produce better outcomes than others? Did an attendance gap create a measurable dip? Did the child actually cross a key functional threshold six weeks before the re-evaluation was even scheduled?
Without structured session-level data linked to domain scores, you simply cannot see any of that. You have two dots. You do not have a trajectory.
Why This Shows Up Differently in Medical vs. School Settings
In a medical outpatient practice, the moment this gap becomes visible is usually an authorization request. You are being asked to justify continued services and the strongest argument is quantitative: here is where the child started, here is the rate at which they are improving, and here is where they are projected to land by the end of this period. Without RTI data, you fall back on clinical narrative. Narrative is defensible. It is not the same as a slope.
In school-based practice, the moment arrives at the IEP table. You are sitting with a team that includes a parent, a classroom teacher, a special educator, and an administrator. They are deciding whether OT services should continue, increase, or be exited. The OT who arrives with a performance slope and a comparison to natural developmental growth is a different professional presence than the one who arrives with quarterly progress notes. Both care about the child. Only one has data that can change a decision.
Exit recommendations are especially difficult without RTI data. Recommending that a child be exited from OT services is a clinical and ethical judgment call. With measurement data showing that the child has reached functional independence or participation and is maintaining gains without direct support, it becomes a defensible milestone. Without that data, it is an opinion.
The Plateau Conversation
Every pediatric OT has been here. A child who was making visible progress has leveled off. The parent is worried. The payer is skeptical. The school team is questioning whether to continue services.
The problem is that a plateau looks the same on paper whether it is stagnation or consolidation. A child consolidating a new skill at a lower level of support may show no numerical gain for several sessions. That is not failure. It is a normal part of skill acquisition. But without session-level performance data, you cannot show anyone the difference. You can only explain it.
With structured data, you can show the team exactly when the plateau began, what changed in the child’s routine or support structure around that time, and whether similar plateaus have resolved in this child’s history. That changes the conversation from “we think this is temporary” to “here is what the data shows.”
What Parents Are Actually Asking
When a parent asks whether therapy is working, the most honest answer most therapists can give without RTI infrastructure is a clinical impression. That impression may be completely accurate. But it is not the same as showing a parent a graph of their child’s performance over twenty sessions and saying: here is where he started, here is the rate at which he is moving, and here is what we project by the end of this period.
For school-based OTs, the parent question arrives at the IEP table, in front of an entire team. The quality of your data shapes what parents understand, what they advocate for, and what they accept when the team recommends a service change. That matters.
The Practical Distinction
Documentation and measurement are not the same system. They are not competing systems either. They serve different purposes.
Documentation records what happened, establishes compliance, and communicates clinical reasoning. Measurement tracks rate of change, identifies what conditions produce better performance, and determines whether progress is sufficient.
Most EHRs were built for the first column. Very few were built for the second. The gap between them is not a failure of clinical intent. It is a gap in infrastructure. EMR’s are typically built by people who are interesting in billing insurance and keeping accounting records. They are not clinical intelligence systems.
One Finding Worth Noting
When the O.T. Wizard re-evaluation data was examined, Participation and Executive Functioning showed flat longitudinal profiles compared to the large gains seen in VMI, ADL, and fine motor domains. The initial interpretation might be that OT did not improve those areas. But there is another possibility worth taking seriously: the first evaluation rating in those domains may not have captured authentic baseline behavior. Children often present their best behavior when meeting a new therapist in a structured evaluation setting. By re-evaluation, the novelty has worn off. The rating that looks flat may simply be more accurate.
That is a question that would not have surfaced without measurement data. It has real implications for how we interpret initial evaluation scores in observation-dependent domains. It is the kind of question that data raises and documentation alone cannot.
A Few Things to Reflect On
At re-evaluation, how do you determine the rate at which a child progressed? Can you identify which sessions produced the most meaningful gains? Can you show a parent the slope of improvement over an authorization period? Can you distinguish a true plateau from a reduction in required support?
If those questions are hard to answer with your current system, the infrastructure gap is real, and it is worth thinking about.
Parts 2 through 4 of this series will move from problem framing to evidence to practice, including composite clinical vignettes, full dataset patterns, and what RTI infrastructure looks like in day-to-day clinical workflow.
Disclosures: The author is the Founder and Clinical Architect of O.T. Wizard and has a financial interest in the platform. All data referenced is de-identified clinical data collected through the O.T. Wizard software platform in routine practice.
About O.T. Wizard: O.T. Wizard is a clinical intelligence system for pediatric occupational therapy professionals. The platform evaluates performance across twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, executive functioning, activities of daily living, and participation. For more information, visit otwizard.com.
More Than a Milestone: What a Brief Scissor Skills Assessment Reveals About Tool Use, Hand Dominance, and Cutting Development in Preschool-Aged Children
Aims: To examine scissor cutting performance across preschool age bands (36 to 65 months) in a clinical sample, identify relationships between cutting accuracy, hand dominance, and hand positioning, assess cross-domain correlations, and evaluate longitudinal progress from initial to re-evaluation.
Methods: Descriptive and correlational analysis of 541 pediatric occupational therapy evaluations (Age Bands G through J) using structured cutting tasks scored on a standardized 14-point rubric. Participants were children referred for OT services, predominantly ages 48 to 59 months, with at least 90% qualifying for Medicaid.
Results: Cutting accuracy followed a clear developmental progression across age bands. Thumb-up dominant hand positioning was a large-effect predictor of cutting accuracy (Cohen d=1.17). Established hand dominance and writing-to-scissor hand consistency were strongly associated with performance. Scissor performance correlated significantly with fine motor, visual motor integration, ADL, visual perception, gross motor, bilateral integration, and praxis domains. Longitudinal gains of 5.21 points over 4.8 months exceeded the expected natural growth rate of 1.51 points.
Conclusions: A structured scissor skills assessment captures clinically meaningful variation in cutting skill and supports response-to-intervention documentation, goal writing, and cross-domain clinical reasoning in pediatric OT practice.
Keywords: scissor skills, hand dominance, fine motor development, pediatric occupational therapy, response to intervention, preschool
Scissors are one of the most commonly targeted skills in pediatric occupational therapy, yet they are rarely assessed with the precision required to drive goal writing, track progress, or demonstrate response to intervention. In most clinical settings, a child either can cut or cannot cut. That binary framing misses the developmental story that unfolds across the preschool years and leaves practitioners without the data needed to communicate clinical value to families, educators, and payers.
The preschool years represent the primary window for scissor skill acquisition. Cutting a straight line with 1-inch tolerance is typically expected by 41 months, precision cutting within a 1/4-inch boundary emerges around 48 months, and smooth curvy-line cutting at a 1/4-inch tolerance is an expectation by 60 months (O.T. Wizard Scissor Skills Assessment, v6.2). Despite this well-established developmental sequence, most standardized pediatric OT assessments address scissor skills with limited granularity, and published outcome data on scissor skill development in referred clinical populations remains sparse.
This article presents findings from 541 pediatric occupational therapy evaluations collected through O.T. Wizard, a clinical intelligence platform for pediatric occupational therapy professionals. Using a standardized scissor skills assessment embedded within the evaluation process, this analysis examines cutting performance across age bands, its relationship to hand dominance and hand positioning, its correlation with multiple developmental domains, and longitudinal gains across re-evaluation. The evidence supports scissor skill assessment as a window into neuromotor organization, tool use learning, and cross-domain functional development.
Methods
Participants
This analysis includes 541 pediatric occupational therapy evaluations representing children ages 36 through 65 months (Age Bands G through J): Band G (36 to 47 months, n=61), Band H (48 to 53 months, n=199), Band I (54 to 59 months, n=256), and Band J (60 to 65 months, n=69). An additional 82 children had paired evaluations (E1 and E2), with a mean interval of 4.8 months between assessments, enabling longitudinal analysis. The sample was 57% male and 43% female. Primary language was English for 90% of participants and Spanish for 9%. At least 90% of children qualified for Medicaid, and approximately 86% were recommended for OT services following evaluation. All evaluations were conducted in North Carolina.
Data were collected through routine pediatric occupational therapy evaluations conducted in clinical practice using O.T. Wizard. Parents provided informed consent as part of standard clinical care. As this analysis represents clinical outcomes data rather than human subjects research, institutional review board approval was not required. All data were de-identified in accordance with HIPAA regulations.
Measures
The O.T. Wizard Scissor Skills Assessment (v6.2) is a structured, standardized tool embedded within the O.T. Wizard multi-domain pediatric OT evaluation platform. It measures cutting performance across a progression of task demands anchored to published developmental milestones: holding scissors with one hand (24 months), snipping paper (34 months), opening and closing scissors (36 months), cutting a 1-inch straight line (41 months), 3/4-inch and 1/2-inch straight lines (43 and 45 months), a 1/4-inch straight line (48 months), a 1/4-inch curvy line (60 months), and smooth cutting quality (72 months). Cutting accuracy was scored using a standardized rubric with 0, 0.5, and 1.0 point values per segment, yielding a maximum score of 14 points per cutting task. Therapists also documented scissor type used, dominant hand position (thumb up versus thumb down or absent), stabilizer hand description, and qualitative hand positioning ratings.
This analysis focuses primarily on FM_SCIS_48 (1/4-inch straight line), selected as the primary analytic item due to near-complete data across all age bands (n=541) and its anchor at the 48-month developmental expectation. FM_SCIS_41 (1-inch straight line) and FM_SCIS_60C (1/4-inch curvy line) were included where sample sizes permitted. Scissor type data indicated that 90.6% of children were assessed with school or safety scissors; adaptive equipment (spring-assisted, loop handle) was used in fewer than 1% of cases and is not analyzed separately.
Data Analysis
Descriptive statistics were calculated for FM_SCIS_48 by age band, dominance level, and thumb positioning group. Pearson and Spearman correlations were computed between FM_SCIS_48 and all available domain and subdomain scores. Between-group comparisons used independent samples t-tests; effect sizes are reported as Cohen d. Longitudinal analysis used paired t-tests comparing E1 and E2 scores among children with paired evaluations. A cross-sectional growth rate was used to estimate expected natural maturation over the mean evaluation interval, following the methodology established in the response-to-intervention outcomes article in this series. Artificial intelligence writing assistance (Claude, Anthropic, version Sonnet 4.6) was used in the preparation of this manuscript for language editing and formatting; all analytical decisions, clinical interpretations, and conclusions are those of the author.
All data is pre-normative. Rasch analysis validation is ongoing and will be reported in subsequent publications.
Results
Developmental Progression of Cutting Accuracy
Across all age bands, mean scores on FM_SCIS_48 increased steadily, with the steepest growth occurring between Bands G and H, the window when this skill is developmentally expected to emerge (Table 1). At Band H (the anchor age for this task), 35.6% of children in this clinical sample scored zero, reflecting the referred nature of the population. The bimodal distribution within Band H is more clinically informative than a pass/fail classification. Of 165 children scoring 10 or above on FM_SCIS_48, 162 (98.2%) were also administered FM_SCIS_60C on the same evaluation day. Among children who scored 14 on FM_SCIS_48, scores on FM_SCIS_60C ranged across the full spectrum (mean 8.61), confirming that the curvy task captures a meaningfully more demanding level of motor control.
Table 1. FM_SCIS_48 (1/4″ straight line) performance by age band. Data from a clinical sample of children referred for OT evaluation, prior to intervention.
Age Band
Age Range
n
Mean /14
% Score
Floor (0)
Ceiling (14)
G
36-47m
45
1.79
12.8%
64.4%
2.2%
H
48-53m
180
4.64
33.1%
35.6%
10.6%
I
54-59m
248
6.52
46.6%
18.1%
21.4%
J
60-65m
68
8.07
57.7%
13.2%
26.5%
Hand Positioning and Hand Dominance
Dominant hand thumb-up positioning was associated with substantially higher cutting performance. Children with thumb-up positioning averaged 8.16 on FM_SCIS_48 compared to 2.74 for children without thumb-up positioning (Cohen d=1.17, p<0.001; median scores 9.0 versus 1.0). Within Band H, 28% of children who scored zero had thumb-up positioning compared to 89% of children who scored 14. No children who scored 14 on FM_SCIS_48 were rated Never for overall hand positioning quality.
Hand dominance status was consistently associated with cutting performance (Table 2). Children with established dominance scored more than five points higher on average than children with inconsistent dominance. The Spearman correlation between dominance level and FM_SCIS_48 was rho=0.360 (p<0.001). Writing-to-scissor hand consistency produced a large effect: children who used the same hand for writing and cutting averaged 7.06 on FM_SCIS_48 compared to 2.25 for children who used different hands (Cohen d=1.05, p<0.001).
Table 2. FM_SCIS_48 mean score by documented hand dominance level (Bands G-J, n=534). Dominance level based on therapist rubric rating at time of evaluation.
Dominance Level
n
Mean /14
% Score
Floor (scored 0)
Emerging
18
1.50
10.7%
56%
Inconsistent
70
2.38
17.0%
49%
Strong preference
233
5.17
36.9%
28%
Established
213
7.76
55.4%
16%
Cross-Domain Correlations
Scissor cutting performance correlated significantly with all major developmental domain scores (Table 3). Fine Motor domain showed the strongest correlation (r=0.772). Visual Motor Integration (r=0.511) and Visual Perception (r=0.461) correlations reflect the visual guidance demands of cutting within a narrow boundary. The ADL correlation (r=0.511) speaks to generalization of tool use skill across daily living contexts. The Gross Motor correlation (r=0.430) reflects the role of proximal stability as a foundation for distal precision. Praxis showed the weakest correlation (r=0.271), consistent with motor planning contributing primarily to early skill acquisition rather than refined execution. The FMSAT Speed and Accuracy subdomain showed no significant relationship (Spearman rho=-0.012, not significant), confirming that scissor performance and pencil speed capture distinct aspects of fine motor control and contribute independent clinical information.
Table 3. Pearson and Spearman correlations between FM_SCIS_48 and domain/subdomain scores (Bands G-J). *** p<0.001. FMSAT Spearman correlation not significant.
Domain / Subdomain
n
Pearson r
Spearman rho
Clinical Interpretation
Fine Motor domain
454
0.772***
0.787
Strong — cutting as fine motor expression
Hand Use subdomain
541
0.541***
0.549
Bilateral tool use as hand use marker
Visual Motor Integration
537
0.511***
0.519
Visual guidance of cutting path
ADL domain
510
0.511***
0.530
Tool use generalizes to daily living
Visual Perception domain
515
0.461***
0.469
Line discrimination supports accuracy
Gross Motor domain
540
0.430***
0.444
Proximal stability drives distal precision
Bilateral Integration
538
0.398***
0.395
Two-hand coordination demand
Praxis domain
539
0.271***
0.281
Motor planning, weaker once program formed
Speed & Accuracy (FMSAT)
459
r=-0.105*
rho=-0.012 (ns)
Distinct skill; independent clinical value
Longitudinal Gains
Among 82 children with paired evaluations, FM_SCIS_48 showed a mean gain of 5.21 points over an average interval of 4.8 months (paired t-test t=8.10, p<0.001; Cohen d=0.895). The curvy task (FM_SCIS_60C) showed a mean gain of 3.75 points over the same interval (t=7.20, p<0.001), with 76.4% of children improving. Using the cross-sectional growth rate as a natural maturation baseline, the expected natural growth on FM_SCIS_48 over 4.8 months was approximately 1.51 points. The observed mean gain of 5.21 points exceeded this expected rate by 3.71 points. Ceiling effects were noted among children who had scored at or near 14 at E1; therapists correctly applied FM_SCIS_60C as the clinically sensitive measure for near-ceiling children in 98.2% of applicable cases.
Discussion
This analysis of 541 pediatric OT evaluations demonstrates that a brief, standardized scissor skills assessment generates clinically meaningful data across multiple dimensions of preschool development. The developmental progression of FM_SCIS_48 scores across age bands aligns with published milestone expectations and provides clinical benchmarks for a referred population that are not currently available in the literature. Floor effects in younger bands and the bimodal distribution at the anchor age band reflect the nature of scissor skill acquisition in children referred for OT services, where skill emergence is delayed relative to normative expectations. These distributions support documentation of functional deficit and medical necessity in ways that pass/fail classifications cannot.
The large effect of thumb-up dominant hand positioning (Cohen d=1.17) elevates hand positioning from a clinical observation to a primary, modifiable intervention target. Establishing correct scissor grip before focusing on path accuracy is supported by the data: children without functional thumb-up positioning have mean scores below three regardless of age band, suggesting that grip orientation is a near-prerequisite for achieving cutting accuracy at the mastery level.
The relationship between hand dominance and scissor performance connects these findings to a broader literature on neuromotor specialization in the preschool years (Scharoun & Bryden, 2014). A child who has not yet organized a consistent preferred hand for tool use is reflecting an underlying developmental state that affects performance across all tool-based tasks. Writing-to-scissor hand consistency findings have direct implications for school-based practice: addressing hand consistency across classroom tool use activities, not only during designated scissor tasks, is consistent with both the data and motor learning principles and supports IEP goal development that reflects the child’s functional performance across settings.
The cross-domain correlation profile challenges the framing of scissor skills as an isolated fine motor task. The significant Gross Motor correlation (r=0.430) reinforces that proximal stability, including trunk support and shoulder girdle control, is foundational to distal precision, consistent with developmental neuroscience frameworks (Stoodley, 2016). The VMI and Visual Perception correlations reflect the visual guidance demands of path-following. Intervention plans that address only the distal cutting task without considering foundational postural, visual, and neuromotor systems may produce slower or less durable gains. The absence of a significant FMSAT correlation confirms that scissor accuracy and pencil precision speed are complementary measures rather than redundant ones, and that both contribute independent information to a comprehensive fine motor profile.
Longitudinal gains of 5.21 points over 4.8 months, exceeding the expected natural growth rate by 3.71 points in a population with limited scissor practice outside of OT sessions, provide meaningful support for OT as a driver of skill development. This above-expected gain pattern is consistent with the RTI methodology established in this article series, which separates natural developmental progress from intervention-attributable change using cross-sectional growth rates as a baseline correction. For medical model practitioners, this framing supports quantifiable medical necessity documentation. For school-based practitioners, it provides data to support RTI tier documentation and progress monitoring language consistent with IDEA requirements.
Several methodological factors warrant consideration. This sample represents children referred for OT evaluation in North Carolina and should not be generalized to typically developing children or other geographic populations. The cross-sectional age-band comparisons reflect group differences rather than individual trajectories. Adaptive scissor data was insufficient for analysis. Without a controlled comparison group, causal claims about OT effectiveness cannot be made; above-expected gains in a referred population with limited community practice are suggestive but not definitive. Future research should include a typically developing comparison sample, session-level dosage data, and expanded age coverage through early elementary years to determine whether scissor precision continues to develop beyond 66 months and at what point the 1/4-inch straight line becomes a floor item for older children (Cameron et al., 2012; Zhang et al., 2025).
Conclusions
A structured scissor skills assessment generates clinically meaningful data for pediatric OT practice. Cutting accuracy in a referred preschool population follows a measurable developmental trajectory, is strongly predicted by thumb-up dominant hand positioning and established hand dominance, correlates significantly with fine motor, gross motor, visual motor, visual perception, ADL, and praxis domains, and shows gains that exceed expected natural growth rates over a 4.8-month evaluation interval in a population with limited community scissor access. These findings support the use of standardized, quantitative scissor skills assessment as a component of comprehensive pediatric OT evaluation and as a practical tool for RTI documentation, goal writing, and cross-domain clinical reasoning in both school-based and medical model practice settings.
Disclosure of Interest
Stephanie Seymore Wick is the Founder and Clinical Architect of O.T. Wizard, the platform from which all data in this article was collected. Data collection is ongoing under clinical quality improvement protocols. All data were de-identified in accordance with HIPAA regulations. The author reports no other competing interests.
Data Availability Statement
De-identified aggregate data supporting the findings of this study are available from the corresponding author upon reasonable request. Individual-level data cannot be shared due to HIPAA de-identification obligations.
Biographical Note
Stephanie Seymore Wick, MSOT, OT/L is the Founder and Clinical Architect of O.T. Wizard, a clinical intelligence platform for pediatric occupational therapy professionals, and the founder of Learning Charms, Inc. Her clinical and research focus is the development of psychometrically sound, computable measurement tools that quantify pediatric OT outcomes across multiple developmental domains. She practices and conducts research in North Carolina.
References
Cameron, C. E., Brock, L. L., Murrah, W. M., Bell, L. H., Worzalla, S. L., Grissmer, D., & Morrison, F. J. (2012). Fine motor skills and executive function both contribute to kindergarten achievement. Child Development, 83(4), 1229-1244. https://doi.org/10.1111/j.1467-8624.2012.01768.x
Scharoun, S. M., & Bryden, P. J. (2014). Hand preference, performance abilities, and hand selection in children. Frontiers in Psychology, 5, 82. https://doi.org/10.3389/fpsyg.2014.00082
Stoodley, C. J. (2016). The cerebellum and neurodevelopmental disorders. Cerebellum, 15(1), 34-37. https://doi.org/10.1007/s12311-015-0715-3
Zhang, B.-F., Lin, Z.-C., & Li, C. (2025). Fine motor skills assessment instruments for preschool children with typical development: A scoping review. Frontiers in Psychology, 16, 1620235. https://doi.org/10.3389/fpsyg.2025.1620235
We Have Always Known Scissor Skills Matter. Now We Have the Numbers.
You already know scissor skills are more than a craft milestone. You know a child who cannot hold scissors with thumb up, who switches hands mid-cut, or who can barely score a line with school scissors is telling you something important about their neuromotor development. What we have not always had is the data to say exactly what they are telling us, and to show that picture to the families, teachers, and payers who need to understand it.
That is what this analysis is about.
O.T. Wizard collected structured scissor skills assessment data across 541 pediatric OT evaluations from preschool-aged children (ages 3 to 5 and a half) in North Carolina. At least 90% of children qualified for Medicaid. All were referred for OT services after failing a developmental screening. For many, OT sessions were the primary or only setting where scissors were regularly available. Teachers limit scissor time in large classroom groups. Parents restrict use at home. That context matters a lot when we look at what the data shows.
What We Measured
The O.T. Wizard Scissor Skills Assessment is a structured, scored tool built into the evaluation process. It uses a developmental progression of cutting tasks, from snipping (around 34 months) through cutting a 1/4-inch straight line (around 48 months) to cutting a 1/4-inch curvy line (around 60 months). Each task is scored on a 0 to 14 point scale using a consistent segment-by-segment rubric. Therapists also document thumb orientation, stabilizer hand, and overall hand positioning quality.
Unlike a standardized evaluation tool that captures a single snapshot in one domain, O.T. Wizard tracks performance across twelve domains simultaneously including fine motor, gross motor, visual motor integration, visual perception, activities of daily living, praxis, and executive functioning. This is what made the findings below possible. We were not just looking at how a child cuts. We were looking at what scissor skill tells us about the rest of the child.
What We Found
Scissor skill development follows a clear developmental path, and most 4-year-olds referred for OT are not yet where we would expect.
All data in this table reflects first evaluations (E1) conducted before OT intervention began. These are the starting points, the picture of where children arrived for their first evaluation. On the 1/4-inch straight line task (developmentally anchored at 48 months), here is how children in this clinical sample performed at initial evaluation:
Age Group
Avg Score /14
% Accurate
Scored Zero (0/14)
Scored Perfectly (14/14)
3:0 to 3:11 (Band G)
1.8
13%
64%
2%
4:0 to 4:5 (Band H)
4.6
33%
36%
11%
4:6 to 4:11 (Band I)
6.5
47%
18%
21%
5:0 to 5:5 (Band J)
8.1
58%
13%
27%
Table 1. Scissor skills performance on 1/4-inch straight line task at first evaluation (E1), before OT intervention. Clinical sample referred for OT evaluation. Data from O.T. Wizard, pre-normative.
At the anchor age for this task, 36% of children arrived at their first evaluation scoring zero on the 1/4-inch straight line, and another 21% scored between 1 and 3 out of 14. That means that at initial evaluation, before OT had even begun, the majority of 4-year-olds referred for services could not yet cut a 1/4-inch line with any consistency. That is not a failure of intervention. That is a description of who we serve and why they need us. And now we can describe it precisely, in numbers, instead of writing “emerging scissor skills” in a narrative.
These are clinical benchmarks for a referred population, not norms for typically developing children. They show where our kids start. That starting point is worth documenting and measuring.
The multi-task design works. Advancing to harder tasks captures the full range of scissor skill ability.
A common concern with developmental assessments is ceiling effects: what happens when a child is too skilled for the task in front of them? The data confirmed that therapists in this sample handled this correctly. Of children who scored 10 or higher on the 1/4-inch straight line task, 98% were also given the curvy line task on the same evaluation day. Among children who scored a perfect 14 on the straight line, the curvy line scores ranged across the full spectrum, averaging 8.6 out of 14 (61% accuracy). The harder task captured real, meaningful variation that the straight line task could not.
Group
n
Straight Line Avg /14
Straight Line % Accurate
Curvy Line Avg /14
Curvy Line % Accurate
Scored 10+ on straight line
165
12.5
89%
6.8
49%
Scored 14 on straight line
90
14.0
100%
8.6
61%
Table 2. Curvy line performance among children approaching or reaching ceiling on the straight line task. Confirms that advancing to the harder task captures meaningful clinical variation.
Thumb up is not just a cue. It is the difference between scissor skills and not having them.
This was the strongest predictor in the entire dataset. Hand positioning was not just associated with better cutting accuracy. It was associated with a more than 3-fold difference in performance.
Thumb Positioning
n
Avg Score /14
% Accurate
Median Score /14
Thumb-up (correct)
~270
8.2
59%
9.0
Not thumb-up
~270
2.7
19%
1.0
Table 3. Scissor skills performance by dominant hand thumb positioning. Bands G through J. Cohen d=1.17 (large effect).
Looking at Band H children specifically: among those who scored a perfect 14 on the 1/4-inch line, 89% had correct thumb-up positioning. Among those who scored zero, only 28% did. Not a single child at zero had positioning rated as “always” correct. The positioning is not the finishing touch on scissor skill development. It is the prerequisite.
For intervention planning, this data strongly supports prioritizing grip and hand orientation before focusing on line-following accuracy. Therapists who have structured treatment this way have been right all along. Now there are numbers to back it up, and to share with families and IEP teams.
Hand dominance predicts scissor skill accuracy, and the connection runs deeper than which hand holds the scissors.
Children with established hand dominance performed dramatically better on the cutting task than children with inconsistent or emerging dominance. The data below reflects all children in Bands G through J at initial evaluation.
Dominance Level
n
Avg Score /14
% Accurate
% Scored Zero
Emerging
18
1.5
11%
56%
Inconsistent
70
2.4
17%
49%
Strong preference
233
5.2
37%
28%
Established
213
7.8
56%
16%
Table 4. Scissor skills performance by documented hand dominance level at initial evaluation. Spearman rho=0.360, p<0.001.
The hand consistency finding was equally striking. Children who used the same hand for both writing and cutting averaged 7.1 out of 14 (51% accuracy). Children who used different hands for writing versus cutting averaged only 2.3 out of 14 (16% accuracy). That is a more than 3-fold performance gap, and it connects directly to the Hand Dominance article in this series.
A child who switches hands between writing and cutting is not just making an inconsistent tool choice. They are reflecting an unresolved neuromotor organization question that affects all tool use tasks. For school-based OTs, this is IEP-relevant data. Documenting hand consistency across writing and cutting tasks as part of the evaluation supports accommodation planning for the full classroom day, not just during scissor activities.
Scissor skills are a whole-body skill. The domain correlation data proves it.
This is where O.T. Wizard’s multi-domain approach made findings possible that are simply not achievable with a single standardized assessment tool that only looks at one domain or captures a single point in time.
Cutting performance was correlated against every other domain scored in the same evaluation. Here is the picture that emerged:
Domain
Correlation with Scissor Skills Score
What This Means Clinically
Fine Motor
Strong (r=0.77)
Expected: distal precision underlies cutting
Visual Motor Integration
Moderate (r=0.51)
Visual guidance is needed to follow a line
ADL (Daily Living)
Moderate (r=0.51)
Tool use skill generalizes across daily tasks
Visual Perception
Moderate (r=0.46)
Seeing and interpreting the line drives cutting accuracy
Gross Motor
Moderate (r=0.43)
Trunk and shoulder stability support distal hand control
Bilateral Integration
Moderate (r=0.40)
Confirmed: cutting is a two-handed coordinated task
Praxis
Weak-moderate (r=0.27)
Motor planning matters early; less so once the program is established
FMSAT (Fine Motor Speed and Accuracy Test)
Near zero (r=0.00)
Pencil speed and cutting accuracy are distinct skills
Table 5. Correlations between scissor skills performance (1/4-inch straight line) and O.T. Wizard domain scores. All correlations p<0.001 except FMSAT which was not significant. Bands G through J, n=454 to 541 depending on domain.
The gross motor correlation deserves a specific callout. Cutting is a distal fine motor task, but proximal stability drives distal precision. Trunk support and shoulder girdle control provide the foundation from which hand precision is expressed. A child with poor postural stability will have reduced arm control, which directly limits how precisely they can guide scissors along a line. This is the body-supports-hand principle that experienced OTs understand clinically. This dataset quantifies it.
The FMSAT finding is equally important in a different way. FMSAT “Bubble-popping” measures open-field pencil speed and precision. Scissor skills assessment measures controlled path-following with a bilateral tool. These two tasks draw on related but distinct aspects of fine motor function, and both contribute independent clinical information. Having both in the same evaluation is not redundancy. It is clinical depth.
This is the kind of cross-domain picture you cannot build from a BOT-2 or a Beery VMI alone. Those tools give you a score in an isolated domain. O.T. Wizard gives you a developmental profile across twelve domains, from the same child, on the same day, every time you complete an evaluation.
Children Made Real Progress, and OT Likely Drove Most of It.
Among the 97 children with two evaluations on file, those with scissor skills data at both time points showed an average gain of 5.2 points on the 1/4-inch straight line over an average of 4.8 months between evaluations. Nearly 70% showed improvement. However, not all of the sample had scissor related goals on their plan of care.
Measure
E1 (Before OT)
E2 (After ~5 months OT)
Gain
% Who Improved
Avg score /14 (straight line)
4.4 / 14
9.6 / 14
+5.2 pts
69.5%
% Accuracy
31%
69%
+38 pts
Avg score /14 (curvy line)
2.5 / 14
6.2 / 14
+3.8 pts
76.4%
% Accuracy (curvy)
18%
44%
+27 pts
Table 6. Scissor skills gains from initial evaluation (E1) to re-evaluation (E2). n=82 for straight line, n=72 for curvy line. Mean interval 4.8 months. Clinical sample referred for OT services.
Based on cross-sectional growth data, we would expect natural developmental growth of about 1.5 points over a 4.8-month window. These children gained 5.2 points. That is 3.7 points above the expected natural rate, representing more than triple the growth that maturation alone would predict.
And remember the context: these children were largely practicing scissor skills only during OT sessions. Teachers avoid large-group scissor time. Parents restrict home use. If nearly all scissor practice was happening in OT, and children gained more than three times the expected natural growth rate, that is meaningful evidence for OT-driven outcomes, even without a randomized controlled trial.
For school-based and medical model OTs alike, this is the kind of data that supports medical necessity, justifies continuation of services, and answers the parent question: Is this working?
What Makes This Different From a Standard Evaluation
Standardized tools like the BOT-2, PDMS-3, or Beery VMI are valuable. They are not being replaced. But they have a structural limitation: they capture performance in isolated domains, on a single day, at a single point in time. They cannot show how a child’s scissor skills score connects to their gross motor stability or ADL function. They cannot track how a child changes from evaluation to re-evaluation. And they cannot build a growing evidence base across hundreds of children that gets more precise over time.
O.T. Wizard was designed to do all of those things. Every evaluation adds to the clinical intelligence base. Scissor skills scores sit alongside fine motor, gross motor, visual perception, ADL, VMI, praxis, executive functioning, and participation data from the same child on the same day. The cross-domain correlations in this article were only possible because the platform was built to be computable, not just documentable.
This is the difference between documenting therapy and understanding it.
What This Means for Your Practice
If you work in a school setting: Scissor skills accuracy scores contextualized within a developmental progression give you IEP-ready language. A score of 4.6 out of 14 at Band H is not just “emerging” and is a measurable starting point with room to define a meaningful, achievable goal. The hand consistency finding connects directly to classroom accommodations: if a child switches hands between writing and cutting tasks, that is relevant accommodation data for the IEP team, not just a therapy note.
If you work in a medical model setting: The 5.2-point average gain over 4.8 months, in a population with almost no between-session scissor exposure, supports medical necessity documentation with concrete numbers. The cross-domain correlations support a whole-child framing in your evaluation report: scissor skills difficulty is not just a fine motor problem. It reflects neuromotor organization, visual guidance, postural stability, and bilateral coordination working together.
For both settings: Thumb-up hand positioning is the most actionable clinical target in the dataset. If thumb orientation is not being documented and targeted as a prerequisite for cutting accuracy, this data makes the case for starting there.
Scissor Skills Are the Child’s Story
A pair of scissors in a preschooler’s hand is a window. It shows you how well the brain has organized a preferred side, how the trunk is supporting the arms, how the eyes are guiding the hands, and how much the child has internalized the motor program for this specific tool. It is one of the richest clinical observations we make, and one of the least quantified.
That is changing. Data from over 500 evaluations now gives us a picture of what scissor skill development looks like in the children we actually serve, what predicts success, what domains are implicated, and what progress looks like over time. This is the beginning of an evidence base that the profession has needed.
Stephanie Seymore Wick is the Founder and Clinical Architect of O.T. Wizard, the platform from which all data in this article was collected. Data collection is ongoing under clinical quality improvement protocols. All data has been de-identified in accordance with HIPAA regulations.
About O.T. Wizard
O.T. Wizard is a clinical intelligence system for pediatric occupational therapy professionals. The platform evaluates performance in evaluations, forms, and daily treatment notes across twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, executive functioning, activities of daily living, and participation. O.T. Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update continuously as the clinical database expands. Learn more at otwizard.com.
89% Improved. Five Domains Exceeded Natural Growth. Here Is What the Data Shows.
Measuring Pediatric OT Outcomes Above the Threshold of Natural Maturation
By Stephanie Seymore Wick, MSOT, OT/L | Founder and Clinical Architect, O.T. Wizard
Introduction
Every pediatric occupational therapist knows that the work they do matters. The harder question is whether the profession can show, in precise and reproducible terms, how much it matters. For decades, OT documentation has been built around goals, progress notes, and clinical narratives. These tools record care. They rarely measure change in a way that separates what the child gained through intervention from what developmental maturation would have produced on its own.
This report addresses that gap directly. Using O.T. Wizard, a clinical intelligence system designed to generate structured, reproducible, multi-domain assessment data for pediatric OT practice, we examined functional performance change across eight domains in 71 preschool-aged children who completed two full evaluations an average of 4.85 months apart.
The central question throughout this analysis is not simply whether children improved. The meaningful question is whether the children in this cohort improved beyond what developmental maturation alone would have produced over the same interval. A natural growth correction applied consistently throughout this report makes that distinction explicit in every finding.
This is the expanded replication of a February 2026 analysis of 44 paired evaluations. The findings across the 27 additional pairs are consistent with and strengthen the earlier report across all domains. The core story does not change with more data. It becomes more precise.
ABSTRACT
Background: Pediatric occupational therapy has well-established standardized tools for point-in-time measurement, including the Bruininks-Oseretsky Test of Motor Proficiency, Beery-Buktenica Developmental Test of Visual-Motor Integration, and Peabody Developmental Motor Scales-Third Edition. Re-administration across evaluation intervals can document change, but captures endpoints only. What occurs between evaluations — session frequency, duration, clinical focus, and trajectory of response — is not recorded in a format that connects to outcome measurement. Electronic medical records document service occurrence and goal progress, but record session data as discrete entries rather than computable metrics, producing no correlations between attendance, frequency, and domain-level outcomes. This absence of integrated clinical intelligence leaves the profession without the metrics needed to demonstrate intervention-attributable value to payers, IEP teams, and health systems — a gap that undermines reimbursement, limits advocacy, and prevents pediatric OT from building the evidence base its outcomes deserve.
Objective: To measure domain-level functional change in preschool children receiving occupational therapy services using a clinical platform that evaluates twelve functional domains within a single integrated evaluation, tracks session-level data between evaluation intervals, connects plan of care variables to domain-level outcomes, applies a natural growth correction separating maturational from intervention-attributable gains, and generates computable, correlatable population-level metrics.
Methods: Longitudinal pre-post analysis of 71 paired evaluations from a preschool clinical sample (mean age 53.5 months; mean interval 4.85 months; at least 90% Medicaid-qualifying). A 9.2% natural growth rate was applied as the maturational baseline. Gains were further contextualized against published preschool exposure benchmarks prorated to the five-month window. All data are pre-Rasch ordinal values.
Results: 89% of children improved in under five months. Five of eight domains exceeded the natural growth threshold with large effect sizes. VMI exceeded the published preschool exposure benchmark by d=0.92, ADL by d=0.90, and Fine Motor by d=0.65. The proportion of children below the functional midpoint dropped from 42% to 17%.
Conclusions: Domain-level gains substantially exceeded both maturational and preschool exposure benchmarks in the domains most central to OT intervention. Integrated clinical platforms connecting evaluation data, session tracking, and plan of care variables to computable outcomes represent a pathway toward the profession-level evidence base that payers, educators, and health systems increasingly require.
Study Sample
Age Distribution
The longitudinal cohort consisted of 71 preschool-aged children, each with two complete O.T. Wizard evaluations separated by a minimum of 30 days. The mean inter-evaluation interval was 4.85 months (approximately 148 days), with a range of approximately 37 to 173 days. Mean age at first evaluation was 53.5 months.
Starting age band distribution: Band G (36 to 47.99 months, n=6), Band H (48 to 53.99 months, n=29), Band I (54 to 59.99 months, n=33), and Band J (60 to 65.99 months, n=3). Bands H and I together represent 87% of the sample and are the primary basis for findings reported here. Bands G and J are included in the data but interpreted with caution given their smaller sizes.
Demographics and Clinical Status
All assessments were conducted in North Carolina through the O.T. Wizard clinical platform. Consistent with the broader software dataset, at least 90% of children qualified for Medicaid, and for many, the structured evaluation environment represented an early introduction to formal educational or clinical settings. Primary language was English for 90% of children, with 9% Spanish-speaking and 1% other. All children had been recommended for occupational therapy services following developmental screening failure.
It is important for readers to interpret these findings within this clinical context. This is not a typically developing population. These are children with identified developmental concerns who were referred for and receiving skilled OT services. Outcome findings therefore reflect the response of a clinically referred, predominantly low-income sample to structured early intervention, not population-level developmental norms.
Natural Growth Framework
Before examining domain-level findings, it is necessary to establish what score change we would expect to observe in the absence of intervention. Children in this cohort averaged 53.5 months of age at first evaluation and were reassessed approximately 4.85 months later. On a well-constructed developmental scale, maturation alone would be expected to produce a gain proportional to that age progression.
The natural growth rate for this cohort is calculated as the mean inter-evaluation interval divided by the mean age at Evaluation 1: 4.85 months divided by 53.5 months equals 9.2%. This figure represents the expected score improvement attributable to developmental maturation alone over the study period.
A gain of 9.2% would be expected from natural maturation alone over 4.85 months.Gains above 9.2% represent intervention-attributable change.
This natural growth rate serves as the reference threshold throughout this report. Domain gains below 9.2% suggest performance did not keep pace with chronological age progression. Gains at 9.2% suggest maturation-equivalent growth. Gains above 9.2% represent functional improvement beyond what age progression alone would predict.
This correction is transparent, reproducible, and requires no external normative sample to apply. It is a direct arithmetic relationship between age progression and scale progression on a fixed instrument. The Gain Above Natural column in each table makes this comparison explicit.
Research Hypotheses
Three primary hypotheses guided this analysis. First, children receiving occupational therapy services would demonstrate composite score gains substantially exceeding the 9.2% natural growth threshold. Second, domains most directly targeted by OT in preschool settings, specifically visual motor integration, fine motor skills, visual perception, and activities of daily living, would show the largest gains above expected growth. Third, Band H children (48 to 53.99 months) would demonstrate greater gains than Band I children, reflecting greater developmental sensitivity at a younger starting point.
Literature Review and Context
The preschool years represent a critical period for fine motor and visual motor development. Between the ages of three and five, neuromotor pathways underlying pencil control, bilateral coordination, and hand specialization undergo rapid maturation, establishing the foundation for academic skill development. Handwriting readiness, scissor use, and self-care independence all draw from skill sets that are most efficiently built during this developmental window.
Visual motor integration has consistently been identified as one of the strongest predictors of kindergarten handwriting readiness. Daly and colleagues (2003) found VMI performance at preschool age predicted handwriting speed and legibility at ages six and seven with effect sizes exceeding those of fine motor or visual perception measures alone. Duff and colleagues (2015) demonstrated that children with developmental coordination difficulties who receive targeted fine motor intervention during the preschool years show significantly better handwriting outcomes at school entry than matched peers without services. It is equally important to recognize that VMI does not operate in isolation as a predictor of handwriting development. Emergent literacy skills, particularly alphabet knowledge, letter-sound awareness, and early orthographic processing, are also well-established predictors of handwriting fluency and transcription accuracy (Gerde et al., 2025; Puranik et al., 2011). The relationship between literacy exposure and VMI development is bidirectional: children who are actively engaged in letter-learning and pre-writing activities in preschool settings are simultaneously building the visual discrimination, directionality, and motor planning foundations that underlie VMI performance. This intersection is clinically relevant and is acknowledged as a study limitation below.
The measurement infrastructure required to track these outcomes longitudinally has historically been a limiting factor in OT outcomes research. Standard evaluation protocols typically capture a single snapshot of performance, and re-evaluation data, when it exists, is rarely structured for computational comparison. O.T. Wizard was designed to address this gap, enabling structured, reproducible, multi-domain measurement within the constraints of a standard clinical evaluation and supporting longitudinal outcome tracking that traditional paper-based protocols do not practically support.
This report also extends a prior O.T. Wizard longitudinal analysis (Wick, 2026) that examined 44 paired evaluations from the same clinical platform. The expanded sample of 71 pairs presented here confirms and extends those findings with consistent direction and strength across all eight domains assessed.
Key Findings
Overall Composite Performance
Across all 71 children with valid paired evaluations, mean composite score increased from 519.8 at Evaluation 1 to 646.0 at Evaluation 2, a mean raw gain of 126.2 points. Against the 9.2% natural growth expectation, the expected gain for this cohort was approximately 47.6 points. The observed gain exceeded the natural growth threshold by 78.6 points, representing 165% above expected developmental progress. To be precise about what that means: for every point of progress that natural maturation would have produced, these children gained 2.65 points. They moved forward at more than two and a half times the rate that developmental aging alone would have driven. That is not incremental. That is intervention doing exactly what skilled, structured, early occupational therapy is designed to do.
The gain was highly statistically significant (paired t-test, t=10.62, p<0.001, Cohen’s d=1.26, large effect). 89% of children showed improvement at the second evaluation. 42% of children began below the 500-point composite threshold; by Evaluation 2, only 17% remained below that threshold. Twenty children crossed the functional midpoint of the scale during the study interval.
89% of children improved. 20 children crossed the 500-point functional threshold.The composite gain exceeded expected natural growth by 165%.
Domain-Level Results
The following table presents results across all eight assessed domains, sorted by magnitude of gain above natural growth. All scores are pre-Rasch ordinal percentage values expressed as points within each domain’s maximum possible score. Natural Gain represents the expected gain based on the 9.2% natural growth rate applied to each domain’s Evaluation 1 mean.
Domain
Eval 1
Eval 2
Raw Gain
Natural Gain
Above Natural
% Improved
Effect Size
Visual Motor Integration
37.6
62.2
+24.6
3.4
+21.1 (614%)
94%
d=1.49 (Large)
Activities of Daily Living
45.5
69.2
+23.7
4.2
+19.5 (468%)
87%
d=1.27 (Large)
Fine Motor Skills
47.7
63.7
+16.0
4.3
+11.6 (268%)
86%
d=1.00 (Large)
Gross Motor Skills
56.8
72.3
+15.5
5.2
+10.3 (198%)
73%
d=0.77 (Medium)
Visual Perception
62.7
75.1
+12.4
5.7
+6.7 (118%)
77%
d=0.74 (Medium)
Praxis
52.1
56.6
+4.5
4.8
-0.3 (-6%)
48%
d=0.15 (ns)
Participation
65.2
69.4
+4.2
6.0
-1.8 (-30%)
62%
d=0.26 (*)
Executive Functioning
63.7
66.2
+2.5
5.8
-3.4 (-57%)
52%
d=0.15 (ns)
Table 1. Domain-level longitudinal comparison. Scores are points within each domain’s maximum possible score. Natural Gain = Eval 1 mean x 9.2% natural growth rate. Above Natural = Raw Gain minus Natural Gain. Effect sizes: Large (d>0.8), Medium (d>0.5). Executive Functioning, Participation, and Praxis findings are addressed in the discussion section. All scores are pre-Rasch raw values.
Visual Motor Integration produced the largest gain above expected growth in the dataset, rising from 37.6 to 62.2 points, a raw gain of 24.6 points against an expected natural gain of 3.4 points. The gain above natural growth was 21.1 points, representing 614% above what maturation alone would have produced. 94% of children with VMI scores showed improvement. The effect size of d=1.49 is considered large by conventional standards.
Activities of Daily Living showed a raw gain of 23.7 points against a natural expectation of 4.2 points, placing the gain above natural growth at 19.5 points (468% above expected). Fine Motor Skills showed 11.6 points above the natural expectation (268% above expected, d=1.00, large). Gross Motor Skills showed 10.3 points above expected (198% above expected, d=0.77, medium). Visual Perception showed 6.7 points above expected (118% above expected, d=0.74, medium).
Praxis, Participation, and Executive Functioning showed gains at or below the natural growth threshold, none with statistically significant large effects. The interpretation of these findings requires clinical context and is discussed in detail below. A dedicated companion analysis of the Participation and Executive Functioning longitudinal findings is forthcoming in this research series, as the novelty effect hypothesis and its implications for clinical documentation merit extended treatment.
Performance by Starting Age Band
The following table presents composite score change by starting age band. Natural growth rates vary slightly by band because younger children have a larger age progression ratio over the same elapsed time. Bands G and J are included for completeness but should be interpreted with caution given small sample sizes.
Age Band
n
Eval 1 Mean
Eval 2 Mean
Raw Gain
Above Natural
NGR
G (36-47.99 mo)
6
394.7
496.3
+101.7
+57.0
11.3%
H (48-53.99 mo)
29
516.6
651.3
+134.8
+84.8
9.7%
I (54-59.99 mo)
33
542.7
670.2
+127.5
+81.9
8.4%
J (60-65.99 mo)
3
549.7
628.0
+78.3
+32.8
8.3%
Table 2. Composite score change by starting age band. NGR = natural growth rate (months elapsed / age at Eval 1). Above Natural = Raw Gain minus (Eval 1 mean x NGR). Bands G and J interpreted with caution (small n).
Band H children (ages 48 to 53.99 months) showed the largest absolute gains above natural growth, averaging 84.8 points above the natural expectation on a composite gain of 134.8 points. Band I children showed 81.9 points above expected on a composite gain of 127.5 points. Both primary age bands show gains well above the natural growth threshold, and the difference between them is modest. This is broadly consistent with the earlier 44-pair analysis, which found Band H slightly outperforming Band I. The 48 to 54 month window continues to appear as a period of high clinical yield for OT service delivery, though both bands show substantial responsiveness.
Understanding the Flat Domains: Praxis, Participation, and Executive Functioning
Three domains showed gains at or below the natural growth threshold: Praxis (-6%), Participation (-30%), and Executive Functioning (-57%). These findings are clinically important to interpret carefully, as they do not simply mean that OT failed to produce change in these areas.
For Praxis, the near-zero gain is more likely a reflection of current measurement sensitivity than true insensitivity to intervention. Praxis is a complex, context-dependent construct requiring the integration of motor planning, bilateral coordination, and sequencing across novel tasks. Detecting incremental praxis development over a five-month interval likely requires either longer measurement windows or more precisely calibrated items. Rasch calibration of the praxis item bank is a priority in the continuing research agenda.
Participation and Executive Functioning tell a more nuanced story that involves the measurement context itself. Both domains are rated by the therapist based on behavioral observation during the evaluation. At Evaluation 1, the child is meeting the therapist for the first time. The novelty of the interaction, the structured environment, and the desire to engage with an unfamiliar adult may produce elevated ratings that reflect situational compliance rather than the child’s authentic behavioral baseline. By Evaluation 2, the therapeutic relationship is established and the child is comfortable enough to reveal their genuine regulatory and engagement patterns, including the variability and difficulty that characterize their daily functioning. If Evaluation 1 ratings are systematically elevated by this novelty effect, the apparent absence of gain at Evaluation 2 reflects a measurement context shift rather than a failure of intervention. A dedicated research blog on the novelty effect hypothesis and its implications for clinical documentation is forthcoming in this series.
Correlation Analysis
A moderate negative correlation was observed between starting composite score and magnitude of change (consistent with the 44-pair analysis). Children who began with lower scores tended to show larger gains. This regression-to-the-mean effect is expected in clinical samples and does not invalidate the findings, but is an important interpretive consideration. No significant correlation was found between inter-evaluation interval length and change score, indicating that the range of intervals in this cohort (approximately 37 to 173 days) did not materially influence the magnitude of observed gains.
Implications for OT Practice
The domain-level findings interpreted through the natural growth framework allow occupational therapists to make specific, evidence-informed decisions about evaluation and intervention priorities. Five domains showed gains ranging from 118% to 614% above the 9.2% natural growth threshold, all with statistical significance and large or medium effect sizes. These gains were achieved in a predominantly Medicaid-qualifying, low-income clinical population with significant developmental concerns, which makes the magnitude of change all the more clinically meaningful.
For insurance authorization and educational planning, the natural growth framework provides a communication tool that is both precise and accessible. Rather than reporting a raw score change, the practitioner can state that the child’s VMI performance exceeded the expected developmental rate by 21.1 points over approximately five months, providing clear evidence that skilled OT intervention, not maturation, drove the observed change. This framing is methodologically transparent and directly responsive to the medical necessity standards that payers apply.
The composite score threshold finding carries particular weight for authorization purposes. A child who begins services below the 500-point composite threshold and crosses it during the authorization period has demonstrated objectively measurable functional change. Of the 30 children who began below 500 points, 20 crossed that threshold during the study interval. That is a two-thirds success rate in moving children from below-threshold to at-threshold performance within a single authorization period.
Serial assessments using a consistent instrument also generate slope data that goes beyond a single outcome comparison. The rate of gain above natural growth, calculated at the domain level, can be used to project whether a child is on track to reach functional goals within a given authorization period, supporting proactive communication with payers and educational teams before a plateau needs to be explained rather than after.
Implications for Intervention Planning
The convergent large effect sizes across VMI, ADL, Fine Motor, and Gross Motor domains point toward a functional skill cluster that is highly responsive to structured OT programming during the preschool developmental window. These four domains share underlying requirements for postural control, bilateral coordination, and visually guided hand movement. Interventions that integrate these components across functional activities are supported by both the data pattern and established OT theory.
The Gross Motor finding is particularly relevant for intervention sequencing. A gain of 10.3 points above natural expectation with a medium-to-large effect confirms that proximal postural and movement foundations are responsive to OT services alongside distal fine motor work. For children showing limited fine motor or VMI gains, postural foundation and gross motor assessment should be considered before concluding that the upper extremity is the primary limiting factor.
For children whose evaluation profiles show strength in Gross Motor relative to Fine Motor and VMI, a proximal-to-distal intervention sequence may accelerate gains across the entire cluster. The strength of the ADL finding (d=1.27) reflects the functional integration that OT uniquely provides: when children gain in fine motor, VMI, and postural control simultaneously, daily living skills follow as a natural downstream effect.
The flat findings for Praxis, Participation, and Executive Functioning should not reduce the clinical attention given to these areas. They reflect current measurement constraints rather than evidence of non-response to intervention. Goal writing in these domains should continue, supported by structured therapist observation and emerging platform tools designed to capture behavioral change over longer intervals.
Study Limitations
This study carries several important limitations that readers should consider when interpreting and applying the findings.
The sample is clinical and geographically restricted to North Carolina. Findings cannot be generalized to typically developing children or to populations in other regions with different demographic profiles, service delivery models, or referral criteria. The absence of a control group means observed gains cannot be causally attributed to OT intervention. Natural maturation, regression to the mean, and test familiarity effects each contribute to observed change scores to an unknown degree.
The 9.2% natural growth correction is a methodologically transparent estimate derived directly from the age progression of this cohort on this instrument. It assumes proportional developmental scaling across the score range, an assumption that Rasch calibration will allow us to test empirically. Future work with a typically developing comparison group will allow domain-specific, empirically derived growth expectations to replace this uniform estimate.
The regression-to-the-mean effect means that domains with the lowest Evaluation 1 scores (VMI, ADL, Fine Motor) also showed the largest gains. The true intervention effect within these domains is likely substantial, but the proportion attributable to treatment versus regression toward the mean cannot be fully separated without a control group. All scores remain pre-Rasch ordinal percentage values. Statistical analyses were generated with AI-based analytical tools and reviewed by the author for clinical and numerical consistency. Final responsibility for interpretation rests with the author.An additional and important limitation specific to the VMI domain is the potential confounding effect of preschool attendance and literacy instruction.
A significant body of research demonstrates that access to quality preschool accelerates cognitive and academic skill development, with effects that are particularly pronounced for children from low-income households (Magnuson & Duncan, 2016; Bailey et al., 2024). Preschool curricula in the four-year-old age range routinely incorporate letter recognition, alphabet knowledge, pre-writing activities, and structured fine motor practice, all of which directly engage the visual-motor and orthographic processing skills that O.T. Wizard’s VMI domain measures. Because O.T. Wizard does not currently collect data on whether a child is enrolled in preschool, how many days per week they attend, or what literacy instruction they are receiving, it is not possible to separate the contribution of preschool-based literacy exposure from the contribution of OT services to the VMI gains observed. The large VMI gains reported here almost certainly reflect the combined influence of OT intervention, natural maturation, and classroom-based literacy and pre-writing instruction. Future data collection that captures school enrollment status and attendance patterns would allow this important confounder to be examined directly.
Band G (n=6) and Band J (n=3) findings should be treated as exploratory only. The Participation and Executive Functioning longitudinal findings are subject to the novelty effect interpretation described above, which cannot be confirmed or ruled out without the prospective study design described in the continuing research section.
Continuing Research Needed
Rasch calibration remains the highest research priority for O.T. Wizard. Transforming ordinal raw scores into interval-level person measures will allow true scale-independent longitudinal comparison, validate the proportional scaling assumption underlying the natural growth correction, and identify items requiring revision. Current analyses are pre-Rasch and should be interpreted as preliminary clinical evidence rather than psychometrically standardized measurement. The O.T. Wizard National Try-Out Team initiative is designed to expand sample sizes needed for stable item calibration across all domains and age bands.
A typically developing comparison group would allow the 9.2% natural growth estimate to be validated empirically and replaced with domain-specific growth expectations calibrated against external developmental benchmarks. Recruiting a non-clinical sample, even a modest one, would strengthen the interpretive framework considerably and provide a more precise foundation for the gain-above-expected metric.
The novelty effect hypothesis for Participation and Executive Functioning requires prospective investigation. A study design capturing therapist-rated engagement and work habits at multiple time points within the first year of services, alongside parent-reported and teacher-reported measures, would allow empirical testing of whether first-evaluation ratings systematically overestimate authentic baseline functioning.
Longitudinal expansion with test-retest intervals of 12 to 24 months would allow examination of whether early VMI and ADL gains are sustained through kindergarten entry, and whether children who make the largest gains above natural growth in the preschool period show measurably better school readiness outcomes. Linking O.T. Wizard composite and domain scores to standardized criterion measures, including teacher-rated school readiness and kindergarten entry assessments, would establish predictive validity and position the platform’s data within the broader early childhood outcomes literature.
Conclusion
This analysis of 71 preschool-aged children with paired O.T. Wizard evaluations, examined through a transparent natural growth framework, extends the domain-level outcome picture established in the February 2026 report. Across a mean interval of 4.85 months and a natural growth expectation of 9.2%, five of eight assessed domains showed gains that were statistically significant, clinically large in effect, and substantially above what maturation alone would produce.
89% of children improved overall. 20 children crossed the 500-point composite functional threshold during the study interval. Visual Motor Integration showed gains of 21.1 points above the natural expectation, 614% above what developmental maturation alone would predict over the same period. Activities of Daily Living and Fine Motor Skills showed gains of 468% and 268% above expected, respectively. These are not marginal differences. They represent functional gains at rates that developmental maturation cannot explain.
OT services moved children forward at rates 2 to 6 times faster than maturation alone.
These findings are preliminary. They require replication with larger samples, validated comparison conditions, and Rasch-calibrated measurement. What they establish is that structured domain-level digital assessment in pediatric OT can generate longitudinal outcome data that clearly and transparently distinguishes intervention-driven change from natural maturation. For a profession that has historically struggled to quantify its impact in terms that payers and educational systems recognize, that distinction is not a minor technical refinement. It is the foundation of evidence-based practice.
A six-part blog series examining what outcome data reveals about pediatric OT, what documentation systems currently miss, and how structured Response to Intervention measurement changes clinical practice is forthcoming in the O.T. Wizard Research Series beginning the week of March 9, 2026.
Disclosures
The author is the Founder and Clinical Architect of O.T. Wizard and has a financial interest in the platform. All analyses were conducted on de-identified clinical data collected in routine practice. Statistical analyses were generated with AI-based analytical tools and reviewed by the author for clinical accuracy and numerical consistency. Final responsibility for interpretation and reporting rests with the author. Data collection is ongoing. All data is de-identified in accordance with HIPAA regulations.
References
Daly, C. J., Kelley, G. T., & Krauss, A. (2003). Relationship between visual-motor integration and handwriting skills of children in kindergarten: A modified replication study. American Journal of Occupational Therapy, 57(4), 459-462. https://doi.org/10.5014/ajot.57.4.459
Duff, S. V., Chow, S. M., & Henderson, S. E. (2015). Developmental coordination disorder and its consequences for children. In A. F. Farrow & P. J. Tremblay (Eds.), Pediatric rehabilitation: Principles and practice (5th ed., pp. 189-215). Demos Medical Publishing.
Zwicker, J. G., Missiuna, C., Harris, S. R., & Boyd, L. A. (2012). Developmental coordination disorder: A review and update. European Journal of Paediatric Neurology, 16(6), 573-581. https://doi.org/10.1016/j.ejpn.2012.05.003Bailey, D. H., Duncan, G. J., Cunha, F., Foorman, B. R., & Yeager, D. S. (2024). Persistence and fadeout of educational-intervention effects: Mechanisms and potential solutions. Psychological Science in the Public Interest, 21(2), 55-116.Gerde, H. K., Zhao, Y., Shu, L., & Gagne, J. R. (2025). Evidence-based instructional support for early writing in preschool and kindergarten: A scoping review. Reading and Writing. https://doi.org/10.1007/s11145-025-10751-8Magnuson, K., & Duncan, G. J. (2016). Can early childhood interventions decrease inequality of economic opportunity? RSF: The Russell Sage Foundation Journal of the Social Sciences, 2(2), 123-141.
About O.T. Wizard
O.T. Wizard is a clinical intelligence system for pediatric occupational therapy professionals. The platform evaluates performance in evaluations, forms, and daily treatment notes across twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, executive functioning, activities of daily living, and participation. O.T. Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update continuously as the clinical database expands. For information about O.T. Wizard research or accessing the platform, visit otwizard.com.
A descriptive snapshot from structured pediatric OT evaluation data (pre-Rasch)
Occupational therapists are trained to interpret evaluations one child at a time. What we rarely get to see is how our observations look in aggregate, across hundreds of evaluations completed using the same structure.
This article summarizes descriptive patterns from the PHI free data of OT Wizard’s preschool OT evaluation data collected since September 2025 using a consistent, structured evaluation framework. The goal is not to draw diagnostic or psychometric conclusions, but to surface real patterns that emerge when data is viewed collectively.
Sample Overview
Number of evaluations: 404
Approximate number of students: 404
Gender: Male: 56%, Female: 44%
Student Age range (two bands): 4:0-4:11.99
Primary Language Spoken: English: 90%, Spanish: 9%, Other: 1%
Primary Diagnosis: F82(Specific Dev Disorder of Motor Function): 95%, Autism: 5%
Sample is from preschoolers referred to OT evaluation; not a normative population sample
Total item-level data points: 19,062 normalized responses
Therapist Experience in this dataset
Average therapist experience: 10.3 years
Range of therapist experience: 3 to 30 years
Composite Scores (raw, non-Rasch)
Sample average composite score:559.7 /1000 (or 55.9%: Emerging Band)
Observed range:137 – 873 /1000 (or 13.7%-87.3%)
This wide range reflects substantial variability within a relatively narrow age span, reinforcing the importance of looking beyond single summary numbers.
Methods (brief and explicit)
Item-level results are based on normalized scores (0–1) , averaged across evaluations and excluding missing values. Domain and subdomain scores reflect the mean of available scores only; missing data were not treated as zero. No Rasch or item-response modeling has been applied at this stage.
Domain Scores (out of 100 points)
Ranked highest → lowest
Participation, 71
Executive Functioning, 67
Visual Perception, 63
Gross Motor Skills, 63
Praxis, 55
Fine Motor Skills. 53
Activities of Daily Living (ADL), 50
Visual Motor Integration, 42
Even when age bands are combined, the ordering remains stable. Participation and executive functioning consistently rank highest, while visual motor integration anchors the lower end of the distribution. However, even the highest score (Participation) is at the low range of Proficient.
Subdomain Scores (out of 100 points)
Ranked highest → lowest
Visual Discrimination, 74
Visual Figure–Ground, 71
Participation, 71
Work Habits (1:1 Therapy Setting), 67
Trunk Stability, 62
Hand Use, 62
Visual Memory, 56
Bilateral Integration, 61
Visual Spatial Relations, 55
Sequencing Praxis, 55
Dressing, 50
Scissor Use, 47
Pre Handwriting ,43
Speed and Accuracy, 43
Complex Visual Motor Representation & Integration, 38
Looking at the full ordering — not just the top or bottom — helps clarify where variation clusters across perceptual, motor, and participation-based constructs.
Item-Level Results (out of 100 points)
Normalized scores, ranked highest → lowest
Highest-ranking items
Match the shape (VP_VD_35), 100
Find the same shape (VP_VD_40), 100
What color is each bear? (VP_VD_47), 91
Dons slip-on shoes (ADL_DRESS_36), 100
Balance on one foot for 5 seconds (GM_TS_40), 86
Snip paper in at least one spot (FM_SCIS_6), 82
Dons simple front-opening coat (ADL_DRESS_33), 87
Transition performance during evaluation (EF_WORKHAB_5), 81
Several of these items show average normalized scores approaching or reaching 1.0.
Lowest-ranking items
Sequencing praxis: 4 steps with symbolic action (PR_SP_72)
Visual Memory (multiple VP_VM items)
Hop on one foot repeatedly (GM_TS_44)
Skip for multiple cycles (GM_TS_60, GM_TS_72)
Draw-A-Person task (VM_DAP)
Letters written legibly (VM_WANDPK4)
Pencil bubble popping (FMSAT) – non-dominant hand (POP_BUBB_11)
Latch a separated zipper (ADL_DRESS_54)
These items consistently fall at the lower end of the distribution across the combined preschool sample.
A necessary note about very high average scores
Items with average normalized scores near 1.0 may reflect several possibilities:
Developmentally easy tasks for much of the sample
Ceiling effects
Limited discrimination at this age range
Potential item misfit
Which explanation applies cannot be determined without psychometric modeling. These patterns are signals to investigate, not conclusions — and they help inform future Rasch analysis decisions.
Why this snapshot matters
None of these findings are surprising in isolation. What is meaningful is seeing how the entire evaluation system orders itself when applied consistently across hundreds of children.
This kind of ranking:
Makes implicit patterns visible
Highlights where variability concentrates
Provides a grounded baseline for future validation work
Most importantly, it reflects how structured OT evaluation data actually behaves — not how we assume it does when cases are viewed one at a time.
As this dataset grows and Rasch analysis is completed, these descriptive patterns will be tested, refined, and in some cases challenged. For now, they offer a clear, honest snapshot of the current data — and a strong foundation for what comes next.
About O.T. Wizard
Data for this analysis was collected through OT Wizard, a clinical intelligence system for pediatric occupational therapy assessment. The platform evaluates performance across up to twelve domains including visual-motor integration, fine motor skills, gross motor skills, praxis, visual perception, visual motor integration, executive functioning, activities of daily living, and participation. OT Wizard is undergoing Rasch analysis validation to establish psychometrically sound, norm-referenced scoring with living norms that update continuously as the clinical database expands.
Unlike traditional checklist-based assessments, OT Wizard converts all observations to continuous metrics that enable progress tracking, cross-domain comparison, and comprehensive reporting. The platform captures all six factors identified in this research as predictive of handwriting success: fine motor skills, visual perception (with subdomain specificity), praxis, cooperation, attention, and task participation. Behavioral regulation is assessed within the context of actual task performance rather than as an isolated rating, providing clinically relevant data about how attention and cooperation affect functional skill demonstration.
For handwriting readiness assessment specifically, OT Wizard provides quantified performance across visual discrimination, visual-motor integration, fine motor control, motor planning, and behavioral engagement during writing tasks. This comprehensive approach addresses the multifactorial nature of handwriting development identified in this research. As the platform undergoes Rasch analysis validation and accumulates longitudinal outcome data, it will establish whether comprehensive baseline assessment across all six predictors improves identification of children at risk for handwriting difficulty and informs more effective intervention planning.
OT Wizard is committed to advancing the occupational therapy profession by collecting de-identified clinical data from real therapist users, building the largest developmental database in pediatric occupational therapy history. This continuous data collection enables research on developmental trends, intervention effectiveness, and response to intervention patterns that elevate practice from perception-based to data-driven decision making and strengthen the evidence base for the entire profession
For OT professionals interested in data-driven assessment tools, visit otwizard.com to learn more about evidence-based pediatric evaluation.
You spent 90 minutes conducting a thorough pediatric occupational therapy evaluation. Another hour and a half writing a detailed report. You submitted it to insurance with confidence. Then the denial letter arrives: “Medical necessity not established.”
Sound familiar? You’re not alone. Insurance denials for occupational therapy evaluations are frustrating, time-consuming, and costly. But here’s the good news: most denials happen because of how the report is written, not whether the child actually needs services.
Let’s fix that.
The Insurance Approval Formula: Medical Necessity + Functional Impact + Skilled Service
Insurance companies don’t deny services because they don’t believe children need help. They deny because the documentation doesn’t prove three critical elements:
1. Medical Necessity: A documented diagnosis or condition that requires intervention 2. Functional Impact: Clear evidence that the condition limits daily functioning 3. Skilled Service: Proof that an occupational therapist’s expertise is required (not just supervision or general instruction)
Your evaluation report must explicitly address all three. If even one is missing or unclear, expect a denial.
What Insurance Reviewers Actually Read (And What They Skip)
Here’s a secret: the person reviewing your report spends about 90 seconds on it. They’re not reading every word. They’re scanning for specific elements. Most likely they are being read by their AI Bot.
The takeaway: Front-load your report with the information they need. Don’t bury medical necessity in paragraph seven.
The 7 Elements Every Insurance-Approved Report Contains
1. Clear Diagnosis at the Top
Wrong: “Johnny is a 5-year-old male referred for fine motor concerns.”
Right: “Johnny is a 5-year-old male with a diagnosis of Developmental Coordination Disorder (ICD-10: F82) referred for occupational therapy evaluation secondary to significant fine motor and visual-motor integration deficits impacting activities of daily living.”
Notice the difference? The second version includes diagnosis code, specific deficit areas, and functional impact in the first sentence.
2. Functional Limitations Stated Explicitly
Insurance doesn’t care that a child scores in the 5th percentile on the Beery VMI. They care that this score means the child cannot complete age appropriate functional activities, such as independently buttoning their shirt, writing their name legibly, or using utensils safely.
For every test score, include the “so what” statement:
Test Result: Visual perception skills measured at 2 standard deviations below age expectations on TVPS-4 (standard score: 70).
Functional Impact: This significant deficit prevents Johnny from independently locating items in his backpack, finding his desk in the classroom, and distinguishing similar letters (b/d, p/q) during early literacy tasks. Parent reports Johnny requires maximum assistance with dressing due to inability to orient clothing correctly.
3. Objective Measurements and Standardized Scores
Clinical observations alone don’t prove medical necessity. You need numbers.
Include :
Standardized test scores with percentiles or standard scores (like Peabody, Beery VMI) or Rasch-calibrated /Criterion referenced assessments (like PEDI-CAT, OT Wizard, or HELP)
Timed performance measures (e.g., “completed pegboard task in 145 seconds; age expectation is 45 seconds”)
Quantifiable observations (e.g., “grasped pencil in fisted grasp 100% of observed writing attempts”)
Measurable functional deficits (e.g., “required 4 verbal cues and 2 physical assists to don shirt”)
Research shows: Reports with comprehensive domain coverage across 8 areas (ADL, Executive Functioning, Fine Motor, Gross Motor, Visual Perception, Visual Motor Integration, Praxis, and Participation) have significantly higher approval rates because they provide objective evidence across multiple functional areas.
4. Medical Necessity Language (Not Educational Language)
If you primarily treat in schools, but are a medical based provider (meaning you aren’t an IEP provider), this is critical. Insurance reviewers don’t understand educational terminology.
Educational Language (Don’t Use): “Johnny requires OT services to access his educational curriculum and participate in classroom activities per his IEP.”
Medical Necessity Language (Use This): “Johnny requires skilled occupational therapy intervention to develop functional grasp patterns, visual-motor integration skills, and bilateral coordination necessary for age-appropriate self-care tasks including dressing, feeding, and personal hygiene.”
Key Differences:
Educational
Medical
Student
Patient
Classroom participation
Functional independence
IEP goals
Treatment goals
Educational benefit
Medical necessity
School activities
Activities of daily living
5. Safety Concerns (When Present)
Safety issues fast-track approvals. If present, state them clearly.
Examples:
“Child demonstrates impulsive behavior and poor body awareness, resulting in 3 falls from playground equipment in past month per parent report. Requires skilled OT intervention to develop safety awareness and motor planning.”
“Significant oral-motor deficits result in choking incidents during meals 2-3 times per week. Skilled feeding therapy required to establish safe swallowing patterns.”
“Decreased proximal stability and postural control result in frequent loss of balance during mobility, with 2 documented injuries requiring medical attention in past 6 months.”
6. Why Skilled OT is Required (Not Just Caregiver Training)
Insurance will deny if they think a parent or teacher could provide the same intervention. You must prove why your clinical expertise is necessary.
Not Skilled: “Child will benefit from practice with buttoning and zipping.”
Skilled Service: “Child requires skilled occupational therapy to analyze specific motor planning deficits preventing successful fastener manipulation, develop individualized strategies to compensate for bilateral coordination limitations, and systematically grade activity complexity while addressing underlying sensory processing difficulties that interfere with tactile discrimination necessary for fastener manipulation.”
See the difference? The second version demonstrates clinical reasoning, assessment expertise, and therapeutic skill that cannot be provided by non-therapists.
7. Concrete Frequency and Duration Recommendations
Vague recommendations get denied. Be specific.
Too Vague: “Recommend outpatient OT services.”
Specific and Justified: “Patient requires skilled occupational therapy 2x/week for 8 weeks (16 sessions) to address bilateral coordination deficits, visual-motor integration delays, and ADL skill development. Frequency based on severity of deficits (2+ standard deviations below age expectations across 4 domains) and need for motor learning repetition to establish new movement patterns. Re-evaluation recommended after 8-week intervention period to assess progress and determine ongoing needs.”
Common Denial Reasons and How to Avoid Them
Denial Reason #1: “Diagnosis not covered”
Prevention: Check the insurance company’s covered diagnosis list before evaluating. If the primary diagnosis isn’t covered, lead with a secondary diagnosis that is covered but still supports the need for OT.
Example: Autism (F84.0) might not be covered for outpatient OT, but Developmental Coordination Disorder (F82) or Sensory Processing Disorder coded as Other Specified Developmental Disorders (F88) often are.
Denial Reason #2: “Educational, not medical”
Prevention: Even if you’re a school-based therapist, emphasize ADL and home function impacts, not just classroom performance.
Include:
Dressing difficulties
Feeding/utensil use challenges
Hygiene and self-care limitations
Safety concerns at home
Community participation barriers
Denial Reason #3: “Not medically necessary”
Prevention: State explicitly in your report: “Skilled occupational therapy is medically necessary to address [diagnosis] which significantly impacts patient’s ability to [specific functional tasks], resulting in dependence on caregivers for age-appropriate self-care and safety concerns during daily activities.”
Denial Reason #4: “Insufficient objective data”
Prevention: Use standardized assessments. Clinical observations alone aren’t enough. Data from 404 evaluations shows that assessments with zero missing data and comprehensive domain coverage provide the objective evidence insurance requires.
The Report Structure Insurance Prefers
Section 1: Demographics and Diagnosis (Top of Page)
Name, DOB, date of evaluation
Primary diagnosis with ICD-10 code
Referring physician
Section 2: Medical Necessity Statement (First Paragraph) One clear paragraph stating diagnosis, functional limitations, and why skilled OT is required.
Section 3: Assessment Results
Standardized test scores
Functional performance observations
Quantifiable data
Each with functional impact statement
Section 4: Clinical Impressions
Summary of findings
How deficits impact daily function
Safety concerns (if applicable)
Section 5: Recommendations
Specific frequency (2x/week)
Specific duration (8 weeks)
Justification for both
Explicit medical necessity statement
Keep it concise: 2-3 pages maximum. Remember, they spend 90 seconds reading it.
Real Example: Before and After
Before (Gets Denied):
“Johnny is a pleasant 5-year-old boy who was referred for OT evaluation. He has difficulty with handwriting and gets frustrated during fine motor tasks at school. During testing, Johnny had trouble copying shapes and his pencil grasp looked immature. He would benefit from OT to work on these skills. Recommend weekly OT.”
Problems: No diagnosis code, no standardized scores, educational focus, vague recommendations, no medical necessity statement.
After (Gets Approved):
“Johnny is a 5-year-old male with Developmental Coordination Disorder (F82) referred for occupational therapy evaluation secondary to significant visual-motor and fine motor deficits impacting activities of daily living and self-care independence.
O.T. Wizard: Composite 550/1000, ADL 50/100, Fine Motor 72/100, Gross Motor 42/100, Sequencing Praxis 27/100, Visual Motor Integration 72/100
Functional grasp assessment: Fisted grasp pattern 90% of observed attempts
Functional Impact: Visual-motor integration and fine motor deficits prevent Johnny from independently managing fasteners (buttons, zippers, snaps), requiring maximum assistance for dressing. Unable to use utensils safely, resulting in frequent spills and parent reports of choking incidents 1-2x weekly. Cannot complete age-appropriate self-care tasks including tooth brushing and hair combing without hand-over-hand assistance.
Medical Necessity: Johnny requires skilled occupational therapy to develop functional grasp patterns, bilateral coordination, sequencing praxis, gross motor, and visual-motor integration skills necessary for age-appropriate self-care independence. Deficits 2 standard deviations below age expectations indicate significant impairment requiring therapeutic intervention. Safety concerns related to feeding and frequent falls during mobility necessitate skilled assessment and intervention.
Recommendations: Skilled occupational therapy 2x/week for 12 weeks to address bilateral coordination, visual-motor integration, and ADL skill development. Frequency based on severity of deficits and need for repetition to establish motor learning. Re-evaluation after 12 weeks to assess progress.”
Why it works: Diagnosis code in first sentence, standardized scores with functional impact, medical necessity explicitly stated, safety concerns noted, specific recommendations with justification.
Special Considerations for Different Settings
School-Based Therapists Seeking Medical Insurance Coverage
You can write reports that work for both IEP teams and insurance, but you need two versions:
IEP Version: Focus on educational impact and access to curriculum Insurance Version: Same data, different framing focused on ADL and medical necessity
Pro Tip: Complete your evaluation once, but generate two reports with different emphasis. Your assessment data doesn’t change, just how you present it.
Outpatient Clinic Therapists
You have an advantage because you’re already documenting medical necessity. Just ensure you’re:
Using covered diagnosis codes
Quantifying functional limitations
Stating skilled service needs explicitly
Providing specific frequency/duration with rationale
Early Intervention Providers
Insurance approval for 0-3 age range requires extra emphasis on:
Developmental delay severity (how far behind age expectations)
Impact on parent-child interaction
Safety concerns
Risk of further delay without intervention
The Bottom Line
Insurance approval isn’t about luck. It’s about documentation. Every denied evaluation report is missing at least one of these elements:
✓ Diagnosis code in first paragraph ✓ Standardized assessment scores ✓ Functional impact statements for every deficit area ✓ Medical necessity language (not educational) ✓ Explicit statement of why skilled OT is required ✓ Specific frequency and duration with justification ✓ Safety concerns (when applicable)
Master these seven elements, and your approval rate will skyrocket.
Stop spending hours appealing denials. Write it right the first time.
Streamline Insurance-Compliant Documentation
Writing insurance-approved reports doesn’t have to take hours. OT Wizard generates comprehensive evaluation reports with all required elements automatically included: diagnosis codes, standardized scores across 8 domains, functional impact statements, and medical necessity /educational eligibility. Choose medical or educational report tone with one click. Stop rewriting reports for insurance appeals.
A data-based look at executive functioning during OT evaluations
Occupational therapy practitioners often hear some version of the same question: “They seemed to do fine with you. Why are they struggling so much in class?”
To explore this, we examined therapist-rated Executive Functioning collected during one-on-one OT evaluations with preschoolers. These ratings reflect how children worked with the therapist during the evaluation itself, not how they function in classrooms, group settings, or home environments.
This article focuses on what the data shows and how therapists can interpret these observations responsibly.
Important context about this sample
Children were referred for OT after failing an occupational therapy screening
400+ preschoolers aged 4:0-4:11 , with about equal % of males and females
This is a clinically referred preschool sample, not a normative group
Ratings reflect performance during a one-on-one OT evaluation
Work habits were rated using descriptive anchors:
Not present or Beginning
Inconsistent or Emerging
Developing
Proficient
Mastered
What the work habit data shows overall
Across all work habit areas, most children in this referred preschool sample were rated in the Developing to Adequate range when working one-on-one with an occupational therapist.
This suggests that many children who struggle in classrooms or group environments are able to participate meaningfully when demands are reduced, expectations are clear, and adult support is individualized.
How children performed across specific work habit areas
Transitions and cooperation
Most children were rated as Developing or Proficient in their ability to transition between activities and cooperate with the therapist during the evaluation.
This indicates that many preschoolers can adapt well to structured tasks when the environment is predictable and supportive, even if transitions are difficult in other settings.
Task initiation and task completion
Task initiation and task completion were commonly rated in the Developing to Proficient range.
Many children required prompting or encouragement to begin tasks, but were generally able to complete them once engaged. This reflects emerging executive functioning skills rather than refusal or lack of effort.
Impulse control and task persistence
Impulse control and task persistence tended to fall in the Developing range.
Children often attempted tasks and remained engaged for short periods but had difficulty sustaining effort across multiple activities. This pattern is typical in preschoolers and becomes more pronounced when task demands increase.
Attention during the evaluation
Attention was the most commonly impacted work habit, even in a one-on-one setting.
Many children were rated between Emerging and Developing for attention. This is an important clinical signal.
When a child struggles to attend during a one-on-one evaluation, they will generally struggle even more in larger, less structured settings such as classrooms, group activities, or busy home environments. Difficulty sustaining attention in a low-demand context often predicts greater challenges when environmental demands increase.
When poor evaluation performance reflects anxiety, not ability
It is also important to acknowledge that not all low work habit ratings reflect true functional ability.
Some preschoolers perform poorly during evaluations because they are:
Nervous
Anxious
Overwhelmed by unfamiliar adults or environments
Uncertain about expectations
In these cases, performance may underestimate a child’s true skills.
Occupational therapists play a critical role in distinguishing between skill limitations and emotional readiness. Establishing rapport is not optional. It is a clinical necessity.
For some children, establishing rapport may include:
Spending additional time building trust
Allowing the child to observe before participating
Using play or preferred activities to reduce anxiety
Delaying formal evaluation tasks until the child appears comfortable with the therapist
Evaluation data is only as meaningful as the context in which it is collected.
What this tells us about one-on-one OT evaluations in our 400+ preschool sample
Taken together, these findings suggest:
Most preschoolers in this referred sample can demonstrate Developing to Adequate work habits in a one-on-one OT evaluation
Stronger performance in this setting does not negate difficulties in classrooms or group environments
Attention difficulties observed in one-on-one contexts often signal more significant challenges in larger settings
Anxiety and lack of rapport can temporarily suppress performance and must be considered during interpretation
Why this matters for interpretation and communication
These data support a critical clinical message:
Performance in a one-on-one OT evaluation reflects what a child can do with individualized support, not what they are expected to manage independently in more complex environments.
When communicating results to families, teachers, and teams, it is essential to clarify the difference between supported performance and real-world participation demands.
Key takeaway for therapists
Based on therapist-rated work habit descriptors:
Most preschoolers in this referred sample demonstrated Developing to Adequate work habits during a one-on-one OT evaluation, while attention and self-regulation remained vulnerable areas, particularly when demands increase or anxiety is present.
This reinforces the importance of careful interpretation, thoughtful rapport building, and contextualized clinical judgment.
📣OTPs : Does this match your experience when evaluating preschoolers? Leave us a comment and let us know.
What the Draw-A-Person Task in O.T. Wizard Is Actually Showing Us
A data-based look at the Draw A Person Task in preschool OT evaluations inside O.T. Wizard.
Important context: The data presented here were drawn from preschool children who did not pass an occupational therapy screening and were subsequently evaluated. This sample is not a normative population and should not be interpreted as representative of typically developing preschoolers.
The Draw-A-Person task is a familiar tools in pediatric occupational therapy. It is widely used, information rich, and often referenced in evaluation reports. At the same time, many therapists find it challenging to interpret, especially when scores are low.
Rather than debating the value of Draw-A-Person conceptually, this article looks at how the task behaves in real evaluation data when it is administered consistently across a preschool sample. You may have heard of “The Good Enough Draw A Person” drawing assessment. The O.T. Wizard uses a similar but more simple version.
This analysis is descriptive only. No Rasch or item response modeling has been applied yet.
Sample overview
Number of evaluations included: 404
Approximate number of students: 404
Evaluations with Draw-A-Person data: 401
Total item-level data points across the evaluation system: 19,062
Referred clinical sample, not normative population
Draw-A-Person was part of the standard evaluation battery and was administered when the child tolerated the task. Missing responses were excluded rather than scored as zero.
Draw-A-Person scoring rubric
Draw-A-Person was scored using the following criteria:
Score 0 (0.0): No approximations
Score 1 (0.2): Approximations emerge
Score 2 (0.4): Head and parts present but no body
Score 3 (0.6): Recognizable person with body and at least four body parts
Score 4 (0.8): Recognizable person with six or more body parts
Score 5 (1.0): Recognizable person with twelve or more body parts
It is important to note that even a score of 1 reflects emerging representational drawing rather than an absence of skill. The score was converted to a normalized score where a raw score of 5 = normalized score of 1.0 (most credit).
Average Draw-A-Person performance
Across the combined preschool sample:
Draw-A-Person scores were analyzed using normalized values derived from the scoring rubric. The average normalized score across the preschool sample was 0.38, corresponding to an average rubric level of approximately 1.9. Clinically, this places the average child between “approximations emerge” and “head and parts present but no body.”
While individual scores varied, this average suggests that most preschoolers in the sample demonstrated emerging representational drawing skills rather than fully recognizable figures
Distribution of Draw-A-Person scores
When responses are mapped directly onto the scoring rubric, the distribution looks like this:
Score 0, no approximations: 290 children, approximately 72 percent
Score 1, approximations emerge: 99 children, approximately 25 percent
Score 2, head and parts without a body: 7 children, approximately 2 percent
Score 3, recognizable person with body and at least four parts: 4 children, approximately 1 percent
Score 4, recognizable person with six or more parts: 1 child, less than 1 percent
Score 5, recognizable person with twelve or more parts: 0 children
This is a heavily floor-weighted distribution.
What this distribution tells us
In this preschool sample, most children did not produce a recognizable person. Nearly three quarters of children showed no recognizable approximations, and only a very small percentage produced a clearly recognizable person with a body.
Draw-A-Person age anchors suggest that by approximately 36 months, children often demonstrate a head with emerging parts, by 48 months a recognizable person with a body and multiple body parts, and by 64 months increasingly detailed human figures. In contrast, the average performance in this referred preschool sample falls between “approximations emerge” and “head and parts present but no body.” This indicates that, as a sample group, these children are demonstrating representational drawing skills that are less mature than would be expected based on age anchors, which is consistent with a population that did not pass an occupational therapy screening rather than a normative sample.
This helps explain why Draw-A-Person often feels like a difficult task for young children and why it frequently stands out in evaluation reports.
Why Draw-A-Person behaves differently than many other tasks
Compared to many visual perceptual, gross motor, or participation-based tasks, Draw-A-Person requires multiple skills to work together at once. These include visual motor integration, visual perception, motor planning, body awareness, fine motor control, and representational thinking.
Because of this high level of integration, Draw-A-Person tends to function as a high-threshold task. It separates children who are beginning to integrate these skills from those who are not yet developmentally ready to do so.
The data supports what many therapists experience clinically. Draw-A-Person is informative, but it should not be interpreted in isolation.
Interpreting Draw-A-Person results responsibly
Based on this data, several points are important for clinical interpretation:
Low Draw-A-Person scores are common in preschoolers
Progress from score 0 to score 1 is clinically meaningful
Scores of 3 or higher represent a small minority of children
Draw-A-Person should be interpreted alongside visual perceptual, fine motor, praxis, and participation data
Whether the task is developmentally ambitious, exhibits a floor effect, or is a candidate for item misfit will be evaluated during future Rasch analysis. At this stage, the data supports careful, contextual interpretation rather than over-weighting the score.
Why this matters for practice
Most therapists have an intuitive sense that Draw-A-Person is hard for young children. Very few have seen how strongly that intuition is reflected in actual data.
Seeing the full distribution helps recalibrate expectations and supports clearer communication with parents, teachers, and teams. It also reinforces the importance of viewing Draw-A-Person as one piece of a much larger evaluation picture.
As this dataset grows and formal psychometric analysis is completed, these descriptive patterns will be tested and refined. For now, they provide a grounded, data-anchored explanation for why Draw-A-Person feels the way it does in real preschool OT evaluations.
I’m thrilled to announce that OT Wizard has officially begun psychometric validation through Rasch analysis – a major milestone in our journey to elevate evidence-based practice in pediatric occupational therapy!
What’s Happening Now
We’ve sent evaluation data for ages 4-5 years (Age Bands H & I) to an independent psychometrician for comprehensive analysis. With over 400 evaluations from 18 therapists across North Carolina, we have robust data to validate that OT Wizard measures what we say it measures & accurately, reliably, and fairly.
What Is Rasch Analysis?
Rasch analysis is a sophisticated psychometric approach that goes beyond traditional test validation. Unlike norm-referenced assessments that simply compare students to each other, Rasch analysis creates an interval-level measurement scale, similar to measuring temperature or weight.
Think of assessments you may know that use Rasch methodology:
HELP (Hawaii Early Learning Profile) – Rasch-validated developmental assessment
Original PEDI – Rasch-based functional assessment
These assessments are considered gold standards because Rasch analysis ensures:
Equal intervals: A 10-point gain at any level represents the same amount of growth
Sample-independent measurement: Item difficulty doesn’t depend on who takes the test
Missing data handling: Scores are valid even when not all items are administered (adaptive testing)
Precise error estimation: Know exactly how confident you can be in each score
Item hierarchy validation: Confirms items are developmentally sequenced correctly
Why Rasch Instead of Traditional Norming?
Traditional norm-referenced tests (like BOT-3, PDMS-3) require testing typically-developing children to create percentile ranks. That’s valuable, but has limitations:
Norms become outdated (tests re-normed every 15-20 years) Percentiles are ordinal, not interval (85th→95th ≠ 15th→25th in actual ability) Can’t track growth accurately across different ability levels Require complete test administration
Rasch analysis provides:
Continuous measurement scale – Track growth precisely over time Adaptive testing – Administer only relevant items, still get accurate scores Sample-independent – Item difficulty stays stable regardless of who’s tested Living calibration – Can update and refine continuously with new data Clinical utility – Scores directly interpretable for intervention planning
We’re building OT Wizard to work like the PEDI-CAT and AMPS. These are tools that OTs trust because they’re built on rigorous Rasch foundations.
What We Already Know About Our Data
Before even sending data to our psychometrician, we conducted preliminary analysis on our 404 evaluations to ensure data quality. Here’s what we’ve learned:
Strong Sample Characteristics
Well-balanced age distribution: 184 evaluations (ages 4.0-4.4) and 220 evaluations (ages 4.5-4.9)
18 therapists contributing: Average of 22 evaluations each, with range of 7-45 per therapist (good for inter-rater reliability)
Diverse language backgrounds: 90% English, 8% Spanish, 2% other languages
Clinical population validity: 96% recommended for OT services
Excellent Data Completeness
Zero missing responses – every administered item was answered
75.5% average completion rate – our adaptive basal/ceiling rules are working perfectly
19,062 total data points across 74 items and 8 domains
Strong Domain Coverage
Visual Perception: 15 items
Activities of Daily Living: 14 items
Gross Motor & Fine Motor: 10 items each
Participation: 9 items
Executive Functioning: 8 items
Visual Motor Integration: 5 items
Praxis: 3 items
Areas for Improvement Identified
Our preliminary analysis flagged several items for the psychometrician to examine closely:
Ceiling Effects in Visual Perception (Preschool evaluation): 56% of responses scored at ceiling (mastered), with 33% at floor (not yet observed). This bimodal pattern suggests we may need more mid-difficulty items to better differentiate students in the middle range.
Rating Scale Consistency: A small number of responses (3.3%) showed raw scores instead of normalized scores, indicating a formula issue we’ve already corrected.
Developmental Anchor Gaps: About 54% of items (primarily Executive Functioning and Participation domains) lack developmental anchors. The Rasch analysis will empirically determine difficulty levels so we can assign appropriate anchors.
Item-Age Alignment: Many items administered are anchored above student age ranges (51-55% above range). This is actually expected and appropriate. Students with developmental delays are working on skills typically seen at older ages. However, Rasch will help us recalibrate anchors based on clinical population performance vs. typical development.
Best-Performing Domains
ADL: Excellent distribution with only 8% floor and 8% ceiling – items are well-targeted
Executive Functioning: Minimal floor effect (0.7%), good spread across ability levels
This preliminary work means we’re sending clean, robust data to our psychometrician. This is maximizing the value of the Rasch analysis and ensuring reliable results.
What’s Being Analyzed
Our psychometrician is conducting comprehensive Rasch analysis across multiple dimensions:
1. Construct Validity (Unidimensionality)
Do items within each domain (Gross Motor, Fine Motor, Visual Perception, etc.) measure a single, coherent construct? This is critical for Rasch – if items don’t “hang together,” they can’t be on the same measurement scale.
2. Item Fit
Which items contribute to reliable measurement? Rasch provides specific fit statistics (infit/outfit MNSQ) showing whether each item:
Is too predictable (doesn’t add information)
Is too unpredictable (confuses the measurement)
Functions optimally (contributes to precise measurement)
Items outside acceptable ranges get flagged for revision or removal.
3. Rating Scale Functioning
Do our 5-point performance bands (Beginning → Mastered) function as intended? Rasch examines:
Are all categories used appropriately?
Do response thresholds advance in the right order?
Should categories be collapsed (e.g., 5-point → 3-point)?
This is similar to how AMPS validates its 4-point scoring scale.
4. Item Hierarchy
Rasch places all items on a single difficulty scale (measured in logits). We’ll see if:
Items anchored at 48 months are empirically easier than 54-month items
Our developmental sequencing matches actual difficulty
Gaps exist where we need additional items
This will be especially important for items currently lacking developmental anchors – the Rasch analysis will tell us where they belong.
5. Measurement Precision
Unlike traditional reliability (one number for whole test), Rasch shows precision at every ability level:
Where is measurement most accurate?
What’s the standard error at our 70% clinical threshold?
Can we distinguish between students with small ability differences?
6. Differential Item Functioning (DIF)
Do items work the same way for:
Boys vs girls?
4-year-olds vs 5-year-olds?
English vs Spanish speakers?
Different diagnoses?
Items showing bias get flagged or removed – ensuring fairness.
7. Person Separation
Can we reliably distinguish between students at different ability levels? Rasch provides a separation index showing how many distinct ability levels we can measure. Higher separation = more precise clinical distinctions.
8. Addressing Known Issues
The psychometrician will specifically examine:
Visual Perception’s ceiling effects : do we need additional mid-difficulty items?
Praxis domain with only 3 items : is this sufficient or should it combine with another domain?
Executive Functioning and Participation rating scales : do they function as separate constructs from performance-based items?
Why This Matters for You
Rasch validation transforms how you can use OT Wizard scores:
Meaningful Progress Monitoring
Because Rasch creates interval-level measurement, you can confidently say:
Traditional percentage scores can’t make these claims and a jump from 40% to 50% isn’t necessarily the same growth as 70% to 80%.
Adaptive Testing Validation
Like the PEDI-CAT and AMPS, OT Wizard uses basal/ceiling rules so students aren’t frustrated with too-hard items or bored with too-easy ones. Rasch analysis confirms:
Scores are comparable even when different items are administered
Our 75% completion rate is optimal
Missing items are appropriately “not administered,” not missing data
Credible Clinical Decisions
When you document that a child’s gross motor ability is at -1.2 logits:
School districts understand the methodology (same as PEDI-CAT/HELP)
You can defend your clinical reasoning with published psychometric evidence
Item-Level Interpretation
Rasch analysis creates item hierarchy maps showing exactly which skills a child has mastered, which are emerging, and which aren’t yet present. This directly informs intervention planning just like how AMPS users identify specific ADL breakdowns or PEDI-CAT shows functional skill patterns.
Why Start with Ages 4-5?
We strategically chose this age range because:
Sufficient sample size: 400+ evaluations provide robust statistical power for Rasch analysis
Diverse representation: Students with various diagnoses, languages, and ability levels
Multiple raters: 18 different therapists ensure inter-rater reliability analysis
Item overlap: Many items in this age range also appear in adjacent ages, so findings inform the entire platform
Strong data quality: Our preliminary analysis confirmed excellent completion rates and coverage
The psychometrician will identify any problematic items, validate our developmental anchors, assign anchors to items missing them, and ensure rating scales function optimally. We’ll implement improvements before these issues cascade into other age bands.
What Happens Next
Based on the psychometrician’s findings (expected in 4-5 weeks), we’ll:
Remove or revise misfitting items that don’t meet Rasch fit criteria
Add mid-difficulty items to Visual Perception domain to address ceiling effects
Optimize rating scales if analysis shows categories aren’t functioning as intended
Recalibrate existing anchors if clinical population performance differs from typical development
Establish measurement precision estimates at different ability levels
Publish validation statistics you can cite in reports and presentations
This refined version becomes the foundation for validating additional age bands.
Expanding Validation: Ages 3 Months to 12 Years
Over the next 12 months, as we reach 200+ evaluations per age band, we’ll validate each additional age group. This will create a comprehensive, linked measurement system similar to how PEDI-CAT links across age ranges where we can:
Track individual students across multiple years on the same logit scale
Provide age-equivalent scores based on item difficulty calibration
Create clinical reference data comparing students receiving OT services
Document growth trajectories with true interval-level measurement
Demonstrate outcomes with unprecedented precision
By Month 12, we’ll conduct a comprehensive linking study that places all age bands (3 months through 12 years) on a single, continuous measurement scale. Items that appear in multiple age bands will “anchor” the scales together, ensuring continuity.
This approach mirrors how major Rasch-based assessments (PEDI-CAT, AMPS, HELP) maintain measurement continuity across ages and versions.
How You Can Support This Work
1. Keep Using OT Wizard
Every evaluation you complete contributes to our growing database. Rasch analysis becomes more robust with larger samples and the more data we collect, the more confident we can be in item calibrations.
Brittany B., administers an eval with OT Wizard
2. Share Your Clinical Insights
If you notice items that seem:
Confusing or ambiguous to score
Too easy or too hard for the age range
Misaligned with what you observe clinically
Culturally biased or inappropriate
Please let us know! Your real-world feedback is invaluable. Rasch analysis will identify statistical misfits, but your clinical judgment helps us understand why items aren’t working.
What This Means for Our Profession
Most therapy documentation tools rely on subjective clinical observation. While tools like PEDI-CAT, AMPS, COPM, and HELP exist and use Rasch methodology, they’re limited in scope and focus on narrow or specific functional domains, requiring specialized training, or covering narrow age ranges.
OT Wizard is different: We’re creating a comprehensive, Rasch-validated clinical intelligence platform that covers:
Birth through early adulthood (3 months – 18 years)
Performance-based, participational based, and functional assessment
Integrated into everyday clinical workflow
By pursuing rigorous Rasch validation, we’re increasing psychometric rigor to comprehensive pediatric OT assessment.
When this validation is complete, you’ll be able to say:
“I use OT Wizard, a comprehensive Rasch-validated pediatric OT assessment platform with published psychometric evidence across 2,000+ evaluations providing the same measurement quality as tools like PEDI-CAT and AMPS, but covering all developmental domains.”
The Vision: Living Calibration and Real-Time Data
Here’s what excites me most about Rasch methodology: Unlike traditional norm-referenced tests that become frozen in time, Rasch-calibrated assessments can be continuously refined.
The PEDI-CAT has demonstrated this – as more data is collected, item calibrations can be updated, new items added, and measurement precision improved all while maintaining the same measurement scale.
OT Wizard will have “living calibration”:
Continuous item refinement as we collect more data
New items added to fill gaps in difficulty coverage
Real-time quality monitoring
Annual recalibration studies
Regional and demographic analyses
Imagine:
Item difficulties that reflect current populations
Outcome analytics showing which interventions are most effective
Predictive data identifying which early skills best predict later success
The world’s largest real-time Rasch-calibrated pediatric development database
Every evaluation you complete contributes to this unprecedented resource.
Thank you for building the future of evidence-based OT assessment with O.T. Wizard.
There are a brand new batch of 43 new O.T. Wizards that were trained at the Charlotte Pediatric OT Conference this past week! It was great to meet the OTA’s and OT’s that attended the training. In the first 1.25 hours , we discussed the struggles and pain points.
We had a 50/50 mix of OTA’s and OT/L’s with about 1/2 from clinics and the other 1/2 from either school systems or home health. I was able to show them how long this process has been in development.
I spent a good amount of time explaining and answering questions about our past (FUNdamental Foundations ) and current scoring philosophy, including Rasch analysis (and plans for the future!).
In the second part of the training, everyone created an account, added a “student” and then we created an evaluation together. I showed everyone how to administer some of the specialized assessments, including the WAND Preschool (pre-handwriting assessment), the Popping Bubbles (dexterity, speed and accuracy), and the scissor skills assessment.
After the evaluation , everyone created their instant report. I think they were pleasantly surprised! I received lots of notes about how excited they are to start using the OT Wizard.