Knowing an accommodation was delivered is only half the picture. The other half is whether it actually helped, and that question is harder, because it asks for a judgment rather than a checkmark. This guide is about answering it honestly: gathering the comparison without ever removing a support the student is entitled to, rating effectiveness in a way that survives a real week, and reading the pattern toward a clear decision.
The comparison question, handled ethically
To know whether an accommodation helps, you need some sense of the difference it makes. The temptation is to run a clean experiment: take the accommodation away, see what happens, compare. Do not do this.
There are two ethical ways to get the comparison. The first is to use naturally occurring occasions where the support simply was not available. A substitute did not know to read the directions aloud. The text-to-speech device was charging. The testing room was booked, so the test happened in the regular room. These gaps happen on their own. When they do, note what you observed, because they are an honest with-and-without contrast you did not create.
The second is to compare across task types where the accommodation applies differently. Extended time matters on a long test and barely registers on a five-question quiz. If a student finishes the quiz fine but struggles to complete the long test even with extra time, that contrast tells you something, and no support was removed to get it.
Both methods keep the accommodation in place at all times. You are reading the difference from situations that arise naturally, not manufacturing a deprivation.
A third honest source is the student's own report. An older student can often tell you whether a support helps, and on which tasks. Asking a ninth grader whether extended time made a difference on a particular test is not soft data; it is information you cannot get any other way, and it costs nothing and removes nothing. Pair what the student says with what you observe and you have a fuller picture than either alone.
One more caution about reading natural gaps. A single occasion without the support is one data point, and one bad day proves nothing. If the text-to-speech device was dead one morning and the student struggled, that is a hint, not a verdict. Wait until you have a few such occasions pointing the same way before you treat the contrast as real. The discipline here mirrors progress monitoring: a pattern of several points means something, a single point does not.
A rating scale that survives a real week
Effectiveness data dies when it is too heavy. A five-page observation protocol will not happen during a normal week. A 3-point rating with concrete anchors will, because it takes about ten seconds and you can do it on the occasions that already matter.
A composite sixth grader we will call Devon. Accommodation: directions read aloud on multi-step assignments. After an assignment where the directions were read aloud, the teacher marks one of three:
- 1, little or no effect: Devon still could not start the task without one-to-one reteaching of the directions.
- 2, partial effect: Devon started the task after the read-aloud but needed one clarifying prompt.
- 3, clear effect: Devon started and worked independently after the read-aloud, no further prompting on the directions.
The anchors are written so two different teachers would land on the same number for the same observation. That is what makes the rating worth recording. A vague "how did it go today" scale where 3 means whatever the rater felt is not data; it is mood. Concrete anchors turn a quick judgment into something the team can actually read.
The anchors do the work. Write them once, in plain behavioral terms, and the rating becomes fast and consistent. You are not grading the student. You are rating whether the support did its job on that occasion.
If more than one adult delivers the accommodation, share the anchors so the science teacher and the math teacher mean the same thing by a 2. A rating scale only produces comparable data when everyone reads the anchors the same way. Five minutes spent agreeing on what each number looks like is what keeps a column of ratings from becoming a mix of five different people's private scales.
How often to rate effectiveness
You do not rate effectiveness every time an accommodation is delivered. Delivery you mark every applicable occasion; effectiveness you sample. How often depends on the accommodation, because some apply many times a week and others only on a handful of occasions a term.
| Accommodation type | Realistic effectiveness sampling |
|---|---|
| Directions read aloud (happens most days) | Rate 2 to 3 occasions a week, not every time; a weekly sample shows the pattern |
| Extended time on tests (a few times a term) | Rate every qualifying test, since there are only 6 to 8 a term |
| Separate testing room (test days only) | Rate each test day it is used; the occasions are few |
| Preferential seating (a standing setup) | Rate once a month; the question is whether the arrangement still helps, not a daily check |
| Text-to-speech for reading (used most reading tasks) | Rate 1 to 2 occasions a week across different task types |
The principle is to rate often enough that one odd occasion does not define the picture, and rarely enough that you can actually keep it up. For a daily accommodation, a couple of ratings a week is plenty. For a rare one, rate each time, because each occasion is precious. A standing setup like seating needs only an occasional check that it is still pulling its weight.
A useful rule of thumb across a term: aim for somewhere between six and ten effectiveness ratings per accommodation before the review. That is enough that a pattern is visible and a single odd day cannot dominate, and few enough that the rating stays a quick habit rather than a second job. For an accommodation used most days, six to ten ratings is two a week for about a month and you can stop once the pattern is clear. For an accommodation used only on test days, six to ten ratings might be the whole term, which is why you rate every one. Match the sampling to how often the occasion actually arises, and let the target be a readable pattern rather than a fixed number of marks.
Reading the pattern toward a decision
A column of ratings is only useful if it points somewhere. Effectiveness data feeds one of three decisions at the review, and the pattern usually makes the decision obvious.
The first decision is keep as is. The ratings are mostly 2s and 3s, the natural with-and-without occasions show the student does worse when the support is absent, and the accommodation is clearly doing real work. You leave it in place and your record says why.
The second is adjust. The ratings are mixed or low, but the accommodation seems right in spirit. Maybe directions read aloud earns mostly 1s, and on the days you noted it, the real problem was that the directions were also too long, not just unread. The fix is to change the accommodation (read aloud and chunk the steps) rather than drop it. Low ratings are a signal to adjust the support, the same way a flat goal line is a signal to change the intervention, not the student.
The third is fade, with data behind it. The ratings have been consistently high and the natural without occasions show the student now does about as well when the support is absent. That can be evidence the student has built the skill the accommodation was scaffolding, and the team may decide to reduce it. Fading should be a team decision supported by a pattern in the data, never a quiet drop because someone forgot to provide it.
The throughline of all three decisions is the same: gather the contrast honestly, rate with anchors so the numbers mean something, and let the pattern, not a hunch, tell the team what to do. An accommodation kept because it works, adjusted because the data showed a gap, or faded because the student outgrew it are all defensible. An accommodation continued out of habit, or dropped out of forgetfulness, is not.