A behavior plan can fail for two completely different reasons, and they call for opposite fixes. Either the plan was never really delivered the way it was written, or it was delivered faithfully and still did not change the behavior. If you look only at the behavior data, you cannot tell which one you are facing. You will be tempted to rewrite a plan that was actually sound but never run, or to keep running a plan that was delivered perfectly and built on the wrong function. The way out is to check two things, not one. You do not need a research team to do it. You need a five-item check, a baseline you already have, and twenty minutes.
Fidelity and outcome are different questions
Fidelity asks: was the plan delivered as written? Outcome asks: did the behavior change? These are separate questions with separate answers, and conflating them is the most common review mistake.
Treatment integrity is the formal name for fidelity, and it matters because a plan is a set of instructions that only works if the instructions are followed. If the break request is supposed to be honored instantly every time but the aide honors it about half the time, the plan on paper was never actually tested. The behavior data from that week is not telling you the plan failed. It is telling you the plan was not run. No amount of staring at the outcome numbers will reveal that, because the numbers look identical whether the plan was poorly designed or simply not delivered.
So before you judge whether a plan is working, you establish whether it ran. Fidelity first, then outcome. A team that reverses the order keeps redesigning plans that would have worked if anyone had followed them.
It is worth being clear that low fidelity is rarely about an uncooperative adult. The usual causes are mundane and fixable: the plan had too many steps to remember, a key step was unclear, a substitute was never trained, the plan lived in a binder no one opened, or two adults were running it two different ways. Treating fidelity as a logistics problem rather than a willingness problem changes the whole tone of a review. You are not auditing people. You are finding the gap between what the plan asked for and what the day allowed, and closing it. That framing is also what makes staff willing to score the check honestly, which you need, because a fidelity check everyone fudges to look good is worse than no check at all.
A five-item weekly fidelity check
You do not need to observe all day. Pull the five steps that matter most from the plan and score them yes or no for a typical week. Five items keep the check honest and fast, and they should be the steps the plan depends on, not every step in it.
A composite fifth grader we will call Andre, whose plan teaches a break request for an escape function. Week of the review, scored by the classroom teacher in about two minutes:
- Writing tasks were broken into three-sentence chunks. Yes
- Andre was given a starter sentence on writing tasks. Yes
- Break requests were honored immediately every time. No (honored most times; the substitute on Wednesday did not know the plan)
- The worksheet stayed present after the problem behavior. Yes
- The break request was prompted before the behavior when Andre tensed up. Partly (prompted some days, missed on the busy ones)
Fidelity read: roughly three and a half of five steps delivered. The weak spot is the most important rule in the plan, honoring the break instantly, and it broke down with an untrained substitute.
That check took two minutes and it already tells the team something the behavior data alone could not: the one non-negotiable rule slipped, and it slipped for a specific, fixable reason. The fix is to train the substitute, not to rewrite the plan.
Tie outcome tracking back to the FBA baseline
Outcome tracking is only meaningful against a baseline, and you already have one. The FBA recorded how often the behavior happened before the plan started. Measure the same behavior the same way now, so the comparison is honest.
Resist the urge to invent a new measure at review time. If the FBA counted instances of tearing the paper per writing block, keep counting instances per writing block. A measure that changes between baseline and review cannot show progress, because you are no longer comparing like to like. Track the replacement behavior too: a plan is working when the problem behavior drops and the break requests rise, because the rise in the replacement is the evidence the student found the better option. If both are flat, that is data, not failure, and it feeds the decision below.
You also do not need elaborate measurement to do this well. A tally of the problem behavior per writing block and a tally of break requests, marked on the same slip the adult already carries, is enough to see a trend over a few weeks. The temptation is to build a richer data system than anyone will keep up, which collapses within a week and leaves you with no data at all. A measure simple enough to be recorded every day beats a sophisticated measure recorded twice and then abandoned. Match the measurement effort to the people who have to sustain it, not to the ideal study you would run with a research team you do not have.
A 20-minute review meeting
A review does not need an hour. With the fidelity check and the outcome data in hand, twenty minutes is enough to reach a decision and assign the next step.
- Restate the function and the plan's one non-negotiable, two minutes. Remind the team why the behavior happens and which single step the plan depends on most, so everyone judges against the same target.
- Read the fidelity check aloud, four minutes. Go through the five items. Name where delivery slipped and why, without blame. A slipped step is a logistics problem to solve, not a person to fault.
- Look at the outcome data against baseline, five minutes. Compare the problem behavior and the replacement behavior now versus the FBA baseline, using the same measure.
- Place the plan in the decision grid, four minutes. Cross fidelity (high or low) with progress (yes or no) and read off the action the cell calls for.
- Assign one change and an owner, three minutes. Change one thing, not five, so the next review can tell what moved the data. Name who does it and by when.
- Set the next review date, two minutes. Put it on a calendar before anyone leaves the room.
The discipline of changing one thing at a time is what makes the next review interpretable. A team that changes the plan, the staff, and the schedule all at once cannot learn anything from the result.
The decision grid
Crossing fidelity with progress gives four cells, and each cell calls for a different action. This is the single most useful tool in a review, because it stops the team from applying the wrong fix.
| Fidelity | Progress | What it means and what to do |
|---|---|---|
| High | Yes | The plan is working and being delivered. Keep going, and start planning how to fade supports as the replacement becomes automatic. |
| High | No | The plan was delivered faithfully and behavior still did not change. This points at the design. Revisit the function, because a plan built on the wrong function will not work however well it is run. |
| Low | Yes | Behavior is improving even though delivery was spotty. Something is helping. Tighten fidelity before you conclude the plan works, so you know what is actually driving the change. |
| Low | No | The plan was not really run, so it has not been tested yet. Do not redesign it. Fix delivery first (train staff, simplify steps, cover substitutes) and review again before changing the plan itself. |
The bottom-right cell is where the most plans are wrongly thrown out. A plan that was never delivered looks exactly like a plan that does not work, and a team that skips the fidelity check will rewrite a sound plan instead of fixing the delivery. The grid forces the question that prevents that mistake.
Keeping delivery and outcome in one place
The reason this is hard in practice is that fidelity lives in one place (a checklist in a folder) and outcome lives in another (a behavior chart somewhere else), so the team reassembles them by memory at the meeting. Logging the plan's delivery and the behavior data in one place, as you can in Evident, means the fidelity check and the outcome trend sit side by side at review time instead of being stitched together from two systems. However you store it, the goal is the same: never look at the behavior data without the fidelity data next to it, because each one is meaningless without the other.
A plan that is reviewed this way is hard to fool yourself about. You will know whether it ran, whether it worked, and which of the two to fix next, which is exactly the knowledge a behavior team needs and rarely has when it looks at outcome alone.
Common questions
What is the difference between fidelity and outcome?
Fidelity is whether the plan was delivered as written. Outcome is whether the behavior changed. A plan can be a sound design but never delivered, or delivered faithfully but built on the wrong function. Checking only outcome leaves you unable to tell which problem you have, so check both.
How do you check fidelity without observers all day?
Use a short weekly self-check or peer check of five concrete items drawn straight from the plan, scored yes or no in two minutes. You are not running a research study. You are getting an honest read on whether the key steps actually happened in a typical week.