Key takeaways
- The verb in the goal usually tells you the measurement type to use.
- Binary, frequency, duration, trial-based, interval, rate, latency, and task analysis each capture a different shape of behavior.
- The same vague goal becomes a different data plan depending on the type you pick.
- Inter-observer agreement is the test that your type is scorable by more than one person.
The measurement type you choose for a goal decides everything downstream: what a data point looks like, how often you can realistically collect it, and whether the goal is even loggable during a normal teaching day. Pick the wrong type and you end up with a goal nobody can score consistently. Pick the right one and the data almost collects itself. This guide walks through eight measurement types, starting with the five most IEP goals use, and shows how the same concern turns into different data plans depending on the type.
The eight measurement types
Each type fits a different shape of behavior. Read the verb in the goal first, then match it to the type whose data point you could actually capture.
Binary (yes or no)
Use binary when the behavior either happens or does not, once per occasion, and there is nothing in between. One data point is a single Y or N. It fits goals about completing a routine, arriving on time, or turning in an assignment.
One-row example: "Will independently start the morning unpacking routine within 2 minutes of entering the room." A data point on Monday is simply Y. On Tuesday, N. Across a week you get a string like Y N Y Y N, which is 3 of 5, or 60 percent of days.
Frequency (count)
Use frequency when the behavior is a discrete event with a clear start and stop that you can tally. One data point is a count over a fixed window. It fits calling out, requesting a break, hitting, or hand-raising.
One-row example: "Will reduce call-outs to fewer than 3 per 30-minute morning meeting." A data point is a tally: 7 call-outs during the 30-minute meeting. The fixed window (30 minutes) matters, because a count without a window is not comparable across days.
Duration
Use duration when how long the behavior lasts matters more than how often it happens. One data point is an elapsed time. It fits time on task, time staying in an assigned area, or how long a tantrum lasts.
One-row example: "Will remain on task during independent work for at least 15 of 20 minutes." A data point is 12 minutes on task out of the 20-minute block. Duration answers a question a count cannot: episodes can stay just as frequent while getting shorter, which is real progress a frequency count would miss.
Trial-based (percent of trials correct)
Use trial-based when you can present the skill as repeated, scorable trials. One data point is a ratio of correct to total. It fits naming, identifying, solving, decoding, or any skill you can probe.
One-row example: "Will identify the main idea in a grade-level passage with 80 percent accuracy." A data point is 8 of 10 prompts correct (80 percent). Keep the number of trials and the difficulty consistent across sessions, or the percentage is measuring the task, not the student.
Interval recording
Use interval recording when the behavior is so frequent or continuous that counting every instance is impossible. You divide an observation window into intervals and mark whether the behavior occurred in each. One data point is the percentage of intervals with the behavior.
One-row example: "Will be engaged in the group activity during at least 70 percent of observed intervals." Over a 10-minute session split into 10 one-minute intervals, you mark engaged in 6 of them: 60 percent. It trades exact counts for a number a busy teacher can actually capture on a near-constant behavior.
Rate (count per minute)
Use rate when you are counting, but the time you watch changes from one session to the next. One data point is a count and the minutes observed, written as a number per minute. It fits reading fluency, math facts, or responses during a lesson that runs long one day and short the next.
One-row example: "Will read at least 60 words correctly per minute from a grade-level passage." A data point is 84 words read correctly in 2 minutes: 42 per minute. Dividing by the time is what lets a 2-minute probe and a 5-minute probe sit on the same graph.
Latency (time to start)
Use latency when the question is how long the student takes to get started after the cue. One data point is the seconds between the direction and the start of the response. It fits following a direction, starting independent work, or answering a question.
One-row example: "Will begin the assigned task within 30 seconds of the direction." A data point is 45 seconds from the direction to the first move. Latency is not duration: duration times the behavior itself, and latency times the wait before it.
Task analysis (step by step)
Use task analysis when the skill is a routine made of steps done in order, and you need to know which steps the student does alone. You write the steps once, then give each step one answer every session: independent, prompted, no response, or not asked. One data point is the percentage of steps done independently. It fits washing hands, packing a backpack, or long division.
One-row example: "Will wash hands, completing at least 6 of 7 steps independently." A data point is 5 of 7 steps independent, or 71 percent. The step-by-step record also shows which steps still take help, which a single yes or no for the whole routine would hide.
A decision table by goal verb
When you are unsure, start from the verb. This table maps common goal language to the type, an example goal, and what a single data point looks like.
| Goal verb | Recommended type | Example goal | One data point |
|---|---|---|---|
| "complete," "turn in," "arrive" | Binary (yes/no) | "Will turn in homework on the day it is due" | Y or N for today |
| "reduce," "request," "raise hand" | Frequency (count) | "Will request a break instead of leaving the room" | 4 break requests in the morning block |
| "remain," "stay," "sustain" | Duration | "Will remain in the assigned area during centers" | 14 of 20 minutes in area |
| "identify," "solve," "decode," "name" | Trial-based (percent) | "Will solve two-step word problems" | 7 of 10 correct (70 percent) |
| "engage," "attend," "participate" (near-constant) | Interval recording | "Will attend to the lesson during group instruction" | Behavior present in 6 of 10 intervals |
| "read," "write," "answer" (per minute) | Rate (count per minute) | "Will read 60 words correctly per minute" | 84 words in 2 minutes (42 per minute) |
| "begin," "start," "respond within" | Latency (time to start) | "Will begin the assigned task within 30 seconds" | 45 seconds from the direction to the start |
| "wash hands," "pack up," "follow the steps" | Task analysis (step by step) | "Will wash hands, doing each step independently" | 5 of 7 steps independent (71 percent) |
The verb is a strong hint but not a law. "Participate" could be a count (number of contributions) or an interval measure (attending across intervals), depending on what you actually want to change. Decide what behavior you care about, then let the data point you can capture settle the type.
The same goal, measured two ways
A vague goal hides the choice. Watch what happens when you make a single concern measurable two different ways. The type you pick changes what you collect every day.
The concern: a composite fifth grader we will call Priya is "off task too much" during independent work.
Measured as frequency: "Number of times Priya leaves her seat during the 20-minute independent block." A data point is a tally, like 5 times today. You are counting discrete events. This works if the problem is getting up and wandering.
Measured as duration: "Total minutes Priya spends off task during the 20-minute independent block." A data point is elapsed time, like 9 minutes off task. You run a stopwatch (or estimate in chunks). This works if the problem is long stretches of disengagement rather than frequent movement.
Same concern, two different data plans. If Priya gets up rarely but stares out the window for ten minutes at a stretch, the frequency count looks fine while the real problem is invisible. Picking duration here is not a style preference; it is the difference between data that catches the problem and data that misses it.
The concern: a composite first grader we will call Devon is not reading along during shared reading.
Measured as trial-based: "Devon will correctly read 8 of 10 targeted sight words when shown flashcards." A data point is 6 of 10 correct. You are probing a discrete skill on demand. This fits if the goal is sight-word acquisition.
Measured as interval: "Devon will visually track and follow along during shared reading in at least 70 percent of observed intervals." A data point is following along in 5 of 10 one-minute intervals. You are sampling attention during a continuous activity. This fits if the goal is participation and attending, not word knowledge.
The two are not interchangeable. The trial-based version tells you whether Devon knows the words. The interval version tells you whether he is engaged in the routine. Choosing the wrong one means you collect clean data about the thing you did not actually care about.
Common mismatches and how to catch them
A few pairings go wrong often enough to name. Spotting them saves a quarter of data you cannot use.
The first is counting a continuous behavior. If a student is off task almost the entire block, a frequency count of "times off task" is meaningless, because off task never really stops to start again. That is an interval or duration behavior. When a tally feels impossible because the behavior never pauses, switch types.
The second is timing a discrete behavior. If a student calls out in quick bursts, a stopwatch tells you little; you want a count. When the behavior has obvious edges (it happened, then it was over), frequency beats duration.
The third is a percentage with no fixed denominator. "Percent of time on task" sounds like duration but is often collected as a vague estimate. Decide whether you are timing minutes (duration) or sampling intervals (interval recording), and hold the window constant either way.
The concern: a composite kindergartner we will call Theo struggles with transitions between activities.
Measured as binary: "Theo will transition to the next activity without a tantrum." A data point is Y or N per transition. Simple, fast, and it answers whether the transition blew up at all. This fits if the team only needs to know how often transitions succeed.
Measured as duration: "Theo will complete transitions within 2 minutes of the signal." A data point is the elapsed time, like 90 seconds. This fits if transitions technically happen but take so long they eat instruction.
Pick binary if the question is "did it work," duration if the question is "how long did it take." A team that wanted to shorten transitions but only collected Y or N would have clean data that never shows the transitions getting faster, which was the actual goal.
When in doubt, write down what a single data point would physically look like in your hand: a tally mark, a number of minutes, a Y or N, a fraction of trials, a fraction of intervals. If you cannot picture capturing that data point during a real lesson, the type is wrong for your setting, however correct it looks on paper.
Inter-observer agreement, in plain terms
There is one test that tells you whether your measurement type is actually scorable: would two people watching the same behavior at the same time write down close to the same number?
That is inter-observer agreement. If you and a paraprofessional both count call-outs during the same morning meeting and you get 7 while the para gets 8, that is close agreement and the measure is solid. If you get 7 and the para gets 2, you have a definition problem. One of you is counting something the other is not, maybe muttering versus full call-outs, and the data is not trustworthy no matter how diligently you collect it.
A type that two people score the same way is a type you can trust, share with a co-teacher, and defend in a meeting. A type that only you can score, because only you know what counts, is a single point of failure. When you are choosing between two measurement types, the more objective one (a clean count, an elapsed time, a correct-or-not trial) usually wins over a judgment call, precisely because more than one person can capture it.
A quick chooser, from goal to data sheet
When you have a goal in front of you and a minute to decide, answer these questions in order and stop at the first yes.
- Is the team's question why the behavior happens? That is not a progress measure yet. Collect ABC data (antecedent, behavior, consequence) to find the function, using an ABC observation log, then choose a type below for the goal itself.
- Does the behavior happen once per opportunity, with nothing in between? Binary. Score yes or no per opportunity on a yes / no data sheet.
- Is the skill a routine made of steps, and do you need to know which steps the student does alone? Task analysis. List the steps in order and score each one every session.
- Is the skill something you can present as repeated trials? Trial-based. Record correct out of total on a trial-based data sheet.
- Does how long it lasts matter most? Duration. Record start and stop times on a duration data sheet.
- Does how long it takes to start matter most? Latency. Use the same duration data sheet, and start the clock at the direction.
- Is it a discrete event you can tally without it taking over your lesson? Frequency. Tally within a fixed window on a frequency data sheet. If the window changes length from day to day, write down the start and end times too and use rate, the count per minute.
- Is it too frequent or continuous to count or time? Interval recording. Mark each interval on an interval recording sheet.
Then ask one more question for any goal about independence: does the amount of help matter? If yes, record a prompt level with every entry, from independent through verbal, gestural, model, partial physical, and full physical. A student who holds at 7 of 10 correct while moving from a model prompt to a verbal one is making progress a plain percentage hides. The prompt level data sheet puts the least prompt of the day in its own column so that fade is easy to read.
| Measurement type | One data point | Free printable |
|---|---|---|
| Binary (yes/no) | Y or N for this opportunity | Yes / No IEP Data Sheet |
| Frequency (count) | A tally in a fixed window | Frequency Data Sheet |
| Duration | Elapsed minutes or seconds | Duration Data Sheet |
| Trial-based (percent) | Correct out of total trials | Trial-Based IEP Data Sheet |
| Interval recording | Intervals with the behavior out of intervals observed | Interval Recording Sheet |
| Prompt level (added to any type) | The least help the student needed | Prompt Level Data Sheet |
The chooser has one override. If staff cannot score the goal during a normal day with the type you picked, change the method before you conclude anything about the student. A goal nobody can log consistently produces a quarter of gaps, not a trend.
Evident records all eight types, with optional prompt levels, on each IEP goal. If you collect on a Chromebook, Evident Capture logs a data point from any browser tab in under ten seconds. Task analysis is the exception: it is scored step by step in the web app.
Where to go next
Once every goal on your caseload has a type tagged, the complete guide to IEP progress monitoring shows how to turn those types into a collection schedule you can actually keep across the year. If the goal itself is too vague to measure, fix the wording first with the measurable IEP goal wording bank. And when the quarter closes, from daily notes to IEP progress report turns the data points into report sentences.
Common questions
Which measurement type works best for behavior goals?
It depends on the shape of the behavior. Count discrete events like call-outs or break requests (frequency), time behaviors where length matters like staying on task (duration), score once-per-opportunity routines as yes or no (binary), and sample near-constant behaviors like engagement (interval). Pick the one whose single data point a teacher can capture during a normal lesson.
Is ABC data a measurement type?
Not in the same sense. ABC data records what happens before and after a behavior so a team can work out why it happens. It supports an FBA. Progress on the IEP goal itself is still measured with one of the eight types.
What are prompt levels, and when should I record them?
Prompt levels record how much help a student needed, usually from independent through verbal, gestural, model, partial physical, and full physical. Record them alongside any measurement type when the goal is about independence, because a skill can stay at the same accuracy while the help it needs fades, and that fade is progress.
Can one goal use two measurement types?
It can, but it doubles the collection load. Usually one type answers the goal's question and the second is context. A common exception is a trial-based or yes-or-no goal with a prompt level on every entry, which adds seconds, not a second system.