Yes, when the routine has a physical sequence the student has never watched anyone perform. The story carries the reasons and the feelings. The video carries the motor steps. Pair them only when both jobs exist, because in a 2024 survey of 16 school SLPs, pediatric OTs, and parents, 94% already spend 30 or more minutes on the story alone.
Why would you add a video to a story that already works?
Because text explains and video demonstrates, and some routines fail on the demonstration side. A student who understands exactly why the class lines up can still not know what lining up looks like from the inside.
AFIRM lists social narratives as an evidence-based practice for autistic learners aged 3 to 22, and lists video modeling separately as a video-recorded demonstration of the target behavior shown to the learner. Two practices, two different mechanisms. You are not doubling up on the same thing.
Parents describe this pattern without the vocabulary. One version that shows up repeatedly in autism parenting threads: the first haircut was a disaster, so the second time they went and watched another kid's haircut twice, just watched, then tried again, and it worked. That is in-vivo modeling. The video is the version you can run in a school building on a Tuesday.
When does pairing actually help?
When you can film the routine and someone watching the film would learn the order of the steps. That is the whole test.
| Target | Story alone | Add video? | Why |
|---|---|---|---|
| Handwashing sequence | Explains why | Yes | Six observable steps in fixed order |
| Walking to the bus line | Explains where and when | Yes | Route and landmarks are visual |
| Opening a lunch container | Weak fit alone | Yes | Pure motor sequence, little social content |
| Asking for a break | Strong fit | No | One action, nothing to demonstrate |
| Coping with a schedule change | Strong fit | No | Internal, no filmable sequence |
| Losing a game without quitting | Strong fit | Rarely | The hard part is the feeling, not the motion |
The rows marked no are the ones where a video gets made anyway because video feels thorough. A clip of a peer sitting calmly through a schedule change shows the outcome, not the process, and the student already knows what calm looks like.
What order do you run them in?
Story, then video, then story again right before the routine. The first reading frames what the video is about, which is the difference between a teaching clip and thirty seconds of screen time.
Concretely, a session looks like this. Read the story. Say one sentence connecting it to the clip, such as "this shows the part where we walk past the office." Play the clip once at normal speed. Play it a second time only if the student asks. Then re-read the story on the day of the routine, as close to the routine as your schedule allows.
Do not narrate over the video with different wording than the story uses. Competing phrasing is the most common way a paired intervention gets less effective than either piece alone.
From the same 2024 community survey: "Getting suitable pictures is 90% of the work." Video makes that worse before it makes it better. If you are already losing an evening to photo hunting, adding a filming step to every story is how the whole practice gets abandoned. Pair on two or three students, not on the caseload.
Who should be in the video?
An adult or a generic model for most school routines, the student themselves only with consent and a clear reason. Video self-modeling, where the student watches an edited clip of their own successful performance, is powerful and is also the highest-friction option in a school.
A video of an identifiable student is an education record. That puts it under FERPA, so your district's photo and video consent process applies before you film, before you store the file, and before you share it with a classroom teacher. Filming your own hands, or a willing staff member, skips the paperwork entirely and works for most motor sequences.
One practical note on storage: keep the file in the district drive, not on your phone. A clip that only exists on your personal device cannot be handed to a paraprofessional, which defeats the point of making it.
How long should the clip be?
Thirty to sixty seconds, one routine, one pass at normal speed. If a clip runs past a minute it almost always contains two routines, and splitting it into two clips is better than trimming it.
Speed matters more than people expect. Slowing the footage down seems helpful and usually is not, because the student then has to translate the slow version into real time. Film it at the pace the routine actually happens.
Shoot from behind or over the shoulder when the routine is something the student does with their hands. That point-of-view angle removes the extra step of mirroring what a facing model is doing.
Does the pairing change how long you run the intervention?
No, run the same schedule you would run for the story alone. The ASSSIST-2 trial, across 87 schools and 249 autistic children aged 4 to 11, delivered the story at least six times over four weeks. Adding a video does not shorten that window.
The 2026 Frontiers in Psychology meta-analysis of 21 single-case studies reported a moderate overall effect, Tau-U of 0.743, with digital formats scoring slightly higher than paper but not by a statistically significant margin. Read that as permission to use whichever format your building supports, not as evidence that screens beat paper.
How do you hand a paired intervention to someone else?
Print a QR code to the video on the last page of the story. Then the story and the clip are one object, and a substitute or paraprofessional can run both without finding you first.
Add the six reading checkboxes next to it. The person carrying the intervention on a day you are at another building needs the sequence, the link, and a place to tick. Anything more elaborate does not survive contact with a school schedule.
Frequently Asked Questions
Should you pair a social story with video modeling?
Yes, when the routine has a physical sequence the student has never seen performed. The story carries the reasons and the feelings, the video carries the motor steps, and AFIRM lists both social narratives and video modeling as separate evidence-based practices for autistic learners.
Which one comes first, the story or the video?
The story first, then the video, then the story again just before the routine. The story tells the student what the video is about, which stops the video from being watched as entertainment.
When does adding a video not help?
When the target is a feeling, a choice, or a rule with no observable sequence. Waiting your turn, coping with a change of plan, and asking for a break have nothing useful to film, and a video of a peer sitting quietly teaches almost nothing.
Do I need permission to film a student for video modeling?
Yes. A video of an identifiable student is an education record under FERPA, so follow your district's photo and video consent process before filming, storing, or sharing it. Filming an adult or using generic stock footage avoids the issue entirely.
How long should the video be?
Thirty to sixty seconds for most K-5 routines, showing the sequence once at normal speed with no narration competing with the story wording. Anything longer usually contains more than one routine and should be split.
Can I just use a YouTube video instead of filming my own?
For community routines like a haircut or a dentist visit, often yes. For school routines it usually fails, because the hallway, the cafeteria line, and the bathroom in the video are not the ones the student has to walk into.
Does adding a video make the intervention harder to hand off?
It can. Put the video link as a QR code on the printed story so a paraprofessional or classroom teacher can run both parts without asking you, and keep the file in the shared district drive rather than on your phone.
One approach for school SLPs short on time is to keep a 5-tool stack: a methodology checklist for the descriptive-to-directive ratio, a slide template you reuse, a folder of stock photos sorted by scenario, an AI text drafter (ChatGPT, Claude, MagicSchool, or Emoquest for one-sentence-in story output), and a delivery format your district already uses (Google Slides or PDF). Video is a sixth tool, and it belongs on two or three students at a time. Film the routines that repeat across your whole caseload first, because a handwashing clip works for every student who needs one.