Why AI Answers Cite Long Videos, Not Shorts.
94 percent of YouTube citations in AI answers come from long form video and only 5.7 percent from Shorts. What that means for a Shorts first channel.
When an AI assistant answers a question and cites YouTube, it almost never cites a Short.
In a study of more than 100 million AI citation instances published by the AI search monitoring platform Otterly.ai in March 2026, 94% of YouTube citations pointed to long-form video and 5.7% pointed to Shorts.
That is not a small skew. It is the difference between being a source and being absent from a surface that has quietly become a route to discovery.
The useful question is not whether to keep making Shorts. It is what a Shorts-only channel gives up, and what specifically makes a long-form video the thing a model reaches for. Neither of those is a volume problem.
What the numbers actually say
Otterly.ai collected more than 100 million citation instances over a 30 day period across six assistants: ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, Microsoft Copilot and Gemini.
Inside that set sat roughly 5.5 million social media citations. YouTube took 31.8% of them, second only to Reddit on 46.4%, which works out at around 1.75 million YouTube citations.
The split inside those YouTube citations is the finding worth acting on:
- Long-form video: 94%
- Shorts: 5.7%
- Playlists, channels and livestreams: 0.3%
Those three figures come from the same pool and sum to 100, so the Shorts number is not an artefact of a different denominator. It is the whole picture.
Length matters less than the headline suggests. Half the cited videos ran under eight minutes, and the largest single cluster was 10 to 20 minutes (32.1%), followed by 5 to 10 minutes (26.1%) and over 20 minutes (17.6%). Long form here means an ordinary video, not a lecture.
Citation is reference, not recommendation
A recommendation engine picks what will hold someone’s attention next. A citation engine picks what will support a claim it is about to make.
The first rewards a hook. The second rewards a complete, checkable answer, which is why the same video performs so differently on each.
The correlation data makes the point bluntly. Views, likes and subscriber counts all showed effectively no relationship with how often a video was cited, sitting within a whisker of zero and slightly negative at that (the study reports about minus 0.03 for both views and subscribers).
More than 40% of cited videos had fewer than 1,000 views at the time of analysis and 36% had fewer than 15 likes. Around 35% of the channels had under 10,000 subscribers. This is one of the few surfaces where a niche channel that never chased views competes with a large one on level terms.
What a Shorts-first channel actually loses
Not reach. Shorts still do what Shorts do, and a channel built on them can grow a real audience.
What it loses is the chance to be the answer. When someone asks an assistant how a thing works or which of two options to pick, the assistant builds a response from sources it can quote and then names them. A channel with nothing quotable is not in that running.
The loss is hard to spot because it does not appear in your analytics as a decline. It appears as an absence, and absences do not show up on dashboards.
There is a second cost. Citations compound. A video answering a durable question can keep being pulled into answers for months, long after its impressions curve has flattened. A Short’s working life is measured in days.
So the practical question is what makes a video citable. Three properties do most of the work.
1. A clear spoken answer, said in full sentences
The spoken word is the text of a video. If the answer is never said out loud in a complete sentence, there is nothing to lift.
Two habits break this more than anything else. The first is answering in pronouns and gestures: “this one is better”, “it’s about that much”, “like I said earlier”. Pulled out of context, none of it means anything.
The second is putting the answer only on screen. A figure in a graphic, a spec in a lower third, a comparison in a table nobody reads aloud. To anything working from the words, that content does not exist.
Both are fixed the same way. Say the question back before you answer it, then answer it once with the nouns intact. “The setup fee on a standard account is charged once, not monthly” survives being quoted. “It’s a one off, not like the other one” does not.
2. A transcript a machine can read
Every long-form video has a transcript, whether you write it or the platform generates one. The generated version is what most channels ship, and it is where the errors live.
Automatic captions cope well with plain conversational speech. They are least reliable on the words that carry the meaning: brand names, product names, technical terms, figures, anything said over music, anything said by two people at once.
A transcript that renders your product name three different ways cannot attribute an answer to you.
The fix is a review step, not a rewrite. Open the transcript, correct the proper nouns and the numbers, upload it. That check belongs in the edit rather than an admin queue: it is the cheapest step in the whole production and one of the few that decides whether the video can be quoted at all.
3. Structure a model can quote in pieces
A cited video is often not cited whole. Google’s AI surfaces treat a chaptered video as a set of separately addressable segments.
In the study, 31% of cited videos carried timestamp signals, and 78% of those were cited more than once, usually across two to five different chapters. One properly chaptered video became several citable units.
That behaviour was confined to Google. Every timestamped citation observed came from AI Overviews (73%) or AI Mode (27%), with none detected in ChatGPT, Copilot, Gemini or Perplexity.
Chapters are not difficult. YouTube asks for the first timestamp to be 00:00, at least three timestamps in ascending order, and a minimum chapter length of 10 seconds, all listed in the description.
The part teams get wrong is naming. Chapter titles written as moods (“The problem”, “Let’s get into it”) describe nothing. Chapter titles written as the question a viewer would type (“How much does an extra approval round cost you”) describe exactly what the segment answers.
The description is documentation, not a caption
Description length showed the strongest correlation the study measured (r = 0.31), with recency close behind at about 0.3 and hashtag presence at 0.20. The average cited video carried a description of 334 words.
A correlation of 0.31 is modest and it is not causation. A long description does not make a video citable on its own. What it signals is that videos with real descriptions tend to be the videos built as references in the first place.
Write the description as a summary someone could read instead of watching: what the video answers, in what order, with the chapter list underneath. Not a keyword block with three social links stapled to it.
Where the evidence stops
This is one study, run by a company that sells AI search visibility monitoring, over a single 30 day window. It has a commercial interest in the finding that AI search matters.
The prompt set is not published, and that matters more than the sample size does. A question mix weighted towards how-to and comparison queries would produce a long-form skew whatever the platform was doing.
The study also reports a pattern rather than a mechanism. It does not explain why long form wins, and nobody outside the model developers can.
Which is why the argument here does not rest on the number. It rests on the fact that the properties long-form video happens to have, a full spoken answer, a usable transcript and addressable segments, are the properties any system picking a reference would need. That bet would still be sensible at 70%.
This is not an argument against Shorts
Cutting Shorts to chase citations would be a poor trade. Shorts put your work in front of people who were not looking for it, and they are the cheapest way to learn which ideas hold attention before you spend a day filming one properly.
The failure mode is a Shorts-first channel with nothing underneath it. Every Short that lands raises a question, and when no long-form video answers that question the interest has nowhere to go, on the platform or in an assistant’s answer.
Moving attention across formats is its own job with its own mechanics, which is the ground covered by turning Shorts views into long-form growth. It only pays off when there is something worth linking to.
What good looks like
A channel with a spine. A small, deliberate set of long-form videos, each answering one question buyers actually ask, with Shorts running at whatever cadence the team can genuinely hold.
The spine gets built by question, not by calendar. Take the ten questions your sales conversations keep repeating and make one video for each. That is a quarter of work that keeps returning value rather than a quarter of output that expires weekly.
Each of those videos ships to the same standard:
- The question is asked out loud and answered in a complete sentence inside the first minute
- The transcript has been reviewed, with names, figures and product terms corrected
- Chapters start at 00:00, run at least three deep, and are named as questions
- The description is a written summary, not a keyword block
- Nothing load-bearing exists only as on-screen text
None of that is a creative decision. It is a production standard, which means it can be written down once and applied every time.
How NBK thinks about citable video
Most teams treat this as a content problem and reach for a better idea. It is a workflow problem, and it is solved at the moment the video leaves the edit.
A transcript check, a chapter list and a written description are three lines on a publishing checklist. Add them once and every video after that is citable. Leave them to whoever happens to be uploading and you get them on some videos and not others, which is close to not having them.
On the study’s own numbers, being cited has little to do with the size of the channel and a lot to do with whether the process makes each video legible by default.
Next step
If your team is publishing consistently and the work still is not compounding, the constraint is usually in the system around the content rather than the content itself. NBK can help you find it.
The NBK Social briefing
Our YouTube coverage, and everything else we publish, by email.