Text tools

How Long Should a Voiceover Script Be? Calculate Video and Presentation Timing

Estimate voiceover script length for videos and presentations. Calculate speaking time from word count, narration pace, pauses, and visual timing.

A voiceover script timing worksheet showing spoken word count, narration pace, and reserved visual pauses for planning a 90-second educational video.

You write a script for a 60-second video. The text looks short. Your word counter shows an estimated reading time of one minute. You start recording. The narration runs for 72 seconds, and you still need time for the introduction, visual demonstrations, and closing screen.

What happened? Nothing necessarily went wrong with the word counter. You estimated silent reading time, but your project required spoken delivery time. Those are different measurements. A useful script-length estimate must account for three things: the number of words actually spoken, the narrator's delivery speed, and the time reserved for visuals or pauses. This guide provides a practical method for budgeting those elements before you record a voiceover, presentation, tutorial, or short educational video.

Three clocks are running in a video

A script-based video effectively has three different clocks.

MeasurementWhat it tells youWhat it does not tell you
Word countHow many whitespace-separated words are in the draftHow long a narrator will take to say them
Speaking timeHow long the spoken words should take at a particular paceHow much non-speaking time the video requires
Total runtimeHow long the complete edited video lastsWhether the narration is clear or appropriately paced

Confusing these measurements is one of the easiest ways to make a script too long. A 200-word document may take approximately one minute to read silently in one context. That does not mean a person can narrate it clearly in one minute. A 60-second video containing several demonstrations may need far fewer than 200 spoken words. The format determines which clock matters most.

Why a word counter's reading-time estimate is different

The Anvil Tools Word & Character Counter estimates reading time using a fixed rate of 200 words per minute, rounding upward to the nearest minute. That is intended as a rough silent-reading estimate. It is not a spoken narration timer. For example, consider a script containing 180 words.

At the tool's reading-time rate, the calculation is 180 divided by 200, or 0.9 minutes. The displayed estimate rounds up to one minute. But if the script is narrated at 145 words per minute, the spoken portion takes approximately 74.5 seconds. Add 15 seconds of non-speaking visuals and transitions, and the planned runtime becomes approximately 89.5 seconds. The difference is substantial. Neither calculation is inherently contradictory. They answer different questions.

Use word count as the input for a speaking-time calculation, not the tool's silent-reading estimate as the video's final duration.

How fast should a narrator speak?

There is no universal ideal speaking speed. A measured technical explanation, a casual product demonstration, an energetic advertisement, and a formal presentation may all call for different pacing. For an initial estimate, you can model several delivery speeds.

Planning paceWords per minuteTypical planning use
Deliberate120Dense explanations or carefully paced instruction
Moderate140General educational narration
Brisk160Straightforward, energetic material

These are illustrative planning settings, not fixed rules or guaranteed averages. Your actual delivery speed may fall outside them. University presentation guidance provides a useful reminder that pace depends on the situation. The University of Kent describes approximately 100–130 words per minute for some scripted presentations, while faster unscripted delivery may reach 130–160 words per minute.

A 2019 research review by Marc Brysbaert reported average English reading-aloud rates of approximately 183 words per minute across the studies it examined. That research describes reading performance, not an ideal narration speed for every video. In other words, published figures can help you choose a starting point, but your own timed delivery is more relevant to your particular production. The University of Kent's presentation guide also emphasizes rehearsal, pauses, and coordinating speech with visual material.

Calculate the speaking time

The basic calculation is straightforward.

Speaking time in minutes = Spoken word count ÷ Words per minute

For seconds:

Speaking time in seconds = Spoken word count × 60 ÷ Words per minute

Suppose you write 210 spoken words and plan to narrate at 140 words per minute. The estimate is: 210 ÷ 140 = 1.5 minutes. That is 90 seconds of speech. If the finished video must also include 12 seconds of non-speaking material, the estimated runtime becomes 102 seconds.

The formula is useful, but notice an important limitation. It does not automatically account for pauses, music-only sequences, demonstrations, or the time needed to read information shown on screen. Those must be included separately, unless your measured speaking rate already incorporates them.

Calculate your word budget from the target runtime

When a video has a fixed duration, work backward. First determine how many seconds are available for speech.

Speaking seconds = Target runtime − Reserved non-speaking seconds

Then calculate the word budget.

Word budget = Speaking seconds × Planned words per minute ÷ 60

Here are four illustrative production budgets.

Total runtimeNon-speaking time reservedPlanned paceApproximate spoken-word budget
60 seconds8 seconds150 WPM130 words
90 seconds15 seconds145 WPM180 words
3 minutes25 seconds140 WPM360 words
5 minutes40 seconds135 WPM585 words

The budgets are calculated from the listed assumptions. They are not measurements from recorded videos. They also assume the reserved time is separate from the speaking time. If the narrator speaks continuously while the visuals change, those visuals do not necessarily require additional runtime. If the narrator stops so viewers can examine a diagram, that pause does consume time. The distinction matters when building an accurate timeline.

A worked example: planning a 90-second explainer

Imagine producing a short video about why a home circuit breaker trips. The goal is to explain a basic concept using an animation. You have exactly 90 seconds. Before writing the narration, decide what the audience needs to see without the narrator speaking. For this example, the production plan reserves 15 seconds for visual holds and transitions.

That leaves 75 seconds for spoken content. At an illustrative narration pace of 145 words per minute, 75 seconds allows approximately 181 spoken words. A practical target is about 180 words. Now divide those words among the message's four essential parts.

SectionWord budgetPurpose
Opening question15Introduce the problem
Basic explanation65Establish the concept
Visual demonstration60Explain the sequence
Closing takeaway40Reinforce the lesson
Total180Narration budget

At 145 words per minute, 180 spoken words take approximately 74.5 seconds. Adding the reserved 15 seconds gives an estimated finished runtime of 89.5 seconds. That leaves almost no spare time. The logical next step is not to squeeze in another paragraph. It is to rehearse the script and check whether the planned visuals genuinely fit. This is a fictional production plan intended to demonstrate the calculation. No actual recording or measured playback is being claimed.

Budget the visuals before polishing the sentences

Writers often finish the script first and leave visual timing until later. That works poorly when the audience must inspect something complicated. Imagine an educational video showing how current flows through a circuit. The narration says: "Notice what changes when the switch opens." The animation then needs time to show the switch opening and the current path changing.

If the voiceover continues immediately into an unrelated explanation, viewers may miss the demonstration. That is why a visual plan should identify where the audience needs time to observe, read, or compare. A simple production worksheet can contain:

SceneNarration needed?Visual eventTiming consideration
IntroductionYesShow the complete circuitBrief opening
Normal operationYesHighlight the current pathMatch explanation to animation
Switch opensBrieflyShow the path breakingAllow the change to register
Final explanationYesShow the resulting stateAvoid rushing the takeaway

Do not assign the same fixed pause to every visual. A simple image transition may need almost no dedicated silence. A detailed chart or safety-critical demonstration may need more time. The appropriate duration depends on the information the viewer must understand.

Do not count production notes as spoken words

A working script may contain more than narration. Consider this illustrative excerpt:

Script elementContent
Scene instructionShow a close-up of the switch
NarrationWhen the switch opens, the circuit is interrupted.
Animation instructionHighlight the broken current path
Editing noteHold the final frame briefly

Only the narration sentence is intended to be spoken. If you paste the entire production document into a word counter, the tool will count the scene instructions and editing notes as well. That inflates the word budget. For an accurate estimate, create a narration-only version of the script.

Count the spoken dialogue separately. Keep scene directions and production notes in another column or document. This also makes it easier to revise the narration without accidentally affecting the timing of unrelated visual instructions.

Numbers, abbreviations, and symbols can distort the estimate

A written token does not always correspond to one spoken word. Consider a script containing: "Set the resistance to 10 kΩ." The actual narration might be: "Set the resistance to ten kilo-ohms." The printed notation and the spoken sentence have different structures. Other examples include dates, currencies, percentages, mathematical expressions, URLs, abbreviations, and technical identifiers.

A narrator may read an acronym as a word or pronounce each letter separately. A mathematical expression may require several seconds to communicate clearly. For this reason, technical scripts are often easier to time when written as they will actually be spoken. Instead of relying on the printed notation, prepare a narration version with the intended pronunciation.

That gives you a more useful word count and helps the narrator deliver consistent terminology. For complex equations, the timing still needs rehearsal. A word count alone cannot capture the time required to explain or emphasize mathematical relationships.

Measure your own speaking speed in under two minutes

The most useful pace is the one your narrator can sustain while remaining clear. To find an initial estimate, select a representative passage and read it aloud in the intended style. Suppose your sample contains 120 spoken words. You record the narration and measure 48 seconds of continuous speech.

The pace is: 120 ÷ 48 × 60 = 150 words per minute. You can now use 150 WPM as an initial estimate for similarly written material. However, that result should not be applied mechanically to every script. A passage containing familiar words may be delivered faster than one containing difficult technical terminology.

A conversational introduction may be faster than a careful explanation of a diagram. A useful approach is to measure one straightforward passage and one technically demanding passage. That provides a realistic working range rather than a single number that looks more precise than it actually is.

Avoid counting pauses twice

There are two valid ways to measure speaking pace, but they should not be mixed carelessly.

Method A: Continuous speech rate

Measure the time spent actually speaking, excluding intentional silent holds. Then calculate speaking time and add planned non-speaking pauses separately.

Method B: Full delivery rate

Measure the complete performance, including its natural pauses. Then use that observed overall rate to estimate a similar performance. If you use Method B and add those same kinds of pauses again, the estimate may be too long. For example, suppose a 150-word passage takes 75 seconds including natural delivery pauses.

Its measured overall pace is 120 words per minute. If you use 120 WPM for another similar passage, that estimate already reflects the measured pauses. Do not automatically add a second allowance for the same pauses. You may still need extra time for separate animation holds, demonstrations, or non-speaking scenes not represented in the test. The key is to know what your timing measurement includes.

Plan presentations differently from narrated videos

A recorded voiceover is usually based on a relatively fixed script. A classroom presentation or live demonstration is more variable. The presenter may point to a slide, answer a question, react to the audience, or explain a concept in an unexpected way. For example, a ten-minute presentation may need time for:

  • Introducing the topic.
  • Explaining the main points.
  • Showing visual examples.
  • Allowing questions, if required.
  • Delivering a proper conclusion.

If all ten minutes are filled with a word-for-word speech budget, the presentation has little flexibility. The University of Kent recommends planning presentation time across sections and allowing for the specific requirements of the event. For a live presentation, a script-length estimate is a useful starting point, but rehearsal with the actual slides is more dependable than word count alone. If the presentation uses cue cards rather than a full manuscript, word count becomes even less predictive. A presenter may say much more or less than the outline suggests.

What changes with AI-generated voiceovers?

Synthetic narration introduces another variable. Some text-to-speech systems allow voice speed to be adjusted. Others produce different pacing depending on punctuation, sentence structure, emphasis, or voice selection. A slower speed setting may increase runtime, but the exact relationship depends on the system. Do not assume that setting a voice to a particular speed produces an exact, universal words-per-minute rate.

Generate a short representative sample. Measure its actual duration. Then estimate the full script using the same voice and settings. If possible, test a passage containing technical terms, numbers, and pauses similar to those in the final production. Once the complete narration is rendered, the audio file's measured duration is more authoritative than the original word-count estimate. At that point, use the measured audio timeline to plan the edit.

A script's timing is not the same as caption timing

Captions need to match the audio that viewers actually hear. An estimated 90-second script duration does not provide accurate caption timestamps by itself. The final narration may include timing changes, re-recorded sentences, or pauses. Captions should therefore be synchronized with the finished audio or video timeline.

They should also represent relevant speech and non-speech audio information rather than treating the original draft as a guaranteed exact transcript. The W3C Web Accessibility Initiative explains how captions support people who are Deaf or hard of hearing and why synchronization matters. For an educational video, it is also worth considering whether important visual information is explained in the narration. For example, saying only "as you can see here" may not communicate the actual change to someone who cannot see the demonstration.

The W3C guidance on describing visual information explains ways to make essential visual details available through audio description or integrated narration. These are accessibility considerations, not merely techniques for adjusting video length.

Does the same formula work for Hindi, Urdu, or other languages?

The arithmetic is universal. If you know the number of counted words and the delivery rate measured using the same counting convention, you can estimate speaking time. But the word-counting rule itself may not behave identically across languages. The Anvil Tools counter uses whitespace boundaries.

That is a practical way to count many space-separated scripts, but it is not a language-aware speech segmentation system. Different languages also express the same information using different numbers of written words. A speaking rate measured for an English script should not automatically be applied to Hindi, Urdu, Chinese, or another language. For multilingual narration, measure a representative passage in the actual language and writing style.

If your narrator is delivering a Hindi script, calibrate using Hindi narration. If the script mixes English technical terms with another language, include that mixture in the sample. The objective is consistent measurement, not forcing different languages into one assumed speed.

How to shorten an overlong script without losing the lesson

Suppose your narration-only draft contains 235 words, but your 90-second plan allows approximately 180. You need to remove about 55 words. Deleting random sentences may damage the explanation. Instead, examine the role of each sentence. Some sentences introduce essential information. Some explain a process. Some repeat a point that is already clear from the visuals.

Others describe details that could be communicated more effectively by the animation. For example, if an on-screen demonstration already shows a button being pressed, the narration may not need a lengthy description of every cursor movement. However, essential visual information should still be communicated accessibly. Do not remove an important explanation simply because viewers are expected to see it. A useful editing principle is:

Keep what the audience needs to understand, not every detail the creator knows.

After shortening the draft, count the spoken words again. Then read it aloud. A shorter script is only an improvement if the message remains clear.

The actual workflow in Anvil Tools

The Anvil Tools Word & Character Counter is a suitable starting point for script budgeting. It reports words, characters, sentences, paragraphs, and an estimated reading time. Its word count uses whitespace-separated tokens. Its reading-time estimate uses 200 words per minute and rounds up. The tool does not provide a dedicated speaking-speed selector, voiceover duration calculator, audio recording feature, or video timeline editor. To use it for narration planning:

  1. Prepare a narration-only draft, without editing instructions.
  2. Paste the spoken text into the counter.
  3. Record the word count.
  4. Choose a preliminary delivery rate appropriate to the material.
  5. Calculate the estimated speaking duration.
  6. Add separate non-speaking time when appropriate.
  7. Rehearse or generate a sample recording.
  8. Update the estimate using the observed pace.
  9. Check the completed audio before finalizing the video timeline.

If you want to understand why the tool's counts differ for emoji, accented characters, or languages without spaces, see How Our Word Counter Counts Emoji, Accents, and Chinese Text. That experiment explains the counting rules. The workflow in this guide explains how to use a word total for spoken production timing.

A reusable script-budget worksheet

Before recording, fill in these fields:

FieldYour value
Target runtime___ seconds
Reserved non-speaking time___ seconds
Available speaking time___ seconds
Planned narration pace___ words/minute
Calculated spoken-word budget___ words
Actual narration word count___ words
Rehearsed speaking duration___ seconds
Estimated or measured final runtime___ seconds

Then answer two questions. First, does the script fit the intended runtime without forcing unnatural speed? Second, have you allowed enough time for the audience to understand the visuals? If the answer to either question is no, revise the script or the production plan before recording everything. That is generally more efficient than discovering the timing problem after the animation and editing work are complete.

The takeaway: write for the clock that matters

Word count is valuable because it turns a vague writing task into something measurable. But it cannot determine video duration independently. A reliable script estimate combines:

Spoken word count + realistic delivery pace + required non-speaking time.

Use the word counter to measure the narration. Use arithmetic to build the initial time budget. Then use a real rehearsal or generated audio sample to replace assumptions with observed timing. For a 60-second video, a 90-second explainer, or a ten-minute presentation, the same principle applies:

The right script length is the amount of spoken content your audience can understand within the time actually available.

Try the relevant text tool

Preparing a voiceover, educational video, short presentation, or spoken tutorial? Start with the Anvil Tools Word & Character Counter to measure your narration draft. Use the spoken-word total to estimate delivery time, and verify the result with a timed recording before final production.

Further reading

← Back to guides & experiments