Reading time, counted not guessed

How a small plugin for Astro’s new Markdown engine counts the words in every post, in any language, and files it as a quick read, an essay or a long read.

“5 min read” is a small promise. It tells you whether a post fits in a coffee break or needs an evening. On most sites that number is either typed by hand, and forgotten when the post changes, or guessed by splitting the text on spaces and dividing by 200.

Both are easy to get wrong. So on this site the number is counted, at build time, by a plugin that runs on every post. Nobody writes it down.

Where it runs

This site is built with Astro 7, which renders Markdown with Sätteri, a Markdown engine written in Rust. Sätteri plugins come in two kinds: mdast plugins see the Markdown syntax tree, hast plugins see the HTML tree. Reading time is about words, so it is an mdast plugin.

It runs once per post when the site is built, and writes its result into the post’s frontmatter, next to the title and the date:

// astro.config.ts
import { satteri } from '@astrojs/markdown-satteri';
import { readingTime } from './src/plugins/reading-time/index.ts';

export default defineConfig({
  markdown: {
    processor: satteri({ mdastPlugins: [readingTime({ wordsPerMinute: 220 })] }),
  },
});

Templates then read readingTime like any other field. There is no JavaScript in your browser doing the counting, and nothing to keep in sync.

Counting words properly

Splitting on spaces breaks in three common ways. It counts punctuation that stands alone, like a dash, as a word. It cannot count languages that do not put spaces between words. And it has no idea what a word is in French or Arabic.

JavaScript has had a real answer for a while: Intl.Segmenter. Ask it to cut text into words and it returns every segment with a flag, isWordLike, that is true for words and numbers and false for spaces and punctuation:

const words = new Intl.Segmenter('en', { granularity: 'word' });
let n = 0;
for (const part of words.segment(text)) if (part.isWordLike) n++;

That one loop handles English, French elisions like “l’intention”, Arabic, and Chinese or Japanese text, which the segmenter splits into words using a dictionary.

One block at a time

The obvious way to get the text of a post is to ask for the text of the whole tree. That glues blocks together: the last word of a heading and the first word of the next paragraph come out as one long word. Over a long post you lose dozens of words.

So the plugin counts each block on its own: every paragraph, heading and table cell. Inside a block, formatting does not split words, so a word that is half in italics still counts once.

What does not count

Two things get special treatment.

Code blocks are skipped. You scan code, you do not read it word by word, and a long snippet would otherwise double the estimate. A post about code can switch this off with one option.

Images add time. Each image adds twelve seconds, a common rule of thumb for looking at a picture.

The arithmetic

The plugin assumes 220 words per minute, a middle-of-the-road pace for adults reading silently. The estimate is rounded up and never goes below one minute:

Words Minutes Label
150 1 1 min read
700 4 4 min read
1,500 7 7 min read

The result also carries the exact word count, which shows up when you hover the little ring next to the time.

Organising the blog with it

Because the number is reliable, the blog can use it. Every post is filed automatically by length:

  • Quick read: under 3 minutes.
  • Essay: 3 to 6 minutes.
  • Long read: 7 minutes and up.

On the blog page you can filter by length as well as by shelf, so “something short about agents” is two clicks away. None of those labels are typed by hand either; a post that reaches seven minutes moves to the long reads on the next build.

Tested

The plugin has no dependencies and comes with tests that run on Node’s built-in test runner. They cover the counting (English, French, Arabic), the rounding, the length labels, skipped code, images, and the case that bit first: words in neighbouring paragraphs must not be glued together.

It also works outside Astro. Handed to Sätteri directly, it writes its result to the compile’s data bag instead of the frontmatter, so it can be reused anywhere Sätteri runs.

That is the whole trick: count the words once, properly, where the Markdown is already being read, and let everything else use the answer.

Keep reading

Notes

A slower home on the internet

Why this site is one calm page, a short bio and a notebook, and why almost everything on it is drawn by hand in SVG.

2 min read Quick read

Agents

Claude Code in your pocket

A single Python file turns a Telegram bot into a full Claude Code session on your own server. What it does, how it stays safe, and the bugs you do not have to rediscover.

5 min read Essay

Stars

Why your birth time matters

The Sun barely moves in a day. The Ascendant goes all the way round the zodiac. A plain guide to what a birth time changes in a chart, and what to do if you do not know yours.

3 min read Essay