Open methodology · 70 lessons

How does A1 news differ from A2 news?

We measured a fixed sample of public LinguaDrop lessons in English, French, Lithuanian, and Indonesian. The result is a descriptive benchmark of text length, sentence size, and vocabulary variety, not a claim that a formula can validate CEFR proficiency.

70

public lessons

4

target languages

10

lessons per language and level

A1 and A2 results

Snapshot: 2026-09-04. Each row is a mean over 10 lessons, collected on the date shown in its row.

Readability measurements for A1 and A2 LinguaDrop news lessons in English, French, Lithuanian, and Indonesian
LanguageCEFR settingLessonsWords / lessonWords / sentenceUnique-word shareCollected
EnglishA11069.58.678.7%2026-08-16
EnglishA21077.67.876.8%2026-08-16
FrenchA110619.179.4%2026-08-16
FrenchA21095.610.975%2026-08-16
IndonesianA21067.97.878%2026-09-04
LithuanianA11041.85.889.1%2026-08-16
LithuanianA21065.67.587.8%2026-08-16

What the snapshot does, and does not, show

A2 lessons are longer than A1 lessons in every language that has both bands. Sentence length does not rise uniformly: the English A2 sample has fewer words per sentence than its A1 sample. That is why we publish the rows instead of compressing them into a universal “readability score.”

Reading the Indonesian A2 row

Indonesian is the one language here with no A1 row, because LinguaDrop publishes no Indonesian A1 lessons. There is nothing to sample, so nothing is shown. The useful comparison is sideways, against the other A2 samples: Indonesian A2 averages 67.9 words per lesson and 7.8 words per sentence, which is the same sentence length as English A2 and noticeably shorter than French A2 at 10.9.

Its unique-word share of 78.0% sits between French and Lithuanian, and that gap is mostly grammar rather than difficulty. We count distinct word forms, so a heavily inflected language such as Lithuanian scores high because one lemma appears in many forms, while Indonesian inflects far less and its forms are closer to distinct words. Compare unique-word share within a language, not across languages.

Read Indonesian news lessons for yourself and judge the numbers against the text.

Reproducible methodology

  1. 1

    Freeze a bounded public sample

    We selected the 10 most recent lessons in each published language and level band, scanning the public API in pages until every band was filled. Bands are added over time and each lesson records the date it was collected, so an added language does not restate figures already published. The rule, lesson IDs, public URLs, and collection dates are checked in.

  2. 2

    Measure only the adapted text

    The parser keeps target-language sentences and removes the interleaved source translation and vocabulary gloss lines. No user data, private corpus, or original full news article is included in the download.

  3. 3

    Use language-aware segmentation

    JavaScript’s Intl.Segmenter counts words and sentences using the target-language locale. Unique-word share is lowercased orthographic types divided by words.

  4. 4

    Generate the page from the snapshot

    A deterministic repository script calculates all six table rows and the CSV. The test suite fails if generated numbers drift from the checked-in source snapshot.

Inspect representative lessons

Every CSV row links to the complete public lesson and its publisher credit. These six links provide one example from each measured band.

Find your starting level

Take the five-question CEFR reading test, then continue into a matching free lesson.

Take the free test

Check your own passage privately

Compare its sentence length, repetition, vocabulary variety, and long-word share without uploading or saving the text.

Explain a passage

Turn a lesson into classroom practice

Use bounded bilingual excerpts, vocabulary, and original comprehension prompts in a printable worksheet. No account or generated answer storage.

Make a free worksheet

Open dataset v1

Vocabulary density by language and level

Aggregate vocabulary variety, repetition, and sentence-length spread across 70 public lessons, derived from the same snapshot as the benchmark above. Aggregates only: no article text, lesson id, or URL is published.

Vocabulary density and sentence-length distribution for A1 and A2 LinguaDrop news lessons in English, French, Indonesian, and Lithuanian
LanguageCEFR settingLessonsUnique-word shareRepetition shareWords / sentence (median)Words / sentence (range)
EnglishA11075.3%24.7%7.26.2–11.6
EnglishA21076.7%23.3%7.56–12.9
FrenchA11078.5%21.5%8.46.5–12.2
FrenchA21074.8%25.2%10.79.8–12
IndonesianA21077.3%22.7%7.35.9–12.6
LithuanianA11089%11%5.94.7–6.8
LithuanianA21087.5%12.5%7.46.3–8.9

Repetition spread

How many lessons in each band fall into each repetition bucket.

Lesson counts per repetition bucket for each band
BandUnder 10%10% to under 20%20% to under 30%30% and above
English A11423
English A20451
French A10550
French A200100
Indonesian A20442
Lithuanian A16400
Lithuanian A22800

Read a lesson from any band

Each link opens the reading challenge with that language and level requested. If no lesson is playable at that exact band today, the next screen says which one it picked.

Download and cite

LinguaDrop bilingual news vocabulary-density dataset (v1), LinguaDrop, snapshot 2026-09-04. https://linguadrop.com/en/news-readability-benchmark
Version v1 · checksum 8f10c208138d020d · CC BY 4.0. Attribute to LinguaDrop and link to the canonical URL.

Limitations

  • Counts come from automated word and sentence segmentation, not from human annotation.
  • Segmentation is only sanity-checked for Latin, Cyrillic, Greek script text; no logographic, abugida, or unspaced-script band is included in this release.
  • Unique-word share counts surface word forms, so inflected languages score higher than analytic ones for grammatical rather than difficulty reasons. Lithuanian and English are not comparable on this measure alone.
  • Sentence-length figures summarise per-lesson averages across the band, not the length of every individual sentence.
  • These are adapted news lessons, not a general corpus, and the bands are not a formal CEFR calibration or any claim about learning outcomes.
  • The smallest published band contains 10 lessons, so a single unusual lesson moves a band noticeably.
  • Coverage is uneven across languages: en (A1, A2); fr (A1, A2); id (A2); lt (A1, A2). A level missing for a language was not collected in the source snapshot; its absence is not a finding about that band.

Cite or share this benchmark

The link opens the public methodology and data on linguadrop.com. Nothing about your visit is attached.

A fixed tool-only marker helps us count useful shares. It never includes your text, answers, or identity.

Part of the Alfred van der Heide platform

Building tools that make life easier