Open methodology · 70 lessons
How does A1 news differ from A2 news?
We measured a fixed sample of public LinguaDrop lessons in English, French, Lithuanian, and Indonesian. The result is a descriptive benchmark of text length, sentence size, and vocabulary variety, not a claim that a formula can validate CEFR proficiency.
70
public lessons
4
target languages
10
lessons per language and level
A1 and A2 results
Snapshot: 2026-09-04. Each row is a mean over 10 lessons, collected on the date shown in its row.
| Language | CEFR setting | Lessons | Words / lesson | Words / sentence | Unique-word share | Collected |
|---|---|---|---|---|---|---|
| English | A1 | 10 | 69.5 | 8.6 | 78.7% | 2026-08-16 |
| English | A2 | 10 | 77.6 | 7.8 | 76.8% | 2026-08-16 |
| French | A1 | 10 | 61 | 9.1 | 79.4% | 2026-08-16 |
| French | A2 | 10 | 95.6 | 10.9 | 75% | 2026-08-16 |
| Indonesian | A2 | 10 | 67.9 | 7.8 | 78% | 2026-09-04 |
| Lithuanian | A1 | 10 | 41.8 | 5.8 | 89.1% | 2026-08-16 |
| Lithuanian | A2 | 10 | 65.6 | 7.5 | 87.8% | 2026-08-16 |
What the snapshot does, and does not, show
A2 lessons are longer than A1 lessons in every language that has both bands. Sentence length does not rise uniformly: the English A2 sample has fewer words per sentence than its A1 sample. That is why we publish the rows instead of compressing them into a universal “readability score.”
Reading the Indonesian A2 row
Indonesian is the one language here with no A1 row, because LinguaDrop publishes no Indonesian A1 lessons. There is nothing to sample, so nothing is shown. The useful comparison is sideways, against the other A2 samples: Indonesian A2 averages 67.9 words per lesson and 7.8 words per sentence, which is the same sentence length as English A2 and noticeably shorter than French A2 at 10.9.
Its unique-word share of 78.0% sits between French and Lithuanian, and that gap is mostly grammar rather than difficulty. We count distinct word forms, so a heavily inflected language such as Lithuanian scores high because one lemma appears in many forms, while Indonesian inflects far less and its forms are closer to distinct words. Compare unique-word share within a language, not across languages.
Read Indonesian news lessons for yourself and judge the numbers against the text.
Reproducible methodology
- 1
Freeze a bounded public sample
We selected the 10 most recent lessons in each published language and level band, scanning the public API in pages until every band was filled. Bands are added over time and each lesson records the date it was collected, so an added language does not restate figures already published. The rule, lesson IDs, public URLs, and collection dates are checked in.
- 2
Measure only the adapted text
The parser keeps target-language sentences and removes the interleaved source translation and vocabulary gloss lines. No user data, private corpus, or original full news article is included in the download.
- 3
Use language-aware segmentation
JavaScript’s
Intl.Segmentercounts words and sentences using the target-language locale. Unique-word share is lowercased orthographic types divided by words. - 4
Generate the page from the snapshot
A deterministic repository script calculates all six table rows and the CSV. The test suite fails if generated numbers drift from the checked-in source snapshot.
Inspect representative lessons
Every CSV row links to the complete public lesson and its publisher credit. These six links provide one example from each measured band.
As Islamophobia rises, Australia's Muslims celebrate Eid
69.5 mean words per lesson in this band
Filial de CK Hutchison en Panamá solicitará arbitraje por adquisición de puertos de Maersk
77.6 mean words per lesson in this band
Police 'missed' Noah Donohoe on CCTV filmed minutes before disappearance
61 mean words per lesson in this band
Almost 500 arrested on suspicion of starting wildfires in France this summer
95.6 mean words per lesson in this band
Iranians Condemn Strike on a Top University
67.9 mean words per lesson in this band
What Is Happening at the Border in Big Bend National Park?
41.8 mean words per lesson in this band
Ann Widdecombe Was Killed in ‘Targeted Attack,’ UK Police Say
65.6 mean words per lesson in this band
Find your starting level
Take the five-question CEFR reading test, then continue into a matching free lesson.
Check your own passage privately
Compare its sentence length, repetition, vocabulary variety, and long-word share without uploading or saving the text.
Explain a passageTurn a lesson into classroom practice
Use bounded bilingual excerpts, vocabulary, and original comprehension prompts in a printable worksheet. No account or generated answer storage.
Make a free worksheetOpen dataset v1
Vocabulary density by language and level
Aggregate vocabulary variety, repetition, and sentence-length spread across 70 public lessons, derived from the same snapshot as the benchmark above. Aggregates only: no article text, lesson id, or URL is published.
| Language | CEFR setting | Lessons | Unique-word share | Repetition share | Words / sentence (median) | Words / sentence (range) |
|---|---|---|---|---|---|---|
| English | A1 | 10 | 75.3% | 24.7% | 7.2 | 6.2–11.6 |
| English | A2 | 10 | 76.7% | 23.3% | 7.5 | 6–12.9 |
| French | A1 | 10 | 78.5% | 21.5% | 8.4 | 6.5–12.2 |
| French | A2 | 10 | 74.8% | 25.2% | 10.7 | 9.8–12 |
| Indonesian | A2 | 10 | 77.3% | 22.7% | 7.3 | 5.9–12.6 |
| Lithuanian | A1 | 10 | 89% | 11% | 5.9 | 4.7–6.8 |
| Lithuanian | A2 | 10 | 87.5% | 12.5% | 7.4 | 6.3–8.9 |
Repetition spread
How many lessons in each band fall into each repetition bucket.
| Band | Under 10% | 10% to under 20% | 20% to under 30% | 30% and above |
|---|---|---|---|---|
| English A1 | 1 | 4 | 2 | 3 |
| English A2 | 0 | 4 | 5 | 1 |
| French A1 | 0 | 5 | 5 | 0 |
| French A2 | 0 | 0 | 10 | 0 |
| Indonesian A2 | 0 | 4 | 4 | 2 |
| Lithuanian A1 | 6 | 4 | 0 | 0 |
| Lithuanian A2 | 2 | 8 | 0 | 0 |
Read a lesson from any band
Each link opens the reading challenge with that language and level requested. If no lesson is playable at that exact band today, the next screen says which one it picked.
Download and cite
LinguaDrop bilingual news vocabulary-density dataset (v1), LinguaDrop, snapshot 2026-09-04. https://linguadrop.com/en/news-readability-benchmark
8f10c208138d020d · CC BY 4.0. Attribute to LinguaDrop and link to the canonical URL.Limitations
- Counts come from automated word and sentence segmentation, not from human annotation.
- Segmentation is only sanity-checked for Latin, Cyrillic, Greek script text; no logographic, abugida, or unspaced-script band is included in this release.
- Unique-word share counts surface word forms, so inflected languages score higher than analytic ones for grammatical rather than difficulty reasons. Lithuanian and English are not comparable on this measure alone.
- Sentence-length figures summarise per-lesson averages across the band, not the length of every individual sentence.
- These are adapted news lessons, not a general corpus, and the bands are not a formal CEFR calibration or any claim about learning outcomes.
- The smallest published band contains 10 lessons, so a single unusual lesson moves a band noticeably.
- Coverage is uneven across languages: en (A1, A2); fr (A1, A2); id (A2); lt (A1, A2). A level missing for a language was not collected in the source snapshot; its absence is not a finding about that band.
Cite or share this benchmark
The link opens the public methodology and data on linguadrop.com. Nothing about your visit is attached.
A fixed tool-only marker helps us count useful shares. It never includes your text, answers, or identity.

