Shanraq.org Shanraq.org
Three Hundred and Eighty-Six People
Culture

Three Hundred and Eighty-Six People

A language was safe as long as enough people spoke it. That rule has stopped working: machines learn from text, and a machine's ability to speak a language depends on how much has been written in it. Iceland, with 386,000 people, has 215 active encyclopedia editors; Kazakhstan, with twenty million, has just 386.

In July 2026 the Kazakh Wikipedia took 13,099 edits. That is its best month in seven years: in January 2019 there were 6,325. Twice as many.

Good news, if there is nothing to compare it with.

Over the same seven years the Uzbek edition went from 3,608 edits a month to 31,244 — a factor of eight and a half. We doubled; our neighbour multiplied ninefold.

But that is still not the number this piece was written for.

The Number

Every Wikipedia edition publishes a line called “active users” — how many people performed at least one action in the last thirty days. The data is open and served from an address anyone can check.

In the past month, 386 people wrote in Kazakh.

Not three hundred and eighty-six thousand. Three hundred and eighty-six. Everyone currently adding to the main open encyclopedia in the Kazakh language would fit into a single lecture hall.

For comparison: the Icelandic edition has 215 active editors. Iceland is a country of 386,506 people, smaller than Aktobe.

The Whole World in One Table

We took fifteen countries whose state language has no second homeland — that is, countries where nobody but themselves will write the corpus in that language. For each we calculated two figures: Wikipedia articles per thousand inhabitants, and active editors per million. The data is Wikipedia’s own statistics and World Bank population; the calculation can be reproduced in half an hour.

Country Population Articles Editors Editors per million
Israel 10,002,200 403,338 6,760 676
Estonia 1,372,341 261,396 904 659
Iceland 386,506 61,230 215 556
Finland 5,619,911 623,439 2,970 528
Slovenia 2,127,400 198,754 535 251
Latvia 1,866,124 145,111 412 221
Armenia 3,033,500 331,239 659 217
Lithuania 2,888,278 226,585 564 195
North Macedonia 1,824,359 163,625 291 160
Albania 2,377,128 106,179 345 145
Georgia 3,812,518 198,136 360 94
Azerbaijan 10,202,830 216,714 879 86
Kazakhstan 20,592,571 244,775 386 19
Uzbekistan 36,361,859 356,701 542 15
Kyrgyzstan 7,221,868 76,524 98 14

The last three rows are the whole of Central Asia, one after another. This is not about us alone.

And one detail worth reading twice: Estonia has written more articles than Kazakhstan. 261,396 against 244,775. Fewer people live in Estonia than in Almaty.

What Does Not Explain It

Poverty? Doesn’t hold. Kazakhstan has $14,155 per capita. Armenia has $8,556 — forty per cent less — and 217 editors per million against our nineteen. Georgia has $8,968 and 94 editors. We are richer than the countries that write more.

Internet access? Holds even less, and this is the most uncomfortable line in the whole calculation. 93% of Kazakhstan’s population has internet access — more than Estonia (92%), more than Israel (88%), considerably more than Armenia (81%). We are among the best-connected countries in the table and near the bottom of it for what has been written. The problem is not that people cannot get online.

The size of the country? Israel is half our size and sustains 676 editors per million — thirty-five times our figure. Azerbaijan has half our population and writes four and a half times more actively.

The alphabet? The theory looked convincing: a country changes its script, the corpus splits in two, and the machine sees two different languages instead of one. But it breaks on the neighbour. Uzbekistan has been moving its language to Latin script since 1993 — thirty-three years, longer than us — and it is precisely Uzbek edits that grew eight and a half times. Changing the alphabet hurts us, but it cannot explain a thirty-five-fold gap with Estonia.

Not one of the obvious explanations survived checking.

What We Are Actually Measuring

The state measures language in percentages. Here is what those percentages say.

From the Concept for the Development of Language Policy in the Republic of Kazakhstan for 2023–2029, approved by government decree:

  • share of state document flow in the state language: 93% in 2020, 96% in 2021, 96.5% in 2022;
  • by sociological survey, command of the state language: 90.5% in 2020, 91% in 2021, 92% in 2022;
  • by the 2021 national census, 13,768,406 people — 80% have command of the state language;
  • and, one line further down in the same document: 8,472,661 people — 49.3% — use it in daily life.

Read those last two lines again. One document, one page, one and the same subject. Command: eighty per cent. Use: forty-nine. And a sociological survey in the same year says ninety-two.

Ninety-two, eighty and forty-nine — three official figures about the same thing. All three cannot be right.

The difference between them is the difference between what a person tells a questionnaire and what they do when there is no questionnaire.

Where the Documents Went

Now put the 96.5% together with what we counted above.

If a state produces almost all of its document flow in Kazakh for nine consecutive years, a mountain of Kazakh text should have accumulated: laws, regulations, decisions, correspondence, reports, expert opinions. That is the kind of text that makes a language a working language — not songs and not poetry, but the dull prose of decision-making. Every modern language rests on it.

The mountain is not there.

Under 0.1% of websites. Two hundred and forty-four thousand articles — fewer than Estonia’s. Three hundred and eighty-six people writing the encyclopedia. A depth of 25.9 — a corpus nobody came back to.

One of two things. Either the documents exist and sit where nobody can see them — in which case, for the language, they might as well not exist. Or the 96.5% is the share of documents issued in Kazakh rather than written in it: prepared in another language and translated before signature.

The second explains everything. A translated document counts towards the 96.5% and adds nothing to the language: it is derivative, formulaic, and often machine-translated. The mountain of reporting grew; the mountain of language did not.

Where a Corpus Is Actually Born

A language lives where decisions are taken in it.

Look at the top of our table. Israel — 676 editors per million, first place. A century ago Hebrew was a liturgical language nobody spoke at home. It came alive not because its status was written into law, but because people began working in it: holding meetings, judging cases, treating patients, fighting wars, writing scientific papers. The elite switched to the language, and the corpus followed.

Estonia, Iceland, Finland — the same. There is no separate programme there for filling up Wikipedia. There is simply no second language in which it would be more convenient to hold a meeting.

We have one, and for thirty-five years it has remained the working language of decision-making. Not by law — by habit.

You can see it without any statistics; it is enough to turn on the television. For thirty-five years the public address of a leader has been built the same way: a few sentences in the state language at the start, and then two or three hours in Russian. The first president spoke like that. They speak like that today — from the head of state down to the akims of regions and districts. The ritual is observed; the work happens in another language.

Those few opening sentences are the 96.5% of document flow and the 80% of command. The two or three hours that follow are the 49.3%.

The gap between the figure and the life, which we examined a paragraph ago, is broadcast every evening.

This is not a question of patriotism and it is not an accusation. It is mechanics. A language in which the decision-makers do not think produces no text — it produces translation. And translation, as we have just calculated, barely enters a corpus at all.

While the top of the country works in one language and reports in another, the figures will look exactly as they look. Nineteen editors per million. Third from the bottom.

Why This Stopped Being a Matter of Pride

For all of history a language was safe if enough people spoke it. A language with thirteen million speakers could not disappear — there were not enough generations for it.

That rule has stopped working, and it stopped within our lifetime.

The tools of the next decade — translation, search, dictation, assistants, speech recognition — do not learn a language. They reconstruct it statistically from text. A machine’s ability to speak Kazakh is a function of how much Kazakh text made it into the training. Not of the number of speakers. Not of the language’s status in the constitution.

According to W3Techs on 21 August 2026, the share of websites with content in Kazakh is under 0.1%. For Estonian it is 0.1%. For Russian, 3.4%; for English, 49.5%.

Estonian, spoken by one million people, is more widely present on the web than Kazakh, spoken by thirteen million.

From which follows a conclusion worth stating without euphemism: machines will speak Estonian better than Kazakh. Not through anyone’s malice. Through arithmetic.

One More Measure, and It Is Worse

Article count is a crude measure: it counts a three-line stub and a ten-thousand-character analysis alike. So Wikipedia has its own measure of how alive an edition is — “depth”: edits per article, adjusted for the share of non-article pages. It shows not how much has been written, but how much of what was written is still living.

Edition Edits per article Depth
Hebrew 108.4 343.5
Azerbaijani 41.7 90.2
Armenian 32.8 87.1
Finnish 38.7 60.1
Icelandic 32.2 52.6
Uzbek 17.5 45.7
Estonian 27.6 37.6
Kazakh 14.9 25.9
Kyrgyz 8.5 5.5

A Kazakh article has been edited on average fifteen times in its whole life, an Estonian one twenty-eight, a Hebrew one a hundred and eight. Our corpus is not merely smaller. It is thinner: a substantial part of it is stubs nobody returned to.

What Comes Next

The decree transferring the Kazakh alphabet to Latin script was signed in 2017. The transition was to be completed by 2025; the deadline was then moved to 2023–2031.

Those are exactly the years in which the models our children will use are being trained.

Everything written in Kazakh over the past thirty years is written in Cyrillic. To a machine, Cyrillic Kazakh and Latin Kazakh are not one language in two scripts but two different sets of symbols, unless a large parallel corpus sits between them. There is no parallel corpus: nobody has compiled one.

So we are splitting a small corpus in two — in precisely the years when it ought to be doubled.

Who These 386 People Are

Nobody knows. They are volunteers. They have no salary line, no department and no reporting. Not one state programme assigns them tasks or counts their work as its own result.

Meanwhile real money is spent on the language: television quotas, compulsory dubbing, signage, examinations, civil-service certification. All of it produces compliance. Not one of those mechanisms produces text that anyone wants to read.

We built an apparatus for compelling the language and not a single institution for producing it.

What to Do About It

Three things follow, and all of them are boring.

First. Start counting volume rather than percentages of command. How many characters of Kazakh appeared in open access over a year, and how many of them are original text rather than translation or reprint. This is measured automatically and costs nothing.

Second. Digitise what has already been written. Over the Soviet and post-Soviet years, tens of thousands of books, dissertations and newspaper runs were published in Kazakh. All of it sits in archives and does not exist for a machine. Digitising the archive is the only way to obtain a large corpus without writing a single new line.

Third, and it is almost indecently cheap. Three hundred and eighty-six people do this work for free. Doubling them is a task costing about one road interchange. Of all the country’s problems this may be the cheapest to solve: it cannot be closed by decree, but it can be closed by paying salaries to roughly a thousand people.

And a fourth, which will appear in no programme.

None of the three will work on its own, because a language’s corpus is the sediment of working in it. It forms where people argue about a budget, conduct an investigation and write a technical specification in that language. Not where they unveil a monument in it.

For thirty-five years we grew percentages — of command, of document flow, of dubbing, of certification. All of them rose: 92%, 96.5%, 80%. And here is the result none of those percentages shows: nineteen people writing per million inhabitants, fewer articles than Estonia, and under 0.1% of the world wide web.

Percentages are compiled by the people they are used to judge. A corpus is compiled by nobody — it simply exists or it does not.

All week the world discussed Nvidia preparing to back some hundred billion dollars of OpenAI financing. Not one model trained on that money will speak Kazakh better than the corpus allows — the corpus that three hundred and eighty-six volunteers are writing today.

And as long as those who take decisions in this country take them in another language, there will be three hundred and eighty-six volunteers.

Sources

If you have found a mistake or a typo in this article, tell us about it

Comments (0)

No comments yet. Be the first.