The Lingua Fortis System

Get better results with data driven language learning.

Jul 20, 2026 By Ryan Ellman
The Lingua Fortis logo on a navy blue background.

The comprehension cliff: why learning a language is so hard

Part of what makes language learning so difficult is just how much we have to learn to be able to understand even the most basic conversations. The reason for this is that comprehension does not degrade gracefully as the number of unrecognized words increases. If even a small percentage of words in a text or conversation are unknown, we can completely lose the ability to understand any of what is being said.

This is why a student can study for an entire year, turn on a film in their target language, and yet hear nothing but noise.

The foundational study on the link between vocabulary coverage (the percent of words in a sample of language content that are known) and comprehension found that when vocabulary coverage was 80%, none of the participants understood what they read.[1] This means that even knowing 8 out of every 10 words was not enough to achieve adequate understanding. Even 90% vocabulary coverage was not found to be adequate for most participants.

An additional study found 95% vocabulary coverage to be the minimal threshold for adequate comprehension and 98% to be the ideal threshold.[2]

Failing to reach these admittedly high thresholds prevents you from using your target language in ‘real world’ situations such as having a conversation beyond a basic level or even just consuming native content such as films and books. This is even more frustrating because engaging with native content is itself one of the most powerful language learning methods available.

So language learning is hard, in large part, because so many learners spend years studying without ever crossing these thresholds, meaning they never graduate to the point where they can learn by using their language in the real world.

The goal is not as distant as it seems

But there is a fundamental feature of language that makes these thresholds far less daunting than they first appear: a surprisingly small number of words makes up the vast majority of all words we use. Whether in spoken language, academic papers, movie scripts or novels, as few as 1,000 words generally comprise about 80% of the words used. This pattern holds true for essentially all languages, despite the fact that any given language contains enormous amounts of words.

The Oxford English Dictionary contains over 500,000 entries.[3] Looking only at root words (also called lemmas) reduces this count to ~60% of the original number, or 300,000. Meanwhile, adult native speakers are estimated to only know approximately 42,000 lemmas by the age of 20.[4] Productive vocabulary, words known well enough that the speaker can use them rather than simply recognizing them, might be only half that, or 21,000 lemmas, less than 10% of the number of lemmas in the dictionary.[4]

But even this much reduced number is enormous compared to the limited vocabulary necessary for daily use.

Nation and Waring estimate that a vocabulary of only 2,000 to 3,000 word families is necessary for "productive use in speaking and writing."[5] Word families are a bit broader than lemmas, so to get a count comparable to the 21,000 lemmas figure I gave before, we can look at the analysis I did when building Lingua Fortis. I found that in a collection of 5,000,000 English sentences, with a combined 103,000,000 words, there were 177,000 unique lemmas. But only 1,000 lemmas represented 80% of all words found in the text and 3,000 lemmas were sufficient to cover 90% of the text.

This skew has been observed across languages and is described as following Zipf's Law, which states that the frequency of an item in a large data set is inversely proportional to its rank.

Put more simply, we could describe this as a Pareto distribution, where a small number of the words make up the vast majority of words that are used.

This is great news for language learners because this drastically decreases the number of words that we need to learn in order to use our target language. While the number of words in a language is massive, the actual number of words necessary to use the language in daily life is both small and attainable.

Even better news for language learners is the finding that spoken language requires fewer words than written.[6] Given that the primary goal of many language learners is be able to use their target language in conversation, this further lowers the number of words needed to reach their goal.

Because of this natural skew found in word frequency, reaching the critical thresholds is not as daunting as it would seem. This does not require you to master tens of thousands of words to start using the language. It only requires you master the right few thousand.

Prioritizing the right words

Therefore, to reach these critical thresholds as efficiently as possible, it is crucial to prioritize mastering this relatively small subset of high frequency words. Mastering the highest frequency words first increases the chances that we have the minimum necessary vocabulary coverage for any situation in which we are using the language.

Prioritizing this way creates a virtuous cycle. Because every high frequency word appears constantly, each one you master produces a noticeable jump in how much you understand. You feel your fluency rising, and that feeling is real, not an illusion of progress. That feedback is motivating in a way that no grade or completed chapter can be.

Additionally, once you know this high frequency subset of words, you cross the threshold necessary for native content to be useful as a learning material. You now know enough of the surrounding words to interpolate those you don't. The same movies and conversations that were noise become a nearly unlimited training ground.

This is not a new discovery. As far back as 1997, Nation and Waring stated that the first 3,000 high frequency words "are an immediate high priority and there is little sense in focusing on other vocabulary until these are well learned."[5] Nation further argues that low frequency words do so little to enhance vocabulary coverage that it is not even worth dedicating classroom time to studying these at all. Instead, by mastering the high frequency words, the students will have the necessary knowledge to be able to learn low frequency words from context when they consume content in there target language.

Therefore, one would expect that language learning materials, particularly those aimed at beginners, would be hyper focused on teaching students the highest frequency words as this is the quickest path to being a functional user of the language.

The problem with traditional language learning materials

But studies have found that many language learning materials in fact do the exact opposite.

They prioritize low frequency words and underemphasize the highest frequency words, leading to an inefficient use of time and lengthening the path to functional fluency.

In the earliest chapters of many language learning textbooks, students are given vocabulary lists that include words that intuitively seem like they would be common, but are in fact surprisingly rare. A study of popular first year German textbooks found that between 29% and 44% of all the vocabulary covered had a frequency rank greater than 4,000, meaning they aren't just uncommon, they are actually exceedingly rare.[7] Not only did these textbooks overemphasize rare words, they did a poor job of covering very common words, with only 53% to 64% of the top 1,000 most common words being included in the text books.[8] This means that a student relying on these very popular textbooks would finish their first year of study knowing barely half the words they would need on a daily basis.

This finding has been replicated in similar studies looking at other languages such as Spanish[9], English as a Foreign Language textbooks targeting the Chinese market[10] and more. Several studies have found the same pattern of outsized coverage of rare words and inadequate coverage of the most common words.

This creates frustration because even the most unrealistically ideal student, one who memorized all the vocabulary in their textbook, would find themselves unable to understand common everyday conversations due to the comprehension cliff, all while having a wide array of specialist vocabulary with neither the opportunity nor ability to use these words in a real world setting.

This creates a vicious cycle. The student puts in more effort, they learn the vocabulary they are taught, they feel like they are making progress, but they do not get much closer to achieving the goal that most language learners have: to use their target language in the 'real world.' The dissonance between how they feel they are doing in their studies and their lack of real world capability is demoralizing. The student never gets the positive feedback of real world language success that would encourage them to go further, and they give up, often thinking they are 'just bad at languages' or 'too old to learn a new one.'

A data driven approach

Therefore students should use a more data driven approach that prioritizes words based on their observed frequency in natural use.

The simplest version of this approach might be to use word frequency lists. There are many such lists available online covering a wide range of languages, and they have existed for a long time. Starting at the top and memorizing the list would theoretically give the learner the necessary vocabulary to use the language in the real world and provide much better prioritization than what many language learning materials offer.

Memorization is insufficient

But simply memorizing frequency lists is insufficient because this misses the mark about what it means to "know a word"

In an ideal world, you would simply be able to memorize a comprehensive list of the most frequent 3,000 words in a language and next thing you know, you would be out using the language in the real world: watching movies, traveling, making friends in a foreign country, etc. All the things we imagine when we set out to start studying a new language.

Unfortunately, things are not so simple. To understand why, we have to ask, what does it mean to know a word?

A naïve response might say that knowing a word means knowing its translation in a language you already speak. We don't usually say that, but this is exactly how many of us attempt to learn new vocabulary in another language. There are two problems with this approach.

The first is that while our existing vocabulary can serve as a helpful bridge for gaining a quick understanding of another word, words generally do not have a one-to-one relationship with a word in another language.

This is especially true of the most common words in a language, which are more likely to represent abstract concepts as they are the building blocks of the structure of the language.

Take the French words de and à, by our counts, the 2nd and 5th most common words in French, respectively. Generally, students are taught that 'de' means 'of' or 'from' and 'à' means 'to' or 'at.' The reality is that, while these words are sometimes used in a way that matches the usage of these English equivalents, they have far more functions in the French language that don't conform to these translations.

  • Je suis à l'heure. → I'm on time.
  • Elle vient à vélo. → She is coming by bike.
  • Elle est en train de lire. → She's reading.
  • Je vais essayer de comprendre. → I'm going to try to understand.

Notice that this works in both directions. A French student learning English will probably learn that 'on' translates to 'sur' or 'by' translates to 'par'. Oftentimes they do, but even these simple examples show just how often that will not be the case.

The second problem is that words do not exist in isolation in their own languages. They are much more likely to be used with certain other words. This is demonstrated by analyses that have found that, in the same way that certain words are much more common than others, certain word clusters are far more frequent than other clusters as well. One study looking at English found that more than half of a given text will consist of prefabricated word combinations.[11] 'Knowing' a word also means knowing the correct words to use with it, something that simple frequency lists do not provide.

As an example, take the French words de and à again in the following sentences:

Il essaie de finir ses devoirs. → He is trying to complete his homework.

Here, de functions as "to" so maybe to can also translate to "de" when connecting auxiliary verbs to infinitives? But then take the following sentence:

Elle cherche à améliorer son français. → She is trying to improve her French.

Here, 'to' is translated as à. And tries is no longer essayer but chercher? Chercher is usually translated as 'to search.'

A native speaker will reflexively know that 'essaie' should be followed by de in the first sentence and 'cherche' should be followed by 'à' in the second. So knowing a word means much more than knowing its definition, or knowing its translation. To truly know a word, you have to have an instinct for its usage, including the correct words to use with it.

Instinct is the goal

Therefore, the goal is not to simply memorize the highest frequency words in a language, but to develop an instinct for how they are used.

While there are sometimes 'rules' that describe the logic behind why certain patterns are considered correct and others are not, I would argue that these rules are nothing more than descriptions of the expectations of native speakers regarding usage. "Il essaie de.." sounds correct and "Il essaie à…" does not, not because there is a rule determining correct usage, but because the native speaker has heard the former hundreds of thousands of times and the latter very rarely (perhaps only from people learning the language). The rule describes the instinct of native speakers and the instinct is developed through massive amounts of exposure to certain patterns and virtually no exposure to others.

This is why memorizing grammar rules often does not lead to better speaking ability. I believe that grammar rules are just formal descriptions of the expectations of native speakers, and that these expectations are themselves formed by being exposed to massive amounts of repetitions of certain "correct" patterns and minimal exposure to "incorrect" patterns.

Memorizing grammar rules overestimates our ability to spontaneously apply these patterns ourselves in new situations. Rather, if we want to reflexively use words "correctly," that is, follow the patterns expected by native speakers, we should seek to develop the same instincts by exposing ourselves to, and repeating aloud, these patterns an enormous number of times.

Lingua Fortis: Data driven and muscle memory focused

I built Lingua Fortis because I wanted a program built around these two fundamental truths about how we use language: that a small subset of the language represents the vast majority of the language we use, and being able to use the language properly comes from developing an instinct for it. My experience with many other apps is that they fall into the same pitfalls as textbooks. They give too much priority to measurably rare words, and they focus on teaching about the language rather than building an instinct for it.

The philosophy of Lingua Fortis is that fluency comes from muscle memory and muscle memory comes from repetition. Prioritizing exposure to the highest frequency structures in the language and developing the muscle memory to produce them spontaneously ourselves, it the most effective path to fluency.

Lingua Fortis offers many different ways to study rather than a single fixed course you follow from start to end. The analogy I always use is that it is more like a gym membership than a fitness class. You decide what equipment to use and how intensely to train. What unites every feature is that each one is data driven and muscle memory focused.

Everything starts with the data. To build the content for a language, we analyze around 5,000,000 sentences drawn from a variety of sources, extract the statistically most important words and patterns, and organize them by frequency.

Standard Courses: Built directly from that analysis, the courses drill the most common words and word combinations through repeated exposure and spaced repetition review. Every focus word is shown in full sample sentences alongside the words it most frequently appears with, never in isolation. This is the first stop for beginners: each pattern is drilled to mastery, starting with the most common, so the fundamental structures of the language are solidified in your muscle memory before anything else is added on top of them.

Rapid Sentence Exposure: This module is designed to give you rapid, broad exposure to the 15,000 most common words in the language through sample sentences adjusted to your current level. The routine is always the same: see the sentence, hear it, repeat it out loud. The goal is volume, encountering the vocabulary that carries you toward the critical thresholds and seeing such a wide variety of grammatical structures that your mind begins internalizing them the way a native speaker's does: through sheer quantity of examples.

Data Center: A set of interactive frequency dictionaries that puts the statistical landscape of the language directly in front of you. Every word comes with definitions, inflection tables, and muscle memory drills built on sample sentences adjusted to your CEFR level. Because the program tracks the words you have mastered, the frequency chart becomes a map of your own knowledge. You can see exactly where you stand relative to the critical comprehension thresholds and exactly which high frequency words you are still missing. Beginners use it to control precisely what they study. Advanced learners use it to find the gaps that years of unprioritized study leave behind.

Interactive Lessons: Even when we teach grammar or specific scenarios, the format stays true to the method. Each lesson includes an article with a full explanation, but the core of the lesson is a drill: a large volume of sentences demonstrating the concept, shown to you and repeated out loud, plus an AI chat where you can practice the concept in open conversation. As I argued above, a rule is only a description of a pattern. The lesson exists to install the pattern, not the description.

High Context Translator: Part of learning a language is understanding the different ways to express a similar idea. Our translator gives multiple translation options and explains the differences between them, precisely because words do not map one to one between languages. And since the things you look up are, by definition, the things you personally need to say, the translator has spaced repetition review built in, so your lookups get committed to long term memory instead of forgotten by tomorrow.

Lingua Fortis works for both beginners and advanced learners. You choose your current level and the content adjusts. For beginners, the goal is to get you as quickly and efficiently as possible past the inflection point where native content becomes usable, by having you master the fundamental structures of the language first. For advanced learners, the goal is to fill the important gaps in your knowledge and develop a truly instinctual relationship with the language.

I used Lingua Fortis myself as an advanced learner of French, and I now find myself spontaneously producing grammatical structures I did not consciously know I knew. I "feel" what is right in the language rather than simply "know" what is right. That shift, from knowing to instinctual feeling, is the goal of the method.

The Lingua Fortis method comes down to this: gather an enormous amount of data on how a language is actually used, break it down into its fundamental structures, prioritize by what is objectively most important, and commit those patterns to muscle memory through high volumes of repetition.

Learning a language is hard no matter what you do. For many people it is one of the most complex long term projects they will ever take on. What Lingua Fortis promises is that your effort will turn into results, because you will be relentlessly training on the most important content in the language. After thousands of repetitions, you will find yourself effortlessly saying things in your target language you didn't even realize you had in you.


References
[1] Marcella Hsueh-chao Hu and I. S. P. Nation, "Unknown Vocabulary Density and Reading Comprehension," Reading in a Foreign Language 13, no. 1 (2000): 403–430.
[2] Batia Laufer and Geke C. Ravenhorst-Kalovski, "Lexical Threshold Revisited: Lexical Text Coverage, Learners' Vocabulary Size and Reading Comprehension," Reading in a Foreign Language 22, no. 1 (2010): 15–30, https://files.eric.ed.gov/fulltext/EJ887873.pdf.
[3] "About the OED," Oxford English Dictionary, Oxford University Press, accessed July 2026, https://www.oed.com/information/about-the-oed.
[4] Marc Brysbaert, Michaël Stevens, Paweł Mandera, and Emmanuel Keuleers, "How Many Words Do We Know? Practical Estimates of Vocabulary Size Dependent on Word Definition, the Degree of Language Input and the Participant's Age," Frontiers in Psychology 7 (2016): 1116, https://doi.org/10.3389/fpsyg.2016.01116.
[5] Paul Nation and Robert Waring, "Vocabulary Size, Text Coverage and Word Lists," in Vocabulary: Description, Acquisition and Pedagogy, ed. Norbert Schmitt and Michael McCarthy (Cambridge: Cambridge University Press, 1997), 6–19, https://www.lextutor.ca/research/nation_waring_97.html.
[6] I. S. P. Nation, "How Large a Vocabulary Is Needed for Reading and Listening?," Canadian Modern Language Review 63, no. 1 (2006): 59–82, https://doi.org/10.3138/cmlr.63.1.59.
[7] To give a sense of just how rare these words are, from Lingua Fortis' analysis of English, the 4,000th most common word appeared 1,238 times in 5,000,000 sentences, or approximately once every 4,000 sentences. For comparison, the 1,000th most common word appeared every 526 sentences, the 500th every 263 sentences, and the 100th every 60 sentences.
[8] Silke Lipinski, "A Frequency Analysis of Vocabulary in Three First-Year Textbooks of German," Die Unterrichtspraxis/Teaching German 43, no. 2 (2010): 167–174, https://www.jstor.org/stable/40961806.
[9] Mark Davies and Timothy L. Face, "Vocabulary Coverage in Spanish Textbooks: How Representative Is It?," in Selected Proceedings of the 9th Hispanic Linguistics Symposium, ed. Nuria Sagarra and Almeida Jacqueline Toribio (Somerville, MA: Cascadilla Proceedings Project, 2006), 132–143, https://www.lingref.com/cpp/hls/9/paper1373.pdf.
[10] Yang Sun and Thi Ngoc Yen Dang, "Vocabulary in High-School EFL Textbooks: Texts and Learner Knowledge," System 93 (2020): 102279, https://doi.org/10.1016/j.system.2020.102279.
[11] Britt Erman and Beatrice Warren, "The Idiom Principle and the Open Choice Principle," Text 20, no. 1 (2000): 29–62, https://doi.org/10.1515/text.1.2000.20.1.29.
Ryan

Ryan Ellman

en flagfr flag

Ryan Ellman is the founder of Lingua Fortis. He has a passion for languages, travel, and fitness. He speaks English and French and is currently learning Spanish.

Subscribe for evidence based language learning insights.

Lingua Fortis
Copyright © 2025 Lingua Fortis | Made in New York