The Data Center

With multiple interactive frequency dictionaries, the Data Center is like a road map to your target language with tools to help you drill into the details.

Learning a language often feels like exploring unfamiliar territory without a map. You study vocabulary lists, complete exercises, and maybe even pass a proficiency exam, but you're left wondering: how much of the language do I actually know? What should I focus on next? Where are the gaps?

The Data Center is your complete atlas to the language.

Just as an atlas can show you how the continents fit together on one page and city streets on another, the Data Center gives you a bird's-eye view of your target language, while enabling you to drill down into the details. It's built on our analysis of 4 to 5 million sentences, revealing exactly which words, patterns, and verb forms matter most for real conversation.

Whether you're a beginner trying to figure out where to start or an advanced learner hunting down blind spots, the Data Center gives you the complete picture.

Statistical Overview

The first page of the Data Center presents the vital statistics of the text corpus (body of language data) behind our analysis.

Statistics interface for Data Center module

In this example from our English analysis, you can see that we analyzed a corpus containing 5 million sentences with a total of 103 million words. Of these, 205 thousand were unique word forms.

You'll also notice that there were 177 thousand unique lemmas. A lemma is simply a root word: all instances of "goes," "go," "went," and "going" count toward the single lemma "(to) go."

Frequency Distribution Chart

Below the statistics summary is an interactive chart showing how word frequency is distributed across the language.

Use the dropdown to switch between three views: number of unique words, number of unique lemmas, and number of unique verbs. The "words" view counts each form separately (so "goes" and "went" are distinct), while the "lemmas" view groups them under their root.

The x-axis shows frequency rank buckets (groupings of words by how common they are), and the y-axis shows what percentage of total language usage each bucket represents.

For example, the English lemmas chart reveals that the first 90 words cover 50% of all usage, and the next 210 lemmas add another 14%. Simply put, the first 300 words represent, on average, 64% of the words used in a sampling of the language. Frequency drops off quickly, which means mastering the most common words delivers enormous returns.

A cumulative frequency line runs above the bars, showing total coverage up to each point. The green "User Mastery %" markers track your personal progress, giving you a clear visual of how your knowledge compares to the language as a whole.

In the example above, this means that the user has mastered 60% of the 40 most frequent words, and 10% of the verbs ranked 41 - 100.

A strong long-term goal is reaching 100% mastery for each bucket up through the 80% to 90% coverage range.

Dictionaries

Use the tabs at the top of the page to switch between the ‘Statistics’ and the ‘Dictionaries’.

On the Dictionaries page, you can select the Top Words, Top Lemmas, or Top Verbs dictionaries. In the future, we will be adding additional lists such as multiword expression and phrase dictionaries.

Use the search bar to look up specific words, or use the page controls to browse the rankings.

The table displays each word's rank (how common it is), a mastery toggle, the word itself, and its raw frequency count. Taking the English list as an example, the most common word "the" appeared over 5.2 million times across 5 million sentences, while "of", "and", and "to" each appeared about half as often.

The mastery toggle lets you mark words you've internalized and committed to muscle memory. You decide when a word deserves this status. As you mark words mastered, you'll see your progress reflected in the frequency distribution chart.

Definitions

Click on the arrow in the table row in order to view the definitions panel.

In the top group, there is information on pronunciation and usage, while in the bottom is a list of definitions along with sample sentences.

Inflection Tables

If there are inflections for this word, such as conjugations, you can see them by clicking the toggle at the top right.

This page lists all inflection patterns for the word. Use the search bar to filter, and note the counter showing your total practice repetitions across all inflection types.

Click the arrow on any row for a quick view of that inflection table. The info buttons next to each form open a popup with audio, pronunciation details, and the option to add that inflection to a custom course (available with a premium subscription).

Practice Interface

Click on ‘Begin study’ in order to go to the practice interface.

In this screen, you can practice listening and speaking the various inflections of the focus word. The goal is to build muscle memory so you will want to do as many repetitions as necessary to be able to effortlessly produce the pattern yourself.

Once you feel comfortable with the inflections, toggle sentence mode in order to practice sample sentences rather than the isolated inflections.

The sample sentences shown are at the CEFR level that you chose in your course settings. For example, if you are at an A1 level, you will see A1 sentences. This enables the practice to become more advanced as you progress. So even if you are an advanced student who has “mastered” a word, it can be helpful to continue to review even basic words, but in more advanced contexts, to keep pushing your muscle memory to new heights.

Subscribe for evidence based language learning insights.

Lingua Fortis
Copyright © 2025 Lingua Fortis | Made in New York