Look Up Every Arabic Word or Guess? What Research Says
Guesses from context are fully right about a quarter of the time, and easy reading needs 98% of words known. When to look up an Arabic word, and when not to.
· Written by Amaan
Should I look up every word when reading Arabic?
No, and do not guess every word either. Look up the words that block the sentence and the words that keep coming back, and skip the rest. In a study of intermediate learners, only 25.6% of guesses from context were fully right, and comfortable reading needs about 98% of words known. Unvowelled Arabic makes guessing harder still.
Look up the words that block the sentence or keep coming back, and skip the rest. Guessing alone fails more often than learners expect: in Nassaji's 2003 study, 25.6% of guesses from context were fully right.
Reading research puts comfortable reading at about 98% of the words on a page known, one unknown word in fifty, and the floor at about 95%. Below that, context has too little to work with.
Arabic makes guessing harder. Printed without vowels, كتب can be “he wrote”, “it was written” or “books”, and only the words around it decide.
The cost of a lookup decides the strategy. With a paper dictionary ordered by root, learners stop looking words up; when a lookup costs one tap, selective lookup is the sensible default.
If you are marking more than about one word per line, the text is above you for now. That is our rule of thumb rather than a research finding, and it applies the numbers above.
How many words do you need to know to read a page?
About 98% of the words on the page, by the most-cited estimate: one unknown word in fifty. Laufer and Ravenhorst-Kalovski (2010) put the minimum for adequate comprehension at about 95% of words known and the optimum at 98%. Hu and Nation (2000) found that most readers needed about 98% to understand a short story without help.
Treat those numbers as proportions, not as a law. They were measured on English, with vocabulary counted in word families. When Kremmel and colleagues repeated Hu and Nation's study in 2023 with 104 adult learners in Sri Lanka, readers at 98% still fell short of adequate comprehension by the original definition. The safer reading is a slope rather than a cliff: every percentage point of known words helps, and there is no line above which reading suddenly becomes easy.
Nation (2006) estimated that 98% of general written English takes a vocabulary of 8,000 to 9,000 word families. Arabic counts differently, because one root carries many words and one word appears in many attached forms, so the English totals do not transfer as they stand. What does carry over is the proportion: a text that hides an unknown word in every line is not yet a reading text; it is a dictionary exercise.
How often does guessing from context work?
Less often than the usual advice suggests. In Nassaji's 2003 study, 21 intermediate learners of English tried to work out unknown words from context: 25.6% of their guesses were fully right, 18.6% partly right and 55.8% wrong. Guessing works best exactly where it is least needed, in a text where nearly every other word is already known.
That is not an argument against guessing. It is an argument against guessing as the only strategy, and against the popular advice to read texts you understand only 80 to 90% of. If that means 80 to 90% of the words, an unknown word arrives on almost every line, and each guess leans on a context that is itself full of holes.
A wrong guess also costs more than it seems. It does not feel wrong at the time, so it can settle in memory as the meaning. A quick check afterwards turns a guess into a word you know, or into a correction you will remember.
Why is guessing harder in Arabic?
Because Arabic is usually printed without short vowels, one written word can be several different words. كتب can be كَتَبَ, he wrote, كُتِبَ, it was written, or كُتُب, books. ملك can be مَلِك, king, مَلَك, angel, مُلْك, dominion, or مَلَكَ, he owned. The words around it decide, and they only help if you know them.
Two more things make a word you know look unknown. Small words attach to it in writing: وَبِكُتُبِهِمْ, and with their books, is a single word on the page. And the word changes shape inside itself: the plural of كِتَاب, book, is كُتُب. A learner who knows كِتَاب can fail to recognise كُتُبِهِمْ, which is why a useful Arabic lookup has to find the dictionary form and the root, not only match letters.
Fluent readers are not guessing blind. They recognise the root and the pattern, and the vowels follow from them; our article on reading Arabic without vowels walks through that process step by step. Until the recognition is automatic, a lookup is how you build it.
Does looking a word up help you remember it?
Yes, when readers actually do it. In a 1996 study by Hulstijn, Hollander and Greidanus, advanced learners who read a story with meanings printed in the margin remembered more of the target words than learners given a dictionary, largely because the dictionary group rarely opened it. Meeting a word three times helped too, but mainly when its meaning had been supplied or looked up.
Luppescu and Day (1993) looked at the same trade from the other side. Of 293 Japanese university students who read an English short story, those allowed a bilingual dictionary scored significantly better on a vocabulary test afterwards than those without one.
Put together: a lookup makes a word more likely to stay, a gloss in the margin makes a lookup free, and the cost of looking up decides how often it happens at all. Of those three, the cost is the one a reader can change.
Why does the cost of a lookup matter so much in Arabic?
Because a paper Arabic dictionary is ordered by root, not by spelling, every lookup starts with finding the root, and weak letters hide it. اتَّصَلَ, he contacted, is filed under و·ص·ل; مُسْتَشْفًى, hospital, under ش·ف·ي. When each lookup costs that search, skipping is the rational choice. When it costs one tap, the calculation reverses.
A tap dictionary is a margin gloss on demand: the meaning is there when you ask for it and invisible when you do not. The honest question is how often the tap finds the right entry. We measured our own: on 214,689 words of real Arabic text, drawn from learner articles, stories, subtitles, a classical book and Arabic Wikipedia, a tap in Fasaha opened a dictionary entry for 96.03% of words (measured 3 September 2026). The misses were mostly words the dictionary does not yet hold, and names or foreign words, with a smaller group of forms the lookup should have found.
No tool reads unvowelled Arabic perfectly. When a written word can be two different words, check the reading the card gives you against the sentence before you save it.
The two-pass read: a routine you can use tonight
Read a paragraph once for the gist and mark the words that stop you. Look up only those, then read the paragraph again, with audio if you have it, and save the words that came up twice. If you are marking more than about one word per line, the text is above you for now: drop a level and come back to it.
This is our routine rather than a finding from the studies above, but it is built on their numbers. It keeps you near the 95 to 98% band where context works, it spends lookups where they change the meaning, and it lets repetition choose what goes into review. The steps below are the same routine in order.
Where Fasaha fits, and where it does not
Fasaha is built for this routine. Every word in its stories, articles, podcasts and videos opens its meaning, root, pattern and family on a tap, the texts are graded from A0 to C2, and the words you save come back for spaced review.
It does not decide which words matter; the two-pass read does that. It also cannot make an ungraded text easier: when a newspaper page has an unknown word on every line, the right move is still to read a level down first.
Eight Arabic words that change with their vowels
Each item is one written form and the words it can be. This is why an unvowelled Arabic word is hard to guess, and why the words around it carry so much of the meaning.
- كتب: كَتَبَ · كُتِبَ · كُتُب — he wrote · it was written · books
- علم: عِلْم · عَلَم · عَلِمَ · عَلَّمَ — knowledge · flag · he knew · he taught
- ملك: مَلِك · مَلَك · مُلْك · مَلَكَ — king · angel · dominion · he owned
- شعر: شِعْر · شَعْر · شَعَرَ — poetry · hair · he felt
- قبل: قَبْلَ · قَبِلَ · قُبِلَ · قَبَّلَ — before · he accepted · it was accepted · he kissed
- حمل: حَمَلَ · حَمْل · حِمْل · حَمَل — he carried · carrying, pregnancy · a load · a lamb
- عقد: عَقْد · عِقْد · عُقَد · عَقَدَ — contract · necklace · knots · he tied, he concluded
- ذهب: ذَهَب · ذَهَبَ — gold · he went
| Words known | Unknown words | On lines of about ten words | What reading feels like |
|---|---|---|---|
| 98% | 1 in 50 | one every five lines | Comfortable: context can do its job |
| 95% | 1 in 20 | one every two lines | Workable, with a few lookups |
| 90% | 1 in 10 | about one on every line | Slow: the sentence keeps breaking |
| 80% | 1 in 5 | about two on every line | Decoding, not reading |
Is it better to guess an unknown word or look it up?
Both, in that order: guess, then check the words that matter. In Nassaji's 2003 study of intermediate learners, only 25.6% of guesses from context were fully right, and a wrong guess can feel right. A quick check turns a guess into a word you actually know.
How many words do I need to know to read Arabic?
There is no reliable single number for Arabic. Research on English puts comfortable reading at about 98% of the words in a text known, which took 8,000 to 9,000 word families for general written English (Nation 2006). For Arabic the practical test is the page in front of you: if you know nearly every word on it, it is at your level.
What is extensive reading, and does it work for Arabic?
Extensive reading means reading a lot of easy text quickly, for meaning, with few lookups. It works when the text really is easy for you, at around 98% of words known. For most learners below B1 that means graded texts written for learners, because in our experience few ordinary Arabic texts are that easy at that stage.
Why can I not find an Arabic word in the dictionary?
Most Arabic dictionaries are ordered by root, not by spelling. To find مَكْتَبَة, library, you look under ك·ت·ب; to find اتَّصَلَ, he contacted, under و·ص·ل, because the و has merged into the ت. Our guide to Arabic roots and patterns shows how to find a root by hand.
Sources
- Hu and Nation (2000), “Unknown vocabulary density and reading comprehension”, Reading in a Foreign Language 13(1)
- Laufer and Ravenhorst-Kalovski (2010), “Lexical threshold revisited”, Reading in a Foreign Language 22(1)
- Kremmel and colleagues (2023), “Unknown vocabulary density and reading comprehension: replicating Hu and Nation (2000)”, Language Learning
- Nation (2006), “How large a vocabulary is needed for reading and listening?”, Canadian Modern Language Review 63(1)
- Nassaji (2003), “L2 vocabulary learning from context”, TESOL Quarterly 37(4)
- Hulstijn, Hollander and Greidanus (1996), “Incidental vocabulary learning by advanced foreign language students”, The Modern Language Journal 80(3)
- Luppescu and Day (1993), “Reading, dictionaries, and vocabulary learning”, Language Learning 43(2)
- Quranic Arabic Corpus, Quran dictionary, root م ل ك
About Fasaha
Fasaha is a reading-first Arabic app. Every word in a story, a podcast, a video or an article opens into its meaning, root, pattern and family, and can be saved for spaced review.
- Available on iPhone, iPad, Mac and Android.
- Interface languages: English, Arabic, Turkish, Russian, Uzbek, Spanish, French, German, Indonesian, Malay and Urdu.
- Free to start; Premium opens the complete library and the stores handle billing.