I tried to organize vocab by difficulty level for an English language-learning app once.
It shocked me how there is absolutely no "right" answer.
If you are teaching English for travel, then you're prioritizing a lot of stuff around bathrooms, transportation, menu items, etc.
If it's for understanding TV, it's a lot of words like "murder", etc. Depending on which TV shows you want to understand.
If it's for reading the newspaper, you don't ever need to know "bathroom", but you sure do need to know words like "congressman".
While if you are living somewhere, it's really important to know a lot of basic supermarket items that you wouldn't prioritize for other usages.
Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed. And the substitutes -- transcribed speech from TV, radio, podcasts, etc. -- is not the same context as the random stuff you say at home and during an average day.
I tried to teach 'magicE vocab', sorting them by difficulty level for an English language plan, and got them easily arranged. For example, sham/shame and slide/slid are for hard to learn level, while ate, pale, kite, are for the easy level.
> While if you are living somewhere, it's really important to know a lot of basic supermarket items that you wouldn't prioritize for other usages.
Why? Having spent a good amount of time living in Shanghai, I found it important to be able to understand menus. But there's no pressure to know the words for supermarket items; you can just go to the supermarket and look for the item.
Otherwise your point is correct; all semantic words are equally difficult and which ones you know depends on the things you like to talk about. Grammatical words are more difficult, and more important, but this is so widely understood that language-learning material already treats them as an entirely separate class of things to learn.
> Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed.
(1) You seem to want COCA, which includes a bunch of transcribed telephone calls.
(2) Word frequencies are still the wrong concept. If you want to understand a particular document, you need to understand almost all of the words that appear in that document. (You'll be able to learn some of them from their use in the document.) If you decide to learn a list of "frequent" words, you're unlikely to be able to understand more than a couple of isolated sentences in any given document.
I tried to build a similar list myself for German and it's not easy as just taking a lot of content and counting frequency. I also haven't found existing curated lists of most useful vocabulary.
There are some databases but e.g. they are biased towards Wikipedia and web which makes some very obscure words at the top of popularity (like some technical words which are present on each wiki page like Datenschutz or Impressum).
> The “Social-Communicative” level barely changed in size. But nearly a quarter of the words in the 1953 list are gone, and 39% of the 2023 words are new. Humble, loyalty, fellowship, generous, polite, and companionship gave way to community, identity, organization, ethnic, gender, and narrative. ...It offers fewer words for the people directly around you, but more for belonging at a distance.
I would blame inequality on this one. In a more unequal world tribalization is a survival strategy and language follows.
When you see everybody else as your equals then focusing on describing that individual person, instead of their group, makes more sense.
Economic inequality affects deeply how we think about others.
Interesting one if you look at Google Ngram Viewer – usage dropped off massively to 70s/80s, and it's picked back up since but not to 1953 or earlier levels, so even that doesn't explain it.
Must just be the combination of that increase as well as other words decreasing in usage I suppose. E.g. perhaps we're a bit less keen, but also much less passionate, so keen ends up making the cut.
As a native English speaker, keen feels like a very 1950s TV word. Don't know if that's actually true but feels like that way. I expect you're more likely to hear like cool today are some more contemporary word.
At the moment this story has 2 upvotes in a half hour and is in the 8th position on the front page. Apparently HN has a fairy godmother algorithm that randomly promotes posts.
It shocked me how there is absolutely no "right" answer.
If you are teaching English for travel, then you're prioritizing a lot of stuff around bathrooms, transportation, menu items, etc.
If it's for understanding TV, it's a lot of words like "murder", etc. Depending on which TV shows you want to understand.
If it's for reading the newspaper, you don't ever need to know "bathroom", but you sure do need to know words like "congressman".
While if you are living somewhere, it's really important to know a lot of basic supermarket items that you wouldn't prioritize for other usages.
Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed. And the substitutes -- transcribed speech from TV, radio, podcasts, etc. -- is not the same context as the random stuff you say at home and during an average day.
Why? Having spent a good amount of time living in Shanghai, I found it important to be able to understand menus. But there's no pressure to know the words for supermarket items; you can just go to the supermarket and look for the item.
Otherwise your point is correct; all semantic words are equally difficult and which ones you know depends on the things you like to talk about. Grammatical words are more difficult, and more important, but this is so widely understood that language-learning material already treats them as an entirely separate class of things to learn.
> Also, while it's easy to calculate word frequencies for stuff like newspaper articles, there aren't any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn't getting recorded and transcribed.
(1) You seem to want COCA, which includes a bunch of transcribed telephone calls.
(2) Word frequencies are still the wrong concept. If you want to understand a particular document, you need to understand almost all of the words that appear in that document. (You'll be able to learn some of them from their use in the document.) If you decide to learn a list of "frequent" words, you're unlikely to be able to understand more than a couple of isolated sentences in any given document.
There are some databases but e.g. they are biased towards Wikipedia and web which makes some very obscure words at the top of popularity (like some technical words which are present on each wiki page like Datenschutz or Impressum).
I would blame inequality on this one. In a more unequal world tribalization is a survival strategy and language follows.
When you see everybody else as your equals then focusing on describing that individual person, instead of their group, makes more sense.
Economic inequality affects deeply how we think about others.
In 1953 people were not exposed as much to different groups of people far away.
Must just be the combination of that increase as well as other words decreasing in usage I suppose. E.g. perhaps we're a bit less keen, but also much less passionate, so keen ends up making the cut.
https://news.ycombinator.com/pool
But not in this case. It's a slow Sunday afternoon and getting a few upvotes quickly is enough