Why Google’s AI Has Trouble with Spelling, Even with ‘Google’
How many instances of the letter P are there in Google? According to Google, the answer is two.
Moreover, Google’s AI Overview claims that there is “precisely 1 ‘r’ in the word ‘poop,’” and identifies two ‘d’s in the word journalism, although it spelled it as: j-o-u-r-n-a-d-i-s-m. At least it recognized that there is one P in the surname of the U.S. president, albeit spelled t-r-p-u-m.
It wasn’t too surprising to predict that Google’s AI-powered Search update would face challenges. We’ve encountered this type of scenario previously. The initial rollout of AI Overviews in Search led to citations from satirical websites like The Onion and Reddit, suggesting ridiculous actions such as eating rocks and adhering pizza.
As Google intensifies its efforts to make generative AI a fundamental component of its long-established products, these missteps are to be expected.
“Counting letters within words has been a known challenge for LLMs, and we are actively addressing this particular issue,” Google conveyed to TechCrunch via email.
These trivial spelling errors may seem familiar. LLMs, the technology behind chatbots and text generators, are not inherently programmed to understand spelling. There’s a prevailing joke that whenever a new AI model is unveiled, one should ask about the count of ‘r’s in the word strawberry. Despite their capability to develop applications in mere seconds or solve intricate math problems, these AI models often exhibit spelling skills comparable to that of a kindergartner.
The issues with Google’s AI overview extend beyond mere amusing misspellings. The company already rectified an incident last week where a search for “disregard” yielded a response resembling a dictionary definition, yet absurdly displayed as: “Understood. Let me know whenever you have a new prompt or question!” Nonetheless, these spelling errors remain comical due to their unyielding presence.
As researchers have previously discussed regarding these spelling issues, AI does not perceive sentences as cohesive entities composed of words and letters. Many LLMs utilize transformer models, which break down text into tokens that can represent whole words, syllables, or individual letters, depending on the model. Rather than “reading” in the human sense, the AI converts text into numerical formats, contextualizing them to produce coherent replies.

“LLMs are built using a transformer architecture, which notably doesn’t actually engage in text reading. When you enter a prompt, it gets transformed into an encoding,” Matthew Guzdial, an AI researcher and assistant professor at the University of Alberta, explained to TechCrunch. “When it encounters the word ‘the,’ it produces a specific encoding that represents ‘the,’ but it lacks understanding of ‘T,’ ‘H,’ and ‘E.’”
The token-based approach that powers LLMs like Google’s AI overview is inherently limiting, and researchers are not optimistic about resolving the spelling issues.
“Defining what constitutes a ‘word’ for a language model is quite complicated, and even if human experts achieved a consensus on an ideal token vocabulary, models would still likely gain from ‘chunking’ data even further,” stated Sheridan Feucht, a PhD candidate studying large language model interpretability at Northeastern University, in a discussion with TechCrunch. “I suspect that a flawless tokenizer doesn’t genuinely exist due to this inherent ambiguity.”
This problem isn’t an urgent concern among researchers, as the primary value of LLMs doesn’t lie in their spelling accuracy. Still, these striking errors remind us that AI is not infallible, even if it occasionally seems to possess extensive knowledge. We should not accept AI outputs blindly without verifying their correctness.
When you make purchases through links in our articles, we may earn a small commission. This does not affect our editorial independence.


