Idioms and AI: why ChatGPT still doesn’t quite get us

Kick the bucket. Monkey business. Elephant in the room. Such idioms are naturally difficult for computers to understand, but large language models were surely supposed to solve that – after all, such cliched language abounds online.

But idioms remain a challenge. Now, if you ask ChatGPT to explain such phrases, it will nail them – so to speak – and even offer a detailed history of the origins of each. But ask ChatGPT to generate an image of “monkey business” and OpenAI’s DALLE-3 will return an illustration of the animals wearing suits. What about “kick the bucket”? A cartoon character kicking a bucket.

ChatGPT interpretation of monkey business... showing monkeys dressed in suits
Ask ChatGPT to come up with an image of “monkey business” and you’ll get stockbroker simians

What’s going on here? That’s what Tom Pickard, a researcher at the University of Sheffield is hoping to find out – and it’s revealing plenty about how these systems work.

More AI and idiom research needed

At the Turing Institute’s AI UK conference in London last week, Pickard said LLMs get better with idioms as the models get bigger – more parameters plus more training data does lead to better accuracy.

ChatGPT generated image of a cartoon character kicking a bucket
ChatGPT could explain the meaning of this idiom, but had trouble creating a relevant image

“But it doesn’t work reliably or 100% and especially those generative models,” he explains. “If you ask them the question a few times, you’ll sometimes get different answers, right? Because it’s stochastically producing the text.”

Humans, on the other hand, will understand that something is an idiom even if they’ve never heard it before and don’t quite get it. If an idiom is rare – in the “long tail of language”, he says – then the LLMs “work a bit, but not as well as you would hope so”.

To help, Pickard is hoping to build a system to assist LLMs with unpicking idiomatic language, perhaps by letting them explore multiple meanings.

That’s something generative systems aren’t very good at. They’re designed to pick a definitive solution rather than sit on the fence, though that’s starting to change.

Elephant in the room when it comes to generative AI idioms

Elephant in the room idiom, kind of, by ChatGPT
ChatGPT struggles with negatives in image generation: ask the AI to make a room without an elephant, and it can’t help but include the animal in the artwork

The challenge is especially evident with images. Pickard pointed to a classic example: ask these AI systems to generate an image of a room without an elephant. Clearly, that’s just a room.

But AI fixates on the word elephant, and places an elephant out the window, waiting just outside the door, or includes two elephants. When I tried to replicate this, ChatGPT gave me a lovely image of a room with artwork of an elephant. It just can’t help itself.

This highlights how AI works: the main sources for training data for images for such systems use the captions for labelling. As all of the images that mention that word in the caption normally feature the animal, AI systems recreate those – rather than understanding that the negative means they should ignore all images with elephants.

“It doesn’t learn negative associations,” says Pickard. “It learns to spot things that are mentioned in the caption.”

Testing questions

Now, to be fair to AI, if you were trying to get an image of a room without elephants, you’d simply ask for a room. By asking in this way, people are trying to sort of stress test the system.

Another example in text is the infamous strawberry conundrum. As smart as these systems are, ChatGPT – for a time – couldn’t count how many instances of the letter “r” were in that word.

While these are perhaps silly examples, they reveal how the system works: LLMs aren’t reading the way we do, but rip apart letters into tokens, destroying any meaning. “The tokens, the words are broken up in training, it doesn’t process text, it processes multi dimensional vectors and numbers,” says Pickard. It has no access to the representation as letters in any algorithm.

OpenAI has since fixed that particular fruit-based example, as any AI developer does when an apparent flaw goes viral. But Pickard notes that the system still doesn’t really understand, it’s just told how to respond to avoid the negative publicity.

Impact of AI when you’re not understood

Betting the farm idiom according to ChatGPT
Don’t bet the farm on ChatGPT’s ability to understand idioms

Why does any of this matter? These models have clear limitations – but they’re capable enough in so many ways, it’s easy to forget that fact. And now that the UK Government is cheerfully foisting AI on the civil service, the fact that AI doesn’t always understand idioms and other multi-word phrases can start to really matter.

What if you use a figure of speech in an application for benefits or in an interview with a civil servant that’s transcribed and interpreted by AI? Or your first-language isn’t English, and translators are traded for translation software? If AI can’t manage English idioms, given it’s the main language they’ve been trained in and most are defined in dictionaries, how well will it do on foreign speech? And in these cases, being misunderstood could have serious implications.

“The UK government, for whatever reason, is ‘betting the farm’ on using generative AI, and I’m not sure they should be,” Pickard says. “It doesn’t understand things – and I think this is where it becomes quite important.”

More examples…

Long tail concept according to ChatGPT

Asked to make an image illustrating the idea “long tail” in reference to language, ChatGPT first responded with a description of the image, clearly understanding the concept and suggesting a winding tail made up of words, more common at the base and less so at the tail extends – great, except it struggles to generate images of words.

Sitting on the fence idiom according to ChatGPT

Remember earlier I was talking about sitting on the fence? I asked ChatGPT to generate an image of a person sitting on a fence, and rather than leap into spitting out cartoon fence sitters, it made a text-based suggestion. This showed that it was seeking to highlight someone looking confused or indecisive. But the end result was still literally a bloke sitting on a fence.

a fierce female action hero kicking ass in a bold, dynamic cartoon style
Thanks ChatGPT!

More kick-ass articles from Nicole

About The Author

Nicole Kobie
Nicole Kobie

Nicole is a journalist and author who specialises in the future of technology and transport. Her first book is called Green Energy, and she's working on her second, a history of technology. At TechFinitive she frequently writes about innovation and how technology can foster better collaboration.

Read more from this author.

We take journalism seriously. To learn more on why you should trust us, head to our editorial guidelines page or meet our team.