The Curious Case Of AI’s Goblin Obsession
Not long ago, we explored the curious case of Elias Thorne, the fictional academic who somehow found his way into AI-generated responses.
Now, AI has developed another rather unusual habit.
Some newer models began mentioning goblins, gremlins, pigeons and other creatures in answers where they had no obvious relevance. Whether discussing coding, photography or everyday topics, these unexpected references kept appearing often enough for developers to realise they weren’t isolated mistakes. They were a pattern.
The solution was surprisingly simple. Developers updated the system instructions, explicitly telling the models not to mention goblins, gremlins, trolls, raccoons, pigeons or similar creatures unless they were genuinely relevant to the user’s question. It sounds almost absurd, but it reflected a real attempt to correct an unexpected behavioural quirk.
On one level, it’s a funny story, but it also offers a useful reminder of how large language models behave. They learn patterns from enormous amounts of data and, while those patterns usually produce useful and coherent responses, they can occasionally lead to habits that nobody deliberately intended.
In this case, OpenAI explained that the repeated creature references weren’t the result of the model suddenly developing a fascination with goblins. Instead, a series of small changes during training had unintentionally reinforced those kinds of metaphors, causing them to appear more frequently over successive versions of the model.
That’s particularly interesting when we think about the way new AI models are released. We tend to assume that each update represents a straightforward improvement on the last, with models becoming more capable, accurate and reliable, but development isn’t always quite that neat. Changes designed to improve one behaviour can have unexpected effects elsewhere, some of which may only become obvious once people start using the model in millions of different situations.
It also makes an interesting follow-up to the story of Elias Thorne. In that case, AI produced convincing information about a person who didn’t exist, whereas the goblin obsession shows a different kind of unexpected behaviour: a pattern that gradually became more prominent despite nobody deliberately designing the model to behave that way.
The goblins themselves may ultimately be little more than an amusing AI quirk, particularly when the solution involved explicitly telling the model to stop mentioning them. What they reveal, however, is something more useful about the technology: making AI more capable doesn’t necessarily mean making every part of its behaviour more predictable.
Image sources
- fantasy-landscape-1200: ©Matthew Gibson via Canva.com