OpenAI: we trained intelligence... ended up with goblins.
It sounds like a joke: AI models randomly talking about goblins and gremlins. After the launch of GPT-5.1, OpenAI noticed something odd - usage of the word "goblin" jumped 175%, and "gremlin" by 52%.
The team started digging and realized this wasn’t a bug but ... behavioral artifact. Basically, the model learned that "goblin" is a valid token to use in many contexts - even when it shouldn’t be.
Small patterns in AI don’t stay small but scale. If it gets reinforced even slightly, it can propagate across millions of interactions.
Like training a junior developer who accidentally picks up one weird habit and then that habit shows up in every file they touch.
ChatGPT had a nerdy personality with a system prompt like "play with language, the world is weird" and the reward model randomly developed a taste for… creatures (+reward => more creatures => goblins).
Nerdy handled just ~2.5% of replies but generated 66.7% of all goblins!
Model trains on its own outputs => goblin population explodes. Boom!
By March, they shut nerdy down, cleaned up the reward signals, filtered the dataset but GPT-5.5 had already internalized it.
So in Codex, they literally added: do not summon goblins
Now imagine what happens after years of AI-generated posts, SEO sludge, synthetic content loops…
What a time to be alive.
Source: openai.com/index/where-th…

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.