As much as we like to write about cutting-edge technology here at Always Scheming, new inventions frequently struggle with commercialization. In such cases, breakthroughs sometimes emerge not from the groundbreaking new creations themselves, but rather from the ingenious repurposing of once-overlooked technologies in conjunction with these novel developments. Reassessed from new perspectives, this tech that was once constrained by its era can discover a second chance at meaningful impact and relevance.
Perhaps the most famous historical examples of this in the games industry have come from Nintendo. The original Game & Watch system was a repurposing of existing calculator technology. Similarly, the motion controls of the Wii were based on patents licensed from a company called Gyration, which was first awarded those patents back in 1999 for use in aviation a full seven years before the Wii would release.
Another technology that we believe may be poised for a similar type of resurgence is speech recognition. The growing capabilities of neural networks, transformer models, and generative AI more broadly will allow this previously under-utilized technology to expand beyond the domain of so-called “voice assistants,” such as Amazon’s Alexa or Apple’s Siri, and into a new frontier of interactive entertainment.
Game designers have tried for years to incorporate voice commands into interactive experiences.
One of the earliest and most famous examples of this was the bizarre Seaman for the Sega Dreamcast. Seaman tasked players with feeding, caring for, and providing accompaniment to a sort of human-fish hybrid creature, voiced by Leonard Nimoy (of Star Trek fame).
Other more mainstream attempts would soon follow. Most of the major consoles boasted at least one title that attempted to leverage voice inputs. Prominent examples include Hey, You, Pikachu! (N64), Lifeline (PS2), and Mass Effect 3 (Xbox 360; PS3), among many others.
The typical use case for speech recognition in these games was to allow the player to provide direction to companions or allies (i.e. “dodge,” “attack,” “take Point B,” etc.). Some games added these features as a type of upsell gimmick, while others were built entirely around voice commands. Lifeline, for example, conducted the vast majority of its gameplay through the PlayStation 2’s microphone, requiring players to guide the main character, Rio, entirely through vocal directives.
While none of these examples have been major commercial successes, that hasn’t stopped game developers from trying new approaches. Even today, speech recognition is still being explored in novel and interesting ways.
One recent example that caught our attention was indie game Don’t Scream, “an Unreal Engine 5 short horror experience inspired by ’90s camcorder found footage.” Made by two developers over just five months, Don’t Scream tasks players with exploring the spooky Pineview Forest for 18 minutes without screaming. If you scream, the game restarts.
Simple though it may seem, Don’t Scream managed to capture the eye of major streamers like Markiplier and Pokimane on its way to racking up more than 130,000 Steam wishlists in just two weeks of pre-launch marketing.1
Despite the lack of a clear commercial breakthrough, there is undoubtedly precedent for incorporating speech recognition into interactive entertainment experiences. But why should we expect the status quo to change any time soon?
As alluded to earlier, it is our belief that this is an area of game development primed for reinvention as a result of the improvements that generative AI tools are bringing to speech recognition technology.
Consider the limitations faced by the prior wave of products built around this tech. Last year, Microsoft CEO Satya Nadella labeled voice assistants like Siri and Alexa “dumb as a rock” — and for good reason. Previous iterations of speech recognition relied on huge arrays of predefined voice commands. These commands could not be used outside of a given product’s established rules, inevitably limiting their flexibility.
Such datasets also grew to be incredibly cumbersome for developers to maintain. In the case of Apple’s voice assistant, Siri, simple updates like adding new commands required the entire database to be rebuilt — a six-week process. Advanced feature updates could easily balloon into months-long slogs.
Even if a sufficiently large set of voice commands could be implemented, they were easily broken by users with strong accents, unexpected dialects (e.g. slang, regional variations, etc.), or different languages altogether. This fragile setup was further complicated by variables beyond developers’ control, such as noisy environments.
“[Voice-driven] products never worked in the past because we never had human-level dialogue capabilities. Now we do.”
– Aravind Srinivas, Founder, Perplexity
Source: New York Times
Today, these problems seem ripe for new solutions.
Tools like NVIDIA’s Riva and OpenAI’s Whisper, trained on hundreds of thousands of hours of data scraped from the web, offer “improved robustness to accents, background noise and technical language.”
Developers will no longer be forced to rely on cumbersome databases of predefined voice commands. Access to huge swathes of data will make AI-enabled speech recognition tools increasingly powerful and comprehensive.
At this point, you might still be wondering if voice inputs are even worth pursuing in the first place. Yes, generative AI makes speech recognition better and more efficient, but do gamers even want voice-enabled games? After all, the vast majority of the gaming audience has grown up using controllers, mice, keyboards, and touchscreens — all far more tactile input devices than microphones.
The most obvious reason is the “hook” factor. Voice inputs are still novel to most gamers — and likely, to most users of consumer technology more broadly.
Unless you’re a Siri power user or have an Alexa in your home, it’s unlikely that you interface with speech recognition frequently in your day-to-day life (outside of the occasional phone call to customer support). This provides voice-driven games with an inherent “hook,” which may be exciting to some…but may also seem gimmicky to others. The Steam reviews for Don’t Scream are illustrative of this, as even negative reviews tend to credit the novelty of the game’s central “idea” or “concept”:
However, if we consider a broader audience beyond the confines of traditional “gamer” stereotypes, voice games suddenly have greater appeal.
For one, voice inputs make games more accessible to players that are differently-abled. Developers that can solve the challenge of translating voice commands to conform with a given game’s control scheme may find an eager market among the accessible gaming crowd whose best option today is to buy expensive adaptive controllers (PlayStation, Xbox) to be able to play their favorite games. Eventually, we may even see automatic speech recognition built into console operating systems directly.
Audiences don’t need to learn any specialized equipment to use their voice, which means voice-driven games might also be more appealing to seniors or young children. In fact, this specific use case has been cited publicly by Max Child, CEO of voice-enabled gaming company Volley, who stated “that grandparents will often play on their own time and text their grandchildren about their progress, and the children will advance their moves when they have a minute.”
Voice inputs may even unlock new forms of games and gaming platforms entirely. One can easily see the applications of speech recognition for Dungeons & Dragons-like TTRPGs, as well as other types of narrative-driven games: mysteries and crime-solving games, dating simulators, and so on. The combination of speech recognition and AI-driven NPCs might even transform social deduction games or other forms of multiplayer-only genres into viable single-player experiences, potentially attracting new audiences altogether.
With all that said, speech recognition is still far from a perfect fit for gaming.
For starters, voice input has inherent speed limitations. There’s a reason many people choose to listen to podcasts at 1.5x or 2x speed, yet no one listens at 5x speed. There are upper limits to how quickly humans can speak and be understood by other humans, making voice input a poor fit for certain types of fast-twitch games.
Indeed, there may even be some evolutionary reasoning behind this, as venture capital investor Chris Paik pointed out in an interview on the Invest Like the Best podcast:
“I think there’s something interesting to think about Darwinistically. If voice had actually been meant as an input mechanism between humans and computers, wouldn’t we have also come up with some tool to advance our data input via the sounds we can make?
We would have come up with some sing-song language or we would make specific sounds that map to certain data, but we haven’t. And I think that begs a question of maybe voice is just not meant to be [an] input mechanism between humans and computers. Maybe we have a structural advantage with our other senses with touch and haptics.”
Additionally, many people are simply not interested in speaking while playing games. Or, put differently, players do not necessarily want their own speech to impact gameplay. People talk all the time while gaming (just watch any Twitch streamer), but their voices rarely impact the game itself.
The incorrect conclusion to jump to would be that everyone has a cell phone with a built-in mic, and that mobile gaming should therefore be a huge opportunity for voice-enabled games. Unfortunately, talking into a microphone just doesn’t fit well with the play patterns of most mobile gamers. No one wants to be voice-gaming while waiting in line at a coffee shop or riding the subway to work.
Similarly, games that do rely heavily on player speech, such as multiplayer titles supporting player-to-player communication, would need a way to differentiate between speech intended for teammates or opponents and speech meant for the AI to ingest. This feels like more of a design constraint than an unsolvable problem, but it will nevertheless impact the types of games that can incorporate this technology.
As it stands today, there are not many companies dedicated to building games that rely on speech recognition and AI as core technologies.
The most prominent studio operating in this space is the aforementioned Volley, a former Y-Combinator startup that describes itself as “the leading developer of voice AI games for smart speaker platforms like Amazon Alexa, smartphones, and connected TVs.”
Others have previously tried their hand at making games for voice assistant platforms, though most appear to have moved on by now. Doppio Games, another voice-gaming startup, was acquired by Fortis Games in early 2022. Warner Bros. published a Batman game for Amazon’s Alexa, largely as a promotional vehicle for the movie Batman v Superman: Dawn of Justice. Long-running Pittsburgh-based developer Schell Games created Baker Street Experience, also for Alexa.
While Volley appears to be one of the few remaining players in the voice games niche, more competition should be expected. Apple, Google, Microsoft, and Amazon are unlikely to walk away from their voice assistants entirely given the rapid technological improvements taking place. These platforms will inevitably be hungry for compelling content and may even be willing to pay established developers to create experiences that showcase their products’ new capabilities.
Similarly, a new wave of natural language-driven AI hardware is set to hit the market soon, such as Humane’s ai pin and Rabbit’s r1 handheld. Even Sam Altman and Jonny Ive are rumored to be developing an AI-powered device of their own. All of this new hardware represents potential future distribution platforms for aspiring developers.
Other entrants may come from AI or voice tech middleware companies as they seek to create proof-of-concept games to showcase their technologies, perhaps as a tentative first step towards vertical integration. Firms like Replica Studios, Inworld AI, or Voicemod will be worth keeping an eye on here.
Finally, we should look to existing game developers in search of new platforms and/or design paradigms. In addition to the many incumbent publishers, AI-forward studios like Latitude.io or Hidden Door may be interested in incorporating speech recognition into their games.
For enterprising game developers seeking a blue ocean category to build in, voice-enabled games may be an interesting wave to ride.
If you’re building something in this space, we’d love to hear from you!
Credit to the excellent GameDiscoverCo newsletter for the data on Don’t Scream.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.