Trends: The Coming of Age for Voice Capture and Voice-Enabled Applications
By Scott Wiley, SVP of Sales & Marketing
It has been a remarkable 30+ years since the Internet of Things began its daily advance into our lives. Who knew three decades ago that smartphones and personal assistants would play such a central and ever-increasing “new normal” role in how we consume information, shop for everyday goods and services, entertain ourselves, and even control our homes and businesses?
As the complexity of interaction with those devices increased, the manual typing of input commands quickly became cumbersome. It was natural for voice capture to emerge as a promising input method for those devices, used mainly as a “supplement” to the primary input method of typing. Early attempts to use voice capture for command input left quite a bit to be desired in the areas of accuracy: the inability to properly capture commands in the presence of loud competing noises, and other annoying shortcomings. Nevertheless, the technology naturally progressed to a point of minimal acceptability. More recently, the emergence of a whole new category of device, the smart speaker, in which the only interaction option is voice control, signaled the true coming of age of voice capture as an input method.
Smart Speakers and the Status Quo of Voice
Today’s smart speaker is thought by many consumers to employ the pinnacle of voice capture capability. After all, some of the most innovative companies are deploying smart speakers that feature voice capture. So, it’s logical to many that the technology that those highly-innovative companies deploy, must be the best performing voice capture capability available, right? Well…actually, no! That is not what is happening and that’s not what’s being deployed. In many cases the smart speakers that find their way into millions of consumer homes are products that are heavily subsidized by the platforms to which they connect.
The end game is a “market share grab” capturing the largest number of customer connections, at the lowest possible cost, while providing a “good enough” solution that people will buy. Few companies are optimizing in the direction of providing the most capable device to enable the “best” user experience. As a result, today’s consumer generally believes that repeating commands, having to move closer to the device, ensuring no obstructions in the audio path, lowering the volume of competing audio sources, and other user adjustments remain a necessary evil of the technology to make voice capture work reliably. That is simply and decidedly not the case!
After having surveyed hundreds of users of voice-enabled devices, we know a key complaint is the “limitation of many of the existing voice capture products to ‘hear’ voice commands accurately and clearly.” That is, the device does not always “hear” what is being spoken or understand what is said. The user may be too far away. There may be too much background noise like streaming music, TV dialog, or multiple competing conversations in the same room. There may be objects in the audio path—between the person issuing device commands and the smart speaker—that obstruct what is being said. Consumers also report frustration in the fact that the device may “hear” a word that sounds like the trigger word and it becomes “confused.” As they say, “garbage in, garbage out.” Bottom line, we know from our research that people are tired of screaming at their devices.
3-D Reverberation vs. Beamforming Technology: A Game Changer
The good news is that many application-focused companies are offering solutions that feature an alternate, significantly more robust, and more technically capable technology. Unlike less capable beamforming technology used by some, the most capable voice capture solutions use 3-D reverberation technology to deliver more noise reduction, significantly extended usable range, and more accurate real world trigger word performance. Reverberation technology does not rely on geometric constraints to define microphone configuration, placement, or orientation. This results in better performance and additionally allows industrial designers the mounting flexibility to achieve their product visions with fewer constraints. Whereas the old beamforming technologies often resulted in false positives and false negatives, or required users to repeatedly shout to have devices hear them accurately, reverberation technology overcomes those problems.
This new technology is also referred to generically as “far-field technology.” But not all far-field solutions perform the same. Taker for example our solution, Ark X Labs’ EveryWord™ Ultra Far-field Voice Capture Technology outperforms all existing OEM solutions. It is truly best-in-class. What does best-in-class performance mean in this context? It means the ability to capture voice commands from across the room (>9 meters), in extremely noisy and reverberative environments, with other loud music or audio playback, and in the presence of competing conversations. It also means that permanent obstructions, like furniture, architectural columns, and other physical barriers directly in the voice path are no problem. Further, in H2H or human-to-human communications (conference speakers, for instance) these best-in-class solutions also provide audio output processing which enhances fidelity and volume of playback resulting in “naturalness” and intelligibility of the person’s voice originating on the other end of the call or from the audio source.
Applying Far-field to Voice-Enabled Devices and Products
Even more exciting, this same far-field technology is not limited to smart speaker applications. Far-field voice technology can be built into nearly any electronic device to provide voice control of that device. Applications like these generally fall into the category of H2M or human-to-machine applications. Companies like Ark X Labs offer the voice modules or building blocks of far-field voice solutions that allow other companies to enable voice control in their products regardless of application, either H2H or H2M.
Today, Ark X Labs is in the testing or deployment phases with multiple global brands in smart hub, smart appliance, TV, audio soundbar, connected exercise, educational, video conferencing, agricultural, industrial, and healthcare products. In these applications, Ark X’s Ultra Far Field Solutions offer voice capture modules featuring the capability of being installed in products, hubs, ceilings, even on-and-in wall applications. In addition, Ark X offers solutions that are platform neutral. The Ark X solution is simultaneously compatible with multiple voice services and trigger-word providers, including Alexa, Google, Siri, Cortana, AliGenie, Baidu/Kitt.ai, Tencent, and Sensory. This allows the user to pick and choose from the best of the skills available from each platform and craft a solution that best suits his or her particular needs.
Regardless of the application, there has been a lot of interest from both industrial and consumer companies in enabling voice control in their products over the last year. In a post-COVID, hands-free world where people don’t want to touch any surface in a public (or, for that matter, private) environment, companies who use voice capture technology can enable great experiences in applications as diverse as a classroom-in-a-box, lobby check-in (hospitality or healthcare), elevator control, and even hands-free point of sale (POS) products and kiosks.
It doesn’t take a great deal of imagination to extrapolate what is developing today in cutting-edge companies in the area of voice capture and voice control to envision a day when most or perhaps all electronic products will have voice control as a primary method of interaction and control. We see endless possibilities. It’s no exaggeration to say that the future is voice and the future has arrived.
If you would like to explore how to build in best-in-class Ultra Far Field voice capture and voice control into your electronic products, feel free to reach out to us at Ark X Labs. We’ll be happy to help you get started and provide as much assistance as needed along the way.