Product Development Trends: The Future of Far-field Voice & Audio Technology
Q & A with Eric Bauswell, CEO at ArkX Laboratories on the future of far-field voice & audio technology.
What precipitated the joint venture?
E.J. Constantine, Founder and Chief Executive of Ark Electronics USA, and our leadership at Surfaceink together found that customers in the U.S. and overseas have had a need for a better voice solution. Brands want more control while mitigating the risks and cost of developing their solutions. This next generation voice solution crosses the chasm for the end-user experience, dramatically improving performance for both real world voice enabling interfaces. Our production-ready solutions are designed and manufactured to help them do just that.
What do you offer the market with Ark-X?
In today’s competitive landscape, time-to-market is everything. With Ark’s manufacturing expertise and Surfaceink’s experience with audio and voice product development, this joint venture offers OEMs and start-ups a better option in performance while reducing development costs and accelerating their time-to-market.
ArkX’s advanced line of ready-made audio/voice products and development kits, powered by Cirrus Logic SoundClear® and FlexArray™ technology, are designed to outperform the market in far-field voice capture. The products are Alexa-compatible, meet or exceed all requirements for the Amazon Voice Services (AVS) premium rating product qualifications and are compatible with other platforms. This includes an acceleration in not only the hardware, but also the software side of solutions. From accelerating the use of multiple triggers to integrating the device software stack with the cloud, we want to help enable your ideas and solutions.
This is a next generation evolution beyond older beam forming technologies. Unlike the 2-dimensional planar view of other technologies, we add a new dimension to far-field voice by analyzing sound in 3D, increase range and precision while capturing voice and suppressing noise. Twelve independent AECs also make it ideal for soundbars and surround sound solutions.
What’s really great is that these can be customized for a company’s ecosystem and applied to a wide range of products including speakers, TVs, appliances, voice controllers and other gadgets. For smart home applications, the modules can be installed in hubs, ceilings, and in-wall. This is something customers have been seeking. We’ve now got it available for them.
This technology enables great experiences for a wide range of applications, like working from home, classroom-in-a-box, lobby check-in (hospitality or healthcare) and even hands-free point of sale (POS) products.
How will this joint venture benefit both Fortune 500 companies and start-ups?
Our combined technical expertise and agile manufacturing processes enable customers to create original products and solutions at scale. ArkX Labs can help get ideas and IOT products to market faster. Simply put, ArkX offers a clearer path. If a company is ready with an idea or product line, we can help get them to market today.
Our next generation of advanced performance far-field Voice/Audio modules and custom solutions, powered by Cirrus Logic and NXP technology, are pre-qualified and ready to go. Any Fortune 500, OEMs, or start-up that wants their own branded, voice-enabled IoT products and smart devices now has the ability to reduce their development costs and accelerate their time-to-market, while mitigating the risk for far field solutions.
What can you share about the product line?
ArkX is dedicated to providing a powerful product line with advanced Voice/Audio product innovations and differentiating features. This includes EveryWord ™ Audio Front End (AFE) Module, EveryWord ™ Voice Module(System-on-Module + AFE), EveryWord ™ Development Kit and EveryWord ™ Software Stack.
All of our products come with clear performance advantages for capturing voice commands from twice the standard distance, around corners, in noisy and reflective environments, and without lowering playback volume. Additionally, they provide the unique ability to identify and suppress speech from TV or other single point noise sources.
They do not require source ducking for reliable interaction and provide linear, circular, square, triangular or 3D mic array geometries, all while requiring fewer microphones. There is ultra-low power battery operation for wake-on-word and the flexibility for placement of microphones allows for in-wall, ceiling, or dashboard solutions. The 3D approach versus linear beam forming enables fewer blind spots with increased performance while incorporating fewer redundant microphone arrays for coverage.
Let’s talk some specifics. First, EveryWord ™ Audio Front End (AFE) Module. As the name suggests, it features Far-Field Voice Audio Front End (AFE). It’s powered by Cirrus Logic technology, has an Amazon “Premium” performance rating, 7 to 9 mic performance using only 3 to 4 mics, and an onboard clock. It’s compatible with most PDM mic arrays, has exceptional capture range of greater than 6 meters and has flexible mic array spacing and orientation (2D, 3D, vertical, horizontal). The hardware accelerates a royalty-free Amazon wake-word engine for applications leveraging existing host/computer board. It requires ArkX F/W integration on host and eliminates need for high density PCB.
The EveryWord ™ Voice Module(System-on-Module + AFE) is powered by Cirrus Logic and NXP technology, has Quad [email protected] + M4@400Mhz host processor and 802.11 a/b/g/n/ac Wi-Fi and Bluetooth 5 (BR+EDR+BLE). Like the EveryWord ™ Audio Front End (AFE) Module, it also has an Amazon “Premium” performance rating, 7 to 9 mic performance using only 3 to 4 mics, has exceptional capture range of greater than 6 meters and has flexible mic array spacing and orientation (2D, 3D, vertical, horizontal). Additionally, the hardware accelerates a royalty-free Amazon wake-word engine for applications leveraging existing host/computer board.
The EveryWord ™ Development Kit is a far-field AVS Reference Design Kit with
FlexArray™ Technology. This was developed at ArkX Laboratories by our expert team of engineers, pre-tested using Amazon AQT and manufactured to exceedingly high standards. The EveryWord ™ Dev Kit allows developers to test and validate their own voice-enabled products using an advanced voice solution. This, of course, leads to mitigating risk, reducing development costs, and accelerating time-to-market.
Finally, EveryWord ™ Software Stack is a highly-integrated vertical software stack developed in collaboration with Cirrus Logic that gives customers the ability to integrate far-field voice capture on the audio front end with Amazon Voice Services and other cloud services on the backend. The hardware and software development are validated in parallel with Amazon’s automated qualification testing along with additional requirements for larger spaces and non-line-of-sight capabilities.
The product features include a Yocto-based toolchain, cloud build support, and pre-configured development VM, and a custom Linux Kernel and root filesystem with debugging, tuning, and update capabilities. It’s a ported, configured, and pre-qualified AVS stack that includes hooks and a Python environment for easy integration of custom behaviors. It has Wi-Fi (client and AP modes) and Bluetooth stacks with advanced USB capabilities, including OTG/host/peripheral and USB-Audio interfaces.
The EveryWord ™ Software is available with additional “starter” components, such as a mobile companion app, sample Alexa Skills, and system configuration console with multiple software trigger options (Amazon, Kitt.AI, Google, Sensory, Baidu, etc.).
They do not require source ducking for reliable interaction and provide linear, circular, square, triangular and 3D mic array geometries. Additionally, there is ultra-low power battery operation for wake-on-word and the flexibility for placement allows for wall, in-wall, ceiling, or dashboard.
Where do you see the future of voice technology going?
In light of the current Covid-19 crisis, it is easy to envision a future where hands-free solutions like far field voice interfaces would become more common. The old beam forming technologies often resulted in false positives and negatives or even having to shout over content repeatedly to have devices hear the user. A better way needs to used moving forward and this is it.
Scott McNeese, the Director of Voice and Smart Solutions at Surfaceink, recently said he sees smart voice, the combination of voice-as-an-interface and AI, in its infancy. However, as its application explodes in the coming years, the impact on the human experience will grow by orders of magnitude.
I’m completely in alignment with Scott because I see clients currently approaching us and they aren’t just the early innovators. We’re now seeing a second group come in who wants to capitalize on that early brand success to integrate voice into their own enterprise and digital solutions. I would agree that the depth and breadth of what we’re going to see in the next five years is probably an order of magnitude larger and with greater significance then we’ve seen in the last five years.
Everyone knows what Amazon and Alexa are as well as, “Hey Google,” and certainly, “Siri.” Moving forward, I believe voice interaction is going to be so common that we will get away from using those trigger words.
Keep in mind, many of the devices are still push-to-talk and they aren’t listening for commands until you physically push the button. We still have that process with voice commands by using “trigger words.” While this is easier than walking across the room and pressing a button, it’s effectively the same thing and it’s still an artificial experience.
As AI becomes smarter, natural language understanding (NLU) gets better and the machine learning gets faster, the value of AI devices to the consumer grows exponentially. As human interaction with AIs becomes more natural and efficient, the appeal of using smart voice grows. At some point, you could have a healthcare version or hospitality version of your voice assistant. Or you could have a single type of platform that knows who you are, where you are, or what you’re doing. It then understands the context of your actions and adjusts accordingly.
With regard to the next generation on the hardware side, I think that we are at the leading edge of that technology. Integrating into different back-ends or existing ecosystems is something we’ve done several times and believe that’s a significant value-add knowing how navigate that software stack and plug into those existing software solutions in the cloud or the back-end.
Being in the business and intimately working with multiple companies and across various industries provides us with a different depth of insight that not a lot of other folks might have. We save companies a lot of time, dollars and effort to market because in addition to knowing what you can do, or could do, or should do, we know what you should not to do.
