Complete content from the Synervoz website. # Untitled > Explore real-world use cases for the Switchboard SDK: consumer, business, metaverse, hardware audio applications. Add engineering and feature flexibility to AI agent solutions. Experiment with and build new consumer experiences. Develop new communications products for your business. Problem solving for a universe of competing activities and sounds. Hardware applications, and more. No need to build an SDK for your audio AI model or DSP algorithms. Got some great audio tech? Let's team up. --- # Voice agents > Build voice-enabled AI agents with a modular audio graph. Real-time speech processing, noise suppression, and voice activity detection for conversational AI apps. Use case Switchboard adds engineering and feature flexibility to AI agent solutions. [Get started free](https://console.switchboard.audio/register) ![Hero image](/_astro/ai-agents_hero-2x-1276x1152px_Z1rrrp8.webp) ![](/_astro/use-cases-audiograph-libraries-native-app_Z1OcGQH.svg) ### Mix and match Switchboard uses nodes—modular containers within the SDK—to make building audio graphs and experimenting with features like speech-to-text, LLMs, and text-to-speech easy. You can try any of the nodes we've already built into Switchboard or request new ones! Switchboard allows you to test and compare results without having to rebuild a graph that connects the features together each time. [See all nodes](/nodes) [![](https://a-us.storyblok.com/f/1008163/0x0/99b2ac7f13/use-cases-ai-agents-mix-and-match-v2.svg)](/nodes) ### Deploy anywhere Gain added flexibility, making it easier to decide where each feature in your graph should live: * in the cloud * on-premise * on-device Run models on your own terms more easily. From prototype through production, Switchboard makes it easier to design, test, and deploy LLM graphs in any configuration. ![](https://a-us.storyblok.com/f/1008163/0x0/6244b1a354/use-cases-on-device-on-cloud-on-premises-reflected.svg) ### Add features easily Switchboard makes it easy to add these into your audio graph and have them run on all of your supported platforms, instantly. * Connect your AI agent into a live call with people * Add music to the background * Experiment with voice changers, languages and accents and effects * Tools to improve audio quality [Play](https://youtube.com/watch?v=Dl4SHFgn4ak) [Play](https://youtube.com/watch?v=TeUxbal67po) [Play](https://youtube.com/watch?v=SvfG45KIlfM) ### Language translation Whether you need a continuous language translation agent for presentations from one to many, or agents that can join phone calls and translate back and forth in both directions, Switchboard can help you build these solutions faster. [Get started](https://console.switchboard.audio/register) [![](https://a-us.storyblok.com/f/1008163/1024x707/95dd523ec0/ai-agents-languagebuddy.webp)](https://console.switchboard.audio/register) ### Customer service Businesses are rushing to have AI Agents solve the problem of poor customer service and waiting on hold. Whether you’re looking for a solution in which AI Agents are able to join human agents on calls, or where the AI can stand alone and replace your Interactive Voice Response system, Switchboard gives you added flexibility and speed to market. [Get started](https://console.switchboard.audio/register) [![](https://a-us.storyblok.com/f/1008163/1024x1024/de5bd5bcf0/ai-agents-customer-service.webp)](https://console.switchboard.audio/register) ### Social and entertainment Building an app to addresses loneliness? Perhaps a watch party, game, or entertainment based app that needs an AI companion to help improve the user experience? Switchboard makes it easy to bring AI Agents into these use cases and to mix the audio alongside other media, as well as other fun features and effects. [Get started](https://console.switchboard.audio/register) [![](https://a-us.storyblok.com/f/1008163/1024x727/b45bcd6699/ai-agents-kosmi-venture-studio.webp)](https://console.switchboard.audio/register) We have multiple options and partners to help you customize solutions. Our partner LiveKit also provides great tools for building AI Agent solutions. Our parent company offers custom engineering services and can build your solution end to end. --- # Team communication > Audio Solutions for Remote Work, Deskless Workers, Creative Industry, and more! Build with Switchboard today Use case Switchboard makes it easy to develop new communications products for business. [Get started free](https://console.switchboard.audio/register) ![Hero image](/_astro/business_hero-2x-1276x1152px_1A5IHV.webp) ### Remote work Go a level beyond Slack Huddles and Teams calls.\ Build casual rooms to stay connected while listening to music. Use voice commands, keyboard shortcuts, and gestures to instantly connect team members together. Embed casual activities like watch parties and games. Switchboard provides a variety of features to support these use cases, and we’ve built numerous apps in this space. ![](https://a-us.storyblok.com/f/1008163/1123x881/111d26cc4e/remote-work.jpg) ### Deskless workers Hands-free walkie talkies for construction, retail, field operations, and more. Use voice commands even in noisy environments. Use trigger words to turn transparency features on and off in noise canceling headphones. Use voice commands to send messages into Slack or Teams. ![](https://a-us.storyblok.com/f/1008163/1123x881/9e1cd957c5/deskless-workers.jpg) ### Creative industry Collaborate while producing music or a film. Share control over a Digital Audio Workstation and listen to the output together while on a voice or video call. Make it sound like you were in the same studio, rather than connected over a Zoom call. ![](https://a-us.storyblok.com/f/1008163/1123x881/3ba068216c/creative-industry.jpg) ### Get started for free Sign up to get free access to Switchboard’s basic prototyping license. [Sign up free →](https://console.switchboard.audio/register) [![](https://a-us.storyblok.com/f/1008163/800x600/9cb3ef94f3/sign-up-free-panel-800x600.webp)](https://console.switchboard.audio/register) --- # Consumer apps > Audio is the hardest part. Switchboard provides the best user experience and the least pain to build! Get started today at no cost! Use case Switchboard is great for experimenting with and building new consumer experiences. [Get started free](https://console.switchboard.audio/register) ![Hero image](/_astro/consumer_hero-2x-1276x1152px_Z1gBjj5.webp) ### Karaoke apps Karaoke apps are full of audio features. You'll need a music player, microphone handling, effects like pitch-correction (autotune), voice changers, perhaps some mechanisms to measure and score your users. And what if you want to do all of this in a live broadcast to an audience, or with a group of friends. Switchboard has you covered. ![](https://a-us.storyblok.com/f/1008163/1123x881/0a8af2c103/couple-singing-karaoke-on-phones.webp) ### Listen with friends Listening to music together is nothing new. But doing it online with integrated voice and video chat is.\ \ Switchboard was the first Listen Party experience involving voice and video chat with smart, voice-activity-based ducking, and our tech remains the most advanced. Control parameters like attack, release, and VAD parameters. Synchronize playback between users.\ \ Whether you’re integrating streaming music services, live shows, radio stations, or podcasts, you can leave the heavy lifting to Switchboard. ![](https://a-us.storyblok.com/f/1008163/1123x881/20c65fd213/listen-with-friends.jpg) ### Watch parties Do you have a video streaming app? Looking for a way to allow your users to watch together? Across all devices, with voice, video, and text chat? Let Switchboard deal with synchronizing, audio issues, mixing media and voice streams, and all the other real time issues you're likely to deal with. ![](https://a-us.storyblok.com/f/1008163/1123x881/b50012a5bd/watch-party-with-kosmi.jpg) [Get started free](https://console.switchboard.audio/register) ### Multiplayer gaming Good sound brings gaming to life. Adding voice chat often ruins that experience, especially when players are connected through multiple platforms (consoles, PC, phones) and aren’t wearing headphones. It doesn’t have to be that way, and you don't have to struggle to provide a better audio solution. ![](https://a-us.storyblok.com/f/1008163/1123x881/873623ceaf/multiplayer-gaming.jpg) [Play](https://youtube.com/watch?v=sPnYIq47LzI) ### Get started for free Sign up to get free access to Switchboard’s basic prototyping license. [Sign up free →](https://console.switchboard.audio/register) [![](https://a-us.storyblok.com/f/1008163/800x600/9cb3ef94f3/sign-up-free-panel-800x600.webp)](https://console.switchboard.audio/register) --- # Consumer electronics > Real-time audio solutions for Smart Speaker, Companion Apps, Augmented Reality, hearable hardware, and much more! Use case [Get started free](https://console.switchboard.audio/register) ![Hero image](/_astro/hardware-and-more-hero-2x-1276x1152px_1JVTly.webp) ### Smart speakers We pioneered the “drop-in” audio experience, listen parties, and other related behaviors. As social listening normalizes, smart speakers are full of untapped potential. We can provide a competitive advantage at the hardware and software level. ![](https://a-us.storyblok.com/f/1008163/1123x881/f4b8968800/echo-listen-party.jpg) ### Companion apps The Switchboard platform includes apps and SDKs that are a perfect starting point for companion apps for headphones and speakers. Leverage all of Switchboard’s features, or select only the pieces you need. We can customize it and build it into your product for you. ![](https://a-us.storyblok.com/f/1008163/200x150/9872859d69/companion-sound-app.svg) ### Augmented reality and hearables Headphones are getting smaller and new form factors have emerged. Audio is the common denominator across all connected devices. * We can embed audio features in resource-constrained hardware * Develop companion apps with one-of-a-kind experiences * Design, develop, and prototype with you ![](https://a-us.storyblok.com/f/1008163/1123x881/898f3a89ed/augmented-reality-hearables.webp) ### Automotive and infotainment Leverage our existing tech stack and expertise to get to market faster. ![](https://a-us.storyblok.com/f/1008163/200x150/b06e2ce730/automotive-and-infotainment.svg) ### Motorcycles and intercom hardware Our technologies enable clear communication, enhancing safety for riders on the road. ![](https://a-us.storyblok.com/f/1008163/1123x881/8d6786124a/motorcycles-intercom-hardware.webp) ### Get started for free Sign up to get free access to Switchboard’s basic prototyping license. [Sign up free →](https://console.switchboard.audio/register) [![](https://a-us.storyblok.com/f/1008163/800x600/9cb3ef94f3/sign-up-free-panel-800x600.webp)](https://console.switchboard.audio/register) --- # Switchboard for Interactive Apps > Open source examples for building interactive audio apps — voice communication, vocal FX, voice changers, and live streaming on iOS and Android. ![](/_astro/hero-landing-bg-gradient_ODlj5.webp) Open source examples for building real-time audio experiences — voice communication, vocal FX, voice changers, and live streaming on iOS and Android. ![Interactive audio app built with Switchboard SDK](/_astro/real-time-voice-apps_1PQRsV.webp) ## Everything you need to build engaging interactive audio Switchboard handles the real-time audio engine — low-latency processing, live FX chains, multi-party communication — so you can focus on creating experiences that keep users engaged. ### Real-time voice communication Build multi-party voice apps with low-latency audio processing, echo cancellation, and noise suppression baked in. ### Live vocal FX chains Apply real-time effects to voice — pitch shifting, reverb, voice changing — with ultra-low latency on device. ### Third-party effect integrations Plug in professional-grade effects from partners like Voicemod to bring studio-quality tools directly into your app. ### Live streaming ready Integrate with Amazon IVS and other streaming platforms to broadcast interactive audio experiences to a live audience. [Explore the docs](https://docs.switchboard.audio) ## Open source repositories Ready-to-run sample apps to start from — clone, build, and make them your own. ## [vocal-fx-chains-sample-ios](https://github.com/switchboard-sdk/vocal-fx-chains-sample-ios) iOS · Swift Experience dynamic audio enhancement with our iOS Vocal FX Chains sample app, utilizing the Switchboard SDK. Elevate vocal performances by crafting intricate FX chains. Vocal FX•Effects•iOS [View on GitHub](https://github.com/switchboard-sdk/vocal-fx-chains-sample-ios) ## [voice-communication-samples-android](https://github.com/switchboard-sdk/voice-communication-samples-android) Android · Kotlin Voice Communication Sample Apps showcasing communication capabilities of the Switchboard SDK. Voice Comms•RTC•Android [View on GitHub](https://github.com/switchboard-sdk/voice-communication-samples-android) ## [voice-communication-samples-ios](https://github.com/switchboard-sdk/voice-communication-samples-ios) iOS · Swift Voice Communication Sample Apps showcasing communication capabilities of the Switchboard SDK. Voice Comms•RTC•iOS [View on GitHub](https://github.com/switchboard-sdk/voice-communication-samples-ios) ## [vocal-fx-chains-sample-android](https://github.com/switchboard-sdk/vocal-fx-chains-sample-android) Android · Kotlin Vocal FX Chains example app for showcasing Switchboard SDK functionality on Android. Vocal FX•Effects•Android [View on GitHub](https://github.com/switchboard-sdk/vocal-fx-chains-sample-android) ## [karaoke-ivs-sample-ios](https://github.com/switchboard-sdk/karaoke-ivs-sample-ios) iOS · Swift Karaoke App with Interactive Audio Effects using the Switchboard SDK and Amazon IVS. Karaoke•Live Streaming•Amazon IVS•iOS [View on GitHub](https://github.com/switchboard-sdk/karaoke-ivs-sample-ios) ## [karaoke-ivs-sample-android](https://github.com/switchboard-sdk/karaoke-ivs-sample-android) Android · Kotlin Karaoke App with Interactive Audio Effects using the Switchboard SDK and Amazon IVS. Karaoke•Live Streaming•Amazon IVS•Android [View on GitHub](https://github.com/switchboard-sdk/karaoke-ivs-sample-android) ## [voice-changer-sample](https://github.com/switchboard-sdk/voice-changer-sample) Cross-platform · C++ A vibe coded voice changer made with Switchboard. Voice Changer•C++•Cross-platform [View on GitHub](https://github.com/switchboard-sdk/voice-changer-sample) The repos on this page were built with Switchboard. They stand alone, but you can also explore the lower level Switchboard repositories [here](https://publicrepos.labs.switchboard.audio/). ### Need to launch faster or build something custom? Our team offers expert consulting and forward-deployed engineers to help you design, build, and ship exactly what you need on your timeline. All of the above can be made available on any platform. ## Get in touch Have a question or want to work together? We'd love to hear from you. [Contact us](https://switchboard.audio/contact/) ## Additional resources ## SDK documentation Full API reference, guides, and tutorials for integrating Switchboard into your app. [Explore](https://docs.switchboard.audio) ## All open source repos Browse every public Switchboard SDK repository — filter by use case or platform. [Browse](https://publicrepos.labs.switchboard.audio/) ## Talk to an engineer Get hands-on help from the Synervoz team to design or accelerate your interactive audio project. [Get in touch](https://switchboard.audio/contact) --- # Metaverse > Your instant voice network. Use case Problem solving for a universe of competing activities and sounds. [Get started free](https://console.switchboard.audio/register) ![Hero image](/_astro/metaverse_hero-2x-1276x1152px_Z2G1US.webp) Audio plays a major role in the metaverse. People talking, shared video screens, games, music, and sound effects can come together to create intriguing new worlds, if done correctly. ![](/_astro/metaverse-3up_140ngk.webp) Spatial audio is only part of the solution. You still have to connect it all together. ### We can help Whether you’re building a web app, VR app, or a device that needs to run embedded audio code, chances are the Switchboard platform contains features that will save you time. And we can help tailor it to your use case. ![](https://a-us.storyblok.com/f/1008163/1123x881/369e6227d2/metaverse-avatar-and-person.jpg) ### Spatial audio Switchboard has extensions that connect to popular spatial audio tools, and we can make custom extensions upon request. ![](https://a-us.storyblok.com/f/1008163/651x557/9819d6d901/ronday-spacial-audio.webp) ### Get started for free Sign up to get free access to Switchboard’s basic prototyping license. [Sign up free →](https://console.switchboard.audio/register) [![](https://a-us.storyblok.com/f/1008163/800x600/9cb3ef94f3/sign-up-free-panel-800x600.webp)](https://console.switchboard.audio/register) --- # Switchboard for Music Apps > Open source music app examples built with the Switchboard SDK — karaoke, DJ, guitar effects, and more for iOS and Android. ![](/_astro/hero-landing-bg-gradient_25yw7H.webp) Open source examples for building music experiences — karaoke, DJ tools, guitar effects, and more — on iOS and Android. ![Music app built with Switchboard SDK](/_astro/music-apps-hero_1UYXWJ.webp) ## Everything you need to build great music apps Switchboard handles the hard audio engineering — real-time FX chains, low-latency mixing, multi-track playback — so you can focus on building experiences your users love. ### Real-time audio effects Apply vocal FX chains, guitar effects, and audio filters with ultra-low latency directly on device. ### Cross-platform support Ship on iOS and Android from a shared audio graph definition — one pipeline, every platform. ### Live streaming ready Integrate with Amazon IVS and other streaming platforms to take your music app to a live audience. ### Third-party integrations Plug in professional effects from partners like Voicemod to add studio-quality tools to your app. ![Switchboard audio pipeline for music apps](/_astro/innovation-our-domain_1Ac2J8.webp) [Explore the docs](https://docs.switchboard.audio) ## Open source repositories Ready-to-run sample apps to start from — clone, build, and make them your own. ## [karaoke-sample-ios](https://github.com/switchboard-sdk/karaoke-sample-ios) iOS · Swift Explore the seamless integration of audio elements through our iOS karaoke sample app using the Switchboard SDK. Effortlessly record your singing voice over a backing track. Karaoke•Recording•iOS [View on GitHub](https://github.com/switchboard-sdk/karaoke-sample-ios) ## [karaoke-sample-android](https://github.com/switchboard-sdk/karaoke-sample-android) Android · Kotlin Explore the seamless integration of audio elements through our Android karaoke sample app using the Switchboard SDK. Effortlessly record your singing voice over a backing track. Karaoke•Recording•Android [View on GitHub](https://github.com/switchboard-sdk/karaoke-sample-android) ## [guitar-effect-sample-android](https://github.com/switchboard-sdk/guitar-effect-sample-android) Android · Kotlin Guitar effect app for Android showcasing Switchboard SDK functionality. Guitar•Audio Effects•Android [View on GitHub](https://github.com/switchboard-sdk/guitar-effect-sample-android) ## [karaoke-ivs-sample-ios](https://github.com/switchboard-sdk/karaoke-ivs-sample-ios) iOS · Swift Karaoke App with Interactive Audio Effects using the Switchboard SDK and Amazon IVS. Karaoke•Live Streaming•Amazon IVS•iOS [View on GitHub](https://github.com/switchboard-sdk/karaoke-ivs-sample-ios) ## [karaoke-ivs-sample-android](https://github.com/switchboard-sdk/karaoke-ivs-sample-android) Android · Kotlin Karaoke App with Interactive Audio Effects using the Switchboard SDK and Amazon IVS. Karaoke•Live Streaming•Amazon IVS•Android [View on GitHub](https://github.com/switchboard-sdk/karaoke-ivs-sample-android) ## [dj-sample-ios](https://github.com/switchboard-sdk/dj-sample-ios) iOS · Swift DJ App for iOS showcasing Switchboard SDK functionality. DJ•Mixing•iOS [View on GitHub](https://github.com/switchboard-sdk/dj-sample-ios) ## [dj-sample-android](https://github.com/switchboard-sdk/dj-sample-android) Android · Kotlin DJ App for Android showcasing Switchboard SDK functionality. DJ•Mixing•Android [View on GitHub](https://github.com/switchboard-sdk/dj-sample-android) The repos on this page were built with Switchboard. They stand alone, but you can also explore the lower level Switchboard repositories [here](https://publicrepos.labs.switchboard.audio/). ### Need to launch faster or build something custom? Our team offers expert consulting and forward-deployed engineers to help you design, build, and ship exactly what you need on your timeline. All of the above can be made available on any platform. ## Get in touch Have a question or want to work together? We'd love to hear from you. [Contact us](https://switchboard.audio/contact/) ## Additional resources ## [SDK documentation](https://docs.switchboard.audio) Full API reference, guides, and tutorials for integrating Switchboard into your app. [Explore](https://docs.switchboard.audio) ## [All open source repos](https://publicrepos.labs.switchboard.audio/) Browse every public Switchboard SDK repository — filter by use case or platform. [Browse](https://publicrepos.labs.switchboard.audio/) ## [Talk to an engineer](https://switchboard.audio/contact) Get hands-on help from the Synervoz team to design or accelerate your music app project. [Get in touch](https://switchboard.audio/contact) --- # On-device speech to text and voice AI > On-device speech recognition SDK for iOS and Android. Process speech locally with no cloud dependency. Supports offline STT, voice commands, and hybrid cloud fallback. Simplify development and unleash creativity * Run high-performance audio graphs entirely on-device * Free models for speech-to-text, text-to-speech, and LLMs * Modular audio processing nodes to build custom workflows * Reduced latency, enhanced privacy, and easy integration ###### TRUSTED BY COMPANIES LIKE * ![](/_astro/bose-logo_Z1coqbj.webp) * ![](/_astro/meta-logo_Z1VUN2T.webp) * ![](/_astro/amazon-logo_Z2eRcQK.webp) * ![](/_astro/superpowered-logo_Z1BanAH.webp) ![](/_astro/on-devive-stt-full-width-hero-v2_1mIANK.webp) Stop paying per minute. * One-time device licensing — perpetual * Zero bandwidth or egress fees * Per device, install, or according to your use case No latency, no network dependency. * Shave off hundreds of milliseconds of latency * Works offline and in poor connectivity * Scales infinitely in any geography Keep audio local and stay in control. * No third-party servers * Simplify compliance and security * Full model and UX ownership ### Are you a developer? This is an iOS example app that uses the Switchboard SDK. It shows you how to combine Whisper STT and Silero VAD into an on-device audio graph for the purpose of building a voice controlled user interface. [iOS Example App](https://docs.switchboard.audio/examples/voice-control-app-ios/) [Play](https://youtube.com/watch?v=2nWcfIaAhx8) *** [Cross-platform STT example](https://docs.switchboard.audio/examples/stt/) ### Reduce costs and improve reliability Cloud STT costs scale with every request. Switchboard doesn’t. With perpetual and per-device licenses you eliminate per-minute billing and network dependencies. [Switchboard SDK Pricing](https://switchboard.audio/pricing/) ### Key benefits * Predictable margins with no surprise usage bills * Perpetual licensing built for OEMs and integrators * Edge-aligned performance in bandwidth-limited environments * Zero downtime during connectivity loss or peak cloud load ### More than just speech-to-text Switchboard isn’t a single-purpose SDK — it’s a full on-device audio runtime designed for the next generation of intelligent products. Build, combine, and scale advanced voice features. Use the same runtime to power: • Voice changers and filters\ • Text-to-speech and LLM integrations\ • Real-time transcription and translation\ • 50+ modular audio features [Explore examples](https://docs.switchboard.audio/examples/) [![](https://a-us.storyblok.com/f/1008163/1024x1024/43bb7e1989/cartesia-cross-platform-dev.webp)](https://docs.switchboard.audio/examples/) | | Cloud STT | On-Device STT | | --------------- | ----------------------------------------------------------------------- | --------------------------- | | **Cost** | Scales with usage | One-time per-device license | | **Latency** | Iterates on speech pipelines for accuracy, latency, and robustness | Instant local inference | | **Privacy** | Builds and tests custom DSP/ML models, audio effects, and signal chains | Fully private and offline | | **Reliability** | Maintains internal SDKs, tooling, or reusable audio frameworks | Works anywhere, anytime | | **Control** | Works in innovation labs to craft new audio-driven experiences | Full model and UX control | *Switchboard gives you enterprise-grade speech performance without dependency, downtime, or data exposure.* ### Why local speech processing matters As AI moves on-device, control is everything. Switchboard lets you deploy voice recognition that’s as private, fast, and scalable as the devices it runs on. **With Switchboard, you get** • A strategic edge in latency-sensitive applications\ • Predictable cost and compliance control\ • Instant scalability from one device to millions.\ • A foundation for local-first voice AI experiences ![](https://a-us.storyblok.com/f/1008163/x/951ebf7743/woman-speaking-toai-on-cell.avif) * Speech-to-text, text-to-speech, language models, voice changers, and more - Deploy to any platform with simple OS-specific bindings * Hybrid options including cloud-first models with on-device fallback - Customization and support are available ### Scale without limits Deploy speech features anywhere — from prototype to mass production — with no cloud dependencies or performance bottlenecks. * Unlimited concurrent users — every device runs its own model * Zero centralized compute — no shared API or latency spikes * Consistent performance in offline and high-latency environments * Hybrid / cloud-connected options are available as needed for larger models. * Auto-failover solutions are available. ![](https://a-us.storyblok.com/f/1008163/x/47c175f834/man-tapping-headphones.avif) --- # On-device text-to-speech and voice AI > On-device speech generation SDK for iOS and Android. Generate speech locally with no cloud dependency. Supports offline TTS, voice commands, and hybrid cloud fallback. Simplify development and unleash creativity * Run high-performance audio graphs entirely on-device * Free models for speech-to-text, text-to-speech, and LLMs * Modular audio processing nodes to build custom workflows * Reduced latency, enhanced privacy, and easy integration ###### TRUSTED BY COMPANIES LIKE * ![](/_astro/bose-logo_Z1coqbj.webp) * ![](/_astro/meta-logo_Z1VUN2T.webp) * ![](/_astro/amazon-logo_Z2eRcQK.webp) * ![](/_astro/superpowered-logo_Z1BanAH.webp) ![](/_astro/on-devive-stt-full-width-hero-v2_1mIANK.webp) Stop paying per minute. * One-time device licensing — perpetual * Zero bandwidth or egress fees * Per device, install, or according to your use case No latency, no network dependency. * Shave off hundreds of milliseconds of latency * Works offline and in poor connectivity * Scales infinitely in any geography Keep audio local and stay in control. * No third-party servers * Simplify compliance and security * Full model and UX ownership ### Are you a developer? This is an iOS example app that uses the Switchboard SDK. It shows you how to combine Whisper STT and Silero VAD into an on-device audio graph for the purpose of building a voice controlled user interface. [iOS Example App](https://docs.switchboard.audio/examples/voice-control-app-ios/) [Play](https://youtube.com/watch?v=2nWcfIaAhx8) *** [Cross-platform STT example](https://docs.switchboard.audio/examples/stt/) ### Reduce costs and improve reliability Cloud STT costs scale with every request. Switchboard doesn’t. With perpetual and per-device licenses you eliminate per-minute billing and network dependencies. [Switchboard SDK Pricing](https://switchboard.audio/pricing/) ### Key benefits * Predictable margins with no surprise usage bills * Perpetual licensing built for OEMs and integrators * Edge-aligned performance in bandwidth-limited environments * Zero downtime during connectivity loss or peak cloud load ### More than just speech-to-text Switchboard isn’t a single-purpose SDK — it’s a full on-device audio runtime designed for the next generation of intelligent products. Build, combine, and scale advanced voice features. Use the same runtime to power: • Voice changers and filters\ • Text-to-speech and LLM integrations\ • Real-time transcription and translation\ • 50+ modular audio features [Explore examples](https://docs.switchboard.audio/examples/) [![](https://a-us.storyblok.com/f/1008163/1024x1024/43bb7e1989/cartesia-cross-platform-dev.webp)](https://docs.switchboard.audio/examples/) | | Cloud STT | On-Device STT | | --------------- | ----------------------------------------------------------------------- | --------------------------- | | **Cost** | Scales with usage | One-time per-device license | | **Latency** | Iterates on speech pipelines for accuracy, latency, and robustness | Instant local inference | | **Privacy** | Builds and tests custom DSP/ML models, audio effects, and signal chains | Fully private and offline | | **Reliability** | Maintains internal SDKs, tooling, or reusable audio frameworks | Works anywhere, anytime | | **Control** | Works in innovation labs to craft new audio-driven experiences | Full model and UX control | *Switchboard gives you enterprise-grade speech performance without dependency, downtime, or data exposure.* ### Why local speech processing matters As AI moves on-device, control is everything. Switchboard lets you deploy voice recognition that’s as private, fast, and scalable as the devices it runs on. **With Switchboard, you get** • A strategic edge in latency-sensitive applications\ • Predictable cost and compliance control\ • Instant scalability from one device to millions.\ • A foundation for local-first voice AI experiences ![](https://a-us.storyblok.com/f/1008163/x/951ebf7743/woman-speaking-toai-on-cell.avif) * Speech-to-text, text-to-speech, language models, voice changers, and more - Deploy to any platform with simple OS-specific bindings * Hybrid options including cloud-first models with on-device fallback - Customization and support are available ### Scale without limits Deploy speech features anywhere — from prototype to mass production — with no cloud dependencies or performance bottlenecks. * Unlimited concurrent users — every device runs its own model * Zero centralized compute — no shared API or latency spikes * Consistent performance in offline and high-latency environments * Hybrid / cloud-connected options are available as needed for larger models. * Auto-failover solutions are available. ![](https://a-us.storyblok.com/f/1008163/x/47c175f834/man-tapping-headphones.avif) --- # Switchboard for Voice AI > Open source tools and examples to help voice AI developers reduce costs, improve latency, enhance privacy, and enable offline functionality. ![](/_astro/hero-bg-gradient_ZraxIg.webp) Open source tools and examples for voice AI developers. ![](/_astro/orb-and-background_1o38DR.webp) Speech Recognition Text to Speech LLM Integration Echo Cancellation Voice Activity Detection Noise Suppression Turn Detection Speaker Isolation Speech Recognition Text to Speech LLM Integration Echo Cancellation Voice Activity Detection Noise Suppression Turn Detection Speaker Isolation ## Built to address real-world constraints Switchboard for Voice AI uses a hybrid on-device + cloud architecture to help you get the best of both worlds. On-device processing with hand-off to cloud only when necessary. ### Reduce costs Process audio on-device to minimize expensive API calls and bandwidth usage. ### Improve latency Local processing eliminates network round-trips for near-instant voice interactions. ### Enhance privacy Keep sensitive audio data on-device and send only processed text to the cloud. ### Enable offline Build voice AI features that work without an internet connection. ![Hybrid on-device + cloud Voice AI architecture](/_astro/voice-ai-hybrid_ZNccNy.webp) [Learn more](https://switchboard.audio/hub/your-voice-ai-bill-is-telling-you-to-go-hybrid-on-device/) ## Open source repositories Production-ready examples and reusable components to accelerate your voice AI development ## [OpenAI Realtime Toolkit](https://github.com/switchboard-sdk/openai-realtime-toolkit) React Native Build voice-first AI agents with OpenAI Realtime and native audio primitives. On-device voice activity detection, turn detection, and barge-in handling make real-time conversations responsive while simplifying the audio pipeline. OpenAI Realtime•VAD•Turn Detection•Tool Calling [View on GitHub](https://github.com/switchboard-sdk/openai-realtime-toolkit) ## [EdgeSpeech](https://github.com/switchboard-sdk/edgespeech) React Native On device speech recognition (ASR / STT) and text-to-speech (TTS) so that you can cut costs and latency while simplifying cloud infra. You only send text to the LLM so don't have to worry about webRTC, sockets, or scaling audio in the cloud. STT (local)•LLM (cloud)•TTS (local) [View on GitHub](https://github.com/switchboard-sdk/edgespeech) ## [Hybrid Voice Agent](https://github.com/switchboard-sdk/hybrid-voice-agent-sample) iOS Build a full voice agent on iOS, running on-device with seamless cloud handoff when you need more intelligence. Switch between local and cloud models without losing conversation context, while keeping audio processing on the iPhone. STT (local)•LLM (local)•LLM (cloud)•TTS (local) [View on GitHub](https://github.com/switchboard-sdk/hybrid-voice-agent-sample) ## EdgeAudio Coming soon Swift, Kotlin On-device preprocessing for speech to speech models (aka S2S or audio models). On device voice activity detection (VAD), echo cancellation, and other audio preprocessing runs locally before connecting to cloud-based speech model (such as OpenAI Realtime API) to optimize performance. VAD•Echo Cancellation•Specific Speaker Recognition•OpenAI Realtime API ## EdgeWhisper Coming soon iOS, Android, macOS, Windows, Linux Run OpenAI's Whisper speech recognition (ASR) model entirely on-device for maximum privacy and offline functionality across mobile and desktop platforms (iOS, Android, mac, Windows, Linux). Whisper (local ASR) ## EdgeAgent Coming soon React Native Run a full STT-LLM-TTS pipeline locally. The STT and LLM components each have optional hand-off (or fallback) to cloud alternatives. STT (local)•LLM (local)•TTS (local) ## EmbeddedVoice Coming soon Linux (& custom upon request) Optimized voice AI components for resource-constrained IoT and embedded systems, including smart speakers, wearables, and edge devices. Embedded•IoT•Edge Computing ### Need to launch faster or build something custom? Our team offers expert consulting and forward-deployed engineers to help you design, build, and ship exactly what you need on your timeline. All of the above can be made available on any platform. ## Get in touch Have a question or want to work together? We'd love to hear from you. [Contact us](https://switchboard.audio/contact/) ## Additional resources ## [Voice AI resource hub](https://switchboard.audio/hub/) Guides, tutorials, and best practices for building production voice AI applications. [Explore](https://switchboard.audio/hub/) ## [On-device speech recognition](https://switchboard.audio/cases/on-device-stt) Learn how companies are implementing local speech-to-text to reduce costs and improve privacy. [Explore](https://switchboard.audio/cases/on-device-stt) ## [Building voice AI agents](https://switchboard.audio/cases/ai-agent) Explore real-world implementations of conversational AI agents with voice interfaces. [Explore](https://switchboard.audio/cases/ai-agent) --- # Let's talk > Contact sales. Get in touch. --- # Demos > Try live demos of the SDK in action: voice-chat, music parties, embedded audio workflows. A collection of demos built using our Switchboard SDK. These examples showcase some of the capabilities and potential applications of Switchboard, but these are a drop in the bucket. Simplify the process of creating a karaoke app. Mix tracks, apply effects and sync beats in real-time. Guitar effect app showcasing Switchboard SDK functionality. Implement the audio pipeline of a simple online radio app, with easy integration. This example plays an audio file and ducks the playback volume based on the user's microphone input. This example adds a reverb effect that simulates the natural reverberation of sound in a physical space. [View more examples ->](https://docs.switchboard.audio/docs/examples) ### Get started for free Sign up to get free access to Switchboard’s basic prototyping license. [Sign up free →](https://console.switchboard.audio/register) [![](https://a-us.storyblok.com/f/1008163/800x600/9cb3ef94f3/sign-up-free-panel-800x600.webp)](https://console.switchboard.audio/register) --- # Switchboard Editor > Build real-time audio solutions for business, consumer, or hardware with a modular SDK! Save time and money. Easy to use - get started for free! Build and test right in your browser. ![](/_astro/switchboard-landing-hero-v3-1920x1180_Z1E3IyR.webp) The Switchboard Editor helps you rapidly experiment, prototype, and design new audio and voice features. Easily connect the latest open-source and proprietary voice and audio tools in unique combinations, and instantly test the results right from your browser. The audio graphs are built on top of the Switchboard SDK (with a C/C++ core) so they can subsequently be deployed across many platforms. The audio engine runs on-device, though individual nodes in the graph can be run on-device or in the cloud, depending on the node. ![](https://a-us.storyblok.com/f/1008163/x/72dab72c83/landing-nodes-library-sidebar-square.avif) [Play](https://youtube.com/watch?v=n0JVCxd1h6Q) ## Voice AI See how to build a local LLM audio engine in the Switchboard Editor. This demo connects Whisper, Silero VAD, Llama, and TTS nodes in the Switchboard Editor, all running locally (**on-device**) in this case. Switchboard makes it easy to use a combination of local and cloud nodes to construct hybrid graphs, depending on the use case. We can do this with any of the latest models (STT, LLM, TTS, S2S, and many other node types). [Play](https://youtube.com/watch?v=r-N-vUbZ_CY) ## Driving Buddy This demo shows a graph we made for our example app called Driving Buddy. There is a live microphone input routed to Open AI for conversation. There is also a music input. The audio streams are mixed, and the music ducks when voice is detected, allowing for casual conversation with driving buddy to help keep you awake on long drives. Similar graphs are also useful for listen parties, watch parties, and other interactive apps. [Play](https://youtube.com/watch?v=pW50tqMboxc) ## Real time source separation This demo brings Audioshake's industry-leading real time source separation into a Switchboard graph. Each stem can have independent effects / processing chains (example uses gain nodes to illustrate). Graphs of this sort can power many creative music, media, and real-time broadcasting workflows (on any platform). Switchboard contains many complementary music, media, DSP, AI, and other nodes. ### Your creative amplifier Think of it as Unity or Unreal, but for voice & audio apps. Ideal for: * Innovation and R\&D teams * Product managers * Engineers * Ideas-people * Rapid prototypers * Anyone eager to explore and compare the latest in voice and audio AI ![](https://a-us.storyblok.com/f/1008163/x/c3aa626b69/creative-amplifier-nodes-visual.avif) ![](/_astro/editor-in-beta_1WcKkE.webp) Currently the Editor is in closed beta Learn more about [our plans](/hub/introducing-switchboard-editor). ![](/_astro/editor-sdk-free-tier_Z2wxX70.webp) If you’re a developer, you can start on your own The [SDK](/sdk) includes a free tier you can use immediately. Fill out the form to join the waitlist and secure early access, request a demo, or if you’d like to discuss consulting and co-development opportunities (see [Switchboard Labs](/labs)). --- # Frequently Asked Questions > Frequently Asked Questions about Switchboard Audio SDK What is Switchboard? Switchboard is a modular audio SDK and real-time runtime that makes it easy to build audio- and voice-powered applications. It helps developers create low-latency, AI-enabled audio experiences that can run on-device, in the cloud, or both. How does it work? Switchboard uses an audio graph model. Each graph is a JSON-defined pipeline of modular audio building blocks like STT, TTS, noise suppression, voice changers, etc. These blocks can be chained together and run in real time. What does it do? Switchboard is a modular audio framework incorporating a large library of audio **nodes** `AudioNode`. These nodes can easily be put together into **audio graphs** `AudioGraph`. Switchboard passes these graphs into natively compiled C++ code that runs *fast,* across multiple platforms including iOS, Android, Mac, Windows, web, and embedded platforms. Audio graphs can be defined in JSON or native languages, making them easy to build while providing complete flexibility at runtime. Switchboard also has a visual (no-code) node-based editor, the Switchboard Editor, which generates JSON automatically for rapid configuration of audio graphs that can easily be designed and tested in the browser and deployed to many target platforms. It also allows you to tune parameters for each node or change the audio graph at runtime, rapidly speeding up development cycles. What are audio nodes? Audio nodes are essentially modular containers that process audio and include a wide range of functions such as: *speech to text, text to speech, large language models, voice changers, media players, streaming, voice* and *video calling*, the ability to *mix* and *split audio streams, DSP effects*, and much more. Switchboard also includes [Extensions](https://docs.switchboard.audio/extensions/) to many other popular audio tools and SDKs, both open and closed source. What platforms does Switchboard support? Switchboard runs on iOS, Android, macOS, Windows, Linux, and the web. It supports edge compute, embedded systems, and hybrid architectures. Why should I use Switchboard? Whether you’re using it for a single node such as *speech-to-text* or a *voice changer*, or stringing multiple nodes together, Switchboard will: * make it faster and easier to build, test, experiment, and get to market * make it easier to maintain and make changes later * save time and money * drive new revenue and growth by simplifying the addition of new features Who is Switchboard designed for? * AI Agent and Voice AI solutions developers * Real Time Communications (RTC) applications * Music and social app developers * R\&D / ML / AI teams looking to commercialize * Hardware projects like headphones, speakers, wearables, etc. * Pro audio industry - hardware and software * Apps & SDKs with voice, VoIP, media players, and other audio features * Product managers looking for a no code prototyping tool (Switchboard Editor) See *Use Cases* in main menu for more details. Can I use Switchboard for building AI voice agents? Yes. Switchboard is ideal for running LLM-powered voice agents on-device or hybrid. It supports real-time STT → LLM → TTS pipelines, and gives you full control over audio routing, DSP, and model selection. How does Switchboard help with voice interfaces in mobile apps? Switchboard lets you embed voice control, transcription, and audio effects into mobile apps without relying on cloud APIs. It supports real-time pipelines optimized for low latency and power usage. Is Switchboard good for noise suppression or voice enhancement? Yes. You can bring your own noise suppression model (or use built-ins), and chain it with compressors, equalizers, or echo cancellation using Switchboard’s modular graph. Can I use Switchboard for building multiplayer audio experiences? Yes. Switchboard is compatible with WebRTC, LiveKit, Agora, and other voice and video chat frameworks. You can use it to build social audio rooms, watch parties, or spatial audio multiplayer environments. How is Switchboard different from Agora or Twilio? Agora and Twilio focus on cloud-hosted communications. Switchboard gives you control over audio processing, routing, and ML—locally or hybrid—making it ideal for more advanced or privacy-sensitive applications. Agora and other VoIP services are available as nodes in Switchboard. They normally function as an audio source (e.g. you can take the audio from a voice / video chat room) or a sink (you can put audio into the room). In short, you wouldn’t use Switchboard instead of either of these services, you’d use it in addition. How is Switchboard different from Vapi? Vapi and Switchboard serve different purposes. Vapi focuses on helping developers build agents, while Switchboard helps developers build audio graphs. You could potentially use both Vapi and Switchboard in your project. For example, you might prototype an agent in Vapi. You might then use Switchboard to optimize an on-device audio graph to save on API costs or introduce more flexibility or to process audio in other ways that are not provided for as part of the Vapi platform. Switchboard is a modular audio SDK and runtime that gives developers full control over the audio pipeline, letting them build custom real-time graphs with components like STT, TTS, voice changers, and noise suppression that run on-device, in the cloud, or in hybrid mode. In contrast, Vapi is a hosted voice agent platform focused on telephony use cases, offering a pre-built stack that integrates APIs like ElevenLabs and Deepgram but limits customization and runs entirely in the cloud. How is Switchboard different from Deepgram or AssemblyAI? Switchboard is an SDK and runtime for building entire audio pipelines, while Deepgram and Assembly AI focus on individual components such as speech recognition. So, whereas companies like Deepgram and AssemblyAI provide their own speech to text (STT), text to speech (TTS), and related nodes, Switchboard lets developers combine these nodes, as well as many others (voice, music, and other audio related nodes) in a real-time audio graphs that can run on-device, in the cloud, or hybrid. Deepgram and AssemblyAI are typically cloud-only APIs that specialize in speech-to-text (and some related features) with little control over the underlying pipeline. With Switchboard, you can bring your own models, run them offline, chain multiple models or effects, and integrate with WebRTC or embedded devices—offering far greater flexibility, privacy, and lower latency than relying solely on hosted APIs. That said, you can run Switchboard graphs with nodes from Deepgram or AssemblyAI. How is Switchboard different from Cartesia, or Eleven Labs? Switchboard is an SDK and runtime for building entire audio pipelines, while Cartesia and Eleven Labs focus on individual components such as speech recognition or text to speech. You can run Switchboard graphs that use Cartesia or Eleven Labs nodes. For example, you can use the [Cartesia extension](https://switchboard.audio/partners/cartesia/) for Switchboard. How does Switchboard compare to LiveKit or Daily.co? Switchboard and LiveKit solve different layers of the real-time media stack. LiveKit is a powerful open-source infrastructure for media transport—handling audio/video routing, SFU/relay, and room-based sessions over WebRTC. Switchboard, on the other hand, is focused on audio processing and orchestration at the endpoint: voice activity detection, STT, TTS, noise suppression, LLM integration, and more. You can use them together—Switchboard for local audio graph processing (on-device or with cloud-connected nodes), and LiveKit to handle routing audio between multiple users. Switchboard doesn't replace LiveKit—it extends it by giving developers real-time, programmable control over what happens to the audio before or after it's streamed (see our [Partners/LiveKit](https://www.notion.so/partners/livekit) ). It’s similar with Daily.co and other webRTC providers. We offer some [Extensions](https://docs.switchboard.audio/extensions/) already and will continue adding more. How does Switchboard compare to JUCE? JUCE is a low-level C++ framework primarily used for building cross-platform audio applications, especially VST plugins and desktop DAWs. It provides powerful building blocks for UI, audio I/O, and DSP, but it requires deep expertise in C++ and lacks out-of-the-box support for real-time voice AI, on-device STT/TTS, or modular graph-based orchestration. Switchboard, by contrast, is a higher-level SDK and runtime focused on real-time audio and AI pipelines—with language bindings in Swift, Kotlin, and JS, plus support for live graph editing, bring-your-own-models, and hybrid cloud/on-device execution. Switchboard accelerates development for teams building voice-controlled apps, agents, or DSP tools—without the boilerplate and low-level threading work required in JUCE. How does Switchboard compare to PortAudio? PortAudio is a low-level cross-platform audio I/O library used to route audio to and from hardware devices. It’s ideal for simple stream management, but it offers no built-in DSP, audio graph management, or support for real-time speech and AI tasks. Switchboard, on the other hand, offers higher level abstractions—providing a modular audio runtime with real-time graph orchestration, support for AI models (STT, TTS, voice changers), and tight integration with mobile, desktop, web, and embedded platforms. While PortAudio is like wiring raw audio cables, Switchboard is like building a fully programmable audio processing studio—letting developers prototype and deploy advanced pipelines without writing audio I/O code from scratch. How do I start using Switchboard? You can start using Switchboard in two complementary ways: ### 1. SDK Libraries Visit our Docs [Downloads](https://docs.switchboard.audio/downloads/) page, and select the individual libraries needed for your platform. ### 2. Switchboard Editor Use the [Switchboard Editor](https://editor.switchboard.audio/), a browser-based tool to visually construct and test audio pipelines without writing code. These methods work well together. You can visually prototype your audio pipeline using the Editor, which generates JSON configurations. Then, implement these configurations with the SDK libraries on your target platforms. This approach ensures your audio experiences sound consistent everywhere. Note that while the Editor covers most functionalities, you may still need to use SDK libraries directly for advanced capabilities. How do I define an audio graph in Switchboard? Graphs are defined using a simple JSON schema. Each node specifies a module (e.g. TTS, media player, effect), its config, and its connections. You can create graphs by hand or with the Switchboard Graph Editor (GUI). What languages does Switchboard support? * Swift / Kotlin / JavaScript (bindings) * C++ core engine * Python * WebAssembly (limited) * Switchboard also supports frameworks such as React Native, Flutter, Unity etc. See our [SDK API reference](https://docs.switchboard.audio/api/) and [integration](https://docs.switchboard.audio/category/integration/). Where can I find docs and examples? All documentation, SDKs, and example graphs are on the [Switchboard Docs Portal](https://docs.switchboard.audio/). You’ll find copy-pasteable starter graphs, pre-built modules, and real-time testing tools. How do I get help from the Switchboard team? For general inquiries and help, submit a request on our [Contact](https://www.notion.so/contact) page. For existing customers, you can contact us directly via email. For enterprise customers we’ll be happy to set up a private Slack channel for live support. How do I install Switchboard? Use npm (for JS), CocoaPods (iOS), or Gradle (Android). Prebuilt binaries are also available for C++, with support for native integration. If you can’t find the answer in our [Docs](https://docs.switchboard.audio/) portal, please [Contact us](https://www.notion.so/contact). How do you mitigate dependency risk? * Some parts of the SDK are Open Source. See [Introduction](https://docs.switchboard.audio/docs/introduction/) page. * We offer world class support and broader [consulting services](https://synervoz-website-astro.pages.dev/services/product-consulting) through our parent company, Synervoz. Having this support interface (e.g. via a shared Slack channel) is a huge help for urgent requests and should help mitigate risk overall. * Escrow is another option we are open to, but only for sizable enterprise deals. Does Switchboard support BYO models? Yes. You can plug in your own models, compiled to ONNX or other formats. Many open-source models already work out of the box (e.g. Whisper, Silero, RNNoise, etc.), and Switchboard provides extension templates for C++ and WebAssembly. How do I use Switchboard with LiveKit or WebRTC? Switchboard can ingest or output audio to WebRTC-compatible frameworks like LiveKit using media streams. Example integrations and a full LiveKit demo are available in the[ Switchboard GitHub repo](https://github.com/switchboard-sdk) and in the [Examples](https://docs.switchboard.audio/examples/) page of our docs portal. Can I use Switchboard offline? Yes. Switchboard supports fully offline audio processing pipelines for mobile, desktop, and embedded. This is perfect for apps that need privacy, low latency, or function in poor network conditions. Many nodes in Switchboard are on-device capable. How much does it cost? Switchboard has a free tier up to 20K activations and Commercial Licenses available thereafter. Consult our [Pricing](https://www.notion.so/pricing) page. How does pricing work for third party extensions? Typically you would contract directly with the third party extension provider. Nevertheless we have partnerships with certain providers so we encourage you to [get in touch](https://synervoz.com/contact) to discuss your use case and the extensions you’d like to use, as we may be able to help. Can I get access to source code? * Some aspects of Switchboard are already Open Source. You can learn more about this on our [Introduction](https://docs.switchboard.audio/docs/introduction/) page. * We also provide [Example apps](https://docs.switchboard.audio/docs/examples) that are Open Source. * We plan to continue our FOSS contributions and welcome partners to [get in touch](https://synervoz.com/contact). * For the closed-source aspects of Switchboard, we may offer partial source code licenses based on your needs. This option is generally cheaper, faster, and more robust than developing it yourself. Pricing typically ranges from $XX,000 to $XXX,000, depending on the modules required. Is there a free tier or trial? Yes. Switchboard offers a free tier with full access to the SDK and limited usage caps. Switchboard supports many independent developers and zero cost prototyping through to launch. Generally limits are only hit as deployed apps start to scale, or if support is required. Can I use Switchboard in a commercial product? Yes, Switchboard’s licensing is designed for commercial deployment. It scales from indie apps to enterprise use cases. See our [pricing](https://www.notion.so/pricing) page. Is Switchboard open source? We open source a lot of example apps, starter templates, and other tools including the bring your own extension framework. The platform specific SDKs are generally closed source, though parts of it may be open sourced in the near future. The full commercial SDK is available under a flexible developer-friendly license. Partial source code licenses are also an option for larger customers. Contact the Switchboard team for custom pricing. Is Switchboard proprietary? Yes, Switchboard is a proprietary technology. However, we offer a free tier — please see our [Pricing](https://www.notion.so/pricing) page and [Master License Agreement](https://www.notion.so/licensing) for more details. We also provide open source example applications. You are responsible for reading the license files contained with any distribution package. We do our best to keep things simple and reasonable, while also supporting the open source community where possible. Depending on what you build, it might also require a patent license. See [Patents](https://synervoz.com/patents) page. Do you offer a warranty or Service Level Agreement? We can provide this option if needed, but it is available only at the enterprise tier and will incur additional costs. Otherwise, we offer support on a best-effort basis, and typically recommend a monthly support package to address this concern. Do you also offer design and development services? Yes we provide a variety of [consulting services](https://synervoz-website-astro.pages.dev/services/product-consulting) through our parent company, Synervoz. Will Switchboard support embedded systems? Yes. Embedded is a key target for Switchboard. Current builds already support ARM platforms and edge deployments. Optimizations for low-power devices are actively being developed. Is Switchboard planning support for RAG pipelines or streaming LLM inference? Yes. Switchboard already supports RAG and hybrid inference using local + cloud LLMs. Switchboard is not focusing on providing tools for prompt orchestration or conversational memory but these features can be obtained from other platforms, and your agents can be connected into your Switchboard audio graph. We are also able to help with custom development as needed. [Contact us](https://www.notion.so/contact) to learn more. Can I run multiple LLMs in parallel in Switchboard? Yes. The modular graph system allows orchestration across multiple models. So, for example, you might run a real time speech to speech pipeline in parallel with a speech to text / LLM pipeline that records and summarizes the conversation. --- # Audio Glossary > Why we built Switchboard: the audio SDK that eliminates the pain of cross-platform audio development. From echo cancellation to voice effects, one SDK handles it all. A practical guide to the terminology behind audio and voice AI. ![](/_astro/glossary-cover-1600x900_ZwUwSa.webp) ## Acoustic Echo Cancellation (AEC) When a device both plays sound and listens at the same time, it risks hearing itself. Acoustic Echo Cancellation (AEC) solves this by removing speaker output from microphone input in real time. Without it, voice systems can spiral into feedback loops, repeating their own speech. In practice, AEC is essential for devices like phones, earbuds, and smart speakers. Modern systems use adaptive filtering to continuously model how sound travels from speaker to mic. However, distortion—especially from small, overdriven speakers—can confuse these models, leaving traces of echo behind. ## Audio Graph Think of an audio graph as a flowchart for sound. It’s a network of small, specialized components (nodes), each handling one task—like capturing audio, reducing noise, transcribing speech, or generating responses. This modular design makes voice systems flexible. Want to swap a speech model or add a new feature? Just replace or insert a node instead of rebuilding everything. Audio graphs are what make experimentation and rapid iteration possible in modern voice AI systems. ## Automatic Speech Recognition (ASR) ASR is what turns spoken words into text. It’s the entry point for most voice systems—everything downstream depends on how well it performs. There are two primary architectural approaches: **batch** (waits until speech ends) and **streaming** (transcribes as you speak). Streaming feels faster but can introduce mistakes mid-sentence. Accuracy is measured using Word Error Rate (WER), and even small errors can cascade into poor responses. In short: better input audio and better models lead to better conversations. ## Barge-In Barge-in is what happens when a user interrupts a system mid-response—and how the system handles it matters more than you’d expect. A good system stops speaking almost instantly and shifts attention to the user. A bad one keeps talking, creating frustration. Technically, this requires constant listening, fast speech detection, and immediate audio shutdown. The key metric here is barge-in latency—the delay between user speech and system silence. Even small delays can make an otherwise fast system feel broken. ## Bit Rate Bit rate measures how much audio data is transmitted per second, usually in kilobits per second (kbps). It directly affects both sound quality and bandwidth usage. Higher bit rates preserve more detail, improving transcription accuracy and speech naturalness. Lower bit rates reduce data usage and latency but sacrifice clarity. Choosing the right bit rate is a balancing act—too high wastes resources, too low degrades performance. In voice AI, this decision often depends on network conditions and application requirements. ## Buffer A buffer is temporary storage for audio data as it moves through a system. It helps smooth out timing mismatches between recording, processing, and playback. Buffer size directly affects performance: * Larger buffers = more stability but more delay * Smaller buffers = faster response but risk of glitches In real-time voice systems, tuning buffers is critical. Too much delay breaks conversation flow; too little causes audio dropouts. It’s one of the most important—and often overlooked—latency controls. ## Call Center Automation Call center automation replaces or augments human agents with voice AI systems that can handle real conversations. Unlike traditional phone menus, these systems understand natural language and can complete tasks end-to-end. They must integrate with telephony systems, backend APIs, and escalation workflows. The real advantage comes from flexibility—teams can update behavior or swap models without rebuilding the system. As AI improves, these systems are rapidly becoming the default for high-volume customer interactions. ## Component Pipeline A component pipeline splits voice AI into three stages: speech-to-text (ASR), reasoning (LLM), and text-to-speech (TTS). Each runs independently. This modularity makes systems easy to update and customize. But there’s a cost: latency adds up across each step, and important vocal cues—tone, hesitation—are lost once audio becomes text. It’s a practical, widely used architecture, but increasingly being challenged by newer approaches that keep audio intact throughout processing. ## Context Window The context window defines how much information a model can consider at once—everything from conversation history to system instructions. As conversations grow, this space fills up, increasing processing time and forcing tradeoffs. Systems often deal with this by summarizing older content or trimming less relevant parts. A larger context can improve understanding but also increases cost and latency. Designing around this limitation is one of the key challenges in building scalable conversational systems. ## Convolutional Neural Network (CNN) CNNs are neural networks designed to detect patterns in structured data. In voice AI, they’re often applied to spectrograms—visual representations of sound. They excel at identifying features like phonemes, keywords, or noise patterns without manual feature engineering. While newer architectures dominate large-scale models, CNNs remain important for tasks like keyword spotting and audio classification, where efficiency and speed matter. ## Cross-Platform Deployment Cross-platform deployment means building a voice system once and running it across devices—phones, desktops, embedded hardware—without rewriting everything. This is harder than it sounds. Each platform handles audio differently, from buffering to hardware access. A good abstraction layer hides these differences, letting developers focus on features instead of platform quirks. Without it, maintaining separate implementations quickly becomes unmanageable. ## Deep Neural Network (DNN) A deep neural network is simply a neural network with many layers, allowing it to learn complex patterns. In voice AI, DNNs power everything from speech recognition to speech synthesis. Their depth enables them to capture subtle acoustic and linguistic relationships that simpler models miss. They’re the foundation of modern AI systems, but their performance depends heavily on training data, architecture, and compute resources. ## Digital Signal Processing (DSP) DSP is the math that cleans up audio before AI models ever see it. It includes noise reduction, echo cancellation, and signal enhancement. Good DSP dramatically improves transcription accuracy—bad audio leads to bad results downstream. In real-world environments (cars, factories, homes), DSP isn’t optional. The challenge is doing all this processing in real time without adding noticeable delay. ## Edge / On-Device Inference Edge inference means running AI models directly on a device instead of the cloud. This reduces latency, improves privacy, and enables offline use. But it comes with constraints—models must be smaller and more efficient. The tradeoff is clear: cloud models are more powerful, but on-device models are faster and more secure. Many systems now combine both approaches. ## Endpointing Endpointing decides when a user has finished speaking. It builds on basic voice detection by adding timing and context awareness. Get it wrong, and the system either interrupts users or leaves awkward silence. Modern approaches combine acoustic signals with language understanding to improve accuracy. It’s a subtle feature, but it has a huge impact on how natural a conversation feels. ## Formant Preservation When you change the pitch of a voice, you risk making it sound unnatural. Formant preservation fixes this by maintaining the voice’s tonal characteristics. Without it, voices can sound cartoonish or distorted. With it, pitch changes feel realistic and human-like. It’s a key technique in voice transformation systems, especially for real-time applications. ## Frame Every audio pipeline chops a continuous sound stream into small chunks—called frames—for processing. Frame size is simply how long each chunk is, measured in milliseconds. Smaller frames mean faster, more responsive processing but demand more compute per second. Larger frames are more efficient but add delay. Most voice AI systems land somewhere between 10 and 30 milliseconds—a sweet spot that balances responsiveness with practical hardware demands. ## Frequency Frequency is the rate at which a sound wave completes one full cycle per second, measured in hertz (Hz). It's essentially what we perceive as pitch—low frequencies sound deep, high frequencies sound bright. Human speech sits roughly between 80 Hz and 8 kHz, though most telephone systems narrow that to 300–3,400 Hz. In voice AI, understanding frequency content informs everything from microphone selection to filter design to why certain ASR models perform better on some audio sources than others. ## Hertz Hertz (Hz) is the unit of frequency—one cycle per second. In audio, you encounter it in two distinct contexts: as a measure of pitch (440 Hz is the musical note A4) and as a measure of sample rate (16,000 Hz means 16,000 audio snapshots captured every second). Voice AI developers run into hertz constantly—when specifying audio formats, configuring DSP filters, or checking that a microphone's output matches what an ASR model expects as input. ## Hybrid Cloud / On-Device Architecture A hybrid architecture splits processing between the device and the cloud, routing each task to wherever it runs best. Latency-sensitive or privacy-critical stages run locally on the device; heavier computation runs in the cloud. This flexibility makes hybrid deployments attractive for consumer hardware and regulated industries alike—but it comes with real design complexity. The boundary between environments needs to be explicit, handoffs need to be fast, and the system needs to degrade gracefully if one side becomes unavailable. ## Inference Latency Inference latency is how long a model takes to go from receiving input to producing its first output. In voice AI, the LLM typically dominates this figure—the gap between receiving a transcript and returning the first token is usually the largest single delay in the system. This number isn't fixed: server load, context length, and model size all affect it. Mean figures can be misleading; p95 and p99 measurements better reflect what users actually experience. Techniques like quantization and speculative decoding are primarily tools for bringing this number down. ## Interactive Voice Response (IVR) IVR is the older generation of phone-based automation—the "press 1 for billing" systems most of us have navigated with varying degrees of patience. Callers follow a fixed decision tree; anything outside the expected inputs hits a dead end. IVR is the direct predecessor to modern voice AI in telephony, and the contrast is stark. Where IVR forces callers to adapt to the system, voice AI agents understand natural language and adapt to the caller. Replacing IVR with conversational voice AI is now one of the primary drivers of enterprise investment in the space. ## Interruption Handling What happens when a user talks over the AI mid-sentence? That's interruption handling—and it's one of the clearest signals of how polished a voice system really is. A well-designed system stops speaking immediately, resets, and listens. A clumsy one keeps talking or stumbles awkwardly. Getting this right requires tight coordination between audio playback, speech detection, and the AI's reasoning loop. It's less about raw speed and more about making the interaction feel respectful and human. ## Jitter Audio arrives in packets—and in real networks, those packets don't always show up on time. Jitter is the variability in that timing: the difference between when audio is expected and when it actually arrives. A little jitter is invisible. Too much causes choppy playback or gaps in transcription. Jitter buffers help by holding incoming audio briefly before playback, smoothing out the bumps—but they add a small delay in exchange. It's another classic latency-versus-stability tradeoff. ## Keyword Spotting Keyword spotting is how always-on devices wake up without burning through battery or compute. Instead of running full speech recognition continuously, the device runs a lightweight model listening for a single trigger phrase—"Hey Siri," "OK Google," and so on. When the keyword is detected, the full pipeline activates. This two-stage approach keeps resource usage minimal during idle periods. The main challenge: reducing false positives (waking up when you shouldn't) without increasing false negatives (missing the actual trigger). ## Language Translation Agent A language translation agent is a voice AI system that performs real-time spoken translation between languages during live conversation—listening in one language and speaking in another without requiring either party to pause. Building a production-quality translation agent requires integrating ASR, neural machine translation, and TTS in a pipeline optimized for speed, while preserving the speaker's intent across languages. Use cases range from multilingual customer service and healthcare intake to live event interpretation. ## Large Language Model (LLM) The LLM is the brain of a voice AI system. Once speech is transcribed into text, the LLM reads it, reasons about it, and decides what to say back. Modern LLMs can handle nuanced questions, multi-turn conversations, and complex tasks—but they're not instant. Inference takes time, and in a voice pipeline, that time adds directly to the delay the user feels. Faster, smaller models reduce latency; larger models tend to reason better but respond slower. Choosing the right model is always a tradeoff. ## LLM Node In an audio graph architecture, an LLM node is a discrete processing unit that receives text input — typically from an ASR node — sends it to a language model, and passes the response downstream for synthesis. It encapsulates the LLM integration within the graph, exposing the same interface as any other node so it can be swapped or tested independently. This abstraction decouples model selection from the rest of the pipeline. Teams can experiment with hosted APIs, open-source models run via frameworks like llama.cpp, or fine-tuned variants without modifying the surrounding graph. LLM nodes using local inference engines enable fully offline voice AI operation. ## Mel-Frequency Cepstral Coefficients (MFCCs) MFCCs are a compact set of numbers that describe the tonal character of a short audio slice in a way that mirrors how humans perceive sound. They were the dominant feature representation in speech processing for decades, and remain foundational context for understanding how audio AI evolved. The process converts a frame of audio into a frequency spectrum, applies mel scaling to match human pitch perception, and produces a small set of coefficients capturing the spectral shape that distinguishes one phoneme from another. In modern voice AI, mel spectrograms fed directly into neural networks have largely replaced them—but understanding MFCCs helps explain why. ## Multi-Turn Conversation A single voice exchange is easy. A multi-turn conversation—where each message builds on what came before—is where things get genuinely interesting and genuinely hard. The system needs to remember context, track what was said, and understand references like "the second option" or "that one." This is managed through the context window, which holds the running conversation history. As conversations grow longer, so does the cost of processing them—making efficient context management one of the quieter engineering challenges in voice AI. ## Natural Language Understanding (NLU) ASR turns speech into text. NLU figures out what that text actually means. It extracts the user's intent (what they want to do) and any relevant details—dates, names, locations, preferences—needed to act on it. In modern LLM-based systems, NLU is often implicit: the model simply understands. But in more structured pipelines, NLU is a discrete step that classifies intent and extracts data before passing it downstream. Either way, it's what separates a system that hears words from one that actually understands them. ## Neural Network Neural networks are the underlying engine behind almost everything in modern voice AI. Loosely inspired by biological neurons, they're computational systems made of layers of interconnected nodes that learn patterns from data. Feed them enough examples of speech, and they learn to recognize it. Feed them enough text, and they learn to generate it. The "deep" in deep learning just means many layers—and more layers generally means the ability to capture more complex patterns, at the cost of more compute and more training data. ## Noise Cancellation Real conversations rarely happen in quiet rooms. Noise cancellation is what lets voice AI function in cars, kitchens, offices, and everywhere else life actually happens. It works by identifying and separating background sounds—fans, traffic, keyboard clicks—from the speaker's voice. Modern approaches use neural networks trained on thousands of audio environments, making them far more effective than older rule-based filters. The goal is to deliver clean speech to the ASR model, because even the best transcription system performs poorly on noisy input. ## Noise Reduction Noise reduction removes unwanted background sounds from an audio signal before it reaches the speech recognition model. In voice AI, it typically runs as a preprocessing step that improves transcription accuracy. The quality of noise reduction has an outsized effect on real-world robustness. A voice AI that performs well in a quiet testing environment may degrade significantly in a vehicle or on a noisy call. Neural noise suppression models significantly outperform traditional DSP-based approaches, particularly for non-stationary noise sources that change dynamically over time. ## Offline Audio Processing Offline audio processing operates on a complete audio file provided before processing begins. Because the full signal is available from the start, algorithms can use future context—making decisions informed by what comes later in the recording. This enables higher-accuracy results for tasks like transcription, speaker diarization, and audio enhancement. The tradeoff is that it can't be used for live or interactive applications. If you're transcribing a recorded meeting, offline processing is the right call; if you're building a live voice assistant, you'll need a different approach. ## Online Audio Processing Online audio processing handles audio as it arrives, working on a continuous stream in chunks without access to future signal content. Each chunk is processed as it comes in, meaning decisions must be made with only past and present context. This approach is necessary for real-time applications but introduces constraints: algorithms must be causal, and accuracy on any given chunk may be lower than what offline processing could achieve. It's the foundation of all live voice AI—the price of real-time is working with incomplete information. ## On-Premise Deployment On-premise deployment means running voice AI infrastructure on servers within an organization's own facilities. Audio and data stay inside the organization's network and never travel to external providers. This is required in environments where regulations, data sovereignty requirements, or contractual obligations prohibit sending audio externally—classified communications, healthcare, financial services, or enterprises with strict data residency rules. Building a fully on-premise pipeline means self-hosting ASR, LLM, and TTS models, which often involves accepting some capability or latency tradeoffs compared to large cloud-hosted alternatives. ## Packet When audio travels over a network, it doesn't flow as a continuous stream—it's broken into small, labeled chunks called packets, each carrying a slice of audio data along with headers that identify its order and destination. In voice AI systems that rely on network transmission—cloud ASR, VoIP calls, streamed TTS—packets are the fundamental unit of transport. Packet loss, reordering, and jitter are constant concerns: even a small percentage of dropped packets can degrade transcription accuracy or introduce audible gaps in synthesized speech. ## Paralinguistics Paralinguistics involves the nonverbal aspects of spoken language. It includes vocal features like tone, pitch, speed, rhythm, emphasis, and pauses that add meaning beyond the actual words. These cues can reveal a speaker’s emotions, level of confidence, uncertainty, or intentions—details that are often lost in a written transcript. In pipeline architectures, this information is discarded at the ASR stage—the LLM receives text only. Speech-native models, which reason over audio directly, preserve these cues, enabling responses that account not just for what was said, but how it was said. That distinction matters most in emotionally sensitive or ambiguous interactions. ## Pitch Shifting Pitch shifting modifies the fundamental frequency of an audio signal—raising or lowering a voice's perceived pitch—without changing its duration. Applied in real time, it transforms how a speaker sounds to others during a live call or recording. Pitch shifting is a core capability in voice modification features across gaming, social, and entertainment applications. Real-time implementations must balance transformation quality against processing delay, since even small amounts of added latency are perceptible in live conversation. When combined with formant preservation, the result sounds natural; without it, the output tends to sound artificially processed. ## Prosody Prosody is everything beyond the words themselves—rhythm, stress, intonation, and pacing. It's what turns "fine." and "fine?" into completely different messages. For voice AI, prosody matters in two directions. On input, detecting it can reveal emotion, urgency, or uncertainty. On output, generating natural prosody is what separates robotic-sounding TTS from speech that actually feels human. It's one of the hardest aspects of speech synthesis to get right, and one of the most noticeable when it goes wrong. ## Real-Time Audio Processing Real-time audio processing is the manipulation of audio signals with low enough latency that the output can be used in live, interactive contexts — a conversation, a call — without perceptible delay. In practice, this means processing audio in small chunks, on the order of milliseconds. Real-time constraints shape every architectural decision in voice AI. They determine the maximum model size usable on given hardware, the buffering strategies available, and the acceptable complexity of DSP preprocessing. Systems that cannot meet real-time constraints introduce latency that disrupts conversational flow — making interactions feel broken even when the underlying responses are accurate. ## Real-Time Factor (RTF) Real-Time Factor is a simple but important benchmark: it measures how long a model takes to process audio relative to the duration of that audio. An RTF of 1.0 means processing takes exactly as long as the audio itself. An RTF below 1.0 means the model runs faster than real time—which is the requirement for live voice applications. If RTF creeps above 1.0, the system can't keep up and delay accumulates. It's a useful early-warning metric for catching performance issues before they become user-facing problems. ## Sample A sample is a single numerical measurement of an audio signal's amplitude at one instant in time. Digital audio is built from a sequence of samples captured at regular intervals; played back at the correct rate, these discrete values reconstruct a continuous sound wave. In voice AI, a sample is the atomic unit of audio data—every downstream operation, from DSP filtering to neural inference, ultimately operates on sequences of samples. It's the smallest building block of everything a voice system hears or produces. ## Sample Rate Sample rate is how many times per second an audio signal is measured and recorded, expressed in hertz (Hz). Higher sample rates capture more acoustic detail—CD audio runs at 44,100 Hz, while many voice AI systems work at 16,000 Hz, which is sufficient for speech and far more efficient. Mismatched sample rates between components can cause subtle audio quality issues or outright failures. Making sure every part of the pipeline agrees on sample rate is one of those foundational details that's easy to overlook and surprisingly painful to debug. ## Speaker Diarization When multiple people are talking, speaker diarization is what labels who said what. Rather than producing one undifferentiated transcript, it segments the audio and tags each segment by speaker—"Speaker 1," "Speaker 2," and so on. This is crucial for meeting transcription, call center analytics, and any scenario where tracking individual voices matters. It's a technically demanding task, especially when voices overlap, speakers have similar tones, or audio quality is poor. ## Speaker Verification Speaker verification answers a specific question: is this person who they claim to be? It compares an incoming voice sample against a stored voiceprint and returns a confidence score. Unlike speaker identification—which picks a voice out of a group—verification is a one-to-one comparison. It's used in banking, healthcare, and secure voice interfaces where identity matters. Performance degrades in noisy environments or when the voice sample is very short, which is why real-world deployments usually require a minimum phrase length. ## Spectrogram A spectrogram is a visual representation of audio—a map that shows how frequency content changes over time. The horizontal axis is time, the vertical axis is frequency, and brightness or color represents intensity. Voice AI models often treat audio as images, feeding spectrograms into visual-style neural networks. This approach has proven remarkably effective: patterns that are hard to describe mathematically are easy for CNNs to learn visually. When you see an AI "reading" audio, it's often quite literally looking at a picture of sound. ## Speech-Native Model A speech-native model processes and generates audio directly, without an ASR transcription step in between. The model receives audio as input and reasons over the full acoustic signal—including tone, pacing, and prosody—rather than a text approximation of it. Eliminating the ASR stage removes both a source of latency and a point of error propagation. The practical tradeoffs: these models are large, infrastructure for them is less mature, and real-world latency gains depend heavily on deployment quality. But the architectural advantage is structural—they preserve information that text-based pipelines permanently discard. ## Streaming Streaming is what makes voice AI feel fast. Instead of waiting for each stage to finish before passing results forward, a streaming pipeline sends partial outputs as they become available—ASR emits a partial transcript while the user is still speaking, the LLM starts reasoning before transcription completes, and TTS begins synthesizing before the full response exists. Each overlap shaves time off the total delay. In a fully streaming pipeline, the time to first audio approaches the LLM inference time alone, rather than the sum of every stage. The tradeoff: partial inputs are inherently less certain, and acting on them too early can mean revising or discarding work already done. ## Text-to-Speech (TTS) Text-to-speech converts written text into spoken audio. In a voice AI component pipeline, it is the final stage: the language model's text response is passed to a TTS model, which synthesizes the audio the user hears. Modern neural TTS produces speech increasingly difficult to distinguish from human recording, though unusual words, strong emotion, and long sentences can still expose artifacts. TTS models vary significantly in naturalness, latency, speaker options, and language support. Streaming TTS—which starts synthesizing before the full response is written—is especially important for keeping perceived response time low. ## Time to First Audio (TTFA) TTFA measures the gap between when the user finishes speaking and when they hear the system's first word. It's the single most important latency metric in voice AI, because it's what users actually feel. TTFA accumulates delay across every stage — VAD, ASR, LLM inference, TTS synthesis, and audio output buffering. The table below reflects generally accepted perceptual thresholds: | TTFA | User Perception | Conversational Impact | | --------------- | ------------------ | --------------------------------------------------- | | < 200ms | Instantaneous | Human-like; users may overlap speech naturally | | 200ms – 600ms | Snappy and natural | Responsive; most users perceive no meaningful delay | | 600ms – 1,000ms | Noticeable delay | Acceptable for task-oriented use; feels AI-like | | > 1,000ms | Disjointed | Users often start repeating themselves or barge in | ## Tool Calling Tool calling is what turns a voice AI from a conversationalist into an agent that can actually get things done. It lets the language model reach out to external systems—databases, calendars, APIs—mid-conversation, retrieve or act on real data, and incorporate the results before responding. Asking the AI to book an appointment or pull up an account balance only works because of tool calling. The catch: each external call adds latency. Well-designed systems manage this with non-blocking calls and natural-sounding filler responses that buy a second or two while waiting for results. ## Turn-Taking Human conversation has a subtle rhythm, we use falling pitch, slowing pace, and brief pauses to signal that we're done speaking. Turn-taking in voice AI is the attempt to replicate that rhythm computationally. Done well, the system feels like a natural conversation partner: it waits the right amount of time, doesn't cut you off, and responds promptly when you've genuinely finished. Done poorly, it either interrupts constantly or leaves uncomfortable silences. Even when the underlying AI is excellent, poor turn-taking makes the whole system feel broken. ## Voice Activity Detection (VAD) VAD is the first thing that runs when you speak—and it keeps running the entire time you're not. It monitors incoming audio continuously, signaling to the rest of the pipeline when speech is present and when it's stopped. Tuning VAD is a balancing act: too aggressive and it triggers on background noise or truncates mid-sentence pauses; too conservative and it adds dead air at the end of every exchange. Natural pause patterns also vary across languages and individuals, which means a VAD tuned on one dataset may not generalize well to others. ## Voice AI / Voice Agent Voice AI is the broad category—any system capable of spoken, natural-language dialogue. A voice agent is the deployed version: a specific system built to handle a real job, whether that's answering customer service calls, supporting healthcare intake, or managing logistics workflows. What separates modern voice agents from older phone menu systems is genuine language understanding. Where IVR forced callers into fixed paths, voice agents can handle open-ended questions, follow conversational threads, and respond dynamically. They can run on virtually any device with a microphone—phones, wearables, embedded hardware, vehicles—and the list of viable use cases keeps growing. ## Voice Changer A voice changer modifies a speaker's voice in real time, transforming pitch, timbre, or character as audio is captured. The applications range from gaming and entertainment to privacy-sensitive communications. In software, voice changers are typically built as audio graphs: microphone input flows through transformation nodes—pitch shifting, formant adjustment, and so on—before reaching the output. Real-time performance is non-negotiable; even a few milliseconds of added latency is perceptible in live conversation. Getting transformation to sound natural while also running fast is the core engineering challenge. ## Voice-Controlled Device A voice-controlled device is any piece of hardware where speech is a primary way to interact—smart speakers, earbuds, headsets, watches, cars, and dedicated embedded systems. Building voice AI for hardware introduces constraints that browser or app development doesn't face: direct audio device access, acoustic tuning for specific form factors, and real-time processing within tight power and memory budgets. The on-device versus cloud question also becomes more pressing here—network connectivity may be unreliable, users may have strong privacy expectations, and even a few hundred milliseconds of cloud round-trip latency can undermine the experience. ## Word Error Rate (WER) WER is the standard way to measure transcription accuracy. It counts how many insertions, deletions, and substitutions are needed to turn the ASR output into the correct transcript, then divides by the total number of words in the reference. A score of 0% is perfect; the metric has no upper bound—a model that hallucinates extra words can exceed 100%. More importantly, a low WER on a clean benchmark doesn't guarantee good performance on real-world audio with accents, background noise, or domain-specific vocabulary. ASR errors propagate: a misheard word becomes wrong input to the LLM, which can send the entire response off track. *** This glossary covers terminology relevant to the design, development, and deployment of voice AI systems. For practical examples and implementation guides, see our [Engineering Hub](/hub/). --- # How it works > Build cross-platform audio apps fast with our modular SDK — real-time, on-device, minimal latency. A modular audio framework that allows you to assemble new features and products in record time. ![](/_astro/how-it-works_hero-2x-1276x1152px_2pT0Va.webp) The Switchboard SDK is a modular audio framework organized into containers called 'Nodes'. ![](/_astro/nodes-1-sdk_2doK31.svg) Without containers, building, customizing, and connecting audio features requires significant time and effort because they don’t fit together easily. ![](/_astro/nodes-2-no-fitment_Z25IRRw.svg) Switchboard packages audio features (e.g. noise reduction, voice changers, WebRTC) as *nodes* (**AudioNode)** making them easy to assemble into *audio graphs (***AudioGraph)**. ![](/_astro/nodes-3-sdk-wrapper-graph_Z1s2v7G.svg) Audio graphs are then passed to natively compiled C++ code that runs them with *minimal latency*, while simultaneously generating production-ready native code for *multiple platforms*, including iOS, Android, macOS, Windows, Linux, web, and more (including some embedded support). ![](/_astro/nodes-4-integration_1ObSiG.svg) Indeed, most Switchboard nodes will run **on-device, in real time, & cross-platform, **and this is unique to Switchboard. We've pre-integrated a huge library of nodes including popular open-source and third party extensions and are always adding more. You can also easily integrate your own nodes. [See all nodes](/nodes) ### Under the hood Most of the code in Switchboard is compiled C++ code Switchboard lets you configure nodes and define audio graphs using JSON. The Switchboard Editor even allows you to do this visually, without writing code. [Go to Switchboard Editor ⬈](https://editor.switchboard.audio) [![](https://a-us.storyblok.com/f/1008163/1024x647/df14ceeac2/how-it-works-sw-editor.webp)](https://editor.switchboard.audio) * **At runtime,** the JSON configuration is passed to natively compiled C++ code which runs your graph fast (unlike Python). * This approach allows you to benefit from optimal performance, cross-platform architecture, and third-party C++ libraries without having to write C++ code yourself. * You can alternatively write your graph in C++ or using a platform specific language, allowing for complete flexibility to change the configuration of your graph at runtime (e.g. to allow user input to dynamically update the graph). ![](/_astro/sdk-extenstions-diagram-from-figjam_qLklo.svg) Either you’ll spend a lot of time building it in C/C++ or you’ll put things together with higher level tools and face performance issues and audio bugs that are time-consuming to resolve. You would also need to create a different pipeline for each platform you support. Switchboard simplifies this by automating the creation of robust, modular and high-performance audio pipelines that are easy to modify and maintain across all platforms with a much smaller team. Karaoke AppsWatch / Listen Parties & Live EventsHardware Device / Companion App Karaoke apps vary widely, but a common feature is music stem separation—splitting audio into vocals, guitar, drums, etc. This example illustrates how the music is split into stems—vocals, guitar, drums—so vocals can be lowered, allowing the user to sing into their mic. Autotune and effects enhance the performance, and the stems are mixed back together to create a truly unique performance. Sync tools, streaming connectivity, and format handling are also required—all supported by Switchboard, which has powered several karaoke apps like this. ![](https://a-us.storyblok.com/f/1008163/0x0/0b33599499/karaoke-app-graph.svg) This setup solves the long-standing issue of mixing VoIP audio with media player audio. Switchboard’s low-level audio engine enables clean integration of voice and video with music or video streams, without compromising quality (e.g., sample rates). Features like voice activity detection, smart ducking, gain control, noise and echo suppression, and spatial audio enhance the experience. These audio graphs are inherently complex—but with Switchboard, they only need to be built once and can run cross-platform. ![](https://a-us.storyblok.com/f/1008163/0x0/4776930517/watchparty-app-graph.svg) This diagram shows a hands-free device—such as headphones, a watch, or a smart speaker—handling both media playback and voice features. A user might say, “Hey Device, what’s James listening to? Listen along with him or just let him know I’m here.” This natural voice interaction is made possible by Switchboard, which supports embedded platforms ranging from TVs and soundbars to in-ear buds, enabling rich, social audio experiences across hardware. ![](https://a-us.storyblok.com/f/1008163/0x0/e7619d9133/hardware-app-graph.svg) ### Why is C++ so common in audio? ### What about the tools provided on iOS, Android, and other operating systems? ### What other audio tools are out there? ### Check out our [Audio Dev Hub](/hub/) for these insights and more. ![](https://a-us.storyblok.com/f/1008163/1040x818/8333b651ca/how-it-works-questions.webp) ### Switchboard + VoIP LiveKit, Agora, Vonage, Chime, IVS, Dolby.io, etc. For projects with simple audio needs, a VoIP / streaming service and its extensions marketplace may suffice. But projects requiring more audio features, more platforms, better performance, and more flexibility need a capable client-side audio graph. Therefore you should use Switchboard together with these services. Unless you want to spend a lot of time writing your own audio graph and dealing with hard-to-pin-down bugs for the next year. ![](https://a-us.storyblok.com/f/1008163/0x0/f25f8369a5/how-it-works-simple-diagram1.svg) ### Switchboard + Media Spotify, Netflix, all streaming audio / video. Switchboard can open up new ways to interact with these services. Listen parties and watch parties with voice and video chat. Easily add features such as audio mixing, gain control, noise suppression, voice detection, and auto-ducking. Or voice changers, an AI friend to chat with, and other fun features. ![](https://a-us.storyblok.com/f/1008163/0x0/c236338b77/how-it-works-simple-diagram2.svg) ### Switchboard + AI models OpenAI, LLama, Audioshake, Voicemod, or any AI model that has audio as an input or output. There’s been an explosion in audio and audio-adjacent AI models. From speech to text and text to speech to LLMs, voice changers, stem separation, and more. If it touches audio, it fits into Switchboard. By integrating it through Switchboard, you’ll save yourself a lot of effort building a wrapper, formatting audio, connecting it to necessary inputs and outputs from the OS, etc. We might already have the extension you’re looking for. In other cases we can build it for you, or help you bring your own model to Switchboard. 

 ![](https://a-us.storyblok.com/f/1008163/0x0/97664dad10/how-it-works-simple-diagram3.svg) ### Get started for free Sign up to get free access to Switchboard’s basic prototyping license. [Sign up free →](https://console.switchboard.audio/register) [![](https://a-us.storyblok.com/f/1008163/800x600/9cb3ef94f3/sign-up-free-panel-800x600.webp)](https://console.switchboard.audio/register) --- # Engineering hub > Stay up to date with how Switchboard is revolutionizing the audio space. Follow our journey or come be a part of it! Learn more about Switchboard and additional audio development tools. Avoid the financial drain and developer strain of on-device audio with EdgeSpeech React Native hook. The OpenAI Realtime API makes building voice agents easy. But bringing that agent into a mobile app is a different challenge. Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) right on device. Run speech synthesis directly on device, with no internet connection required at runtime. Why the next generation of voice products won’t run entirely in the cloud. Many teams overuse cloud mixing for interactive audio. For personalized experiences, assembling audio on the user’s device is simpler, faster, and more scalable. Choosing the right on-device voice AI stack: comparing Switchboard and Run Anywhere. Patching together AI voices, background music, and audience mics creates chaos. Learn how Switchboard’s shared audio timeline perfectly syncs your interactive live streams. Reliability isn’t claimed; it’s earned in production where unpredictable users and shaky networks meet reality. At Switchboard, reliability isn’t a one-time feature—it’s built over time. Decomposing the speech-to-speech voice AI latency stack. Covers STT/ASR inference, TTS, NLU processing, and audio I/O latency on cloud and on-device, with practical optimization techniques. Comparing the cost of cloud speech APIs against on-device voice AI at scale. Covers per-request pricing, bandwidth, scaling economics, and when on-device STT/ASR and TTS eliminate recurring fees. What it takes to run voice AI offline. Covers on-device speech recognition (ASR/STT), text-to-speech, model packaging, and offline-first architecture for mobile and embedded deployments. How on-device voice AI eliminates the privacy risks of cloud speech processing. Covers data residency, GDPR and HIPAA compliance, on-premise deployment, and architectures where voice data never leaves the device. A developer's guide to how acoustic echo cancellation works, with a deep dive into WebRTC AEC3's architecture. Covers the full pipeline, common debugging scenarios, and platform-specific integration on iOS, Android, and embedded devices. Most mobile developers don’t decide to build a cloud-only Voice AI product. They arrive there accidentally because those are the tools available. I built this project to show how app builders can replace an end of life voice SDK by composing open source audio tools inside a specialty runtime instead of becoming audio developers. A visual web tool for creating real-time audio or data engines—connect nodes from a large library or get started with one of our templates. Switchboard is a modular, real-time audio SDK and orchestration layer built to empower developers who are pushing the boundaries of audio, AI, and real-time media experiences. Switchboard streamlines real-time voice AI projects across all phases. From building audio graphs to tuning latency, deploying cross-platform, and monitoring live performance, it empowers teams to deliver faster, smarter voice experiences. Building production voice AI applications requires different strengths at different stages of development. Switchboard's native C++ runtime excels at mobile production deployment, providing the on-device performance and cross-platform consistency that modern applications demand. On-device voice AI guarantees low latency, privacy, and reliability that cloud-only assistants can’t match. Explore why modern voice interfaces must run their critical path locally, with examples across consumer, enterprise, wearables, and automotive, plus how hybrid approaches and hardware advances make it practical. A concise guide for mobile developers on pairing Stability AI’s Stable Audio Open Small model with the Switchboard SDK to build real‑time, on‑device voice filters and smart messaging features that run privately and with minimal latency on everyday smartphones. A walk through of the benefits of offline processing, the architecture of a voice processing pipeline, and code snippets so developers can build and extend a hands-free field service demo. Tuning audio DSP on embedded devices requires full rebuilds for small changes and doubles effort when ported. Our fixed-memory engine runs the same graph on any chip with live next-block updates, cutting weeks to hours. Switchboard and AudioKit have some key differences in design, scope, and cross-platform capabilities. This article outlines the strengths of both to ensure you make the right choice. WebAudio supports basic audio flows, but fails under real-time pressure. For advanced use cases, only native layers offer the control and consistency serious audio apps require. How a hybrid approach to processing the Switchboard SDK can help in developing high-quality Voice AI solutions. See what makes Switchboard stand apart from other audio tools on the market. Tools include iOS Core Audio / Audio Units, Android’s Oboe / AAudio / OpenSL ES. Audio programming mistakes can produce very interesting sounds. In this talk we are going to look at these mistakes and even listen to them. Amazon’s Interactive Video Service (IVS) is a managed live streaming service for live streaming video and audio at scale. But what if you want to do more with that audio on a device... In this video tutorial, Synervoz VP of Engineering Balazs Kiss shows viewers step-by-step instructions for how to build a simple guitar-effect app for iOS using the Switchboard SDK. In a recent presentation at ADCx, Kieran Coulter, Senior Engineer and Lead Architect at Synervoz, delves into neural audio digital signal processing (DSP)... Exciting new features we have planned in the near future. ###### TRUSTED BY CUSTOMERS AND PARTNERS LIKE * ![](/_astro/amazon-logo_Z2eRcQK.webp) * ![](/_astro/meta-logo_Z1VUN2T.webp) * ![](/_astro/unity-logo_2edYer.webp) * ![](/_astro/slack-logo_Z16SpDc.webp) * ![](/_astro/bose-logo_Z1coqbj.webp) * ![](/_astro/dash-logo_Z1Kg5np.webp) * ![](/_astro/agora-logo_YRb78.webp) * ![](/_astro/superpowered-logo_Z1BanAH.webp) ### Get started for free Sign up to get free access to Switchboard’s basic prototyping license. [Sign up free →](https://console.switchboard.audio/register) [![](https://a-us.storyblok.com/f/1008163/800x600/9cb3ef94f3/sign-up-free-panel-800x600.webp)](https://console.switchboard.audio/register) --- # Acoustic Echo Cancellation: How WebRTC AEC3 Works > A developer's guide to how acoustic echo cancellation works, with a deep dive into WebRTC AEC3's architecture. Covers the full pipeline, common debugging scenarios, and platform-specific integration on iOS, Android, and embedded devices. ![](/_astro/acoustic-echo-cancellation-web-rtc-aec3_Z20qYpb.webp) Every voice and video call you make through a browser uses acoustic echo cancellation to keep your conversation clean. Without it, your own voice would bounce back at you from the other person's speaker, creating an unbearable feedback loop. WebRTC AEC3 is the echo canceller built into Chrome, Edge, and every WebRTC-based application, and it handles this problem in real time with remarkably low latency. Despite being one of the most widely deployed audio processing algorithms in the world, WebRTC AEC3 has almost no accessible documentation explaining how it actually works. This guide covers how acoustic echo cancellation (AEC) works and examines how WebRTC's AEC3 implementation handles each stage of the pipeline. Whether you're integrating echo cancellation into a voice SDK, debugging echo issues in a real-time audio application, evaluating AEC solutions for a mobile platform, or simply want to understand how this technology works under the hood, this is the reference that WebRTC's own documentation doesn't provide. ## What Is Acoustic Echo Cancellation? Acoustic echo cancellation is the process of removing the sound that loops back from a loudspeaker into a microphone during a two-way audio conversation. When a remote speaker's voice plays through your device's speaker, it travels through the room and arrives at your microphone as an echo. Without AEC, this echo gets transmitted back to the remote speaker, who then hears a delayed, distorted copy of their own voice. The core idea behind echo cancellation technology is deceptively simple: if you know what signal was played through the speaker, and you can model how the room transforms that signal before it reaches the microphone, you can predict the echo and subtract it from the captured audio. What remains (ideally) is only the near-end speaker's voice, with the echo removed. In practice, this is significantly harder than it sounds. The room's acoustic characteristics (its impulse response) change constantly as people move and furniture shifts. The speaker and microphone introduce non-linear distortions. Both people may talk at the same time (double-talk). And all of this must be handled in real time, with processing latency measured in milliseconds. ## How Echo Cancellation Works: The Adaptive Filter At the heart of every acoustic echo canceller is an adaptive filter. This filter maintains an estimate of the room's impulse response, which is the acoustic path between the loudspeaker and the microphone. Using this estimate, the filter predicts what the echo will look like and subtracts that prediction from the microphone signal. The process works in a continuous loop: 1. The far-end signal (what the remote speaker said) is played through the loudspeaker 2. That signal bounces around the room and arrives at the microphone, mixed with any near-end speech 3. The adaptive filter takes the far-end signal as input and produces an echo estimate 4. The echo estimate is subtracted from the microphone signal 5. The residual (the difference between the actual signal and the estimate) is used to update the filter, improving future predictions The most common algorithm for this adaptive filter is the Normalized Least Mean Squares (NLMS) algorithm. NLMS adjusts the filter coefficients after each sample (or block of samples) to minimize the error between the predicted echo and the actual microphone signal. The "normalized" part scales the update step by the energy of the input signal, which prevents the filter from diverging when the input is loud and from stalling when it's quiet. A key parameter is the filter length, measured in milliseconds. This determines the maximum echo delay the canceller can handle. A typical room might have a reverberation tail of 100 to 300 milliseconds, so the filter needs enough taps to cover that duration. Longer filters can handle more reverberant spaces but require more computation. ## The Acoustic Echo Cancellation Pipeline A production echo canceller involves far more than just an adaptive filter. The complete acoustic echo cancellation pipeline has several stages, each addressing a specific challenge. Here's how the signal flows through a typical AEC implementation, including WebRTC AEC3: ### Delay Estimation Before the adaptive filter can work, the system needs to know the time offset between the far-end reference signal and the echo in the microphone signal. This delay comes from several sources: the audio playback buffer, the DAC (digital-to-analogue converter), the speaker-to-microphone acoustic path, the ADC, and the capture buffer. On mobile devices, this total delay can range from 20 to 200 milliseconds and can vary during a call. WebRTC AEC3's render delay controller continuously estimates this delay by cross-correlating the reference signal with the capture signal. Getting the delay estimate right is critical: if the adaptive filter is looking at the wrong time offset in the reference buffer, it cannot converge on a useful echo estimate. ### Linear Adaptive Filter Once the delay is estimated, the linear adaptive filter does the heavy lifting. WebRTC AEC3 uses a partitioned block frequency-domain adaptive filter (PBFDAF). Instead of processing one sample at a time in the time domain, it works on blocks of samples in the frequency domain using the FFT. This frequency-domain approach has two major advantages. First, convolution in the time domain becomes multiplication in the frequency domain, which is computationally much cheaper for long filter lengths. Second, partitioning the filter into blocks allows the system to update the filter incrementally, reducing latency compared to processing the entire filter length at once. The linear filter typically removes 20 to 40 dB of echo. For many scenarios this is sufficient, but for challenging conditions (reflective rooms, non-linear speaker distortion), the residual echo can still be audible. ### Double-Talk Detection Double-talk occurs when the near-end speaker talks at the same time as the far-end speaker. This is the hardest problem in echo cancellation because the adaptive filter's update mechanism assumes the error signal (microphone minus echo estimate) represents only the filter's estimation error. During double-talk, the near-end speech is also present in the error signal, and if the filter adapts to it, the filter diverges: it starts trying to model the near-end speaker as part of the echo, which corrupts the echo estimate. WebRTC AEC3 handles double-talk by monitoring the coherence between the reference signal and the capture signal, along with the energy levels of both. When double-talk is detected, the filter adaptation rate is reduced or halted entirely, protecting the filter coefficients from corruption. Once the double-talk ends, normal adaptation resumes. ### Residual Echo Suppression (Non-Linear Processing) After the linear filter has done its work, some echo usually remains. This residual echo has two main sources: the linear filter's inability to perfectly model the room (especially in changing conditions), and non-linear effects that the linear filter cannot capture by design. Non-linearity comes from the loudspeaker itself (which distorts at high volumes), the amplifier, dynamic range processing in the audio path, and clipping at the ADC. WebRTC AEC3's residual echo suppressor estimates the power of the remaining echo in each frequency band and applies a frequency-dependent suppression gain. Bands with more estimated residual echo get suppressed more aggressively. This stage acts as a safety net that catches what the linear filter misses. The suppression gain calculation is a balancing act. Too much suppression removes the echo but also degrades the near-end speech (creating a "hollow" or underwater sound). Too little suppression leaves audible echo. AEC3 uses the quality of the linear filter's estimate (measured by the echo return loss enhancement, or ERLE) to calibrate how aggressively the residual suppressor should act. ### Comfort Noise Generation When the suppressor removes residual echo, it can create unnaturally silent gaps in the audio. These gaps feel jarring to listeners because background noise (room tone, ventilation hum) suddenly disappears during suppression. Comfort noise generation fills these gaps with synthetic noise that matches the spectral characteristics of the room's background noise, preserving a natural listening experience. ## WebRTC AEC3: Architecture and Implementation WebRTC AEC3 (the "3" denotes the third iteration, replacing the older AEC and AECM modules) was introduced into the WebRTC codebase around 2017-2018. It represents a significant architectural improvement over its predecessors, and it runs in Chromium-based browsers, Android WebRTC applications, native desktop apps, and any other platform that links the WebRTC audio processing module. ### Why AEC3 Replaced the Older Modules The original WebRTC AEC module used a time-domain adaptive filter with a fixed delay estimator. It worked adequately in controlled environments but struggled with several real-world conditions: rapidly changing echo paths, devices with variable audio latency (common on Android), non-linear speaker distortion at high volumes, and Bluetooth audio routing changes. The AECM (AEC Mobile) module was a lighter-weight alternative for constrained devices, but it sacrificed quality for efficiency. AEC3 addressed these limitations with a redesigned architecture: * A robust, continuously-adapting delay estimator that tracks changing latencies * A frequency-domain partitioned block filter that converges faster and handles longer reverberation tails efficiently * Improved echo path change detection that recovers quickly when someone moves or a door opens * A more sophisticated suppression gain calculation that balances echo removal against speech quality * Sub-band processing for better frequency resolution in the suppression stage ### Block Processing in AEC3 AEC3 processes audio in blocks (typically 64 samples at 16 kHz, or 4 ms per block). Each block passes through the full pipeline: the render delay buffer is updated with the latest far-end audio, the delay controller adjusts the alignment, the linear filter produces an echo estimate and adapts its coefficients, and the residual echo suppressor applies its gains. This block-based architecture aligns well with how audio hardware delivers data (in buffers, not individual samples) and enables efficient use of SIMD (Single Instruction, Multiple Data) instructions on modern CPUs. AEC3 can process a block of audio well within the 4 ms budget on mobile ARM processors, leaving headroom for other audio processing stages. ### Source Code For developers who want to trace the implementation, the [WebRTC AEC3 source code](https://webrtc.googlesource.com/src/+/refs/heads/main/modules/audio_processing/aec3/) lives under `modules/audio_processing/aec3/` in the WebRTC repository. The main entry point is `echo_canceller3.cc`, which orchestrates the full pipeline. The code is C++ with hand-optimized SIMD paths for x86 (SSE2) and ARM (NEON). ## Common Echo Cancellation Problems and How to Debug Them Understanding the AEC pipeline makes it much easier to diagnose echo problems in practice. Here are the failure modes developers encounter most often, along with what causes them and how to investigate. ### Echo Breakthrough Echo breakthrough, the most common AEC complaint, is when the remote party hears clearly audible echo despite AEC being enabled. **Possible causes:** * The delay estimator has locked onto the wrong offset. If the estimated delay is off by even a few milliseconds, the linear filter operates on misaligned data and cannot converge. This happens most often on Android devices with unpredictable audio latency. * The filter length is too short for the room's reverberation time. The echo tail extends beyond what the filter can model. * The echo path changed suddenly (someone moved the device, a Bluetooth audio route switched) and the filter hasn't re-converged yet. **How to debug:** Log the estimated render delay and the ERLE (echo return loss enhancement). If ERLE is consistently low (under 10 dB), the linear filter isn't converging. If the delay estimate is unstable (jumping between values), the delay controller is struggling with the audio path. ### Half-Duplex Behaviour Instead of allowing both parties to speak simultaneously, the system suppresses the near-end speaker whenever the far-end speaker is active. It feels like a walkie-talkie. **Possible causes:** * The residual echo suppressor is being too aggressive, treating near-end speech as echo during double-talk. * The double-talk detector is not recognizing simultaneous speech, allowing the filter to diverge and then over-suppressing to compensate. **How to debug:** Reduce the suppression aggressiveness if your implementation exposes that parameter. In WebRTC, the `EchoCanceller3Config` struct contains tuning parameters for the suppressor. Check whether the issue correlates with the far-end signal level (worse at higher volumes suggests non-linear distortion is fooling the suppressor). ### Filter Divergence Echo cancellation works initially but degrades over time as ERLE drops and echo becomes audible again. **Possible causes:** * Double-talk is corrupting the filter. The adaptation isn't being paused correctly during simultaneous speech. * A feedback loop exists in the audio path. If the cancelled output is accidentally being fed back as the reference signal, the filter chases its own tail. * Numeric instability in the filter coefficients, usually from an excessively high adaptation rate. **How to debug:** Check your audio routing. The reference signal must be the far-end audio before it's mixed with any near-end audio. Verify that the adaptation rate (step size) is within expected bounds. If you're using WebRTC's AEC3 directly, the defaults are well-tested, so divergence usually points to an audio routing problem. ## Echo Cancellation on Mobile Platforms Mobile devices present unique challenges for acoustic echo cancellation. Variable audio latency, non-linear speaker behaviour at high volumes, diverse hardware configurations, and power constraints all complicate the task. ### iOS and AVAudioSession On iOS, echo cancellation is tightly integrated with the AVAudioSession system. Setting the audio session mode to `.voiceChat` enables the system's built-in AEC along with other voice processing (automatic gain control, noise suppression). This is the simplest path for most iOS voice applications. However, the built-in iOS echo cancellation has limitations. It assumes a standard phone-call-like scenario and may not perform well for applications with custom audio routing or music mixed with voice. If your application uses `.measurement` or `.default` mode for other reasons, the built-in AEC is not active, and you'll need to provide your own. Switchboard's iOS SDK provides a WebRTC AEC3 node that you can insert into your audio graph, giving you echo cancellation without requiring `.voiceChat` mode. This is particularly useful for applications that need AEC alongside other audio processing that `.voiceChat` mode would interfere with. ### Android and AudioEffect Android provides the `AcousticEchoCanceler` class in its `android.media.audiofx` package. This wraps the device manufacturer's AEC implementation, which varies significantly across devices and Android versions. Some devices have excellent echo cancellation; others have noticeably poor implementations. Because of this inconsistency, many voice applications on Android bypass the platform AEC and use WebRTC AEC3 directly (or through an SDK like Switchboard that wraps it). This provides consistent behaviour across the fragmented Android device ecosystem at the cost of slightly higher CPU usage. The biggest challenge on Android is audio latency. The round-trip latency between playing audio and capturing it varies from 10 ms on flagship devices to over 100 ms on budget hardware. AEC3's delay controller handles this variability, but extreme or rapidly changing latencies can still cause problems. Using Android's low-latency audio path (AAudio with AAUDIO\_PERFORMANCE\_MODE\_LOW\_LATENCY) helps stabilize the delay. ### Embedded and Edge Devices For edge AI applications running on embedded Linux boards, Raspberry Pi devices, or custom hardware, there is no platform AEC to fall back on. You need to run an echo canceller as part of your audio pipeline. WebRTC AEC3 is a strong choice here because it's pure C++ with no platform dependencies and runs efficiently on ARM CPUs, having been battle-tested across billions of WebRTC calls. Switchboard's C++ API provides AEC3 as a node that you can integrate into any audio graph on embedded platforms. ## Measuring AEC Performance When evaluating or tuning an acoustic echo canceller, several metrics matter: * **ERLE (Echo Return Loss Enhancement):** The ratio of echo power before and after cancellation, measured in dB. Higher is better. A well-performing linear filter typically achieves 20 to 40 dB ERLE. Below 10 dB indicates the filter isn't converging. * **Residual echo level:** The absolute level of remaining echo after both linear cancellation and non-linear suppression. This is what the remote listener actually hears. * **Near-end speech degradation:** AEC can damage the near-end speech, especially during double-talk. Measuring with PESQ (Perceptual Evaluation of Speech Quality) or POLQA gives an objective quality score. * **Convergence time:** How quickly the filter adapts to a new room or a changed echo path. AEC3 typically converges within one to two seconds in normal conditions. ## Integrating Echo Cancellation in Your Application If you're building a real-time voice application, you have several options for echo cancellation: **Use the platform's built-in AEC.** On iOS (`.voiceChat` mode) and in WebRTC-based browser applications, this is the simplest option. The trade-off is limited control over tuning and behaviour. **Use WebRTC AEC3 directly.** If you're already using the WebRTC native library, AEC3 is available as part of the audio processing module. You feed it the render (far-end) signal and the capture (near-end) signal, and it returns the cleaned audio. You need to manage the audio routing yourself. **Use an audio SDK with AEC built in.** Switchboard provides WebRTC AEC3 as a node in its audio graph architecture. You connect it alongside your other audio processing nodes (noise suppression, voice activity detection, speech-to-text, text-to-speech) and the SDK handles the signal routing, buffering, thread management, and sample-rate conversion. This approach gives you AEC3's quality with less integration effort, and it works cross-platform across iOS, Android, desktop, and embedded Linux. Whichever path you choose, the critical integration requirement is the same: the echo canceller must receive the far-end reference signal and the near-end capture signal with accurate timing. If these signals are misaligned or if the reference signal doesn't match what was actually played through the speaker (e.g., because of post-processing or mixing after the reference tap point), AEC performance will suffer. ## Available Now Acoustic echo cancellation is one of those technologies that works so well, most people never think about it. From browser-based calls to voice chat in games to telehealth consultations to enterprise conferencing, AEC runs silently in the background. WebRTC AEC3 handles this for billions of calls, and understanding how it works gives you the foundation to debug echo problems when they arise and make informed decisions about your audio pipeline. Switchboard's audio SDK provides WebRTC AEC3 as part of its modular audio graph, alongside noise suppression (RNNoise), voice activity detection, speech-to-text, and text-to-speech nodes. If you're building a voice application that needs echo cancellation across platforms, [check out the AEC documentation](https://docs.switchboard.audio/aec/) to get started. --- # Bridging the Gap: Developing Audio AI Applications > How a hybrid approach to processing the Switchboard SDK can help in developing high-quality Voice AI solutions. ![](/_astro/developing-ai-applications-w-hybrid-processing-1200x800_ZNE8rr.webp) However, developing high-quality Voice AI solutions remains challenging due to platform limitations and the shortcomings of current frameworks and tools. Android's fragmented ecosystem and outdated real time communication (RTC) stack create inconsistencies in performance, making it difficult for AI-driven voice applications to deliver reliable and low-latency interactions. Meanwhile, Apple's tightly controlled environment limits developer flexibility, restricting the ability to customize and optimize AI processing pipelines. Existing solutions like WebRTC, platform-specific APIs, and purely cloud-based approaches fail to resolve these issues fully, leaving developers with suboptimal performance, high latency, and limited adaptability. Many tradeoffs depend on whether the AI model lives on-device or in the cloud. For example, latency and offline operation favor the former, whereas accuracy and breadth of capabilities favor the latter. While on-device models continue to become more capable, the most advanced models will likely require cloud compute in many use cases. Hence, the future of AI is hybrid. By intelligently balancing local and cloud-based processing, developers can create responsive, high-performance audio applications that work seamlessly online, offline, and across various devices, including those with limited compute and memory available. Therefore, tooling is needed to help build seamless, cross-platform, low-latency audio pipelines with modular, hybrid processing capabilities. ## Challenges on Android and iOS Despite Android's dominance in the mobile market, its fragmented device ecosystem makes it difficult to ensure consistent audio AI performance. With thousands of Android devices featuring different microphones, audio processing chips, and software configurations, developers face unpredictable behavior in AI-driven voice applications. Additionally, Android's reliance on WebRTC for real time audio presents challenges. WebRTC's default acoustic echo cancellation (AECM, AEC3) and voice activity detection (VAD) were not designed for modern AI-driven applications, leading to suboptimal performance. On top of that, the platform's audio stack introduces unpredictable latencies that negatively impact user experience, making real time responsiveness difficult to achieve. Meanwhile, Apple's ecosystem provides more consistency but introduces other constraints that affect AI-driven audio applications, particularly in adapting AI models to handle real time processing demands efficiently. The company's strict system controls mean that developers have limited access to lower-level audio processing, and while Apple's audio pipeline (VPIO) is highly optimized, it does not allow for easy replacement or modification of key audio components. Furthermore, Apple prioritizes battery efficiency, imposing strict limits on background processing and resource usage, complicating the optimization of real time AI audio applications. ## When Purely On-Device or Cloud Processing Falls Short Given OS limitations on iOS and Android, one might wonder how to develop AI audio apps best. Android's open ecosystem and diverse hardware tend to favor cloud-based solutions, as device constraints often limit the feasibility of high-performance on-device AI. On the other hand, Apple's tightly integrated hardware and software ecosystem offers more robust on-device processing capabilities. Still, its strict system controls make it less flexible. So if you want to develop a mobile application with similar functionality on both platforms, which approach should you take? There are cases where a purely on-device approach makes sense. When privacy is the top priority, such as in voice-controlled healthcare devices or secure voice authentication, keeping all processing local ensures data never leaves the device. On-device AI also shines in low-latency, always-available applications that must function even without an internet connection. Conversely, cloud-based AI is beneficial when models require immense processing power or access to vast, evolving datasets, such as large-scale transcription services or AI assistants that improve through aggregated learning. However, both approaches have inherent trade-offs, and neither platform fully supports a single-method approach across all AI use cases. Hardware constraints limit on-device models, while cloud-based solutions suffer from latency, network dependency, and potential privacy concerns. This leaves a wide range of use cases where neither method alone is sufficient. ## The Hybrid AI Approach The best solution isn't necessarily a choice between on-device processing or cloud-based AI—it's likely a combination of both. A hybrid approach ensures low-latency responsiveness by handling immediate, time-sensitive tasks on-device while leveraging the cloud for computationally intensive processing. This method allows developers to assign tasks based on performance needs, ensuring fast, natural user interactions while maintaining flexibility and scalability. With intelligent load balancing, AI applications can deliver the best possible experience while optimizing for privacy, efficiency, and real time interaction. ## Limitations of Current Solutions Developers attempt to use existing tools to execute this strategy but often fall short. WebRTC, while widely used for traditional VoIP applications, was not built for AI-powered speech processing. It struggles with accurate voice detection, real time noise suppression, and smooth conversational turn-taking. Meanwhile, platform-specific APIs from Apple and Android provide native audio processing tools. Still, they either lack flexibility (as seen on iOS) or are inconsistent across devices (as seen on Android). Relying solely on cloud-based AI is not a viable alternative, as sending all audio processing to the cloud introduces latency, leading to unnatural delays that degrade real time interactions. Without a more adaptable solution, developers grapple with high latency, poor real time detection, and platform-specific limitations that ultimately diminish the user experience. ## New Tools Built for AI The Switchboard SDK circumvents these limitations by allowing developers to construct flexible, high-performance audio pipelines. Unlike platform-restricted APIs, Switchboard provides a comprehensive library of modular first- and third-party audio nodes, enabling developers to create custom audio graphs without requiring deep expertise in audio programming. By abstracting platform-specific details, Switchboard ensures that the same pipeline works consistently across Android and iOS (among other platforms), eliminating the need to build and maintain separate solutions for each platform. Switchboard also provides hybrid processing flexibility, allowing developers to determine whether tasks should run on-device or in the cloud. This capability is crucial for optimizing performance, privacy, and resource efficiency, allowing developers to fine-tune solutions based on specific use cases and hardware constraints. ## How Switchboard Solves These Challenges Diving in more deeply, Switchboard directly addresses the key challenges that Android and iOS present by optimizing latency, improving real time voice interactions, and ensuring platform consistency. For Android, where fragmentation creates unpredictable audio behavior, Switchboard abstracts away hardware differences by providing a unified, high-performance audio pipeline that works across diverse devices. It eliminates the need for developers to manually optimize for various microphones, audio chips, and software configurations. On iOS, where strict system controls limit access to lower-level audio processing, Switchboard offers a flexible framework that works within Apple's constraints while allowing advanced customization of AI-driven audio features. Switchboard can significantly reduce delays and enhance conversational flow by enabling real time, on-device audio processing. It allows for natural back-and-forth interactions without noticeable lag, adds more flexibility around noise suppression and echo cancellation and allows audio pipelines to be configured for AI-driven speech applications. Switchboard also enables developers to fine-tune hybrid AI workflows by seamlessly integrating cloud-based processing while keeping time-sensitive tasks local. This allows AI models to leverage powerful cloud-based machine learning for deep speech analysis while ensuring real time responsiveness—such as interruption detection, latency-sensitive audio transformations, and local speech enhancement—remains fast and efficient. By giving developers precise control over how and where audio processing occurs, Switchboard makes it possible to build AI-powered voice applications that are both high-performing and adaptable to platform constraints. Beyond basic audio processing, Switchboard enables advanced Voice Activity Detection (VAD) to distinguish between primary speech, background chatter, and incidental noises. Switchboard can support VAD designed for AI-driven interactions, interpreting speech intent more accurately. This ensures that AI agents, for example, can respond naturally without being falsely triggered by background noise or momentary interruptions, improving real time communication and enhancing the the user experience. ## Conclusion Developing AI-powered audio applications across Android and iOS is challenging due to fragmentation, outdated RTC stacks, and platform limitations. A hybrid AI model that balances on-device and cloud-based processing is the best way forward. Switchboard makes this possible by providing a flexible, powerful SDK that gives developers control over their audio pipelines without the constraints of WebRTC or platform-native APIs. For developers looking to build high-performance, real time AI audio applications, Switchboard is the tool that makes it easier, faster, and more reliable. --- # Building a Voice Changer with Claude Code and Switchboard > Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) right on device. [Play](https://youtube.com/watch?v=xW39tEDZ6fc) I built this project to provide that pattern. It shows how to build a real time voice changer with Switchboard and open source libraries, and it makes a practical point about agentic coding. Agentic tools work best when they target a specialty runtime. For audio, that runtime is Switchboard. ## The problem and the runtime Most app builders ship UI and backend services. Audio requires a different set of skills. You must handle devices and real time buffers without glitches, keep latency low, and make the same logic work across platforms. Most teams just want a working audio feature in their app, not a crash course in DSP. A specialty runtime makes that possible by handling the hard parts for you. Switchboard provides a graph runtime for audio. You build features by composing nodes, each with a single job. The runtime handles scheduling, buffering, and device I O so the developer can focus on the feature instead of the plumbing. This gives agentic coding tools a stable target where they can assemble known building blocks instead of guessing at audio code. The audio pipeline in this demo is simple and explicit. ![](https://a-us.storyblok.com/f/1008163/2345x1420/b45865e363/voice-changer-diagram.png) Two custom nodes handle the core transformation. Pitch shifting with formant preservation uses an open source library, and ring modulation produces robotic tones. Everything else is composition using existing Switchboard effects. The presets match common voice changer behavior and are built entirely from open components. ## Agentic coding and assembly Claude Code did not design the audio graph. Designing the effects chain still required human judgment. Claude Code assembled the system by wiring nodes together, integrating the pitch shifting library, and structuring the demo application and preset system so the code could ship. Agentic coding accelerates the assembly of known pieces into a working system, and with a specialty runtime like Switchboard, non audio developers can build audio features. ## Portability and reuse This demo runs on Linux, but the graph is the important part. The same audio graph runs on mobile or desktop with no changes to the audio logic. Only the shell around it changes. You build an audio feature that can live inside a real product on multiple platforms, which gives teams coming off the Voicemod SDK a pattern they can own and extend. ## Try it yourself Clone the repo and change it. Replace the nodes, adjust the presets, and build a different audio feature. This is a pattern you can reuse for your own app. [Here's the Repository](https://github.com/switchboard-sdk/voice-changer-sample) --- # Building OpenAI Realtime voice agents in React Native > The OpenAI Realtime API makes building voice agents easy. But bringing that agent into a mobile app is a different challenge. ![](/_astro/building-openai-realtime-voice-agents-in-react-native_13BYN3.webp) Putting that same agent into a mobile app is a little different. Once you get beyond the demo, you start running into problems that don't have much to do with the model itself: * The agent hears its own voice through the phone speaker. * Background noise gets mistaken for speech. * A short pause gets mistaken for the end of a sentence. * The user interrupts, but the agent keeps talking. * Your React Native app needs to manage a native audio session across iOS and Android. * Tool calls need to get routed back into the app and its current state. We built [@synervoz/openai-realtime-toolkit](https://github.com/switchboard-sdk/openai-realtime-toolkit) to handle that layer. It's an MIT-licensed React Native package for building OpenAI Realtime voice agents on iOS and Android. It manages audio capture and playback, the Realtime session, echo cancellation, tool calls, and barge-in. It can also run voice activity detection and semantic turn detection locally on the device when you want more control over when the agent should listen, respond, or stop talking. The API is intentionally small: one provider and two hooks. OpenAI handles the intelligence of the agent and the realtime protocol. But if you're connecting directly from a mobile app, there is still a fair amount of plumbing around it. ## Mobile audio You need to capture the microphone, play the model's response, maintain the correct audio route, and deal with things like headphones being connected or disconnected. Then there's echo cancellation. If you're using WebRTC, echo cancellation is generally part of the audio stack. If you're connecting directly to Realtime over WebSocket, you need to configure the native platform audio correctly yourself. On iOS that means using Apple's voice-processing audio path. On Android it means configuring the appropriate communication audio mode and source. Get it wrong and the model can hear its own output coming back through the microphone. ## Turn-taking OpenAI already provides server-side turn detection, including semantic turn detection, and for many applications it works well. But mobile environments vary enormously. Someone using your app in a quiet bedroom is a very different case from someone using it in a car, gym, warehouse, kitchen or café. Sometimes you want tighter control over questions like: How loud does something need to be before it counts as speech? How long should we tolerate a pause? Does what the user said actually sound like a completed thought? How quickly should playback stop when the user starts talking? Those decisions can also be useful to make locally. For example, if the agent is already playing audio from a buffer on the phone and the user starts speaking, you don't need to wait for a server round trip before turning that playback down. That immediate response makes interruption feel much more natural. ## What the toolkit handles Without a library, a direct React Native integration typically needs some combination of: * native microphone capture and audio playback * audio-session and route management * acoustic echo cancellation * the OpenAI Realtime WebSocket session * reconnect and lifecycle handling * voice activity detection * turn detection * interruption and barge-in behavior * tool-call dispatch * iOS and Android native integration None of these problems is particularly exotic. There's just a lot of them. And they're mostly infrastructure rather than whatever makes your application unique. That's the main reason we built the toolkit: we'd rather have developers spend their time on what the voice agent actually does than on getting Android audio routing or interruption timing right. ## A voice agent in one screen A basic agent looks like this: ``` import { OpenAIRealtimeToolkitProvider, useOpenAIRealtimeToolkit, useTool, } from '@synervoz/openai-realtime-toolkit' export default function App() { return ( ) } function Screen() { const { isRunning, start, stop } = useOpenAIRealtimeToolkit() useTool({ name: 'get_time', description: 'Get the current time.', handler: async () => ({ time: new Date().toLocaleTimeString(), }), }) return ( {isRunning ? 'Stop' : 'Start talking'} ) } ``` That gives you a working OpenAI Realtime voice agent with a tool call. Underneath, the toolkit creates the audio graph, connects the microphone to OpenAI Realtime, plays the response, enables echo cancellation, and dispatches tool calls back into your application. Because the tools are just JavaScript handlers, they can interact directly with your app. For example: **"Show me the cheaper options."** could call a tool that changes a filter in your UI. Or: **"Book the 3:30 appointment."** could call the same application logic your booking screen already uses. This is one of the reasons we like running the agent in the client. Voice becomes another interface to the application rather than a separate service sitting beside it. ## Local turn detection Local turn handling is optional. ``` const { localTurnHandling } = useOpenAIRealtimeToolkit() localTurnHandling.setEnabled(true) ``` When enabled, the toolkit uses two stages. First, on-device voice activity detection decides whether someone is actually speaking. Then a semantic model estimates whether what they said sounds complete. Those are different questions. Consider: **"What's the weather in..."** versus: **"What's the weather in Berlin?"** Both can be followed by exactly the same amount of silence. But only one sounds like a finished thought. Separating speech detection from semantic completion makes it possible to tune the interaction more intelligently than simply waiting for a fixed silence timeout. There's also a separate fast path for interruption. When the user starts speaking while the agent is talking, playback can duck immediately, then pause or cancel according to the configured timing. For most applications you don't need to tune all of this manually. The toolkit includes three presets: ``` localTurnHandling.setConfig(QUIET_CONFIG) localTurnHandling.setConfig(BALANCED_CONFIG) localTurnHandling.setConfig(NOISY_CONFIG) ``` You can switch them while a session is running. That makes it practical to use different behavior for, say, a quiet indoor assistant versus a hands-free app intended for a gym or car. ## When this architecture makes sense This toolkit is mainly for developers who: * are building a React Native app * want to use OpenAI Realtime directly * want the voice agent integrated tightly with their existing application * care about mobile audio behavior and interruption * don't want to build and maintain a separate agent runtime It's particularly useful when voice is intended to control the application itself. Your tools execute directly in the React Native application, so they can work with application state, local storage, device APIs and your existing backend calls. There isn't another server-side tool protocol required just to get an agent action back into the UI. ## When LiveKit is probably a better choice LiveKit solves a broader problem. Its agent architecture runs server-side, with the mobile client connecting over WebRTC. That gives you capabilities this toolkit isn't trying to reproduce: telephony and SIP, multi-party rooms, server-side agent workers, observability, provider flexibility, and the ability to build STT → LLM → TTS pipelines using different vendors. If you're building an agent that needs to work across phone calls, browsers and mobile apps, or you want the agent runtime centralized on your infrastructure, LiveKit is probably the better architecture. The tradeoff is that your agent is no longer running inside the application. For mobile applications where voice is primarily another way of controlling the app, we think the client-side architecture is considerably simpler. A tool handler can just be: ``` useTool({ name: 'set_background_color', handler: ({ color }) => setBackgroundColor(color), }) ``` There's no data channel or application protocol required to get that action from a server-side worker back into the component tree. The other difference is interruption. Because the toolkit owns playback locally, it can react to detected speech immediately rather than waiting for an interruption signal to make a round trip through the agent infrastructure. And there is no agent runtime to deploy or scale. You should still use a small backend endpoint to mint ephemeral OpenAI credentials before shipping. But that's very different from operating the agent itself on your servers. ## Why WebSocket instead of WebRTC? OpenAI recommends WebRTC for client applications, while this toolkit currently connects to Realtime over WebSocket. That's a tradeoff. WebRTC is better suited to unreliable networks. Its media transport handles packet loss and network changes more gracefully, which matters when someone moves between Wi-Fi and cellular or has a poor connection. WebSocket gives us a simpler direct audio path and avoids adding a WebRTC stack to the application. For applications expected to operate frequently on unreliable mobile networks, WebRTC is worth considering and may be the better choice. For applications primarily running on good Wi-Fi or cellular connections, we've found the direct architecture attractive enough to make that trade. ## Current limitations The toolkit is deliberately fairly narrow today. It's for React Native on iOS and Android, and it currently targets OpenAI Realtime. It doesn't give you a provider-independent STT → LLM → TTS pipeline, telephony, multi-party rooms, or a server-side agent runtime. The package also contains native code, so Expo users need a development build rather than Expo Go. React Native's New Architecture is required. And while the wrapper is MIT licensed, the underlying audio engine uses prebuilt Switchboard SDK binaries. Those constraints are worth understanding before choosing the architecture. ## Try it ``` npm install @synervoz/openai-realtime-toolkit ``` The [READ ME](https://github.com/switchboard-sdk/openai-realtime-toolkit) contains the quickest path to a working agent. There's also documentation for: [Getting started](https://github.com/switchboard-sdk/openai-realtime-toolkit/blob/main/docs/getting-started.md) [Turn detection](https://github.com/switchboard-sdk/openai-realtime-toolkit/blob/main/docs/turn-detection.md) [Tools](https://github.com/switchboard-sdk/openai-realtime-toolkit/blob/main/docs/tools.md) [Example app](https://github.com/switchboard-sdk/openai-realtime-toolkit/tree/main/example) Shared test credentials are included so you can run the example before setting up your own. They're intended for evaluation only; use your own ephemeral credentials before shipping. The project is built on the [Switchboard SDK](https://switchboard.audio/). If you're building something adjacent — another model provider, another platform, or a more complicated realtime audio pipeline — [open an issue](https://github.com/switchboard-sdk/openai-realtime-toolkit/issues). We're interested in seeing where people take it. --- # Effort and Challenges in Building Embedded Audio DSP Software Across Platforms > Tuning audio DSP on embedded devices requires full rebuilds for small changes and doubles effort when ported. Our fixed-memory engine runs the same graph on any chip with live next-block updates, cutting weeks to hours. ![](/_astro/effort-challenges-building-embedded-audio-dsp-software_Z1vn0TI.webp) Embedded audio DSP development is notoriously time-consuming and complex, especially when firmware needs to be tuned for high-quality audio and reused on multiple hardware platforms or in different form-factors. Those inefficiencies are primarily the result of only a handful of challenges, which are, as yet, poorly addressed by existing tooling. ## Iteration Cycles: High Cost and Slow Turnaround Developing and tuning audio DSP firmware often requires many iterative cycles of coding, compiling, and testing. Each adjustment to an audio parameter typically means modifying code, rebuilding the firmware, and re-flashing the device, which is time-intensive and hampers quick experimentation. As a result, audio engineers cannot easily perform instantaneous A/B comparisons of different tunings.[ One industry whitepaper ](https://dspconcepts.com/sites/default/files/embeddedaudioproductcreation_whitepaper_final.pdf)notes that "often, DSP engineers must rebuild and compile code for different sound settings", so by the time an engineer listens to a second or third tuning iteration, they've lost the fresh reference of how the first one sounded. This slow turnaround makes it costly to refine algorithms to optimal sound quality. What makes iteration especially high-cost in audio is the subtlety of human hearing; tiny differences in filter coefficients or EQ settings can be discernible, meaning many fine-tuning passes are needed. Without real-time adjustments, each fine-tune is a full software cycle. In [**How to Shorten and Simplify Embedded Audio Product Creation**](https://dspconcepts.com/sites/default/files/embeddedaudioproductcreation_whitepaper_final.pdf), Dr. Beckman emphasizes that real-time tuning capability would greatly streamline this process, since it would eliminate the need to "change code and re-compile before it can be heard again," allowing the audio engineer to tweak parameters live and immediately hear the result. In current workflows, this capability is often lacking, stretching development over long debug/tune cycles. ## Complexity of Reuse Across Multiple Hardware Platforms The effort required to build an audio DSP stack multiplies when that software must run on different chipsets or DSP cores. Porting and generalizing DSP code across hardware is a major challenge: audio algorithms are frequently optimized for a specific processor architecture (sometimes even in hand-written assembly for performance), and these optimizations don't directly transfer to a new platform. A modular design is not common in traditional audio DSP firmware. Historically, "an audio post-processing algorithm was developed considering a specific DSP architecture", meaning the code was heavily tied to one chip's features and instruction set. When a new product uses a different DSP or a new SoC version, engineers often must re-write or re-optimize large portions of the code, effectively rebuilding the audio stack for each platform. Moreover, audio DSP libraries have often been delivered as monolithic blocks combining many signal processing features. This monolithic approach hurts reusability. As one engineer noted, if a customer or new product only needs a subset of the features, the entire library might need to go through a full development cycle again to be adapted and retested for that subset. In other words, lack of modularity means code reuse across products is limited, leading to duplicated effort. Maintaining separate codebases for different chips also increases engineering overhead and risk of bugs. All of this adds complexity and time: teams must debug and tune on each platform's unique toolchain and hardware quirks. ## Lack of Real-Time Configurability and Visibility A common pain point in embedded DSP development is the lack of real-time configurability and internal visibility during development. Unlike software on a PC where developers can often tweak parameters on the fly, embedded audio firmware typically runs without a rich UI or console. Gaining insight into the DSP's behavior usually involves using hardware debuggers or adding instrumentation code. However, embedded systems have tight real-time constraints: even printing debug values can disturb timing. In fact, adding just a few printf statements can significantly affect performance (cache usage, timing, etc.), to the point that such instrumentation is often not usable for real-time audio code. Thus, developers operate with limited visibility into what the audio algorithms are doing in real time, making debugging and tuning akin to a "black box" process. Not having real-time control is equally problematic for tuning audio performance. There is typically no live GUI to adjust filter coefficients or mixer levels on an embedded DSP in real time, so audio engineers must rely on slow compile-flash-listen cycles as described earlier. Beckman emphasizes that *real-time tuning* is highly desirable so that engineers could tweak multiple parameters live without full rebuilds. The lack of such interactive control not only slows down finding the best sound but also reduces confidence; if a change degrades audio, one might not catch it until much later. Similarly, visibility into internal states (like CPU load, memory use, or intermediate audio signals) is often limited. Traditional tools like logic analyzers can't easily be applied inside a modern audio SoC where ["many of the signals of interest are buried deep within the chip"](https://www.embedded.com/the-case-for-real-time-visibility/#:~:text=Unfortunately%2C%20monitoring%20transactions%20among%20system,developers%20interpret%20the%20data%20collected). All these factors make the development and tuning process laborious, requiring cautious trial-and-error with insufficient feedback. ## Long Development Cycles: Real-World Examples Because of the challenges above, it's not uncommon for audio DSP firmware projects to stretch over many months or even years. In some cases, teams spend years iterating on audio algorithms to meet quality or performance targets. For example, Karlheinz Brandenburg, one of the inventors of the MP3 audio codec described their development process as highly iterative; each new idea was implemented and tested, uncovering new issues that prompted further refinement, and "[it took years before we reached a point where quality met our expectations](https://www.thehansindia.com/tech/deep-dive-audio-plunging-headphones-into-a-new-era-of-spatial-sound-962550)". This underscores how even with a focused algorithm, achieving robust, high-quality audio required a long cycle of tuning and testing. In consumer audio products, we see similar multi-year efforts. A notable case is Apple's AirPods. Apple had been engineering AirPods since 2016, yet early models were only "good" in sound quality. It was only after several generations and continuous improvements that the flagship earbuds achieved excellent sound. The 2022 AirPods Pro 2 finally delivered a best-in-class audio experience that rivaled top competitors, earning five-star reviews, [a result of refining the acoustics and DSP over the prior years](https://www.whathifi.com/features/i-spoke-to-apple-to-find-out-the-secret-behind-the-airpods-pro-2s-audiosound-success#:~:text=Apple%20has%20been%20making%20AirPods,been%20good%2C%20but%20not%20great). This implies *multiple years of R\&D and tuning* went into perfecting the audio firmware and hardware synergy for that product. [Another industry anecdote comes from headphone manufacturer V-Moda](https://www.3ders.org/articles/20161103-v-moda-introduces-forza-in-ear-headphones-with-luxury-3d-printed-custom-caps.html), which admitted that it took years of engineering to develop a new tiny driver without sacrificing sound quality. While that example is about transducer hardware, it parallels the timeline for complex audio DSP features like adaptive noise cancellation or spatial audio, which often require several product generations to mature. These examples illustrate that without the right tools, bringing an audio product to "flagship" level performance is a long haul. The high-end earbuds and speakers we see on the market are usually the result of multi-year development cycles, where teams painstakingly tune algorithms (and sometimes continue to fine-tune via firmware updates post-launch). This long cycle directly impacts time-to-market and costs, tying up engineering resources across iterations. ## Impact of Better Tools and Abstraction on Time-to-Market Given the above pain points, it's clear why the industry is searching for better DSP development platforms. Improved abstraction, modular design, and real-time tooling can dramatically reduce time-to-market and tuning overhead. For instance, when development is done with a graphical audio tool that allows on-the-fly adjustments and reuse of ready-made modules, teams can cut down iteration time from days to minutes. A notable claim from DSP Concepts (the makers of Audio Weaver) is that [using their end-to-end audio DSP platform enabled development "up to 10×" faster than traditional methods](https://www.cadence.com/ko_KR/home/solutions/automotive-solution/infotainment.html). This acceleration comes from multiple efficiencies: parallel development by audio engineers and firmware engineers, drag-and-drop assembly of pre-optimized algorithm blocks, and the ability to tune parameters in real time without writing new C code for each change. In such an environment, an audio engineer can focus on sound design and instantly hear tweaks, while the system handles low-level optimization, a stark contrast to the slow compile cycles of the conventional approach. Cross-platform abstraction is another benefit. A well-designed DSP execution framework can provide a hardware abstraction layer where the same audio processing design runs on different chipsets with minimal changes, saving the effort of re-implementing code for each new device. In other words, the platform handles the hardware differences (data format, CPU optimizations, etc.), allowing developers to reuse algorithms across products. For example, Sound Open Firmware (an open-source audio DSP framework) is built to be *modular and portable* so that it "can be ported to different DSP architectures or host platforms" easily. The promise is that by writing to a common API or using portable data-driven configurations, a team could avoid duplicating their work for each chipset,  a huge time saver when a product line includes, say, a Bluetooth earbud on one SoC and a smart speaker on another. Early adopters of these advanced tools have reported significantly shorter development cycles. In general, an effective prototyping and tuning system "can go a long way in taking the product development process out of the Stone Age". By providing real-time insight and letting engineers iterate quickly and safely, modern DSP platforms help teams get to a good sound faster and with fewer resources. This can translate to launching products in months instead of years, or freeing up engineering time to add new features rather than fighting platform-specific bugs. In summary, better tooling and higher-level abstraction directly address the pain points of traditional embedded DSP development, enabling companies to deliver high-quality audio products with a fraction of the effort that was once required. ## Our Contribution Our early access engine boots, loads one audio graph, and allocates its memory pool once. From that moment on it pushes every block through the chain before the next interrupt fires. Parameters sent over USB or UART land in the very next block, so a change in gain or EQ is audible right away. The execution layer hides word length, byte order, and cache quirks, which means the same graph runs on ARM, Xtensa, or RISC‑V without code changes. Each node tracks its own cycle count and buffer headroom, giving honest performance numbers without risky print statements. Because the engine never allocates after start, RAM use is fixed and glitches disappear. Drop the binary on new hardware, wire up the I/O driver, press play, and your mix sounds the same day. We're opening early access to Switchboard Embedded. If you're building something serious with embedded audio and you want faster iteration times, this is for you. It gave us back the control we needed to make our hardware reliable, and we think it'll do the same for you. [Sign Up for Early Access](/signup/embedded) --- # Comparing Switchboard and AudioKit > Switchboard and AudioKit have some key differences in design, scope, and cross-platform capabilities. This article outlines the strengths of both to ensure you make the right choice. ![](/_astro/switchboard-vs-audiokit-600x400_2kvzyw.webp) When it comes to Switchboard and AudioKit, there are some key differences in their design, scope, and cross-platform capabilities. We lay out aspects of both in this article to make your choice easier. ### Switchboard: A Versatile Audio Toolkit Switchboard stands out as a comprehensive, cross-platform audio SDK designed to streamline the development of audio features for a wide range of applications. Its modular architecture empowers developers to construct custom audio pipelines with ease, using a combination of JSON-based configuration and a visual editor. This flexibility makes it ideal for projects spanning from real-time communications (RTC) to AI-driven audio features and professional audio development. ### Key Advantages of Switchboard * **Cross-Platform Compatibility:** Supports iOS, Android, macOS, Windows, and web, ensuring broad reach and reusability of code. * **Modularity and Flexibility: **Allows developers to build complex audio pipelines using interchangeable audio nodes, tailoring solutions to specific needs. * **Easy to Use:** The visual editor and JSON-based configuration simplify the process of creating and modifying audio graphs, even for those without extensive audio programming experience. ### AudioKit: A Swift-Focused Framework for Apple Platforms AudioKit, while a powerful tool, is primarily focused on audio synthesis, processing, and analysis within the Apple ecosystem (iOS, macOS, tvOS). Its tight integration with AudioUnits and Swift-based API makes it a popular choice for developers building music apps, synthesizers, or audio analysis tools on Apple devices. ### Key Features of AudioKit: * **Swift-Native:** Offers a seamless integration with Swift, making it a natural choice for Apple developers. * **AudioUnit Integration:** Leverages AudioUnits for efficient audio processing and effects. * **Focused on Apple Platforms: **Primarily designed for iOS, macOS, and tvOS, providing deep integration with Apple's audio technologies. ### Choosing the Right Tool The optimal choice between Switchboard and AudioKit depends on your specific project requirements: **Cross-Platform Development:** If your application needs to function across multiple platforms, Switchboard's broad compatibility and modularity make it the clear choice. **Apple-Focused Projects:** For projects targeting exclusively Apple devices and leveraging AudioUnits, AudioKit's simplicity and tight integration can be advantageous. ### Conclusion Both Switchboard and AudioKit offer robust capabilities for audio development, but their strengths lie in different areas. Switchboard excels in cross-platform flexibility and modularity, which can improve the reach to a broader range of users, while AudioKit provides a streamlined experience for building audio applications on Apple platforms. By carefully considering your project's needs, you can select the most suitable tool to achieve your audio development goals. --- # Complementary Frameworks: How Pipecat and Switchboard Work Better Together > Building production voice AI applications requires different strengths at different stages of development. Switchboard's native C++ runtime excels at mobile production deployment, providing the on-device performance and cross-platform consistency that modern applications demand. ![](/_astro/pipecat-and-switchboard-work-better-together_ZPQlLM.webp) Developers working on real-time voice and multimodal AI agents are discovering that [Pipecat](https://github.com/pipecat-ai/pipecat) and [Switchboard](https://switchboard.audio/) form a powerful combination rather than competing alternatives. [Pipecat makes it easy to prototype conversational pipelines](https://medium.com/@cloudiafricaa/pipecat-the-easiest-way-to-build-voice-and-multimodal-conversational-ai-7072860ed05a) by orchestrating speech recognition (STT), language models, and text-to-speech (TTS) services in Python, while Switchboard provides the native runtime architecture for deploying these concepts on mobile platforms (iOS/Android) as on-device features. Rather than viewing these as competing options, understanding how they complement each other enables teams to leverage both frameworks' strengths throughout the development lifecycle. ## Pipecat's Complementary Role: Rapid Prototyping and Server Deployment **Server Architecture as a Prototyping Accelerator:** Pipecat's client/server architecture serves a complementary role to native mobile solutions by excelling at the stages where iteration speed matters most. The core bot logic [runs in a Python server process](https://www.daily.co/blog/build-a-voice-agent-for-android-with-gemini-multimodal-live/) (e.g. on a PC or cloud), and client SDKs connect to it over the network. In a typical setup, a mobile app uses Pipecat's iOS/Android SDK to stream microphone audio and receive responses, while [the Python Pipecat backend handles the AI processing in real time](https://www.daily.co/blog/build-a-voice-agent-for-android-with-gemini-multimodal-live/). This separation complements mobile-native solutions by enabling rapid server-side iteration before committing to on-device architectures. Communication via WebRTC or websockets ([the Pipecat **RTVI** protocol](https://www.daily.co/blog/build-a-voice-agent-for-android-with-gemini-multimodal-live/)) for low-latency streaming allows teams to prototype and test conversation flows with real mobile devices while keeping all the complexity on the server. This architecture complements rather than competes with native deployment frameworks because it addresses a different phase of development: concept validation and rapid experimentation. **Python's Complementary Strengths:** Pipecat's Python foundation complements lower-level native solutions by providing the high-level flexibility that experimentation requires. A CPython 3 runtime and numerous Python packages and native dependencies (e.g. for audio processing, networking, ML APIs) that are normally installed on a server enable capabilities that would be cumbersome in compiled languages during the experimentation phase. In fact, [Pipecat's production deployment model expects you to containerize the Python app (e.g. as a Docker image) and host it in the cloud](https://www.zegocloud.com/blog/daily-co-vs-zegocloud-comparison), which complements mobile native solutions by providing a server-side option for scenarios where centralized deployment makes sense. The footprint (dozens of megabytes of interpreter and libraries) and Python's rich ecosystem are assets for server-based experimentation. Pipecat's architecture as a server-side Python framework complements native mobile frameworks by optimizing for flexibility and integration ease, leveraging Python's rich ecosystem (async IO loops, web frameworks, AI service SDKs, etc.) to enable developers to [quickly prototype integrations and experiment with different services](https://www.zegocloud.com/blog/daily-co-vs-zegocloud-comparison). **Complementary Design Decisions:** Several specific architectural choices in Pipecat complement rather than compete with native mobile solutions: 1. **Language/runtime:** Pipecat is written in Python, which complements compiled native solutions by enabling rapid development with simple syntax, dynamic typing, and extensive library support. When you need to quickly validate whether a voice interaction pattern works, Python's flexibility complements the precision of compiled languages used in production frameworks. 2. **Real-time audio handling:** Pipecat's [WebRTC streaming infrastructure between client and server](https://www.daily.co/blog/build-a-voice-agent-for-android-with-gemini-multimodal-live/) complements native audio frameworks by abstracting away transport complexity during prototyping. The Pipecat client SDK handles platform-specific audio capture, complementing your eventual native implementation by allowing you to test the conversation logic separately from low-level audio concerns. 3. **Concurrency and flexibility:** Pipecat's use of async Python and multi-threaded I/O complements production mobile solutions by making it easy to test with multiple concurrent users during the validation phase. This complements single-user focused mobile testing by enabling multi-user scenario validation early. 4. **External service integration:** Pipecat pipelines seamlessly integrate external AI services (cloud APIs for STT, LLM, TTS), complementing eventual native implementations by making it trivial to compare different AI providers during experimentation. This service-agnostic approach complements the production phase where you'll commit to specific providers based on testing data. 5. **Security design:** [Pipecat's design assumes a server will hold API keys and secrets](https://docs.pipecat.ai/getting-started/web-mobile), complementing native mobile development by establishing secure patterns during prototyping that can inform production security architecture. All these factors show that Pipecat was built to complement production deployment tools by excelling at *rapid prototyping and server-based deployment*. It's intended to work alongside rather than replace native mobile solutions. As the documentation notes, the recommended approach is to keep the bot logic on a server (or Pipecat Cloud) and have clients connect for testing. This architecture complements mobile native deployment by handling the prototyping phase and also serving scenarios where server-based deployment is preferred (IVR systems, call centers, web applications). Teams can prototype with Pipecat, test with web and mobile clients, and then either deploy the Pipecat server to production (for server-based use cases) or transition to a native solution like Switchboard for mobile deployment. ## Switchboard's Complementary Role: Native Mobile Production **Mobile Native Architecture Complementing Server Solutions:** When Pipecat's server-based prototyping has validated the concept, Switchboard provides the complementary architecture that production mobile deployment requires. It is an SDK and runtime for real-time audio pipelines built in high-performance C++ and designed to be cross-compiled to every major platform. Instead of or in addition to a Python process running on servers, Switchboard provides a lightweight audio engine that runs inside your mobile app. Developers define an audio/voice processing graph (via a JSON config or using provided APIs), and the Switchboard engine executes it locally in real time. This architecture complements server-based prototyping by enabling the voice agent to live within the iOS or Android application process when that's what the product needs. Crucially, Switchboard's runtime is implemented in C++ for speed and compiled directly for each platform (iOS, Android, Windows, macOS, Linux, etc.), so it runs natively. From the mobile OS perspective, it's a native library that complements rather than replaces server-side logic. In fact, one of Switchboard's goals is to let developers *"build without being constrained by the limited use cases of iOS, Android, and other native platforms"*, meaning it abstracts away low-level audio handling while providing fully native execution that complements server-based AI services you may have prototyped with tools like Pipecat. **Cross-Platform Consistency Complementing Platform-Specific Development:** A key way Switchboard complements Python-based prototyping frameworks is through its language and runtime compatibility across platforms. The core engine is C++, but Switchboard offers idiomatic bindings for multiple languages and environments, including Python for desktop prototyping. It provides a Swift API for iOS, a Kotlin/Java API for Android, a JavaScript API (for web or Node integration), Python bindings for desktop prototyping, and C++ APIs for integration into games or other engines. This means mobile developers can call Switchboard's functions from Swift or Kotlin while maintaining the same underlying engine that could be prototyped in Python, creating a complementary workflow across languages. The runtime architecture is consistent: whether on iPhone or Android or PC, the same Switchboard engine runs, invoked through appropriate language bindings. This complements cross-platform prototyping by providing production-grade consistency. On iOS, you include the Switchboard SDK (as a framework or Swift Package) and it runs within the app's sandbox. On Android, it uses JNI to expose the C++ engine to Java or Kotlin. Because it's compiled native code, it complements web-based prototypes by providing the App Store-compatible architecture that production requires. Switchboard bridges the gap between high-level app code and low-level audio processing, complementing what server-based prototyping tools validate by providing the native execution layer. **Production Architecture Complementing Prototype Validation:** Switchboard's runtime complements the validation work done during prototyping with Pipecat or similar tools. It uses a modular audio graph architecture: you build a pipeline of nodes (microphone input node, STT (speech-to-text) node, LLM inference node, TTS node, output to speaker, etc.) and the engine executes the graph with highly optimized C++ code. This node-based architecture can mirror the pipeline structure validated during Pipecat prototyping, making the transition smoother. It supports both on-device modules and API-based modules in the same pipeline, complementing prototype work by enabling you to keep using cloud services you validated with Pipecat while adding on-device optimizations. For example, one node could run an on-device wake-word detector or VAD (voice activity detection) model, then another node calls the same cloud API for STT that you tested during Pipecat prototyping, then another runs a local filter, etc. This *hybrid* approach complements server-based prototyping by letting you incrementally move processing on-device while maintaining connectivity to validated cloud services. Switchboard is built for real-time performance with sub-10ms processing latency per module and stable execution even under load, complementing prototype validation by delivering the production performance users expect. Unlike a Python loop, the C++ engine can leverage multiple threads efficiently and is tuned for audio processing without garbage collection pauses. Switchboard also offers features like dynamic graph reconfiguration and event callbacks to adapt to network or user changes on the fly, complementing static prototype configurations by providing production flexibility. Furthermore, Switchboard has a library of pre-built nodes for common functions (STT, TTS, effects, mixers, etc.) and allows custom extensions in C++ or via its API, complementing the prototype phase where you identified which components matter most. In short, Switchboard's architecture complements prototyping work by providing a production-ready mobile runtime: compiled, optimized, and integrated with platform audio APIs, enabling voice AI logic to run on-device (even offline, if using local models) while still supporting the cloud services validated during prototyping. The table below shows how Pipecat and Switchboard complement each other: | **Aspect** | **Pipecat (Prototyping & Server)** | **Switchboard (Mobile Production)** | **How They Complement Each Other** | | ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Core Runtime** | Python 3 engine (CPython interpreter) running bot logic. Server process (e.g. FastAPI/WebRTC pipeline in Python). | C++ audio engine running in-app. Compiled to native code for each platform (no interpreter needed). | Pipecat enables rapid server-side validation; Switchboard takes validated concepts to native mobile. Both can call the same cloud APIs for consistency. | | **Deployment Model** | Client-server architecture: bot runs on a server (cloud or local PC), clients connect over network (WebRTC/WebSocket). | Embedded library architecture: voice pipeline runs on-device inside the mobile app process. | Pipecat prototypes and hosts server-side features; Switchboard handles on-device deployment. Hybrid architectures can use both (Pipecat for cloud components, Switchboard for device). | | **Platform Support** | Server-side Python runtime with client SDKs for testing. Mobile SDKs stream data to Pipecat server. Also server-based production (cloud, containers). | Built for iOS, Android, Web, desktop native deployment: unified engine with language bindings for Swift, Kotlin, JavaScript, C++, and Python. | Pipecat enables testing on all platforms during prototyping; Switchboard provides native deployment on each platform. Switchboard's Python bindings even enable some prototyping, creating overlap. | | **Integration Method** | Network-based architecture: backend service (Pipecat server) plus transport. Mobile integration via networking (API endpoints, WebRTC). | Native SDK integration. App links Switchboard library and invokes pipeline via Swift/Kotlin APIs. No network dependency for core operation. | Pipecat validates cloud service integrations; Switchboard can call those same APIs or run models locally. Complementary rather than competing approaches. | | **Performance & Latency** | Network transport enables server-side processing; relies on server resources. Real-time over network. Connection quality affects mobile experience. | Ultra-low latency on-device processing (sub-10ms per module). Can operate even with no internet (for local nodes). Optimized C++ for stable real-time mobile performance. | Pipecat's server processing complements Switchboard's on-device work. Hybrid architectures can use Pipecat for heavy processing, Switchboard for low-latency local components. | | **AI Modules & Flexibility** | Highly flexible in Python: easy to call many cloud AI services or swap providers (supports dozens of STT/TTS/LLM APIs). Mostly cloud-based. | Modular graph of nodes: mix on-device ML models (e.g. Whisper, local TTS) with cloud API calls. 50+ prebuilt nodes, supports custom ones. | Pipecat rapidly tests different AI providers; Switchboard implements the winners as production nodes. Services validated with Pipecat can be called from Switchboard nodes. | | **App Store Compliance** | Server-based deployment, no app store concerns (runs in cloud/containers). Mobile clients are thin networking layers that easily pass review. | Meets app store requirements via compiled native code. No external interpreter or JIT. Runs within app sandbox like any media engine. | Pipecat's test clients pass app store review during prototyping; Switchboard provides production architecture for final app. No conflict. | | **Offline Capability** | Requires connectivity to server. Perfect for prototyping with cloud AI services and server-based production where connectivity is available. | Can be fully offline-capable: on-device nodes (local speech recognizer, offline TTS) allow functionality without network. Developers choose cloud vs. local per feature. | Pipecat validates cloud approaches; Switchboard can implement offline alternatives or use the same cloud services. Complementary capabilities for different scenarios. | | **Use Case Fit** | Perfect for rapid prototyping and cloud-based voice bots. Quickly assemble voice agent logic and test via web or mobile clients. Ideal for server-side production (IVR, web demos, call centers). | Built for production deployment in native apps and devices. Ideal for voice/chat AI features inside mobile apps, games, or IoT devices with native performance. | Different use cases that complement each other: Pipecat for server-based features and prototyping, Switchboard for on-device mobile features. Can be used together in the same product. | ## How Teams Use Both Frameworks Together Many successful teams use Pipecat and Switchboard together in a complementary workflow. During the proof-of-concept stage, they leverage Pipecat to validate concepts quickly. With Pipecat, a small team can script a voice bot prototype (for example, connecting a Twilio phone call to an LLM and TTS in a few dozen lines of Python, or building a web demo where users talk to an AI agent). The Python environment enables rapid experimentation with different APIs and conversation logic. This prototyping work with Pipecat validates which AI service providers work best, what conversation flows users respond to, and whether the concept is viable at all. The server-based architecture means teams can iterate without worrying about mobile-specific constraints during this exploration phase. When prototyping validates the concept and it's time to deliver a production mobile app, Switchboard complements the work done with Pipecat by providing the native deployment architecture. Switchboard is explicitly designed to ease this transition: its creators emphasize that the SDK "saves time from prototyping through production while providing lasting flexibility," especially for cross-platform products. In practice, this means developers take the validated logic from their Pipecat prototype and implement it as a Switchboard audio graph. The AI services tested with Pipecat can often be called from Switchboard nodes, maintaining continuity between prototyping and production. For instance, if a Pipecat prototype used Deepgram's API for transcription and ElevenLabs for TTS via cloud calls, the production app with Switchboard might use those same services via API call nodes, or it might use a Whisper model node on-device for transcription with an offline TTS node, depending on what the prototype testing revealed about requirements. This flexibility shows how the frameworks complement rather than constrain each other. As a practical example, consider a startup building a voice-enabled customer support bot. They might first use Pipecat to prove out the conversation flow and multi-service integration (because Pipecat excels at hooking into different NLP APIs in Python). The prototype runs on a cloud server during user tests. The team rapidly refines the conversation design based on testing feedback, validating the concept with stakeholders via web interfaces and mobile test clients connecting to their Pipecat server. This validates both the conversation logic and the AI service choices. When moving to production mobile app deployment, the team uses Switchboard complementarily: they embed the voice assistant logic into the app itself for better responsiveness and reliability. Thanks to Switchboard's cross-language support, the team implements the voice pipeline in Swift for iOS and reuses the same pipeline definition for Android in Kotlin, with the C++ engine handling execution consistently. The cloud API services they validated with Pipecat can still be used via Switchboard API nodes, or they can transition to on-device models where appropriate. The product may even use both frameworks simultaneously: Switchboard handles on-device voice interactions, while a Pipecat server handles complex multi-turn conversations or integrations with backend systems. This complementary approach is increasingly common in voice AI products. Several developers in the voice AI community have adopted this complementary pattern: using Pipecat or similar Python frameworks for prototyping and server-side features, while using Switchboard for mobile-native components. This approach leverages the strengths of each framework without forcing an either/or choice. Switchboard's documentation targets AI voice assistant developers and real-time communications apps, positioning it as complementary to prototype-focused tools rather than a replacement for all use cases. In summary, Pipecat and Switchboard form a complementary pair that serves different stages and deployment models in voice AI development. Pipecat offers flexibility and speed for assembling conversational AI prototypes and deploying server-based applications, leveraging Python's ecosystem to enable rapid experimentation with AI services and conversation flows. Switchboard provides the native mobile deployment architecture that production apps require, with a C++ real-time engine, multi-language SDKs, and hybrid on-device/cloud processing. Projects benefit from using both frameworks where each excels: Pipecat for rapid prototyping and server-based features, Switchboard for mobile native deployment. They can even be used together in the same product architecture, with Pipecat handling server-side components and Switchboard handling on-device mobile features. This complementary approach harnesses the best of both frameworks: the rapid iteration and service integration flexibility of a high-level prototype framework, combined with the efficiency, performance, and native integration of a production-ready mobile runtime. The result is voice AI features that move efficiently from concept to production, delivering seamless experiences across server-based and mobile contexts. --- # What makes Switchboard different > Your instant voice network. ![](/_astro/engineering-hub-what-makes-switchboard-different_Z17GLG4.webp) Switchboard is a truly unique audio SDK. There are many other audio SDKs out there, but they generally serve a different purpose than Switchboard. There are great audio SDKs besides Switchboard, but they are designed for different use cases. Below we have listed a number of popular audio SDKs to help you understand how they are positioned relative to Switchboard.\ \ It’s worth bearing in mind that we designed Switchboard to solve problems that arose time and time again across the many audio software development projects we’ve been involved with, despite the existence of other audio SDKs. These other SDKs often require: * Audio-specific expertise * C++ expertise * Lots of custom audio handling code * Writing native code for each platform / OS * Spending a lot of time writing glue code to get different DSP modules to work together Switchboard gets around all of that in a unique way. Moreover, it targets different use cases than most other audio SDKs. Most other audio SDKs focus on use cases such as: * Developing DAW plugins * Adding specific DSP effects into an existing music production application * And in many of these cases, using other SDKs independently of Switchboard might make sense. But where audio pipelines are more complex and involve multiple audio processing modules, multiple platforms, or multiple SDKs coming together in a single application that’s rich with audio features, chances are Switchboard can be used together with these other SDKs, or on a standalone basis depending on the use case. Let’s explore some examples below: Largely for individual audio processing nodes, rather than building audio graphs. Requires C++. Optimized for low latency and performance (e.g. minimizing CPU usage on mobile and embedded devices). Switchboard offers a Superpowered Extension to make it easier to use and to construct graphs containing Superpowered nodes. Largely used by audio software developers building plugins for DAWs (e.g. VST plugins). Switchboard targets a different set of use cases. *See ****Use Cases**** in the main menu.* A programming language. Primarily targets music composition / music production use cases (e.g. DAWs, plugins). Switchboard targets a different set of use cases. *See ****Use Cases**** in the main menu.* Web-based audio library. Primarily targets music composition / music production use cases (e.g. DAWs, plugins). Switchboard targets a different set of platforms and use cases. *See ****Use Cases**** in the main menu.* Good for certain projects in the Apple ecosystem. Not cross platform. Switchboard offers a more comprehensive set of nodes, is cross platform, has a visual editor, and enables a broader set of use cases. *See ****Use Cases**** in the main menu.* Music focused modules including stem separation, lyrics transcription, chord recognition, beat detection, mastering, vocal synthesis, and more. Typically used more for offline processing of music files. Switchboard targets a different set of use cases. *See ****Use Cases**** in the main menu.* Creators of Audio Weaver offer a visual interface where DSP nodes can be assembled into signal processing chains, targeting embedded and automotive applications. In contrast, the Switchboard Editor is web-based, user-friendly, and free, with broader node types, use cases, and platforms. While Switchboard is expanding embedded platform support, it doesn’t match Audio Weaver's out-of-the-box capabilities. However, we can enhance support as needed through a services agreement. Focused mainly on wireless headphones market. Different set of nodes and use cases targeted than Switchboard, though there is partial overlap. Switchboard focuses more on novel features and use cases for headphones, speakers, and other embedded devices, with and without a companion app. Such use cases would require a support / services agreement. Get in touch to learn more. [Get in touch](https://synervoz.com/contact/) --- # EdgeSpeech > Avoid the financial drain and developer strain of on-device audio with EdgeSpeech, our easy to use hook for React Native ![](/_astro/why-voice-ai-needs-to-be-on-your-device_29I6tJ.webp) You want to develop an app that your users can talk to; what are your options? You could rely on high level cloud AI to process your audio or you could rely on low level audio processing on the device; each has its tradeoffs. ## The Hidden Costs of Cloud AI To keep things simple, you might consider using a cloud AI provider that can already handle audio. However, this simplicity comes with hidden costs. Chief among them are privacy, latency, and service fees that only grow when your user base increases. ### Privacy By its very nature, cloud-based voice AI requires sending data to a third-party server for processing. This can raise privacy concerns, especially in applications that handle sensitive information. *On-device voice AI processes data locally, ensuring that user data remains private and secure.* ### Latency Audio data requires significant bandwidth to transmit to the cloud. This can result in latency issues, especially in real-time applications where quick responses are critical. *On-device voice AI has a much shorter round trip meaning your app can respond sooner.* ### Service Fees Audio data sent to voice AI cloud services requires processing at the provider which incurs high token usage. *On-device voice AI can reduce this by processing data locally, allowing much smaller payloads to be sent to the cloud requiring fewer tokens to interpret.* ## The Hidden Complexity of On-Device Audio If the tradeoffs of cloud AI are too severe you may consider on-device audio processing. However, those cloud AI providers are hiding a lot of complexity in their service; complexity you’ll need to implement locally all by yourself. Setting up a low level real-time audio pipeline comes with a number of challenges. ### Voice Activity Detection (VAD) First, you need to know when someone is speaking. Processing audio consumes resources so it’s important to avoid processing silence. This involves integrating a voice activity detector (VAD) and tuning it for potentially noisy environments. ### Speech-to-Text (STT) Next, you need to convert the speech to text (STT). This involves integrating an STT model that’s both small enough to run on-device and efficient enough to run in real-time. ### Text-to-Speech (TTS) Finally, you need to convert text back into speech (TTS). This requires integrating a TTS model that can run efficiently on the device but also sounds natural; robotic voices aren’t going to cut it. ## The Best of Both Worlds EdgeSpeech is the best of both worlds. It dramatically simplifies on-device audio so you can focus on building your app, while at the same time saving you money by avoiding sending large payloads to cloud providers. ## The 96% Savings of On-Device AI On-device voice AI can save about 96% of the costs associated with cloud AI services. Consider a voice AI assistant handling 1,000 conversations per day, each lasting 5 minutes. If you used OpenAI Realtime API to do speech-to-speech, this would cost roughly $7,200 per month. Here’s the breakdown: | Component | Calculation | Cost | | --------------------------- | -------------------------------- | ------------ | | Audio input | 150 sec × 10 tokens/sec × $32/1M | $0.05 | | Audio output | 150 sec × 20 tokens/sec × $64/1M | $0.19 | | **Per conversation** | | **$0.24** | | **1,000 conversations/day** | | **$240/day** | | **Monthly (30 days)** | | **$7,200** | Meanwhile, using EdgeSpeech to do speech-to-speech on the device means you can send lightweight text-only messages to the ChatGPT API instead. This would cost roughly $281 per month; a savings of more than 25x. | Component | Calculation | Cost | | --------------------------- | ----------------------- | ------------- | | Text input | \~750 tokens × $2.50/1M | $0.002 | | Text output | \~750 tokens × $10/1M | $0.008 | | **Per conversation** | | **$0.01** | | **1,000 conversations/day** | | **$9.38/day** | | **Monthly (30 days)** | | **$281.25** | ## Integrate EdgeSpeech EdgeSpeech uses a provider and hook pattern. All you need to do is wrap your app in the `EdgeSpeechProvider` and then call the `useEdgeSpeech` hook. The hook returns actions to listen and speak as well as reactive values of the speech-to-text (STT) transcripts. ``` import { EdgeSpeechProvider, useEdgeSpeech } from '@synervoz/edgespeech' // use the hook to communicate with the local model function VoiceChat() { const { listen, speak, onTranscriptComplete } = useEdgeSpeech() onTranscriptComplete(async (text) => { const response = await chat(text) await speak(response) }) return