saa-sdk
View on GitHubAddressee detection for voice agents: device-directed speech detection that runs before STT, so background speech, side conversations, and the agent's own TTS echo never trigger it. No wake word, model-agnostic, drop-in for LiveKit, Pipecat, ElevenLabs, Twilio, and OpenAI. The layer your VAD and turn detection are missing.
Python and JavaScript SDKs for a hosted, real-time classifier that determines whether speech is directed at a voice agent. It filters audio before STT and integrates with LiveKit, Pipecat, ElevenLabs, and Twilio without requiring a wake word.
Use Cases
Filter background speech before it reaches a voice agentPrevent an agent's TTS echo from triggering its pipelineGate speech input for robots and kiosksRoute only device-directed speech to STT and LLM servicesReduce unintended agent interruptions in telephony callsSelect which agent responds when multiple devices share a room
Built With
- Language
- Python
- Frameworks
- LiveKit · Pipecat · ElevenLabs · Twilio
Tags
voice agents · addressee detection · speech processing · real-time audio · VAD · barge-in · turn detection · wake-word alternative · audio gating · WebRTC