I built the voice assistant workflow in NuroFlow by extending the existing text pipeline. Audio recorded in the browser is sent to the FastAPI /voice endpoint, transcribed, processed through the same LangGraph planner‑executor‑validator graph, and finally spoken back with Edge‑TTS. The design mirrors the core concepts described in nuroflow key features.
Audio is captured in the browser as a .webm file. "User speaks in browser"
The transcription step uses the Faster‑Whisper model. "User speaks in browser"
Speech synthesis is performed by Edge‑TTS, producing an MP3 that is Base64‑encoded for transport. "User speaks in browser"
The JSON response from /voice includes conversation_id, user_message, response, and audio_base64. "User speaks in browser"
The planner can emit a pending_question and the validator tracks retry_count against max_retries. "pending_question = "Required information"" "retry_count max_retries"