innovator · talks

Voice Assistants: Building Bilingual Conversational AI

Master the architecture of voice assistants by building speech recognition, synthesis, and dialogue systems. Design your own bilingual assistant with safety guardrails.

20 modules·Difficulty: ★★★★★· 5 Free
Start module 1

Modules

  • 1
    Free8 min
    From Sound Waves to Digital Signals
    Discover how microphones capture vibrations and ADCs sample them into numeric arrays that computers process 🎵
  • 2
    Free10 min
    Phonemes, Spectrograms, and Feature Extraction
    Learn how Mel-frequency cepstral coefficients transform audio into visual patterns that neural nets can read.
  • 3
    Free12 min
    Acoustic Models and Language Models in ASR
    Explore how acoustic models map sounds to phonemes while language models predict likely word sequences.
  • 4
    Free9 min
    Text-to-Speech Synthesis Fundamentals
    Understand how TTS engines generate waveforms from text using concatenative and parametric methods.
  • 5
    Free11 min
    Neural Vocoders and WaveNet Architecture
    Dive into autoregressive models that predict audio samples one at a time for natural-sounding speech.
  • 6
    Paid10 min
    Prosody Control: Pitch, Duration, and Stress
    Manipulate fundamental frequency and phoneme timing to make synthetic voices sound expressive and human-like.
  • 7
    Paid12 min
    Intent Classification with Transformer Encoders
    Train BERT-style models to map user utterances to predefined intents like SetAlarm or PlayMusic.
  • 8
    Paid11 min
    Slot Filling and Named Entity Recognition
    Extract structured data like dates, locations, and times from free-form speech using sequence tagging.
  • 9
    Paid9 min
    Building Dialogue State Trackers
    Implement finite-state machines and belief trackers to maintain context across multi-turn conversations.
  • 10
    Paid10 min
    Reinforcement Learning for Dialogue Policies
    Optimize which action an assistant takes next by rewarding successful task completions in simulated chats.
  • 11
    Paid12 min
    Handling Code-Switching in Bilingual ASR
    Adapt acoustic and language models to recognize when users mix Arabic and English mid-sentence.
  • 12
    Paid11 min
    Streaming ASR and Real-Time Decoding
    Optimize beam search and chunk processing so transcriptions appear instantly as users speak.
  • 13
    Paid10 min
    Voice Activity Detection and Endpoint Detection
    Use energy thresholds and neural classifiers to distinguish speech from silence and background noise.
  • 14
    Paid9 min
    Content Moderation Filters for Safe Outputs
    Deploy keyword blocklists and toxicity classifiers to prevent assistants from generating harmful responses.
  • 15
    Paid11 min
    Privacy-Preserving Wake-Word Detection
    Run lightweight models on-device to listen for activation phrases without sending audio to the cloud.
  • 16
    Paid12 min
    Personalization Through User Embeddings
    Learn speaker-specific accents and preferences by fine-tuning models on individual usage history.
  • 17
    Paid10 min
    Evaluating Voice Assistants: WER and BLEU
    Measure transcription accuracy and response quality using automatic metrics and human annotation studies.
  • 18
    Paid11 min
    Deployment Architecture: Edge vs Cloud
    Compare latency, privacy, and cost tradeoffs when running models locally versus on remote servers.
  • 19
    Paid12 min
    Capstone: Design Your Bilingual Assistant
    Synthesize ASR, TTS, dialogue policies, and safety filters into one working prototype supporting Arabic and English.
  • 20
    Final exam25 min
    Voice Assistant Mastery Exam
    Prove your understanding of speech recognition, synthesis, dialogue systems, and safety mechanisms across ten questions.