Back to courses
IT & ENGINEERING Advanced

Voice AI and Realtime Multimodal Agents

Read the first lesson free — in full No account, no card · plus the interactive platform demo and the AI Professor Start now

A premium, advanced and complete course on building production voice AI and realtime multimodal agents, updated for 2026. You will master the modern voice stack — cascaded pipelines versus speech-native models — including speech-to-text (Whisper, Deepgram, streaming ASR), text-to-speech (ElevenLabs, Cartesia, OpenAI voices), and speech-to-speech via the OpenAI Realtime API and GPT Realtime. You will learn bidirectional streaming over WebSockets and WebRTC, latency budgeting, voice activity detection, turn detection and barge-in, orchestration with Pipecat and LiveKit Agents, function calling in voice, multimodal agents that combine audio, text and vision, realtime video, telephony and SIP integration with Twilio, rigorous evaluation and testing, deployment, scaling and cost modeling, and real-world case studies. The course carries a strong, practical focus on legal and ethical voice AI: consent to record, biometric voiceprints under GDPR, responsible voice cloning, and AI-disclosure transparency under the EU AI Act. Includes a comprehensive final assessment.

10 modules
30 lessons
~25h duration
v1.0 version
AI professor An AI agent built into every lesson — ask questions and get instant answers based on the course content
Hands-on exercises Real scenarios and practical exercises directly on the platform, with instant feedback
Progress & analytics A personal dashboard with statistics, streaks, scores and structured learning paths
Interactive AI quizzes Questions generated by AI and adapted to your level, with detailed explanations
Individual course access
€49
+ VAT / month
Get started
All lessons AI quizzes AI professor included Cancel anytime
Or read the first lesson free
or
Recommended
IT Pro bundle
€399
+ VAT / month
See the IT Pro bundle
  • Every IT Pro courseA full library, not just this course
  • AI professor in every lessonAnswers when you need them, included in your subscription
  • Quizzes, progress, streaks & statistics
  • Content updated regularly
Cancel anytime
Secure payment
Updated regularly
Content in English
Built-in AI agent Exclusive Ask anything about the lesson and get an instant answer — the agent knows the course content
Interactive AI chat Automatic summaries Personalized quizzes

What you will learn

Practical skills you gain by completing this course

Foundations of Voice AI in 2026
Speech-to-Text: Modern ASR
Text-to-Speech: Neural Voices
The Realtime API and Speech-Native Models
Building a Voice Agent: Orchestration
Function Calling and Multimodal Voice
Telephony and SIP for Voice Agents
Evaluation, Testing, and Safety
Deployment, Cost, and Case Studies
Final Quiz — Voice AI and Realtime Multimodal Agents

Who it is for

Developers Software engineers Solution architects CTOs / Tech Leads Data Scientists ML Engineers DevOps Engineers

Recommended level

Advanced

Assumes hands-on experience with AI and complex scenarios.

Updates

Regular

Last update: Aug 8, 2026. Content kept up to date.

Category

IT & Engineering

A technical course for IT professionals — available with individual course access or the IT Pro / All Access bundle.

Advanced level

Hands-on experience required

Assumes practical experience with AI. Covers complex scenarios and advanced strategies.

Always up to date

Last update: Aug 8, 2026

The course is updated regularly with the latest information, tools and practices from the industry.

Practical and applied

30 lessons with real examples

Each lesson includes practical scenarios, actionable checklists and quizzes to check your understanding.

Curriculum

10 modules, 30 lessons — structured to learn step by step.

10 modules
30 lessons
~25h of content
Interactive quizzes
Free preview available The Voice AI Stack: Cascaded Pipelines vs Speech-Native Models
Read the preview
1 Free preview lesson The Voice AI Stack: Cascaded Pipelines vs Speech-Native Models
Read the preview
2 Conversation Anatomy: Latency, Turn-Taking, and Endpointing
51 min
3 Consent, Privacy, and AI Transparency: The Legal Foundation
50 min
1 How Modern Speech Recognition Works: Whisper, Deepgram, and the ASR Landscape
51 min
2 Streaming Transcription: Partial Results, Interim Hypotheses, and Endpointing
50 min
3 ASR Accuracy: WER, Diarization, and Handling Real-World Audio
50 min
1 Neural TTS in 2026: ElevenLabs, Cartesia, and Hosted Voices
51 min
2 Streaming TTS: Latency, Chunking, and Voice Consistency
50 min
3 Voice Cloning: Ethics, Consent, and the Law
49 min
1 Speech-to-Speech Models: The OpenAI Realtime API and GPT Realtime
51 min
2 Bidirectional Streaming: WebSockets, WebRTC, and Audio Framing
51 min
3 Sessions, Events, and Managing Realtime State
50 min
1 Orchestration Frameworks: Pipecat and LiveKit Agents
51 min
2 Turn Detection, VAD, and Barge-In in Practice
51 min
3 The Latency Budget: Engineering Natural Conversation
50 min
4 Conversation Design and System Prompts for Voice Agents
50 min
1 Tool and Function Calling in Voice Agents
51 min
2 Multimodal Agents: Combining Audio, Text, and Vision
50 min
3 Realtime Video and Vision Agents
49 min
1 Connecting Voice Agents to the Phone Network: Twilio and SIP
51 min
2 Media Streams, DTMF, and Call Control
50 min
3 Building a Production Phone Agent
50 min
1 Evaluating Voice Agents: Metrics and Methods
50 min
2 Testing, Simulation, and Regression
50 min
3 Safety, Guardrails, and Responsible Voice AI
51 min
1 Deployment Architecture and Scaling
50 min
2 Cost Modeling and Optimization
50 min
3 Case Studies: Real-World Voice Agents
49 min
4 Observability, Monitoring, and Incident Response for Voice Agents
50 min
1 Final Assessment — Voice AI and Realtime Multimodal Agents
52 min
Access this course from €49 / month

Ready to start learning?

Create an account and choose how you want to learn — just this course, or the full IT Pro bundle.

30 hands-on lessons Content updated regularly AI professor included in your subscription