All videosOther topics0:36

Voice architectures

Where text appears and why full-duplex is a separate axis.

Links containing ?t= open the video at a specific second.

Video summary

The ideas to retain

01

Three modality architectures

02

1. Full cascade: audio → STT → LLM → TTS → audio

A conventional full cascade exposes a contract at every stage:

03

2. Audio-native input with text output + external TTS

Half-cascade is not a formal standard, but the term is now used in production frameworks. LiveKit, for example, defines a half-cascade as a realtime model that understands speech and…

Key moments

Jump directly to a section

  1. Observable boundaries
  2. Modality
  3. Interaction