Where text appears and why full-duplex is a separate axis.
Links containing ?t= open the video at a specific second.
Video summary
The ideas to retain
01
Three modality architectures
02
1. Full cascade: audio → STT → LLM → TTS → audio
A conventional full cascade exposes a contract at every stage:
03
2. Audio-native input with text output + external TTS
Half-cascade is not a formal standard, but the term is now used in production frameworks. LiveKit, for example, defines a half-cascade as a realtime model that understands speech and…