Interactive conversational avatars for training present a compelling real-world architecture challenge combining low-latency audio pipelines, LLMs, and real-time video rendering over WebRTC. This session covers the practical engineering trade-offs required to keep round-trip latency low enough for natural human-AI roleplay. It is a helpful case study for anyone building real-time multimodal AI applications.