Short utterances expose brittle edge cases in voice pipelines where VAD sensitivity, minimum speech duration, and turn-taking heuristics easily collide. When building real-time audio agents over WebRTC, misconfigured silence detection or STT filtering often leads to awkward agent silence. This post walks through practical troubleshooting steps across transcription and endpointing layers to ensure quick confirmations aren’t lost.