Streaming typed every phrase as soon as a 250ms pause came along,
which made long dictations trickle in slowly and double-emitted the
last phrase (once as the phrase, once in the final flush — the source
of the phantom "Thank you."). With phrase_silence_ms at 5s, phrases
only flush at the very end, so the whole text lands in one piece
about 0.1s after the key is released. The 20s force-split still
emits during very long continuous speech, so nothing is lost there.