upvote
latency, instruction adherence, reliable tool cools, conversationality are all in tension.

it's great when you can get a 170ms ttft. but if you have 700 ms endpointing on the stt side and 300ms ttfb on the voice side, then you haven't really made something super snappy.

reply