upvote
basically he is feeding the same input to multiple models, taking their outputs and dumping it int an LLM to sort out what the reason transcription probably is. expensive but effective.
reply