It's weird to think of these kinds of models as having "output tokens". Cross-encoder approaches like Laya add a [MASK] marker per option, but nothing is generated the way an autoregressive transformer generates. It's one bidirectional pass over your input, then a small head scores each option, so you wouldn't really pay for output as much as only input