upvote
With models there are a bunch of other dials that can be tuned even if the model itself remains exactly the same.

Are those dials set the same across all hardware configurations and clusters? Does model behavior average out the same across different hardware?

There are just too many different buttons that can be set to really trust a provider either not to directly commit fraud, or indirectly commit fraud with system complexity affecting the output.

reply
To my understanding, with batched inference and other "optimizations" you wouldn't get the exact same token predictions even with temp=0.0.
reply