i mean we use max benchmarks because they have the best coverage, which sucks because very few people use max day to day, but it's what we have. the performance curve is generally pretty similar across models and effort levels, weird outliers are pretty weird. max is generally a big cost bump from most providers (less so from OpenAI)
reply