upvote
But training LLM's is also a task one can do whenever you have a spare GPU-minutes.

I wonder why they don't have some kind of scheduler which makes sure there are never any idle minutes. One would imagine they at least would have autoscaling on their production serving workload and use the freed compute capacity for model training for example.

reply
I doubt they're inferencing on their training hardware
reply
Is the electricity cost far greater than the marketing value?
reply
The marginal electricity cost is zero.
reply
Or specifically, electricity was already paid for with the pre-paid capacity.

Not using it would not save them any money, they already paid for it.

reply
The first is a physical quantity that can be written down.

The second is approximately no better than astrology.

reply
The second point is, sadly, true of quite a lot of aspects of software, including "design" and "quality"
reply
[dead]
reply