upvote
Last time I checked it costs closer to 500K to run SOTA open models at any usable speed.
reply
How do I know that a private LLM system won't have a degradation of quality similar to this or worse? The only thing I can think of for the proposed scenario is some sort of homomorphic encryption system? But not sure.
reply
>How do I know that a private LLM system won't have a degradation of quality similar to this or worse?

You use benchmarks, you test, and because YOU'RE the sysadmin you know what weights are running at what time, it's very visible. You can airgap the hardware and be guaranteed it won't change over time. And, frankly, degradation over time doesn't seem to be what's happening with the open models.

reply