40% overhead is quite typical, so you'd be looking at $120k/year. In the USA I'd consider that a competitive salary for an admin capable or keeping a $6M rack of specialized hardware running 24/7.
A company of the size that this is worthwhile for, probably has dedicated devops on staff already and can add this rack to the inventory with no additional staff.
Who is going to fix it when the api does something weird ?
Who is going to proactively make sure it’s not overheating?
Chat GPT has enterprise contracts for a reason.
I didn't say it's fire-and-forget. I'm saying all that is maybe a day of work every 3 months.
> A company of the size that this is worthwhile for, probably has dedicated devops on staff already
Adding another rack beside the VMWare cluster, managing any storage/networking issues, etc. will be incremental costs; they already have a pager (probably not a pager not anymore just an app on their phone) like rotation schedule etc.
Why not? Do you trust AWS with things that need to be private?
I'd update this to
'I don’t LLMs for anything that needs to be private'
What's to prevent the LLM from sliding a heavily obfuscated binary blob into the application that does nefarious things? If you aren't creating the LLM itself from scratch, I don't feel it can be trusted.
Or that perhaps you have a perfect alignment algorithm which you are unwilling to share with the broader research community (evil)?
A NN per se is a file... The executable that runs it can "act"...
Having these models in the open caps the inference margin.
Not a devops but I'd say one full time is already too many.
Add to this number another $1.5m/yr in opex, so not sure I’d call such an enterprise wealthy enough to spend those kinds of sums on LLMs an “SMB”.
(managed an entire data center building with thousands of servers a lifetime ago with ~2-3 other people, it’s only gotten easier over the last two decades imho)