As long as you can run on a CSP, and can engineer around the high-ish latency (most business cases can), it's extremely expensive to try engineer around it.
The range of things you can do with blob storage and a (very simple) auth model are surprisingly broad.
We recently replaced our Docker container registry with S3 using a tiny tool [1] we built in-house. I think that even with current capabilities, we can still model a lot services as a very thin layer over object storage.
So you need something to wrap the layout into something Docker understands. Because S3 is not a server, you have to construct the tarball on the fly as the image is pulled and stream it straight into docker load.
We use R2 (not affiliated) which doesn't have egress fees.
Alternatively you can stand up your own object store services, but that's not something I would like to do.
Huh? We're talking about the same S3 right? At list pricing, 1TB is $23/TB to store for one month, and about $90/TB (plus request fees) to send it out to the internet.
While hard drive prices are roughly 3x what they were a year ago, the price per TB of a new hard disk averages around $30/TB -- assume 2x overhead for other hardware and extra space for parity etc, a disk pays for itself in less than 3 months and lasts 5 years or more.
If you assume it takes 1 month to download that 1TB (about 3Mb/s) that's $29.16/Mbps. When I first started buying internet transit in Europe back in 2008 I think I was paying under $10/month. It's now under $0.10/Mbps pretty much anywhere in the US or Europe at the big datacenters.
None of the "big" object storage services are cheap. They're somewhat reasonable if you only access the data from within the same region, but absolutely insane if you ever want to ship that data outside of that cloud vendor (or to another region etc). The pricing of storage and egress has not changed in a decade (I believe AWS last lowered the price of either in 2016) and in fact it costs even more now due to things like NAT Gateways etc.
It's definitely not "shockingly cheap". It's just cheaper than $80/TB of gp3 or $45/TB of st1, and while sc1 is $15/TB it has a baseline performance of only 12 MB/s. There's quite a few companies out there that have $6/TB/month object storage plans with similar performance, rising to about $15-$18/TB/month for SSD backed storage with far lower latency figures.
Whatever AWS is selling you - except Deep Archive - I'll figure out a way to sell you for half that price, if you want, and it'll still be 80% profit for me. Your only downside will be that I don't know what I'm doing so it might not be reliable - but neither is AWS.
Well because your provider assumes you are not using all your bandwidth constantly. Cloud bandwidth is only billed for you actually use
And if you have two specific endpoints you need to transfer data between at a high rate, you can get stupidly cheap cost per GB on a leased line in exchange for making all that commitment upfront.
> I do wonder if we will see an expansion of the s3 api to support more of these use cases
This is actually an area where I think we have a big leg up on folks building on top of S3. My team (which built K2) sits next to the R2 team, and we have the opportunity to co-evolve the products in mutually beneficial ways.
I think the opportunity extends below 100ms too, particularly given the existence of faster object storage tiers like S3 express or more recently GCS rapid bucket (both only offering single-zone durability, so still need to do quorum writes to get region-level durability as with standard tiers).
One of the tensions of course is how long to linger before flushing to object storage - you have to trade off directly between latency and cost of your API ops for PUTs.
When building the serverless offering of s2.dev (which is in a similar space, full disclosure!), we designed around stateful backend processes capable of constantly flushing multi-tenant objects (i.e., containing records from many streams), allowing streams to offer low ack latencies (~50ms p99 from same region) without blowing up the unit economics.
Congrats on the launch btw!
Haven't tried it yet but looks nice for simpler K8s deployments
[0] https://docs.aws.amazon.com/vpc/latest/privatelink/vpc-endpo...
I am currently using Cloudflare R2 right now and if you see their forums, there's always the occasional post about objects going missing.
Don't go for the "object gateway" compatibility layer - just use raw Ceph if you're writing your own app.
Consider your cloud costs.
Backblaze's famous reports put HD failure rate is ~1.39%, so for the 11 9s you get as a guarantee from S3. Assuming nothing else fails, you'd need at least 6 independent copies to get that, plus all the effort to engineer recovery, and constant upkeep.
Suddenly, S3, even when considering bandwidth costs, seems like a steal.
If you have two sets of hardware the server can run on (cold standby), and RAID, and are competent at physical IT work, you can have faulty hardware replaced in an hour. Drive fails - replace it. Anything else fails - swap the drives to the other machine, boot it up and then troubleshoot the original.
Most likely you don't even need that. If the app server runs on a standard platform like Windows you can shuffle it over to some spare tower PC. You hear horror stories of dusty towers that nobody knows what they're doing - the horror there isn't from running server software on a tower, but from unmaintained servers no matter the form factor