Is their uptime acceptable? No. But personal attacks aren’t necessary or constructive.
Core parts of the product, like navigating to individual files in a code review, are broken
I think this is a good argument to underline "It's not _just_ the scale". Adding to this, the Github Code Review experience is kind-of broken, the way comments/threads are stacked in the PR overview has not improved, pagination isn't really a thing, and these issues are age old. Hopefully, one day, Github will mature.
The final step is challenging; likely the most difficult part is changing all the code references and imports. Shadowing changes would be straightforward. Training your likely 25-50 engineers to use the new code review UX would likely not take that long.
Considering the wasted engineering velocity during Github outages, it's worthwhile to do even a partial migration. Github's action runners have in my experience, been the most fragile part of the platform. Given the ease of moving build and merge queue runners to alternates, it's a no-brainer.
https://this.weekinsecurity.com/microsoft-wins-lamest-vendor...
I gather that you have intimate and deep knowledge on the teams and the problems they try to solve there.
And 99% of the bread that I get is good but 1% of the loaves, they forgot to add flour. Consistently, for years, they always have loaves missing a key ingredient that I still end up paying for.
I can be pretty sure that BreadHub have a pretty major internal issue, and should probably be questioning their competence, regardless of their “scale”, and without any knowledge of the “problems they’re solving”
I’m saying that GH is operating at a huge scale with (probably) lots of technical debt and a forced migration to new infrastructure.
I would be (and am) highly critical of leadership. I’m not going to make strong assertions about ICs without knowing their context. I’ve worked at a company with a sterling reputation for engineering excellence where brilliant ICs were kneecapped by poor leadership.
I think a lot of us, at one point or another in our careers, have worked with potato leadership that can be short-sighted or political. It isn’t a comment on the engineers.
An outage is more like a shipping issue with the supplier, if it's owned wholly by them.
I don't think this is commendable at all. I give GitHub a lot of money and I'm tired of it being wasted with downtime.
I don't know the architecture or any of that, but I feel like there could be (and it's not like they would've really known this until the last year or two with the massive spike) separate infrastructure for paid users/orgs vs free the same way they make the distinction with enterprise.
I get the massive load changes that they are under over the last two years, but why does a bunch of vibe coded slop take down the same resources that my company pays for every single month and has for years? I imagine properly splitting that out would be an absolute headache and not worthwhile for them vs stabilizing the rest of the service, but damn it sucks when I get blocked at work because GH is down.
Of course it is not free of all management but for our use case it is working.
There are hiccups with the CI runners from time time but nothing major and we have another machine in another rack that serves as a backup which can be brought up ~< 20 minutes.
I know companies have long since tossed their expertise for hosting their own stuff in favour of SaaS but at some point its hard to beat the up time of a single machine.
It's a lot more expensive and has a good bit more limitations to the regular SaaS product.
1. Github has enterprise users who paid for the service, their day job requires Github to be available and working
2. Github has generous free tier which is the one which is exploring a lot more with the AI generated code.
It is a complexity in itself but the traffic should have been separated, the free users should not be allowed to bring down Github for enterprise customers (Just to clarify, I am free user myself). And if they do not have capacity it would have been perfectly fine to push back or throttle new users/repositories.
But I think they’re in a tough spot. GitHub has historically been a huge supporter, proponent, and provider for open source projects. Engineers are difficult customers, to say the least, and the community would likely freak tf out of the segmented traffic.
The logical, pragmatic, and justifiable answer doesn’t always align with your market.
They are in a tough spot cause they want to continuously ingest all the data to train their LLMs. Any fork or decentralization at large scale of git is going to impact the training pipeline.
They decided they needed to capture the whole open source ecosystem by turning open source work into social networking... on a proprietary platform (because open source is great, especially when it's others' software). That was before they joined Microsoft.
And then Microsoft pushed AI everywhere, including on GitHub itself with copilot.
I would have liked if they had left the open source projects alone and didn't create that FOMO for not using them.
I have no sympathy.
Ding ding ding, we have a winner. I like AI. I work for an AI company. Still, Microsoft aggressively pushed GitHub users toward Copilot. They don't get to do that and complain about increased volume from AI-generated changes.
No Copilot + reasonable operation: the way things were
Copilot + reasonable operation: Well done!
No Copilot + being overwhelmed by AI commits: Sympathy.
Copilot + being overwhelmed by AI commits: "Where did that petard come from that's hoisting us?"
I can have sympathy for the humans caught in the crossfire but only managing one nine of availability on a commercial service is not acceptable.
That is an active choice they are making, over and instead of any reliability for the rest of their users.
It is a problem they are embracing, and actively encouraging, for themselves.
No.
Their leadership went all in on AI and in the last blog post essentially admitted some missing test coverage for a critical path. Time to learn lessons and fix your vibe coded shitslop and stop using "user graph go up" as some kind of excuse.
But we pay enterprise license and GitHub is a big dependency in our software flow.
If this continues to be a problem as an enterprise product they need to do something. Otherwise theyre are going to to start losing business
Yes, that is scale. And yes, that's not actual requests per second. But it's the sort of scale that big (and even mid-sized) tech has known how to deal with for decades. Microsoft doesn't have an excuse.
As I said before (and it was not a popular comment), it's easy to be an "armchair quarterback," with these services, as I think the brittleness was already baked in, and just waiting for the right time to crack. The only true way to have a robust platform, is to design something that will scale, from the start, and many startups don't do that, because they are feverishly trying to get out an MVP; even if it is a mess of bubblegum and baling wire.
They always say "We'll get it done right, once we get funding," but that never happens.
But my sympathy is limited by the fact that MS paid a lot of money for this, quite a while ago, and that is one company that knows all about issues of scale. They should have seen this coming.
We certainly do not.
A restaurant makes pizzas. They suddenly get 100x more popular. They can't make 100x more pizzas. But they are still taking orders from 100x more people. Not only are they not getting enough pizzas delivered that they took orders for, but in their rush to make and deliver more pizzas, they set the kitchen on fire, which makes an even longer wait for pizzas.
When the pizza you ordered doesn't get delivered, do you have sympathy for the restaurant? Or do you tell them to stop taking orders they can't fill and try not to set the kitchen on fire?
Now consider the pizza restaurant has 21 billion dollars in cash, is taking your money, and not giving you pizza.
I don't have sympathy for them. They have disrupted my work so much in the past month that it's ridiculous, and my company likely pays them millions of dollars. And you can see from this site's plotting that there is a trend towards more frequent and more critical outages. It has been extremely bad the past month.
This is a multi billion dollar corporation. History of robbing and stealing from others of their labor or IP. History of enshittifying once great services.
I used to be on-call in a high-traffic environment where single customers pushed more bits than entire nations. I chose the role. I didn't want people's sympathy, if anything, I wanted them to complain to management.
If it gets too bad they can quit. Maybe that would be for the best, just wear the thing down until it outright fails and no one wants to touch it. One less bullshit service sucking all of the oxygen out.