Since yesterday a lot of users in Europe found their workflows failing due to Github seemingly randomly throwing HTTP/401 on git clone/git pull when interacting with public repos without authentication.
It was now confirmed by staff member that this indeed is intentional and no further steps are planned at this point.


The word is “scrapers,” as in to scrape.
no disassemble!
Well given the waves of “buy book, scan, scrap” scrappers might become an acceptable version as well soon.
We’re talking about AI bots scraping GitHub in this post. Nothing to do with books.
Yes, and I was talking about the practice of some AI companies buying books in bulk to feed to their LLMs en masse which sometimes/often destroys the book in the process.
I was trying to build a bridge between those two related topics to “justify” the mistake and make it an intentional remark instead.
Wooooooooosh
Are you claiming that “AI scrappers” is a valid term and not just an (all too common) misspelling? If so, I’ll need to see a source on that, please.
Hah. No. I’m claiming there’s a joke that went over your head. And apparently over the heads of quite a few more readers. That’s OK.
What was the joke?
That the misspelling is amusing because if you look at it in a certain way it would still apply to AI/LLM companies. Whether it was intended or not.
Do we always have to be so tight and literal about everything?
Yeah except in the context of the post/article, it seems pretty far fetched for someone to be referring to the “scrapping” of books (assuming that’s what you meant; it’s strange that neither you nor the person who posted “whoosh” was willing to clearly state what the supposed joke was).
Misspellings annoy me, especially when they’re common like this one. Trying to explain away the misspelling as some kind of “joke” annoys me even more.
I made a typo, so no joke was intended 🤷
otoh bots scrapping the internet infra with massive traffic… 🤔
Thanks for clarifying! cc @aev_software@programming.dev
did microsoft just lock up a good chunk of open source behind a (stochastic (for now)) login wall?
this feels like it should violate the gpl, but i bet it doesn’t. truly devious.
Public GitHub repositories remain public and can still be accessed without a GitHub account, including repositories owned by paying customers. However, a subset of unauthenticated clone or fetch requests may now be asked to authenticate as part of GitHub’s protections against abusive traffic. If you receive a 401, update your application or script to use GitHub credentials.
Uh okay.
I wonder why people still keep up with this bullshit and not switch to some better public Git hosting provider.
Could you suggest an alternative?
I know Codeberg exists, but they had reliability problems recently IIRC?
sr.ht and Gitlab are paid products.
Technically one can self-host a forge, but my attempts at setting up CI were unsuccessful (IMO that’s way more complicated that setting up the forge itself).
Codeberg’s couldn’t be much worse than github’s reliability since 2019
The problem with Codeberg now is that there is a spectrum of AI usage between manually coding and AI slop that Codeberg chooses to ignore with their new policy.
I get it that’s their platform and their choice, and at some point I even looked for ways to support them financially if that meant I could migrate all my repos from GitHub - I wouldn’t mind paying for my use, “vibe coded” or not - but an intransigent LLM ban creates uncertainty for many of my active projects, and I can’t consider them a viable alternative to GitHub anymore.
So sure, I’ll take 3h of GH actions downtime a month over a platform ban. I’m sure I’m not alone there.
So, you should be fine with the measures against scraping, then?
That doesn’t affect me, but you can’t eat your cake and have it too. I’d rather have open source content actually open and easy to access, but GitHub has an availability issue and they’re trying to address it, so I don’t see this as a hostile move.
I’m sure if codeberg was nearly as popular as github, they’d be forced to restrict some clones too. Their current load is unprecedented, even for their standards.

you can self host gitlab too, and its free. yes, you CAN buy a license, but you can run it free forever. you can also use their SaaS free forever too.
at least right now, I think it’s the best alternative, though I completely understand people wanting to favor OSS.
disclaimer: I have contributed code to gitlab, but I am NOT an employee.
Their SaaS is pretty good but of course you are running the same chain trust. You’re betting that gitlab doesn’t enshittify within the next 5 years which is hardly a guarantee.
Self-hosting gitlab is very resource-intense and complex from what I tried, though I did only try it two or three times.
I did set up forgejo which was way easier and less heavy but I haven’t tested it much so who knows.
I hosted GitLab Community Edition on-prem as a trial for a very small team, but switched to a Microsoft offering. Part of it was it being demanding, part of it was not. The point is you can host CE yourself and get an experience that’s very similar to their service, “for free,” where “free” translates to your hardware requirements and responsibility. I’d even say it’s worth it to go with GitLab, in that context, for the familiarity you can provide your team. “You’ll have to learn gitonator9000” turns into “you’ve used GitLab, right?”
Personal use? You could run it in a container and periodically backup your data. Is it proportionally more demanding than other things? Probably.
but they had reliability problems recently
Like GitHub, yes. But if you’re not going to selfhost Forgejo they’re the best option.
Codefloe is very nice and fast, and their CI can do Jsonnet instead of only dumb YAML.
Sounds interesting, thanks for the tip!
Yes, they had reliability problems because so many fucking people are suddenly switching to them because they’re so much better overall and not evil. Those are the kind of reliability problems that it is genuinely nice to see someone having. Having a little bit of a bumpy road when scaling due to significant rapid adoption is normal. So what?
Imagine not wanting to use Linux because it recently had a bunch of security flaws. And it did. But again, so what? Does that mean Linux has always been insecure? Well for those things it was. But is it still insecure? Maybe, who knows, nothing is perfect. Are you going to refuse to use it because you’re not sure? Why? Past performance is not an indicator of future success.
If some minor reliability issues are your foremost concern to the point that the other things codeberg provides for free are not valuable to you because of it, I question the depth of your priorities.
That said, it is much healthier and better for people to self-host or use smaller less centralized providers if possible. I do not wish Codeberg to become a victim of their own success, and a healthy ecosystem is a diverse one. But it is not for everyone, and if all you need is a minimal fuss alternative to Github, Codeberg is right there.
Can you help me find a solution that:
- I don’t have to self host
- Provides SSO
- Provides local and remote build agents and actions
- Allows me to store private proprietary code
- Supports static IPs for runners
Number 4 rules out Codeberg. The only other one that really supports that level is Azure DevOps, and well….
Holy crap this is huge!
Open source is no longer open source on github.
Be ready to supply your fingerprints using Microsoft Authenticator app and your free source access subscription for only 49 USD per month! /s
deleted by creator
Our aim is to make public repositories accessible without authentication as much as possible. However,…
It’s new reddit then. I was still using Github as a shitty backup for my projects, but an alternative may be required faster than expected.
Embrace. Extend. Extinguish.
Time to leave github! Boycott is the only language companies understand!
AI bot problem is real and there is no good solution to it. Look, I very much dislike GitHub for various reasons BUT currently there is no good way to throttle AI bots that literally trash web. They are like that geeky classmate who can never hold his liquors: it’s nice having them around for some answers, but they ramble a lot and shit/puke in random places of the house making it unlivable.
They are the ones who created this crap
I believe that it is some fuckup and they don’t know where exactly problem is and while they are looking for a way to fix it they’ve created plausible lie. When they’ll fix it or believe that it’s fixed there would be a public announcement like ‘we heard the community and reversed our decision’ and users will be happy. There is a serious need for github mirror, they are becoming less and less stable every year.
I’m 100% sure this is damage control on their part, they refused to acknowledge the incident and are looking for their way out.
What users found in this thread is
- problem is limited to EU
- problem is limited to subset of
gitbuilds (gix version x TLS lib x TLS lib version in place) - problem goes away if you switch back to HTTP/1.1 for some reason
If these are LLM scrapper mitigation steps then apparently fighting LLM scrappers is 7D chess game or something 🤷
The Microslop marketing department was always their best department.
Our aim is to make public repositories accessible without authentication as much as possible. However, like much of the Internet, we continue to see significant increases in the volume of robot traffic recently which has increased the need for verification, for example CAPTCHAs.
As I expected, it’s about combating (excessive) bot traffic.
Hmm… I don’t know whose building all those bots to scrape all of internet
“We’re all trying to find the scraping LLM who did this!”
First time ever seeing color in a post titleRendered
fixed fontwith the use of one and only bacticks (`)













