Hacker Newsnew | past | comments | ask | show | jobs | submit | gavinhking's commentslogin

Seems like you’re part of the group represented in this dataset trend then, many of these visits are also from (compromised) Google servers in that same ASN.


Yup, in fact most of them are already. That's one of the ways this data is verifying whether the visits are spoofed or not: https://knownagents.com/insights#spoofing-and-security


Basically, unless CF is counting static asset network requests etc. For what it's worth, GA also miscategorizes some bots as humans as well.


(Insert spider man meme)


Good to know, thank you. Would you do this by fully blocking particular ASNs? Or something more granular?


You can block entire ASNs. If you are frustrated with bots, blocking Tencent's entire IP address space would have very few downsides.

If you have fail2ban or NGINX logs, you can use our CLI to summarize those IPs and identify the ASNs you want to block. But before you block entire ASNs, make sure they are not classified as "ISP" type. For that, visit our website's ASN page first.

I have quite a few community posts around this approach. https://community.ipinfo.io/

If you have raw logs, you can send them to me as well, and I can review them and provide some guidance.


Appreciate it, I'll check out your posts.


What's your strategy?


Another interesting thing here is the paths they're targeting, many are for newish AI coding tools


People or their agents must be accidentally committing or publishing their repository level secrets and configs with enough regularity that it’s worth scanning.


Totally. I'm sure this campaign was inspired by sloppy vibe coding


There are a few novel ones but I’ve been seeing most of them in my logs for longer than generative AI has existed. This isn’t remotely new, the vector is just getting bigger.


Definitely, the point here though is there is a stat sig surge in the chart in the last week, across thousands of websites. At least a surge in this particular spoofing pattern.


It's still not really anything special. Thousands isn't even large scale.

Any random bozo can trigger that.


This is a random sample of completely unrelated websites, which indicates that the total scale is much larger. This is not saying that it is difficult to make thousands of requests.


Is this your company? If it is, your cheerleading makes you very hard to trust. If it’s not, I’m sure that everyone gets the point - you adore everything about this research and can’t see any possible problems.


Not asking for trust, just sharing the data/math


You forgot the "Yes, that's my company" part in your reply (https://ghking.co)


That doesn't change the data/math brother


It's still good etiquette to disclose affiliation in online discussions.


Fair enough


It changes our evaluation of the likely reasons that you are ignoring the reasons that you are wrong.


Yeah, that's exactly what these visits are: faked user agents that fail IP verification or Web Bot Auth. What's interesting is the surge across so many websites in the last week.


There are many possibilities but one of them could be some new vuln was released and they are looking for it. That would require looking at the URL's they are requesting. Botters run their own purpose built campaigns. Do you also have a summary of URL's requested by unique counts?


Looks like many of the paths relate to AI coding tools. There are some examples below the chart


You keep repeating this about a small minority of the tools that were posted.


I've had a similar bump in scanners in the past week, more than half of it is coming from MS and Google owned IPs and all of them are spoofing AI agents.


I made this, let me know if you have questions or feedback.


Thanks for the work on this!

I automated my site's robots.txt[0] by scraping your site. It would be extra nice if darkvisitor.com exposed a plain text version or JSON representation of the list.

[0] https://tbeseda.com/blog/automating-my-robots-txt-to-block-a...


That was one of the points of the new repo. A plain text version of the file is https://raw.githubusercontent.com/ai-robots-txt/ai.robots.tx...

The other point was to make this community maintained rather than rely on one source to provide all the inputs.


Definitely! And I'll likely use that raw URL in my tooling going forward. So thanks for starting the repo.

I'm thinking it also helps to bring up a feature request on the source material so we can all limit the drift. ie. if darkvisitors.com had a sort of plain text API, your repo could check for new entries via GH Actions and create issues or even PRs.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: