Forum Moderators: open

Message Too Old, No Replies

referer spam - "free music"

         

lucy24

12:49 am on Sep 24, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



Moderators, I've intentionally left the offender's real name out of the subject and meta slots, since publicity is exactly what they want.

Can anyone shed any light on this referer spam that's been vexing me for the last month or so? Humanoid UA-- assorted versions of Chrome-- with all supporting files including analytics. I'm assuming infected machines. Mainly but not exclusively Brazil.

http://musica.descargar-musica-gratis.net/descargar-musica-gratis.php?u=http://example.com
http://musicas.baixar-musicas-gratis.com/baixar-musicas-gratis.php?u=http://example.com


example.com is my sitenames (three sites to date).

Current lockout:
SetEnvIf Referer musicas?-gratis keep_out

(using mod_setenvif so it will cover all affected domains)

Free lookup says these two are the only sites living on their current server (Worldstream in the Netherlands, if anyone cares), but I remain apprehensive that they will tweak the referer yet again and I'll have to re-tweak the lockout :( Free lookup also strongly suggests that the most recent registration change was made specifically to enable referer spam, since the dates match a little too well.

:: irritably wondering if I should just throw Brazil into the shoot-on-sight bin alongside China ::

wilderness

11:32 am on Sep 24, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



lucy,
looks to me like their just a log spammer. (same as all those RU's).

I've had these for the past two months on two different websites with two different widget topics.

FWIW, on my own widget site (and my narrow market share)?
I've had one valid visitor from either Argentina or Brazil in fifteen years. That visitor was a request-referral from a widget friend and within hours attempted to run a bot on my entire site and was denied access as a result.
I've nothing against Argentina or Brazil, there just not the market for my widgets.

Angonasec

12:32 pm on Sep 24, 2014 (gmt 0)



South America is on China's shopping list, next to Africa, so you might as well get ready Lucy.

I have :)

not2easy

1:59 pm on Sep 24, 2014 (gmt 0)

WebmasterWorld Administrator 10+ Year Member Top Contributors Of The Month



I have had the same experiences, some coming from ES & FR server farms but also from BR and other S. American LACNIC IPs. Referer spam. Not from essential markets, I can't offer a better solution than what you're using. I guess enough hosts still have crawlable logs to make it worthwhile?

engine

2:29 pm on Sep 24, 2014 (gmt 0)

WebmasterWorld Administrator 10+ Year Member Top Contributors Of The Month



I'm with wilderness on this one: Referrer spam.

dstiles

6:21 pm on Sep 24, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



I've been looking at semalt lately and gratis seems to follow the same characteristics in UA and other major headers. Use of a botnet and the countries involved also seem very similar.

I suggest they are either the same "company" or closely connected (DNS is different but that's easy).

The problem with detecting these is that they are targeted via referers. There is a pattern: it's just working out all the variations.

I can easily reject them but I'm trying to prevent them from hitting my web sites at all. I'm trying to reject them using IIS's Request Filtering at the server level (not per site) but it's a total failure. If anyone has any ideas (that do not involve a linux-like .htaccess)...

lucy24

12:31 am on Sep 25, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



I agree they're very semalt-like. One thing they have in common is their humanoid behavior. They never ask for the favicon, but they get all other supporting files including analytics (even when it lives on a different site). I've been working on the assumption that they're infected human machines, so the browser "thinks" it's making an ordinary request. It makes the lockout extra important because it cuts them back from 10-20 requests for an ordinary page, to 3-4 for the 403 page. Normal robots don't ask for anything but the page, so I often just ignore them.

It now occurs to me that I could exclude referer-spammers from analytics using the same code I already use to bar entities like the plainclothes bingbot. "piwik.js" is a pretty large file, so the server would probably prefer to hand over the 403 page again.

Hm. Should I be using absolute links on the 403 page, or would this cause other problems? Humans and humanoids who ask for the wrong hostname and then meet a 403 don't get the www redirect, of course, so instead their requests for stylesheets (referenced by the 403 page) get redirected.

keyplyr

12:42 am on Sep 25, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month




I ignore most log spam; too much bother to block it. What would be the point anyway? It still gets in the logs.

Trouble comes when these hacked/infected accounts are used for other, more invasive purposes. Then they get blocked or redirected. Lately I've been sending them to an E. Euro SPAMMER resource where these accounts are bought/sold. I figure that's the place for them :)

@ lucy24 - I was told years ago by Jim Morgan to never use any internal links on a 403 page. If you think about it, the reasoning becomes obvious.

lucy24

3:54 am on Sep 25, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



What would be the point anyway?

The point is to reduce the number of requests from 10, 20 or more-- however many supporting files a typical page uses-- to 2 or 3-- the supporting files used by your 403 page.

keyplyr

1:05 pm on Sep 25, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



That isn't why I said "what would be the point?"

And my 403 page doesn't use any "supporting files." I keep it at the absolute minimum... about 200 bytes, no css, no files, no links... only a few words saying they've been blocked and why:


Forbidden

Permission to access this server is denied

Possible Reasons:
• You are in violation of copyright.
• You are hiding your browser or user agent.
• You are using a tool or method not allowed.
• Your host/ISP has been banned for bad behavior.


Your IP Address: ##.##.##.## has been logged.

dstiles

6:52 pm on Sep 25, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



I've analysed this month's trap logs for these; how many I've missed I don't know without checking all site logs.

There are only three UAs, varying only in the google version number. Since it looks as if the UAs may be updated when G's browser updates this is not helpful. The significant headers also vary slightly and I've found a handful of HTTP/1.0 accesses (most are 1.1).

I was hoping to pre-block these and similar but (for me) that looks impractical.

In connection with these hits I've blocked referers for

\.php\?u= and crawler
semalt, reevoo, kambasoft, gratis

The middle two are reputed to be next year's domains.

Lucy: "infected human machines"

Yes, although I THINK some of the IPs I've trapped were servers - or at least DSL with more than one set of open ports (possibly a virus'd machine that moved IPs?). Either way, I think it's likely just a hired botnet centred on LACNIC with a few RIPE thrown into the mix.

It would be nice if registrars and hosting services dumped these domains asap but there are some nasty money-grubbing ones out there! :(

Apropos which: http:[//]threatpost[.]com/researchers-work-to-predict-malicious-domains/108531 - can they be predicted?

How can these scammers get analytics logs? I'm sure they can't on my server...

lucy24

6:43 am on Sep 26, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



How can these scammers get analytics logs?

Are they getting analytics logs?

In my case I just meant that the full humanoid request includes piwik.js (a big, fat file, probably comparable in size to what GA uses) and, from there, a final request to piwik.php?something-or-other. I don't think there's even a file at the other end of the request; maybe something like an administrative gif. The request itself carries all the information that piwik uses. Since piwik lives on my own site, all requests-- including me looking at the data-- are shown in ordinary access logs. I haven't seen anything fishy.

I don't think my own logs are accessible. At least, I've never figured out how to get at them in a browser. Logs live in a different physical directory, not in my own userspace, so they may not even be accessible by http at all.

kambasoft, yup, that was #2 in the list that began with semalt. Had to go consult the mod_setenvif docs to see how to block by referer, because I normally do it in mod_rewrite, site-specific. Haven't met reevoo yet.

not2easy

3:04 pm on Sep 26, 2014 (gmt 0)

WebmasterWorld Administrator 10+ Year Member Top Contributors Of The Month



OT - Side note about bots accessing logs - I have sometimes wondered how many "spammy backlinks that I did not create" seen over in the Google SEO forum are due to badly configured hosts that have access logs in the userspace that allow bots to read all that referer spam and hence multiply it? I know they still exist because they show in the serps..

wilderness

4:30 pm on Sep 26, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



Don't recall which host it was (was only with them for a short while), however they claimed it was not possible to have log access via FTP because that would allow open access.

I've never had any such explanation from hosts on both side of the aforementioned host. in fact in most hosts the log directories are in place above the html directory.

dstiles

6:19 pm on Sep 26, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



Lucy - Ah. My misunderstanding. :)

wilderness - "log directories are in place above the html directory"

The default for windows IIS is everything in a sub-directory-per-site of a common, predefined directory. Having just set up several dozen sites on a new server that was a pain, because I kept forgetting to change the log location for each site. :(

lucy24

7:38 am on Sep 27, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



they claimed it was not possible to have log access via FTP because that would allow open access

Mine changed a couple of years ago. You used to be able to ftp into logs; now you have to sftp (or assorted other protocols that I don't deal with, including one involving Terminal, shudder, that I had to use briefly until I got Fetch to play nice).

My userspace has one directory for each domain, one more for mail (I do not pretend to understand this, since each domain has its own email), and one for logs. But the logs aren't really inside my userspace; it's aliased from somewhere else. And then there's a further nest of aliases once you get in. I was once plagued by a robot asking for nonexistent example.com/logs/something-or-other files, and it was utterly immune to my own lockouts. I finally had to beseech to host to apply a separate block.

I have to say I really like the userspace structure, as opposed to primary/addon, because it means no one site's requests ever have to pass through another site's htaccess. And I can have a top-level htaccess that covers everything, like IP-based lockouts that are the same for everyone.

dstiles

6:43 pm on Sep 29, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



Further to my mention of blocking...
\.php\?u=

There has to be a whitelisting for facebook-originating "real" visitors along the lines of...

\.facebook\.com/lsr\.php\?u=

Obviously, both examples above are formulated for regex tests.

Since April l.facebook and lm.facebook have apparently become common referers.

dstiles

8:15 pm on Oct 5, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



I'm looking at this in more depth. I've set a trap for a specific condition and recorded all who enter thereunto. Interesting. There seems to be a LOT of possible referer domains.

One that is prominent is aosheng-tech.com. I cannot conceive that an Asian site that manufactures springs would ever have legitimate links to my sites, especially the med sites. Looking online for that domain I came upon the site below, which has a few more domains.

http[://]clicky[.]com/forums/?id=16360

I have, I think, almost completed an algorithm to trap these but I'm loathe to reveal too many details, as you may imagine. I'll report back in a few days, when I've completed-ish the current tests.

lucy24

7:33 am on Oct 7, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



Sometimes I get the impression the mere existence of a referer is the point. "Can I look around? John sent me." "Yeah, sure, go on in." Only later do you realize that you don't know anyone named John. It's one step over from auto-referers, which are too easily blocked.

It now occurs to me that maybe the reason this type of botnet doesn't ask for the favicon-- even when they're otherwise humanoid-- is that they're set up to request only supporting files that are explicitly named by the html page (like css and images). Most sites don't specify the favicon that way.

dstiles

6:58 pm on Oct 7, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



In the case of my current investigation I'm finding the referer is often "my" domain without any page (I already block http-less referers). In some circustances, depending on other headers, I sometimes block on empty referers.

Since I always specify a non-root favicon in the header of every page I can't use that one. :(

I was asked the other day if I couldn't just send back something that would blow up the offending computer... A nice thought! :)

lucy24

7:39 pm on Oct 7, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



You could code a popup that says YOUR MACHINE IS INFECTED in huge red letters.

dstiles

8:05 pm on Oct 8, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



That would only work with "real" browsers. And even then only if they were using javascript. :(

I implemented the algorithm today. Now to see how many real people get injured.

lucy24

7:31 am on Oct 9, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



And even then only if they were using javascript.

Remember we talked about analytics earlier? These particular spammers don't only request piwik.js (referenced on the html page). They follow it up by requesting piwik.php in its <script> form, recognizable by its several miles of query string. So yup, they're acting on javascript.

dstiles

6:57 pm on Oct 9, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



But that assumes you have javascript on your pages. On most of my sites JS is only served on forms pages. :)

I used to use JS for all sorts of minor embellishments, until I discovered NoScript and began visiting sites with JS turned off. A reappraisal of my sites indicated JS was seldom necessary so I stopped including it on most pages.

I take your point, though: it would be easy enough to include it for "specials".

dstiles

7:53 pm on Oct 12, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



There are a lot of referer domains being used in this project. I have blocked four common ones in the IIS server itself - semalt, gratis-types, aosheng-tech and simplysoaps - but new ones keep popping up in the log generated by my test algorithm.

All I am prepared to say in public is: look at a common header that is not all it should be. In conjunction with that there should be a series of qualifications / exemptions as to UA types etc. Or you could simply block all chrome UAs (I wish!). :)

So far, in the past three or four days, I have, I think, trapped at least most of the hits on my couple of dozen sites using my derived algorithm. There was only one seemingly false trap which, on investigation, turned out to be an Indian server farm - or at least a range with several open-port IPs: the referer was a bare google india. I have about 24 trapped IPs since 9th October (and several unlisted in my db before that). They are fairly evenly split in origin between server and dsl ranges and I suspect, as stated above, that they belong to a botnet. OR - and this is only speculation - the bots were loaded onto computers through dodgy marketing of some kind ("Install our toolbar!" kind of thing).

I logged about 500 unique IPs (after removing duplicates) at the start of this month, several of which turned out to be baidu CN and similarly unwanted bots. I ended up with about 150 samples of which about a third were semalt types, the others providing me with whitelisting parameters or being already blocked as server farms. IP ranges were from many parts of the world including (mostly) lacnic and ripe.

It is possible that some of the domains used in this project are genuine ones trying for referer-log recognition, probably conned into this method by (probably) the people behind the semalt scam. It's only a possibility, though, based on a couple of referer domains with long registration dates. They could have been just pirated. Either way, the .COM registration system is so lax and even, in some cases, corrupt that it's probably easier to just create new domains with one-year dates.

If anyone would like a copy of my criteria then let me know via the PM system here.

not2easy

8:36 pm on Oct 12, 2014 (gmt 0)

WebmasterWorld Administrator 10+ Year Member Top Contributors Of The Month



I'll vote for: "the bots were loaded onto computers through dodgy marketing of some kind" because when I first started trying to figure out what it was, that is all that I found in the serps for their name - How do I uninstall "blahblah"?

lucy24

5:29 am on Oct 13, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



but new ones keep popping up

A few days ago I found a fake referer that could not possibly have been intended as referer spam. How can I be so sure? Well, I know that some people get strongly attached to a favorite dentist ... but not even my father would go to Norway for the sole purpose of getting his teeth cleaned. So they're just pulling names out of a hat in order to put something-- anything!-- in the referer slot.

:: wandering off to examine headers, which I've been logging diligently for a few years and could just about count on the fingers of one hand the number of times I've actually used the logs for any purpose ::

Fr0mCha0s

11:14 am on Oct 13, 2014 (gmt 0)

10+ Year Member



We have been getting slammed by this garbage. All of these domains you have posted are the same people. semalt! mostly coming from brazillian ISP ranges.

Whats worse is now they are actually resorting to simply using random websites as the referrer! That means they could actually be using your domain name as the referrer in some requests around the globe!

I am seeing it with my own eyes in logs. These bozos are pissing off the entire internet.

I have been thinking of redirecting them to FBI or something else drastic. LOL

[edited by: Ocean10000 at 5:42 pm (utc) on Oct 13, 2014]
[edit reason] Removed URL [/edit]

dstiles

4:15 pm on Oct 13, 2014 (gmt 0)

WebmasterWorld Senior Member 10+ Year Member Top Contributors Of The Month



I do not think they are using real, working sites in the referers. This would be counter-productive unless the spammers had a trojan installed on the web site. It is much easier to set up a suitable web site on a specially-bought domain name.

I think the primary reason here is to entice site owners to click on the referer when it appears in their analyzed stats. The site would then serve up some kind of exploit or gather your browsing data. Again, this is speculation but I can see no legitimate or non-intrusive use for this project.

The latest domain I've stuffed into the IIS server blocking gizmo is website-errors-scanner[.]com. This domain was registered for one year on the 2nd of this month and has appeared three times in my security logs in the past 24 hours. All others over the past few days have been single hits, which means there is no point in blocking them specifically.

It is impractical to block all of these domains but the more frequent visitors should, I think, be terminated with extreme prejudice at the earliest point of entry - firewalls would be favourite if that's possible.

Fr0mCha0s

11:50 pm on Oct 17, 2014 (gmt 0)

10+ Year Member



I do not think they are using real, working sites in the referers. This would be counter-productive unless the spammers had a trojan installed on the web site. It is much easier to set up a suitable web site on a specially-bought domain name.


Could be, but I have looked at many of these random websites. They are real people, like active housewives blogs(these I could see being easily infected lol) and some other random businesses. One of the sites you mentioned here I have seen too.
This 38 message thread spans 2 pages: 38