Login Register






Tutorial Focused Crawler Part 1 - Basics filter_list
Author
Message
Focused Crawler Part 1 - Basics #1
Okay, I always wanted to write a tutorial about crawlers and recently found this forum, so may I just contribute a little? If there is a single person who appreciate my views on focused crawlers, I am happy (also, I am not an experienced writer, so feel free to ask if something is not clear, not on point, engrish...etc.)

Focused Crawler Part 1 - Basics

First I will just give a quick overview about what I am planning to talk about, so nobody wastes time. Also this won’t be a full on beginner tutorial, I don’t want to write novels about how to use basic tools (there is so many tutorials out there so I will just mention them and anyone can probably find something for it easily. Of course, if someone asks I can help them).

The tutorial is about focused crawlers / web scrapers, which is different from a regular crawler (easier at some point). A regular crawler scans the web and tries to index as many content as it can find, for example a search engine’s crawler. A scraper tries to get certain data from a small group of targets, for example you try to build a database about recent gaming news or you are building a webshop where you sell t-shirts and you would like to know other webshop’s prices regularly, so you scrape them every now and then.

What I would like to talk about here  (part 1 at least):
• Required tools
• Main goal of a scraper
• Anonymity of the crawler and not DDoSing the target
• Error / exception handling
• Resumable process
• Not scraping the same content twice
• Bruteforcing

Required tools
Mainly you have to choose a platform you would like to write your crawler on. I personally wrote crawlers for C#, Java and JavaScript (NodeJS/CasperJS). These should be perfectly fine for the task. I am not sure about C/C++, because you have to rely heavily on HTML Parsers and text tools (don’t even think about regex) and it is hard to find an easy to use one (you have to write a lot’s of selectors and be flexible, most site generates incomplete html, not-so-html, escape xml’s inside json and similar), but I am not a c/c++ developer, maybe someone who has experience about the topic could write some tips. Another language (which I don’t use, but I would like to learn at some point) is Python, hands down perfect for the job.
Links for the tools I recommend:
.NET/C#
HTML Parser: Html Agility Pack http://htmlagilitypack.codeplex.com
Client to grab content: System.Net.WebClient / System.Net.WebRequest

Java
HTML Parser: Jsoup https://jsoup.org/
Client to grab content: JSoup can do it well, or https://hc.apache.org/
I wrote my own framework, but there are highly known crawler frameworks for Java: https://github.com/yasserg/crawler4j ; http://nutch.apache.org/ (Nutch is BigData though)

JavaScript
HTML Parser: You don’t need one. (You could even use jquery)
But you need to work within NodeJS https://nodejs.org/en/ http://casperjs.org/

Python
A Compelte scraper framework: https://scrapy.org/
HTML Parser: https://www.crummy.com/software/BeautifulSoup/

Common tools
Something that would help you figure out which request gives you the content. Obviously a browser with developer tools (Firefox/Chrome) and a Http request tester, I prefer Postman (you can send a request with specific parameters and get the result as text or RENDERED <-- I think this is a big thing, because you actually can read the content better when it is rendered Biggrin):
https://www.getpostman.com/
I also use curl , so If you use windows you just install git bash and you have it https://git-scm.com/downloads (also ssh).

Main goal of a scraper
This is just to make it clear how I think a scraper should operate and what the result should be.
You basically have a start url or urls (which is called as the seed in some articles), you request the content and if it is needed, you make a next request. You usually decide based on the currently requested document, for example, if you have a site which list the products with pagination, you check at every request if there is a next link on the page, if there is, you request the next page.
The pages you get from the request are a big text documents, you could store them if you want, but if you want to run queries against them (like average prices) you should get the data out of the text and store it in a database (RDBMS or NoSQL).
So the scenario would be:
Figure out parameters -> run http requests -> parse html -> insert the data to database

Anonymity of the crawler and not DDoSing the target
This is the part which separates good and bad scrapers. Can you download a large chunks of data from a target without being seen? Can you even download it at all? A properly setup and administered website will block you if you make too many request within a short period of time (or better, give you captchas.. amazing, now you need some kind of service which figures out the captchas). On the other hand, a badly written site (with a low performance hosting) will just break if you hit it too hard, which is not good for anyone, you won’t get the data and they can notice something is wrong. Maybe they change the website in some way which is more work for you.
You have to balance out the sent requests, the speed of how fast you want the data and how much money you would like to spend on the project. The basic is a cloud VPS, but If you need millions of request, you have to note, that with a VPS you get only one IP address and if the enemy (Smile) blocks it you are screwed, also they can trace you. I think you have the following options:
• Buy more vps and do slow craw (sounds expensive)
• Build a botnet (sounds easy)
• Pay money for a service to shady companies that gives you IP addresses (crawlera/scrapinghub for example, but I don’t trust them, neither should you).
• Use Tor as a middleman which can route you requests to a different node after a certain threshold. Note that with Tor you still can’t do all the requests you could send (common sense though)!

The problem is that IPv4 addresses are expensive, nobody going to give you tons of them for a low price! My personal choice is doing it as slow as google would (Google’s official limit is 5 request per second if I am remember it correctly), which is still a little too fast sometimes. I usually hop on a test VPS and run test requests to the index page as fast as I could and check if I got an error or not. I measure time, the number of requests, and save the error pages. (Important: Well behaving sites use the HTTP Retry-After header if they block you, quite useful, but all of them gives you different measurements but here is the standard: https://developer.mozilla.org/en-US/docs...etry-After )
[Image: hnhrxl.png]
Using Tor:
1. Install Tor ( https://www.torproject.org/docs/debian.html.en )
2. Check if the service is running: sudo service –status-all
3. Note that the default proxy address of tor is 127.0.0.1:9150 or 127.0.0.1:9151 which you can easily check by taking a look at /etc/tor/torrc
4. Tunnel your request through the Tor proxy. Example Jsoup:

Code:
Jsoup.connect("https://www.target.com/sadpanda.html)                .proxy(“127.0.0.1”,9151)                .userAgent(UserAgentStrings.UA_GOOGLEBOT_1)                .validateTLSCertificates(false)                .maxBodySize(0)                .get(); // Reason why I like Jsoup, easy isn’t it ?
Maybe you noticed the UserAgent, which sets the request’s browser identifier (it is a simple http header), never forget to set it to something, I usually set it to a recent firefox.
If you use Windows, you can just install Tor Browser and launch it, it will create the same proxy for you and you can access it with 127.0.0.1:9150
And please don’t forget that you need to limit the number of request, otherwise you will run out of IPs at some point.
Another tip, always check the robots.txt! No crawler should ever stumble upon urls disallowed there, so if you request pages from there and they can detect that you are not a real user, they can block you more easily.

Error / exception handling
Getting the result as a string is one thing, parsing is another. It is certainly possible, that you simply cannot check every possibility of how the page is generated for you. Maybe somewhere along the products there is an extra <table> tag which contains the information and not the <div> you expected it to be in.  You have to define clear rules for your data. What is the absolute necessary for the data to be accepted? If your selector couldn’t return the products name, it should throw an error, shouldn’t it? But if the missing part of the data is not important, it we could still accept the remaining data.  Also, the program shouldn’t stop just because one request or parsing failed. Log the error, save the full HTML which gave you the error (so you can check the structure later) and go forward (if you can).

Resumable process
The state of your crawler shouldn’t be stored only in memory, you can’t recover from an error without storing where you are currently and knowing what is the next job in the queue (I mention queue, but you don’t necessarily need a literal fifo queue implementation). Also you have to know beforehand if you want multi-threaded / multi-client crawler or a single threaded one.
What I do is define specific type of job names (usually a job is a single request, for example: a job would get a list from the index page another would request all the list elements and parse the result). Define parameters for these jobs (like a function, they take something, and give a result). I can store the job type and the parameters for the jobs in a database easily and can also store a state for it (Important: Implementing a Queue in a RDBMS is an anti-pattern or hard to do. Quick google http://mikehadlow.blogspot.hu/2012/04/da...ttern.html ; https://blog.jooq.org/2014/09/26/using-y...otally-ok/ ).
I do the database saving in an async way, doing everything inside in-memory stores and polling buffers to write to the database. I don’t really care if a few job is lost after a crash, but mostly I would like to be able to recover. (I don’t know how much I should write about this topic, because it is quite big, maybe I should write a more detailed article about this separately).

Not scraping the same content twice
https://en.wikipedia.org/wiki/Spider_trap
Other than traps, nobody wants to do the same work twice or more. Identifying if you are just about to request the content you already did at some point is crucial.
One way is to check the URLs, if you already did a request to an url you can just skip it. End of story? Not really. First, URLs are large and there are many of them (takes lots of space). You have to check if your current one is in the database, but look at these:
http://target.com?p=14874144&sort=username
http://target.com?sort=username&p=14874144
http://target.com/14874144?sort=username
First, you have to optimize your urls, than HASH them with md5 or sha1 and store it somewhere, but it is not enough. You should also cache the results of the check you run, because looking up the hashes takes a lot of time.
What about data? Do you want duplicated data in your database or not? Depends on what kind of data we are talking about, but I usually don’t want any dupes. If it is possible I define certain fields for the data that should be unique, for example the product’s ID or name. Then when I save the data, I can say that if it is already exists, skip it (now, you mostly filter out duplicates by optimizing urls but it can still happen, if you run your program multiple times and store the data in the same database without versioning, which can be totally okay, if you do what I mentioned).    
I use PostgreSQL to store data and an example data table could look like this:
Code:
CREATE TABLE public.pdataexample (    id bigserial NOT NULL,    pid character varying(255) NOT NULL,    data jsonb NOT NULL,    PRIMARY KEY (id),    CONSTRAINT u_pid UNIQUE (pid) )
I only put fields as columns that are absolutely necessary, I store everything else as Jsonb.
Now I can insert data without worrying:
Code:
INSERT INTO public.pdataexample(pid, data) VALUES ('SO-471', '{"dataf1":123245}') ON CONFLICT DO NOTHING;
This skips the current data if it already exists, which could be a problem, but I mentioned versioning, depends on the data, if you could compare them after different runs or delete duplicates later based on user choice.
Note: Why I don’t use MongoDB or other NoSQL? I don’t need them, but you could totally use them. This is the kind of data that totally okay to store in nosql (quick to change).

Bruteforcing
You can either be lucky and could gather every link you need to visit from the source of the target easily or be forced to try to guess. Example for the first is when you get a categories on the website, you can grab them, visit them, visit all the pages in the pagination.
What if all you get is a search box? Type in something and you get a result list maximized to a specific number (like maximum 20 result). You are forced to try every combination of A-Z0-9 (there are escape routes, if you are lucky. If the search allows ‘%’ or ‘*’ to give all the result or displays a pagination or tells you how many entries are there overall).
I had to run 36^4 and 10^6 number of request for a few sites to get all the results because all it outputs is a fixed number of result. I used a simple bruteforcing algorithm:
Let’s say we have an alphabet: 0-9A-Z
Start with length 1, ( ‘0’), run the search, if the returned number of rows is the maximum number of rows the site returns, proceed to length 2 (‘00’), if it is less, proceed with the next character in our alphabet (‘1’). So if ‘00’ returns 19 and our max is 20, our next query is ‘01’ and so on (and if the query ‘01’ returns 20 results, we go with ‘010’ as our next search). You need to walk back characters when you are at the end of the alphabet. Like ‘0AZ’ when the result is smaller than max should proceed to ‘0B’.
Choosing the right alphabet is also important, if the site searches in specific product numbers, you could potentially lower the number of request needed. (What if all the product numbers contain a number? You can use only numbers).

Plans
I would like to start with a simple example of how to code a scraper, from planning to concrete implementation. Than I should delve into some of the problems I mention here. Feedback helps. Also I am planning to use Java as an example, but if everyone hates it I can show you C# too.

Reply

RE: Focused Crawler Part 1 - Basics #2
This is an awesome guide assuming you didn't plagiarize it. You should include detection evasion methods.

Reply

RE: Focused Crawler Part 1 - Basics #3
I've never worked with web crawlers so that was an interesting read.
But I have delt with scraping, I usualy write scrapers in python and use (TOR -> proxy), because the site can detect and bock you if you make loads of requests form a single IP (just as you mentioned).
I just find a list of proxies and circle trough them one by one so each proxie has a cooldown and the site thinks that they are normal users.
I have one question, is there a reason to use PostgreSQL insted of MySQL?

Reply

RE: Focused Crawler Part 1 - Basics #4
Nice work, thank you for sharing.

Reply

RE: Focused Crawler Part 1 - Basics #5
(01-16-2017, 07:20 AM)Pikami Wrote: I have one question, is there a reason to use PostgreSQL insted of MySQL?

PostgreSQL has full support for JSON. Literally it is almost as good as a NoSQL database. Info here: https://www.postgresql.org/docs/9.6/stat...-json.html
You can write quires that select data from json and from normal columns  in the same query. Using the example I posted:
Code:
SELECT id, pid, data->>'dataf1' as dataf1clmn FROM public.pdataexample;
Would give you:
Code:
id pid dataf1clmn 1 SO-471 123245

You can use the fields in the WHERE statement, check if exists or have a certain values. With MySQL all you can do is use it as a string in a LIKE query. Also, with PostgreSQL you can define powerful CHECK constraints, further protecting your data without a single line of program code (like defining the pid 's length has to be more than 4) https://www.postgresql.org/docs/9.6/stat...aints.html

Also, you define INDEXES for fields inside json! http://clarkdave.net/2013/06/what-can-yo...-and-json/
And you can use PostgREST for instantly provide a REST API for your data: http://postgrest.com/api/reading/
(This post was last modified: 01-16-2017, 10:48 AM by d763. Edit Reason: extra )

Reply

RE: Focused Crawler Part 1 - Basics #6
This is actually a very nice formatted tutorial. I can actually use this. I have made (a sort of) crawler that crawls a website for links and checks if there are links that you can use for SQLi. This made me think I have some recoding to do. Looking forward to part 2.
~~ Might be back? ~~

Reply