You can decompose a "search engine" into multiple big components and figure out what you want to look at first:
(1) web crawler/spiders
(2) database cache of web content -- aka building the "search index"
(3) algorithm of scoring/weighing/ranking of pages -- e.g. "PageRank"
(4) query engine -- translating user inputs into returning the most "relevant" pages
Each technical topic is a sub-specialty and can be staffed by dedicated engineers. There are also more topics such as lexical analysis, distributed computing (for all 4 areas), etc.
If you're mainly focused on experimenting with programming another ranking algorithm, you can skip part (1) by leveraging the dataset from Common Crawl:
https://index.commoncrawl.org/
As someone who works in this space, ^ this. I would say don't overthink any component, use common crawl (https://commoncrawl.org/) to build your initial index, use a pagerank implementation that's been thoroughly researched and published, and use off the shelf components from the apache foundation when you can.
You are implying web. If you eliminate web from your major components, the concepts are the same for building a search engine for any data. Like a product search on a company website might need to crawl through a bunch of internal data sources to build an index of all the products, metadata, etc. and you may want some products ranked higher in results for whatever reason.
I think a major effort, if its not web, is defining a common data structure to represent the varying source data sets.
Edit to add: OP if you want to study this, start smaller than "search the web" and get the major component concepts down. Expanding to the web will involve distributed systems stuff to scale but there's only like a handful of companies doing that yet there are many more companies that need a search engine on in-house data.
Hi thank you for your answer. Actually my original idea was to build it for something else rather than the web in particular. Can you please advice me on which masters I should pick between distributed systems , machine learning and theoretical computer science/algorithm if my end goal is this?
I don't think any specific MS is going to cover everything for building a search engine. You may want to look at people that work on Lucene [0], for example, and connect to see what their background might be. My guess is there is much more algorithms (types of searches and their big-O tradeoffs, sorts, etc.) involved than dist systems/ml but I don't work on search engines. Its open source so you could just download it and tinker. You may find that a degree isn't necessary to build one.
I would include (5) the presentation of search results - document snippets or surrogates and affordances for query refinement like faceting. One salient acronym is SERP, for search engine results page.
A different "information retrieval" perspective overall would be to take a deep dive into the career of Susan Dumais, http://susandumais.com/
Hi, between my options to go for a masters in Machine Learning, Distributed systems, Theoretical computer science/algos, Which would you pick if your goal was to build a search engine?
You can decompose a "search engine" into multiple big components and figure out what you want to look at first:
(1) web crawler/spiders
(2) database cache of web content -- aka building the "search index"
(3) algorithm of scoring/weighing/ranking of pages -- e.g. "PageRank"
(4) query engine -- translating user inputs into returning the most "relevant" pages
Each technical topic is a sub-specialty and can be staffed by dedicated engineers. There are also more topics such as lexical analysis, distributed computing (for all 4 areas), etc.
If you're mainly focused on experimenting with programming another ranking algorithm, you can skip part (1) by leveraging the dataset from Common Crawl: https://index.commoncrawl.org/
Here are some videos about PageRank:
https://www.youtube.com/watch?v=JGQe4kiPnrU , https://www.youtube.com/watch?v=qxEkY8OScYY
... but keep in mind that the scope of those videos omits all of (1), (2), and (4).