Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

>search engines

You can decompose a "search engine" into multiple big components and figure out what you want to look at first:

(1) web crawler/spiders

(2) database cache of web content -- aka building the "search index"

(3) algorithm of scoring/weighing/ranking of pages -- e.g. "PageRank"

(4) query engine -- translating user inputs into returning the most "relevant" pages

Each technical topic is a sub-specialty and can be staffed by dedicated engineers. There are also more topics such as lexical analysis, distributed computing (for all 4 areas), etc.

If you're mainly focused on experimenting with programming another ranking algorithm, you can skip part (1) by leveraging the dataset from Common Crawl: https://index.commoncrawl.org/

Here are some videos about PageRank:

https://www.youtube.com/watch?v=JGQe4kiPnrU , https://www.youtube.com/watch?v=qxEkY8OScYY

... but keep in mind that the scope of those videos omits all of (1), (2), and (4).



I would also require reading on the major OSS search engines, e.g. Lucene, Solr and then Elasticsearch.

https://en.wikipedia.org/wiki/Apache_Lucene

https://en.wikipedia.org/wiki/Apache_Solr

https://en.wikipedia.org/wiki/Elasticsearch

those who dont know history will repeat it, etc

then maybe the newer stuff like https://typesense.org/

to be clear i dont know these things but thats what I would do and I'd happily read fly.io style blogposts drip feeding out knowledge over time


As someone who works in this space, ^ this. I would say don't overthink any component, use common crawl (https://commoncrawl.org/) to build your initial index, use a pagerank implementation that's been thoroughly researched and published, and use off the shelf components from the apache foundation when you can.


abadger9, nice username, do you have any cool portfolios ref the same kinda work?


You are implying web. If you eliminate web from your major components, the concepts are the same for building a search engine for any data. Like a product search on a company website might need to crawl through a bunch of internal data sources to build an index of all the products, metadata, etc. and you may want some products ranked higher in results for whatever reason.

I think a major effort, if its not web, is defining a common data structure to represent the varying source data sets.

Edit to add: OP if you want to study this, start smaller than "search the web" and get the major component concepts down. Expanding to the web will involve distributed systems stuff to scale but there's only like a handful of companies doing that yet there are many more companies that need a search engine on in-house data.


Hi thank you for your answer. Actually my original idea was to build it for something else rather than the web in particular. Can you please advice me on which masters I should pick between distributed systems , machine learning and theoretical computer science/algorithm if my end goal is this?


I don't think any specific MS is going to cover everything for building a search engine. You may want to look at people that work on Lucene [0], for example, and connect to see what their background might be. My guess is there is much more algorithms (types of searches and their big-O tradeoffs, sorts, etc.) involved than dist systems/ml but I don't work on search engines. Its open source so you could just download it and tinker. You may find that a degree isn't necessary to build one.

[0] https://lucene.apache.org/


I would include (5) the presentation of search results - document snippets or surrogates and affordances for query refinement like faceting. One salient acronym is SERP, for search engine results page.

A different "information retrieval" perspective overall would be to take a deep dive into the career of Susan Dumais, http://susandumais.com/


Hi, between my options to go for a masters in Machine Learning, Distributed systems, Theoretical computer science/algos, Which would you pick if your goal was to build a search engine?


distributed systems or cs / algo




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: