If you find value in my newsletter and want to support the work that goes into it consider —> [☕Buy Me a Coffee☕]💡 Want more from me? [SEO Strategy Course] | [Build Your Own SEO Tools (with Python & Agentic AI)] | [Premium SEO Mastermind]
Alternatively you can become a paid subscriber:
Intro
I shared the other week my new project, building a search engine. You can read about part 1 of this journey here. I’ve also started other projects since then 😄 but that’s a chat for another day! Today I’m sharing the next step in building the search engine:
Main Content Extraction
The term “main content” has come to light again recently with Google’s update to their helpful content documentation. You see, your content quality is assessed based on the “main content” of the page. Things like top navigation, footer, sidebars, popups, etc… are not part of your main content, and would not influence your content quality score.
So the next step after crawling and collection the URLs, is extracting the main content. This is part of the indexation process. The first step basically.
Because in indexation, what you want to do is build a table mapping each page and a set of keywords. You’re basically saying for these keywords, this page provides a relevant answer.
The table will be used to find and serve pages when a user searches.
Ya… so let’s dive in!
Main Content Extraction
Using my dear friend Claude Code, I decided to extract text from pages in the following way:
I decided to ignore all non 200 pages. [Only 541 pages returned 200, and those are the ones worth extracting.]
Also, I decided to ignore all pages that require JS for their content, to simplify my process.
I build a different set of extraction rules. One is tailored specifically for my substack pages, one for my website pages (and body did I find space for technical optimizations there 😄), one for YouTube and an open source pre-existing tool for other pages called trafilatur.
Here’s a summary of the extraction methods:
As a general rule, make sure to keep headings and paragraphs, as simple Markdown (# headings, paragraphs, lists), not one flat block of text.
Finally I generated a report to review the extraction process:
looks good right?
Limitations
Few things I had to set aside at this stage to focus on delivering vs perfecting:
I would’ve loved to build a better extraction mechanism for YouTube that pulls the transcript.
The goal is not a perfect search engine, but rather a good enough one. I think I was able to do that at this point.
Next Steps
Clearly indexation is not done yet. I still need to normalize, and tokenize, and decide on the mechanism to assign keywords to pages. I like what I’m doing and that all that matters. Will keep you updated!
And That’s a Wrap (Almost 😄)
Building tools is great, but building a search engine is something else. It’s definitely a passion project. I guess that’s why we do SEO 😄
I hope you found this helpful and inspirational.
That’s that for today folks and see you in the next newsletter!
Support the Riddler!
Sign up for my newsletter if you’re not already. (Pssst, you can also become a paid subscriber)
Share the newsletter and invite your friends to signup. Help me reach 2k signups on Substack by end of 2026 please 🙂
Provide feedback on how I can make this newsletter better!!!
If you’re an SEO tool or an SEO service provider, consider sponsoring my newsletter. I’m also open to other partnership ideas as well.
Disclaimer: LLMs were used to assist in wording and phrasing this blog.




