Search Engines

Search Engines

This transcript is a countdown-style video presenting eight search engines or tools that surface what Google deliberately or structurally obscures. Below is a summary of each entry, followed by consolidated lists and citations.

Click to watch - you will find something useful - list and summary below

The Eight Search Engines

#8 — DuckDuckGo ("The Anti-Bubble")

Google personalises results based on location, search history, and clicks — two people searching the same term can receive entirely different results. Writer Eli Pariser named this phenomenon the "filter bubble" in 2011. DuckDuckGo's distinguishing feature is that it does not build a user profile or track searches; everyone searching the same term gets roughly the same results. It is less mind-reading than Google, but that is deliberate. The key takeaway: a search engine simply not personalising results now counts as noteworthy, which reveals how deeply personalisation has become the default.

#7 — Marginalia Search ("The Small Web")

Built largely single-handedly by Swedish software engineer Victor Lövgren, who started it during the pandemic. Marginalia prioritises text-heavy, non-commercial, ad-free websites — what its own crawler calls the "small, old, and weird web". It surfaces content that Google buries under listicles and AI-generated content. Its results are described as "the kind of grass-fed, free-range HTML your grandma used to write." It is slower, quirkier, and not for booking flights, but it finds the genuinely interesting corners of the internet that stopped ranking.

#6 — Million Short ("The Spite Engine")

An experimental search tool from Toronto whose founding idea is to deliberately remove the most popular sites from results. Users can strip out the top 100, 1,000, 10,000, or up to the top 1 million most popular websites on Earth. By removing Wikipedia, Amazon, Reddit, and SEO-dominated sites, it surfaces what sits underneath — the "weird sediment" on page 40. It is a favourite of researchers and open-source investigators. Not an everyday driver, but it answers the question: "What's under the couch cushions of the internet?"

#5 — BASE ("The Locked Library")

The Bielefeld Academic Search Engine, run by a university library in Germany. Instead of crawling the commercial web, it harvests directly from thousands of academic repositories and institutional archives. It has indexed hundreds of millions of scholarly documents — papers, theses, reports, and conference proceedings. Critically, roughly 60% of what it indexes is open access, meaning full texts are free and legal to read. It lets users pull the actual primary source rather than relying on blog posts that summarise (and often misrepresent) the findings.

#4 — Wayback Machine ("The Undelete Button")

Run by the Internet Archive, a nonprofit that has been saving snapshots of web pages since the late 1990s — hundreds of billions of them. It freezes what each page looked like on a specific day, forever. Deleted content usually still exists in the archive, timestamped and waiting. You can view old company websites, track article edits across multiple saved versions, or resurrect entire dead websites. Journalists and researchers treat it as essential because "the receipts basically never disappear." Deleting something can paradoxically make it more permanent, because the deletion is what prompts people to check the archive.

#3 — The Right to Be Forgotten ("The Forgotten Files")

A legal right stemming from a 2014 ruling by the EU's top court allowing people to force search engines to remove certain results associated with their name. However, the transcript emphasises that "forget" is misleading: Google does not delete anything — it cannot, as it does not own the underlying webpages. It simply stops linking to the page when someone searches a person's name, and only on its European versions. A 2019 ruling confirmed this does not have to apply worldwide. The original article remains on its website, visible on other search engines and to anyone outside the EU. It is described not as deletion but as "a voluntary, geographically limited, entirely reversible blindfold."

#2 — Yandex ("The Face Finder")

Russia's biggest search engine. Google deliberately throttles facial matching for privacy reasons, showing a "results for people are limited" message. Yandex made the opposite choice. Its reverse image search is described as "unnervingly good" at taking one photo of a face and finding other photos of that same person across the internet. Investigators and open-source researchers have relied on it for years. The transcript frames this as a stark contrast: the only thing standing between a stranger's photo and their identity is a corporate policy that one company enforces and another skips. Useful for catching catfish and verifying photos, but a reminder that "anonymity in public was never guaranteed by physics — it was guaranteed by a guardrail."

#1 — Shodan ("The Internet's Backdoor")

Built by John Matherly around 2009. Instead of crawling websites, Shodan crawls devices — servers, routers, webcams, printers, industrial control panels, building systems, and infrastructure quietly connected to the open internet. In 2013, CNN called it "the scariest search engine on the internet." The horror is not that Shodan is a hacking weapon, but that vast amounts of infrastructure are left exposed with little or no security. Shodan is overwhelmingly used for good — security researchers use it as a floodlight to find vulnerable systems and warn their owners. But the floodlight shines for everyone.


Consolidated Lists

All Eight Tools/Engines

Rank Name Nickname Origin / Key Person What It Does
8 DuckDuckGo The Anti-Bubble Non-tracking search; pops the filter bubble
7 Marginalia Search The Small Web Victor Lövgren (Sweden) Prioritises text-heavy, non-commercial, ad-free sites
6 Million Short The Spite Engine Toronto Removes top 100 to 1,000,000 most popular sites from results
5 BASE The Locked Library Bielefeld University Library, Germany Indexes hundreds of millions of scholarly documents; ~60% open access
4 Wayback Machine The Undelete Button Internet Archive (nonprofit) Archived snapshots of the web since the late 1990s
3 Right to Be Forgotten The Forgotten Files EU Court ruling (2014, 2019) Forces Google to delist name-based results in EU only
2 Yandex The Face Finder Russia Powerful reverse image / facial recognition search
1 Shodan The Internet's Backdoor John Matherly (c. 2009) Searches internet-connected devices and infrastructure

Key Facts and Figures Mentioned

  • Eli Pariser coined the term "filter bubble" in 2011
  • Marginalia Search started by Victor Lövgren during the pandemic
  • Million Short allows removal of up to the top 1 million most popular websites
  • BASE has indexed hundreds of millions of scholarly documents; ~60% are open access
  • Wayback Machine run by the Internet Archive, snapping the web since the late 1990s; hundreds of billions of snapshots
  • Right to Be Forgotten: 2014 EU court ruling; 2019 ruling confirming no worldwide obligation
  • Yandex: Russia's biggest search engine
  • Shodan: built by John Matherly c. 2009; CNN nicknamed it "the scariest search engine on the internet" in 2013

Citations (Named References in the Transcript)

  • Eli Pariser — named the "filter bubble" in 2011
  • Victor Lövgren — Swedish software engineer; sole creator of Marginalia Search
  • Million Short — experimental search tool based in Toronto
  • BASE (Bielefeld Academic Search Engine) — run by a university library in Bielefeld, Germany
  • Internet Archive / Wayback Machine — nonprofit archiving the web since the late 1990s
  • EU's top court — 2014 ruling establishing the right to be forgotten; 2019 ruling limiting it to European versions of search engines
  • Yandex — Russia's biggest search engine
  • John Matherly — creator of Shodan (c. 2009)
  • CNN — gave Shodan the nickname "the scariest search engine on the internet" (2013)

Overall Significance

The transcript's central thesis is that Google's dominance has created structural blind spots — not through conspiracy but through defaults: personalisation, commercial optimisation, and deliberate policy choices. Each tool on this list fills a gap Google leaves by design or neglect. The closing message reframes the entire countdown: the unsettling truth is not that clever search engines can find hidden corners of the internet, but that so much of the world was quietly connected to the internet by people who assumed no one would ever bother to look — and every day, something proves them wrong.