Gateway Search

A fast, private, multi-source search engine made from scratch.

Open Gateway View on GitHub

A search engine built from scratch

Gateway is a fast, private, multi-source search engine by XPDevs. It searches a bundled local web index of 50,000 real websites across 50,000 domains, plus live Wikipedia and Wikidata open data, with no API keys, no build step, and no generative AI.

Live behaviour: main results paint from the local index first (typically milliseconds), while Quick Wiki, spell-check and web enrichment resolve asynchronously afterwards.

Your searches are recorded to a local-only search history you can replay, filter, switch off or delete. It never leaves the browser and never affects ranking.

50,000 indexed websites No API keys No build step No generative AI 25-language UI Dark mode

Screenshots

The screenshots below are taken directly from the XPDevs/gateway repository.

Homepage

Gateway Search homepage
Clean, minimal interface with instant local search.

Search results

Gateway Search results for Costa Coffee
Results for “Costa Coffee”, featuring the Quick Wiki box, result stats and favicons.

Features

Instant main results

Two-stage search: Stage 1 renders gatewaySearchLocal() (local IndexedDB/memory index only, no network) inside requestAnimationFrame; Stage 2 enriches with gatewayCrawl() without replacing the visible list. Result stats show measured time, e.g. 120 ranked results (0.03 seconds). Uses content-visibility: auto, 20-per-page pagination, memory + localStorage + IndexedDB caches and request dedup.

Quick Wiki

Deterministic answers from Wikipedia extracts plus Wikidata claims (P577, P571, P36, P1082…) with date-intent parsing and never generative AI. It renders after the main list behind a skeleton placeholder, and failures show no box at all without touching results. Can be disabled entirely via Settings.

XPDevs profile

Searching XPDevs, or any query containing the token (xpdevs, xp dev, xp-dev, xpdev, gateway search), is answered entirely from a local deterministic profile with no network call. It pins the official site, the GitHub organisation, Gateway and all 59 public repositories above every ranked result.

Favicons on every URL

Each result row and each Quick Wiki source link shows a favicon: Google S2 primary, DuckDuckGo icons.duckduckgo.com fallback, letter-avatar final fallback, so no row is icon-less. Images load lazily and asynchronously.

Unique descriptions

All 50,000 index.json descriptions are unique (verified 50,000 / 50,000), as are all 50,000 URLs. Every entry is a real website, one per domain, sourced from Wikidata and confirmed to answer over HTTP at build time.

Local search history

Every search is recorded to localStorage under gw-history as {q, t, n, ms}, newest first, capped at 200 entries. Grouped by day, filterable, replayable, deletable and exportable as JSON. Nothing is uploaded, never attached to a search, and never used for ranking.

BM25 hybrid ranking

IDF with additive smoothing, IDF^1.5 term weighting, Okapi BM25 and a static authority prior. See Ranking below for the full formula.

Plus

25-language UI, dark mode, autocomplete (getSuggestions), Did you mean? spell-check, I'm Feeling Lucky, uploadable custom index, safe-domain filtering and tracking-param stripping.

Project structure

gateway-main/
├── index.html       # Home + results UI, settings modal
├── ui.css           # Theme, results, Quick Wiki, favicon, responsive styles
├── main.js          # UI, staged search orchestration, rendering
├── crawl.js         # Index engine, Wikipedia/Wikidata fetchers, ranking, caches
├── index/
│   ├── index.json       # Bundled local index: [{t, u, d, s}] - 50,000 unique-description entries
│   └── index-meta.json  # Index version + entry count (cache-busting)
├── about/ advertising/ business/ how-search-works/ history/ privacy/ terms/  # Static pages
├── favicon.ico
└── 404.html

Setup

No build, no dependencies. Any static server works:

cd gateway-main
python3 -m http.server 8000
# open http://localhost:8000/

Or npx serve ., GitHub Pages, Nginx, etc. Use http(s)://, not file:// (fetch and IndexedDB require it).

Custom index (optional): Settings → Upload Index → select a JSON array of {t, u, d, s}. It is validated, persisted to IndexedDB and versioned, so old caches are invalidated. Settings (gear icon) also cover language, dark mode, Quick Wiki on/off (persisted as gw-quickwiki, default on) and index upload.

Architecture

User query
  ├─ Stage 1 (sync-fast): gatewaySearchLocal() → prepared normalised index → score → _boostXpdevs → paint (rAF) + timing
  ├─ Stage 2a (async): gatewayCrawl() → local + Wikipedia search + official URLs + extlinks → merge new URLs only, re-sort, update count
  ├─ Stage 2b (async, lazy): gatewayQuickWiki() → parse intent → wiki search → best-page match → Wikidata claims → image → renderQuickWiki()
  └─ Stage 2c (async): gatewaySpellCheck() → renderDidYouMean()

Key modules

Ranking

Every result (bundled index, live Wikipedia, official sites, cited web links) is scored with one hybrid function, so all sources compete on identical terms.

FinalScore = (BM25_Score * 0.8) + (Authority_Score * 0.2)
  1. IDF with additive smoothing: IDF(q_i) = ln(1 + (N - n(q_i) + 0.5) / (n(q_i) + 0.5)), where N = 20,000,000,000 (web-corpus scale, RANK_CORPUS_SIZE; override with window.GATEWAY_RANK_CORPUS) and n(q_i) is document frequency, counted at index load.
  2. Dynamic query term weighting: W(q_i) = IDF(q_i) ^ 1.5, so rare entities are boosted and common terms are damped.
  3. Okapi BM25: BM25(D, Q) = Σ W(q_i) · (f(q_i,D) · (k1 + 1)) / (f(q_i,D) + k1 · (1 - b + b · (|D| / avgdl))), with k1 = 1.2 and b = 0.75. Documents are title + description; |D| is the word count; avgdl is the mean across the whole index (12.73 words for the bundled set: 548,586 tokens over 50,000 documents, 79,765 distinct terms).
  4. Authority: a pre-calculated static float in [0.0, 10.0], PageRank-style, resolved per domain: exact table → parent-domain walk (en.wikipedia.org → wikipedia.org) → public-sector/academic host shapes (.gov, gov.uk, .edu, ac.uk, …) → deterministic 2.0–6.0 hash band for everything else. It never varies with the query, and an index entry may ship its own value in a, which always wins.

Since authority contributes at most 2.0 points, relevance dominates and authority only breaks ties, which is the intended 80/20 split.

Tokenisation

Query terms are stopword-filtered; single digits are kept (PlayStation 5, World War 2) while lone letters are dropped. Document text keeps every token, so stopwords still count toward f and |D|. Matching is whole-token, not substring.

Verification

IDF, W, BM25 and FinalScore were re-derived from the raw index for 150+ results across 6 queries with a maximum drift of 0.0. Corpus statistics (avgdl = 12.7286, vocabulary = 79,765) also match an independent count.

Performance notes

Data notes

index/index.json format: t title, u URL (tracking params stripped), d unique description, s normalised domain, and optional a pre-calculated authority (0.0–10.0, falling back to the domain table when absent).

Regeneration check:

python3 -c "import json; d=json.load(open('index/index.json')); print(len(d), len(set(x['d'] for x in d)))"
# → 50000 50000

Local keys written by the browser (all removable via site data): gw-lang, gw-dark, gw-quickwiki (preferences); gw_v6_* (10-minute search cache, max 30 keys); gw-history and gw-history-optout (search history, max 200 entries). Nothing else is persisted, and none of it is ever transmitted.

Security & privacy

Repository details

Repositorygithub.com/XPDevs/gateway
Live sitexpdevs.github.io/gateway/
DescriptionA small search engine made from scratch
LanguageJavaScript, HTML and CSS, with no build step and no dependencies
Index size50,000 entries · version 2026-09-29-gateway-index-v6-50000-verified
DocumentationREADME · How search works · About · Privacy · Terms

Contact Us

For more information about Gateway or to discuss how our technology can support your needs, please contact us at xpdevs.github.io.com@gmail.com.