Skip to content

Search & Ranking

Implementation: internal/search/search.go. Two phases: hard-filter, then fuzzy score.

MatchesDoc — hard filters (bare words ignored)

Section titled “MatchesDoc — hard filters (bare words ignored)”
  • Tags: subset, case-insensitive. NotTags: exclusion.
  • DocType: exact, case-folded.
  • Title / generic key:value: fuzzyContainsFold = case-insensitive substring or sahilm/fuzzy hit.
  • Status: exact-fold, plus status:archived matches IsArchived().
  • Path: substring-fold.
  • Dates: Created/Updated via matchDate (inclusive/exclusive From/To); date/before/after use either-semantics (matchEither). Nil timestamp never matches a present filter.
  1. Filter via MatchesDoc.
  2. No bare words → sort by Path.
  3. Else scoreDoc per doc: every bare word must match — AND semantics, miss drops the doc. A word matches by case-insensitive substring of SearchBlob (fallback Title+" "+Body); only if that misses does it fall back to sahilm/fuzzy against the Title.
  4. Sort: score desc → most-recent UpdatedAt (nil = oldest) → Path.

Per bare word, scoreDoc adds a base plus a graded title boost:

ContributionConstantValueWhen
Match length—len(word)word is a substring of the blob (body/frontmatter)
Fuzzy score—fuzzy.Find scorenot in blob, but fuzzy-matches the title
Exact titleexactTitleBoost+200title equals the word exactly
Whole-word titletitleBoost+100word appears as a whole word in the title
Partial titlepartialTitleBoost+50word appears inside a longer title word

The three title boosts are mutually exclusive, so whole-word/word-boundary hits are preferred over longer words that merely contain the query. A search for butter therefore ranks Butter (200 + 6) above Buttermilk (50 + 6), which in turn edges out a body-only mention (6). Boundaries are UTF-8-aware and treat any non-letter/digit/_ character (spaces, hyphens, punctuation) as a separator.

Rank never hides; FilterArchived(docs, include) / TUI ctrl+a layer applies hiding. archived:true / status:archived still match hidden docs.

See also: Query Syntax for token semantics, and Backend Evaluation for the indexed-alternative decision.