←all writing
0130 Sept 2026/12 min read

Building a help search that finds the right article, without a search server

How the search on help.happypet.tech went from 76% to 100% of test searches in the top 3, using MiniSearch, a synonym list, a test file and no search server. With the code.

SearchNext.jsMiniSearch

The help centre for the Happy Pet Tech SaaS has 104 articles. Until this week, its search was the kind most small sites have. It checked if the typed words appear in the title, the summary or the keywords of an article.

That search failed in four ways. One example of each:

What they typedWhat happenedWhy
invioceNo resultsOne typo
billDid not find the invoice articlesDifferent word for the same thing
collarNo resultsThe word is only in the body of an article, and bodies were not searched
bill kaise banayeDid not find the invoice articleHindi words mixed in (“how to make a bill”)

Our users are groomers and kennel staff. Many type in a mix of English and Hindi. They search while a customer is waiting.

This post is how we rebuilt it. It runs inside the Next.js app, with no search server and no paid API. I will show the real code, and you can copy the approach for any site with up to a few thousand documents.

This is the step that made everything else possible, so it comes first.

I wrote a JSON file of searches, each with the article it should find:

[
  { "q": "invioce", "expect": ["create-an-invoice-and-take-payment"], "kind": "typo" },
  { "q": "collar", "expect": ["belongings-drop-off-pick-up"], "kind": "body" },
  { "q": "check in", "expect": ["check-a-pet-in-and-out"], "kind": "plain" },
  { "q": "how to cook rice", "expect": [], "kind": "none" }
]

The last row is a search with no answer. Rows like it are there to catch a search that returns something for everything.

A small script runs every search and counts how often the right article is first, and how often it is in the top 3:

let top1 = 0, top3 = 0;
for (const row of golden) {
  const slugs = engine.search(row.q).hits.map((h) => h.slug);
  const rank = slugs.findIndex((s) => row.expect.includes(s));
  if (rank === 0) top1++;
  if (rank >= 0 && rank < 3) top3++;
}
console.log(`top 1: ${pct(top1)}  top 3: ${pct(top3)}`);
if (top3 / golden.length < 0.9) process.exit(1);

The target is the right article in the top 3 for 90% of searches. The script exits with an error below that, so a change that makes search worse cannot slip through.

The script also runs the old search on the same file. The old search scored 76% in the top 3. Now I had a number to beat, and every change after this was judged by the number, not by my feeling.

Step 2: the engine is a pure function

We use MiniSearch. It is a small JavaScript library that builds a search index in memory. For 104 articles the index is small, builds quickly and lives inside the server process.

The engine takes articles in and gives ranked results out. It never touches the database.

export function createSearchEngine(articles, categoryLabels, synonymGroups) {
  const index = new MiniSearch({
    idField: 'slug',
    fields: ['title', 'keywords', 'summary', 'headings', 'body'],
    storeFields: ['context_keys', 'service_types'],
    tokenize,
    processTerm,
    searchOptions: {
      boost: { title: 6, keywords: 4, summary: 2, headings: 2, body: 1 },
    },
  });
  index.addAll(docs);
  // ...
  return { search };
}

Two things to notice.

The boost numbers say how much each field counts. A word in the title counts six times more than the same word in the body.

And because the engine is pure, the API route and the test script use the exact same function. The route feeds it from the database. The test script feeds it from the public API. What I measure is what runs in production.

Step 3: clean every word the same way on both sides

Every word goes through one function, both when the article is indexed and when the user searches.

export function processTerm(raw: string): string | null {
  const term = raw.toLowerCase().replace(/['’]/g, '');
  if (term.length < 2 || STOP_WORDS.has(term)) return null;
  if (term.length > 3 && term.endsWith('s') && !/(ss|us|is)$/.test(term)) {
    return term.slice(0, -1);
  }
  return term;
}

STOP_WORDS are words that carry no meaning in a search: “how”, “to”, “the”, “my”. Ours also has Hindi fillers: kaise, kare, hai, banaye. So invoice chahiye (“I want an invoice”) becomes a search for invoice.

Two details here were learned the hard way.

“not” and “no” are not stop words. “Not getting alerts” and “no show” mean something different without them.

The stemming is tiny on purpose. It only removes a trailing “s”, so “invoices” finds “invoice”. I decided against a smarter stemmer. Classic stemmers turn “invoice” and “invoicing” into different stems for some words. A wrong stem is worse than no stem, because then the word finds nothing at all.

Step 4: typos and half-typed words

index.search(query, {
  // only the word still being typed is matched as a prefix
  prefix: (_term, i) => i === lastIndex,
  // one wrong letter from 5 letters, two from 7, none for short words
  fuzzy: (term) => (term.length >= 7 ? 2 : term.length >= 5 ? 1 : false),
});

Both limits are there because of real wrong results.

If every word is a prefix, the word “in” in the middle of a sentence matches “invoice”. So only the last word, the one still being typed, is a prefix. groo finds grooming.

If short words allow typos, “pet” matches “put”, “per” and “set”. So words under five letters must match exactly.

Step 5: three ranking problems, and the rule that fixed each

After steps 2 to 4 the test score was already much better. The remaining misses were ranking problems. The right article was found, but not first.

Searching “invoice” put “Cancel or delete an invoice” above “Create an invoice”. Shorter titles score higher in the ranking formula, and that is all it was.

Our articles have a “search words” field that the writer fills in. If the whole search is equal to one of an article’s search words, the writer has told us that this article answers it. That beats statistics:

if (entry.keywordPhrases.has(typedPhrase)) r.score *= 1.8;

Searching “no show” found the wrong articles. “No” and “show” are both common words. An article with “no” in one place and “show” in another scored nine times higher than an article that says “no show”.

So for searches of two or more words, articles that contain the words side by side, in order, rank ahead of all the others:

if (words.length > 1 && entry.flat.includes(` ${typedPhrase} `)) phraseHits.add(r.id);

results.sort((a, b) =>
  Number(phraseHits.has(b.id)) - Number(phraseHits.has(a.id)) || b.score - a.score
);

Searching “check in” found “What to check before you go live”. “In” is a stop word. So “check in” was searched as “check”.

We joined a small list of two-word actions into one word, on the article side and the search side:

const JOINED_ACTIONS = {
  check: new Set(['in', 'up']),
  clock: new Set(['in']),
  log: new Set(['in']),
  sign: new Set(['in', 'up']),
  pick: new Set(['up']),
  set: new Set(['up']),
};

Now “check in” is indexed and searched as checkin. “Check in” went from 2nd to 1st, and “someone new needs to log in” went from 7th to 1st.

Step 6: synonyms, and the mistake I made first

A synonym list is groups of words that mean the same thing to our users:

["invoice", "bill", "billing"],
["boarding", "hostel", "pet hotel", "overnight stay"],
["refund", "money back", "paisa wapas"],
["vaccination", "vaccine", "jab", "rabies", "tika"],
["upi", "gpay", "google pay", "phonepe", "paytm"]

We have 40 groups. They live in the database and are edited in our CMS. The Hindi words are there because that is what people type.

My first version added all the synonyms to the search. A search for “invoice” became a search for “invoice bill billing”.

The test score went down. The share of searches with the right article first fell to 78%. The reason: an article that mentions “bill”, “billing” and “invoice” collected points for all three, and beat the main invoice article on a search for “invoice”.

The fix is that a synonym replaces the word. It does not add to it. Each synonym gets its own search, the results count for half, and each article keeps its best score:

for (const r of index.search(typed)) keep(r, r.score);

for (const variant of synonymVariants) {
  for (const r of index.search(variant, { prefix: false, fuzzy: false })) {
    keep(r, r.score * 0.5);
  }
}

Half weight means an article that uses the reader’s own word still wins a tie.

One more rule we added: keep groups tight. A broad word in a group drags in every article about it. I tried adding “cost” to the price group and one test search got worse, so it came out.

Step 7: “Did you mean”, and the dictionary it needed

If a word is in no article, the engine looks for the nearest word it does know:

function correct(term: string): string | null {
  if (term.length < 4 || vocabulary.has(term)) return null;
  const max = term.length >= 7 ? 2 : 1;
  // find the known word with the smallest edit distance, up to `max`
  // ties go to the word that more articles use
}

The edit distance counts two swapped letters as one edit. Plain Levenshtein counts invioce as two edits away from invoice, and swapped letters are the most common typing slip there is.

Then I found a bug. The engine was “correcting” real English words that our articles happen not to use:

  • “buzzing” became “buying”
  • “died” became “tied”
  • “upfront” became “front”

A search with the word “died” in it would come back with “Did you mean: tied”. That is not acceptable on a site about pets.

The fix is an English word list. We use the word-list package, about 270,000 words. Before correcting a word, the engine checks the list. A real word is never corrected.

if (isWord(raw)) return null;   // a real word we do not use is not a typo
return best;

This check runs last, only when a correction was found. Loading 270,000 words takes memory and time, and most searches never need it. The test script now fails if any correctly spelt search gets a “Did you mean”.

Step 8: open the article at the right section

Article bodies are split at their headings before indexing. When the match is inside a section, the result links to /article#section and shows a line of that section.

Choosing the section has one trick. A word that appears in every section says nothing about which one to open. So each matched word is worth one divided by the number of sections it appears in. In “walk in customer bill”, the rare word “walk” picks the section, not “customer”.

What did not work: search by meaning

Some users type whole questions, like “big dogs cost more to wash”. No article uses those words. So I tried the popular answer: embeddings. A small open model (bge-small-en) turns each section and each search into a vector, and the search can match by meaning.

On 20 test questions it looked great. The top-3 score went from 65% to 85%.

But I had picked the model and its settings using those same 20 questions. So I wrote 12 new questions blind, after the settings were frozen, and ran them once. The result: 75% in the top 3 with meaning search, and 75% without it. No gain.

It also had two other costs:

  • It always finds something “close”. A search for “best pizza near me” returned help articles. We could only let it reorder results, never add them.
  • It added about 470 MB of packages to every deploy.

I removed it and kept the code as a patch file.

The three questions it missed were fixed another way, with content and no code. I added the users’ words to the “search words” of three articles: “big dogs” on the pet sizes article, “overpaid” on the wallet article, “check up” on the consultation article. All three moved into the top 3.

If you remember one thing from this section: test on questions you did not tune on. Without the 12 blind questions, I would have shipped 470 MB for nothing.

Learn from what people really type

Test searches that I write are guesses. So the site now logs real searches: the words, how many results came back, and which result was clicked.

The log is anonymous. It has no company, user or role. Numbers of three or more digits are replaced, and emails are removed, because staff type phone numbers and invoice numbers into search boxes. Rows delete themselves after 180 days.

Our CMS has a screen with two lists: searches that found nothing, and searches where nothing was clicked. Each row has a button to add a synonym or to start a new article with those search words already filled in. That screen is where the next improvements will come from.

The numbers

Old searchNew search
Word searches: right article in the top 376%100%
Word searches: right article first66%92%
Whole questions in plain language: top 3not measured65%

The last row is the honest one. Plain-language questions are still the weak spot, and the fix for them is better search words on the articles, found from the real search log.

If you want to build this

  1. Write 40 test searches with their expected results. Include typos, other words for the same thing, and searches that should find nothing.
  2. Score your current search with it. That is your baseline.
  3. Put MiniSearch behind one API route. Index title, keywords, summary, headings and body, with title weighted highest.
  4. Clean words the same way for the index and the query.
  5. Add prefix on the last word only, and typo tolerance on long words only.
  6. Run the test. Look at each miss and ask why. Most of our fixes were one rule each.
  7. Add synonyms as replacements at half weight, not as extra words.
  8. Log real searches, and read the ones that found nothing.

The whole engine is about 600 lines of TypeScript in one file.

← all writing nextHow I made help.happypet.tech load in 0.1 seconds →