help.happypet.tech is the help centre for the Happy Pet Tech SaaS. It has 104 articles in 18 sections and a “What’s new” page. It is a Next.js app on Vercel, with its content in MongoDB. The same site also feeds the help panel inside our app, through four small APIs.
This week we put a new search on it. Right after that, the whole site felt slow. Pages took a moment to open. And when you typed in the search box, nothing happened for two or three seconds, and then the results appeared.
This post is how I found the causes and fixed them. I’ll give the numbers first, so you know where it ends.
| Before | After | |
|---|---|---|
| Pages | 0.3 to 0.7 s | 0.10 to 0.18 s |
| A page nobody opened since the last deploy | 1.3 to 1.7 s | 0.18 to 0.22 s |
| Article list API | 1.2 to 1.6 s | 0.12 to 0.17 s |
| What’s new API | about 1.1 s | 0.12 to 0.18 s |
| Search | up to 1.4 s | 0.11 to 0.15 s |
| Sitemap | 1.5 to 2.8 s | 0.08 to 0.14 s |
Measure before touching anything
I did not start by guessing. I requested every public page and every API three times from the command line and wrote down the time to first byte. That is the time until the server starts to answer.
This matters because “the site feels slow” can mean many things. A slow server, a heavy page, a big image and a slow script all feel the same to a user. The timings showed that our pages were light, and the waiting was all on the server side.
So the question became: what is the server doing for a full second?
Cause 1: the server and the database were on different continents
Our database is in Mumbai. Our Vercel functions were running in Washington DC, which is Vercel’s default region.
Every database query had to cross the world and come back. That is about 200 ms each way. A page made between one and four queries. So a page could spend more than a second only on travel, before doing any real work.
The fix is one small file:
{
"regions": ["bom1"]
}
bom1 is Vercel’s Mumbai region. Now the function and the database are in the same city.
You can check where a function ran from any response. Look at the x-vercel-id header. Before, it showed iad1 (Washington). Now it shows bom1::bom1.
If you take one thing from this post, take this one. Check which region your functions run in, and check where your database is. It took one line to fix and it was the biggest part of the problem.
Cause 2: nothing was cached
Every visit asked the database again, even though help articles do not change often.
Now every public page and every public API reads from a cache. The database is asked at most once a minute. Pages are served from Vercel’s CDN, so most visitors get a ready copy and no function runs at all.
The usual worry with caching is stale content. Someone fixes a typo in the CMS and the old text stays on the site. We handled that with cache tags. Every content-changing action in the CMS clears the cache right away, so a save shows on the next request.
A simplified version of the idea:
// read: cached, tagged
export const getArticles = unstable_cache(
() => db.collection('help_articles').find({ status: 'published' }).toArray(),
['help-articles'],
{ tags: ['help'], revalidate: 60 }
);
// write: in every CMS handler that changes content
revalidateTag('help');
While doing this I found that the article page made four database reads, one after another. With the cached data it makes one.
Cause 3: search rebuilt its index again and again
This was the “nothing happens for two or three seconds” problem.
Our search keeps an index of all articles in memory. On Vercel, memory is not shared between server instances, and new instances start often. Each new instance had to build the index first, and that took five to seven slow database queries. The unlucky user who hit a new instance waited for all of that.
Three changes fixed it:
- The index is built from the cache, not from the database. So building it is fast.
- A warm-up request. When someone clicks into the search box, the site sends a small request that makes the server build its index. By the time they have typed a word, the index is ready. The help panel in the app does the same when it opens.
- Searches are remembered in the browser. If you type “invoice”, delete two letters and type them again, the second result is instant.
A typical search now answers in 0.11 s. On a brand new instance it is 0.4 to 0.5 s, and the warm-up usually takes that hit before the user types.
I also added a loader: a spinner in the search field and a thin progress bar. It only shows if an answer takes more than 200 ms. If it showed on every letter, a fast search would flicker and feel slower than it is.
Cause 4: the first visitor after a deploy built the page
After a deploy, two pages took 1.3 and 1.7 seconds. When I asked for them again, they came back in 0.1 s.
The reason: those pages were built on the first request. Whoever opened an article first after a deploy paid for building it. Everyone after that got the cached copy.
Now every article page and every section page is built during the deploy. No visitor is ever first.
This has a cost. The build now reads all 104 articles and 18 sections from the database, so deploys take a little longer, and the build machine must be able to reach the database. If it cannot, the deploy fails and the old version keeps running. I think that is the right way to fail.
The 17 requests that looked like a problem
After the fixes, I reloaded the home page with the Network tab open and saw about 17 extra requests, each with ?_rsc= in the URL. My first thought was that the page was calling 17 APIs.
They are not API calls. They are Next.js prefetches. When a link scrolls into view, Next.js quietly fetches a small preview of the page behind it, so a click opens instantly. The home page shows about 17 links, so it made about 17 requests. The search results did the same, up to 8 more each time the list changed.
They did not slow the page, because they start after the page is shown. But most of them were wasted. Nobody clicks 17 links.
So links now prefetch on intent. The page is fetched when the pointer moves onto a link, when the link gets keyboard focus, or when a finger touches it. That signal comes 100 to 300 ms before the click, which is enough.
// simplified
export function IntentLink({ href, children, ...rest }) {
const router = useRouter();
const warm = () => router.prefetch(href);
return (
<Link
href={href}
prefetch={false}
onPointerEnter={warm}
onFocus={warm}
onTouchStart={warm}
{...rest}
>
{children}
</Link>
);
}
A fresh load of the home page now makes 12 requests in total and zero prefetches. The HTML arrives in 24 ms and the page is fully loaded at 0.57 s. Hovering one card fetches that one page and nothing else.
One thing confused me while testing. I reloaded and still saw a few prefetch requests. My mouse was resting on a link while the page loaded, and that counts as intent.
A bug I found on the way
The site counts article reads and logs searches. These writes happened after the reply was sent to the user, so they would not slow anyone down.
On Vercel, a function can be frozen as soon as it has replied. So those writes could be cut off in the middle and lost.
Next.js has after() for this. It tells the platform to keep the function alive until the work inside it is done.
import { after } from 'next/server';
export async function GET() {
const result = await search(query);
after(() => logSearch(query, result.length));
return Response.json(result);
}
If you write to a database after sending a response on a serverless platform, check that the write really finishes.
What was already right
The pages themselves were light before any of this, which is why the fixes were all on the server side. The help site loads no third-party scripts: no analytics, no chat widget, no cookie banner. It uses the device’s system font, so there is no font file to wait for. The only image on the home page is an SVG logo.
The app’s help panel got faster for free
The help panel inside the Happy Pet app reads the same four APIs. So its requests became fast without a change on the app side:
| Panel action | First request | Repeat |
|---|---|---|
| Open panel: article list | 0.32 s | 0.14 s |
| Open panel: What’s new | 0.10 s | 0.11 s |
| Open an article | 0.15 s | 0.10 s |
| Search | 0.15 s | 0.11 to 0.13 s |
Keeping it fast
The last thing I added is a small script, perf-check.sh. It requests every page and API a few times and prints the timings, the region and whether the CDN answered. I run it after a deploy.
Good looks like this: the region is bom1, pages are well under 0.3 s, search is under 0.15 s, and pages report a cache HIT.