Skip to content
Sudip KC writing notes in a journal at a desk
SK.
← All articles

Development · · 7 min read

How Service Workers Actually Work

  • PWA
  • Performance
  • Next.js
  • Tutorial

A service worker is a small script that sits between your page and the network, deciding what happens to every request. Once that clicks, offline pages, instant repeat visits and "new version available" toasts stop feeling like magic.

When I turned my portfolio into a PWA, I wrote about the what: the offline page, the install button, the update toast. That's in I Turned My Portfolio Into a PWA — Here's How. This post is about the how. I'll walk through the mental model and use real code from the service worker running on this site.

The mental model: a proxy you program

A service worker is JavaScript that runs separately from your page, on its own thread, with no access to the DOM. Its main job is to listen for fetch events. Every request your page makes within the worker's scope, for HTML, scripts, images, fonts, can be intercepted and answered however you like:

  • Pass it to the network as usual.
  • Answer from a cache.
  • Try the network, fall back to a cache.
  • Build a response from scratch.

That's the whole idea. Everything else, offline support, caching strategies, updates, is built on top of that one capability.

A few constraints shape how you write one:

  • HTTPS only (localhost is allowed for development).
  • Scope is set by where the file lives. A worker at /sw.js can control the whole site. That's why mine sits in public/.
  • It's event-driven and short-lived. The browser starts it when there's an event and may stop it when idle. Don't keep state in global variables and expect it to survive.

The lifecycle: register, install, activate

This is the part that confuses people most, and it's where most bugs come from.

Register

The page tells the browser about the worker. On my portfolio, registration adds the build version to the URL, so each deploy counts as a new worker:

navigator.serviceWorker.register(`/sw.js?v=${buildVersion}`);

Inside the worker, I read that version back and use it to name the page cache:

const VERSION = new URL(self.location.href).searchParams.get("v") || "dev";
const PREFIX = "sk-";
const CACHES = {
  pages: `${PREFIX}pages-${VERSION}`,
  static: `${PREFIX}static`,
  assets: `${PREFIX}assets`,
  media: `${PREFIX}media`,
};

Install

The install event fires once per new version. It's the moment to precache what the app needs to work offline. I precache the home page and the offline page, plus icons and the manifest:

self.addEventListener("install", (event) => {
  event.waitUntil(
    (async () => {
      const pages = await caches.open(CACHES.pages);
      await pages.addAll(PRECACHE_PAGES.map((url) => new Request(url, { cache: "reload" })));
      const assets = await caches.open(CACHES.assets);
      // Best effort: a missing optional asset must not block installation.
      await Promise.allSettled(PRECACHE_ASSETS.map((url) => assets.add(new Request(url, { cache: "reload" }))));
    })(),
  );
});

Two details worth copying. event.waitUntil tells the browser not to consider installation finished until the promise settles. And cache: "reload" bypasses the HTTP cache, so you don't precache a stale copy of the page you're trying to refresh. The pages use addAll, which fails the install if anything is missing, because a worker without an offline page is worse than no worker. The icons use allSettled, because a missing icon shouldn't block everything.

Waiting

Here's the part tutorials skip. After a new worker installs, it doesn't take over straight away. If an older worker is still controlling open tabs, the new one sits in a waiting state. This is deliberate: you don't want two versions of your app's logic fighting inside one tab.

Many guides call self.skipWaiting() in install to force the takeover. I don't. Instead, the new worker waits until the user agrees:

self.addEventListener("message", (event) => {
  if (event.data?.type === "SKIP_WAITING") self.skipWaiting();
});

The page notices a worker in the waiting state, shows a "New version available" toast, and posts SKIP_WAITING only when the user taps it. Then it reloads once the new worker takes control. No surprise reloads in the middle of reading something.

Activate

Once the new worker takes over, activate fires. This is where old caches get cleaned up:

self.addEventListener("activate", (event) => {
  event.waitUntil(
    (async () => {
      const keep = new Set(Object.values(CACHES));
      const keys = await caches.keys();
      await Promise.all(
        keys.filter((key) => key.startsWith(PREFIX) && !keep.has(key)).map((key) => caches.delete(key)),
      );
      if (self.registration.navigationPreload) await self.registration.navigationPreload.enable();
      await self.clients.claim();
    })(),
  );
});

The prefix check matters: it only deletes caches this site created, so it won't touch anything else on the origin. clients.claim() lets the new worker control open pages right away instead of waiting for the next navigation. And navigation preload lets the browser start fetching the page in parallel while the worker boots, which hides the worker's start-up time.

Most service worker bugs aren't in the caching code. They're in not understanding which version is in control right now.

The fetch event: one router, several strategies

The fetch handler is a router. It looks at each request and picks a strategy. Mine starts by deciding what not to touch:

self.addEventListener("fetch", (event) => {
  const { request } = event;
  if (request.method !== "GET") return;

  const url = new URL(request.url);
  if (url.origin !== self.location.origin) return;
  if (url.pathname === "/sw.js" || url.pathname.startsWith("/api/")) return;
  // React Server Component payloads vary by router-state headers — never cache them.
  if (request.headers.has("RSC") || url.searchParams.has("_rsc")) return;
  // ...route to a strategy
});

Returning without calling event.respondWith means "let the browser handle it normally." Leaving things alone is the safest default. POSTs, API calls and cross-origin requests go straight to the network.

The RSC line is specific to Next.js and it bit me. App Router navigations fetch Server Component payloads from the same URL as the page, with different headers. Cache those like normal pages and you'll serve the wrong payload for a route. So they're never cached.

Then the actual strategies:

  • Page navigations: network-first. Fresh HTML when online. If the network fails, serve the cached copy of that page, or redirect to /offline.
  • /_next/static/*: cache-first.** These files have content hashes in their names, so a given URL never changes. Once cached, there's no reason to ask the network again.
  • Images, fonts, icons: stale-while-revalidate. Serve from cache instantly, update the cache in the background. That includes the scroll-animation frames, which would otherwise be a lot of repeat downloads.
  • Videos: cache on first play, then serve from cache with byte-range support.
  • Everything else: network only.

Stale-while-revalidate, line by line

This one strategy explains most of the Cache API, so it's worth reading closely:

async function staleWhileRevalidate(event, cacheName, limit) {
  const { request } = event;
  const cache = await caches.open(cacheName);
  const cached = await cache.match(request);
  const network = fetch(request)
    .then(async (response) => {
      if (response.ok && response.type === "basic") {
        await cache.put(request, response.clone());
        await trim(cacheName, limit);
      }
      return response;
    })
    .catch(() => undefined);
  if (cached) {
    event.waitUntil(network);
    return cached;
  }
  return (await network) || Response.error();
}

The network request starts immediately either way. If there's a cached copy, it's returned at once, and event.waitUntil(network) keeps the worker alive long enough to finish updating the cache. The response.clone() is necessary because a response body can only be read once: one copy goes into the cache, the other back to the page. And response.type === "basic" makes sure only same-origin responses get stored.

The offline fallback has a Next.js twist

My first version served the offline page's HTML directly when a navigation failed. It looked fine for a split second, then Next.js tried to hydrate /offline markup on a different route and showed its error screen. The fix was to redirect instead:

return Response.redirect(`${OFFLINE_URL}?from=${encodeURIComponent(url.pathname + url.search)}`, 302);

The browser lands on the real /offline URL, hydration matches, and the from parameter lets the offline page offer a retry back to where the user was going.

Caches need limits

Runtime caches grow forever unless you stop them. I keep a limit per cache and evict the oldest entries:

async function trim(cacheName, limit) {
  const cache = await caches.open(cacheName);
  const keys = await cache.keys();
  await Promise.all(keys.slice(0, Math.max(0, keys.length - limit)).map((key) => cache.delete(key)));
}

It's not a perfect LRU, since keys() returns entries in insertion order. But it's predictable, and for a portfolio it's plenty.

Debugging tips that saved me time

  • Open DevTools, Application, Service Workers. You can see which version is active, which is waiting, and force an update.
  • "Update on reload" is useful in development, but turn it off when testing the real update flow.
  • Use the Cache Storage panel to see exactly what's stored, per cache.
  • Test offline with the Network panel's offline toggle, then for real with airplane mode on a phone.
  • When things get weird, unregister and clear site data. Then fix the bug that got you there.

The MDN Service Worker API docs and web.dev's offline guidance are the two references I go back to.

The short version

A service worker is a programmable proxy with a strict lifecycle. Register it, precache the essentials on install, clean up on activate, and route each fetch to a strategy that fits the kind of request. Be careful about what you cache, especially in Next.js, and let users decide when a new version takes over.

On networks like the ones I build for in Nepal, this is more than polish. It's the difference between a site that disappears when the signal drops and one that keeps working. If that's your situation too, Designing Apps for Slow Internet Connections covers the rest of the picture.