Workflow8 min read

How to Save a Web Page So It Is Still There in Five Years

Bookmarks rot. Five ways to keep an actual copy of a web page — single-file HTML, PDF, the Wayback Machine, and automatic archiving — and which one to use for what.

A bookmark is not a copy. It is a promise that somebody else will keep their page online, unchanged, for as long as you need it — a promise they never made and have no reason to keep.

Most of the time this does not matter. You bookmark a news article, read it, never return. It matters when the page was the answer to something: the migration guide for the version you are actually running, the documentation for the API before they rewrote it, the client brief that is now three revisions out of date on their site.

The five methods, and what each survives

These are not ranked. They fail differently, which is the whole point — pick by which failure you are trying to avoid.

MethodSurvives page going offlineSurvives your laptop dyingSearchable later
BookmarkNoYesTitle only
Cmd/Ctrl+S, single fileYesNoBy file search
Print to PDFYesNoUsually
Wayback MachineYesYesNo — by URL only
Automatic archivingYesYesYes, full text

1. Save it yourself: Cmd/Ctrl+S

The most underrated option, and free in every browser you already have. The setting that matters is the format:

  • Webpage, Single File (.mhtml in Chrome and Edge, .webarchive in Safari) bundles the HTML, CSS and images into one file. This is what you want.
  • Webpage, Complete writes an .html file plus a folder of assets beside it. Move one without the other and you have a broken page — which is exactly what happens the first time you tidy your downloads folder.
  • Webpage, HTML Only saves the markup with no styling or images. Fine for text, useless for anything visual.

The catch is that it is manual, and manual archiving is archiving you will not do. Nobody presses Cmd+S on the fortieth tab of a research session. It works for the handful of pages you already know are important, which is not the set that turns out to matter.

2. Print to PDF

Better than its reputation for documents, worse than its reputation for applications. A PDF is universally readable, will still open in a decade, and preserves a text layer you can search and quote.

It also mangles anything with a sticky header, an infinite scroll, or a layout that assumes a viewport. Use it for articles, specifications and anything you may need to send to somebody else. Do not use it for dashboards or documentation with a sidebar, which is precisely the material most worth keeping.

3. The Wayback Machine

The Internet Archive lets you submit a URL and have a copy taken and stored on infrastructure that is not yours. This is the only method on the list that survives both the page going offline and your own hardware failing, which makes it the right choice for anything you might need to prove.

Its limitations are real. It cannot capture anything behind a login. It respects site opt-outs, so some pages simply refuse. And you can only find a capture if you still have the exact URL — there is no searching inside what you archived. It is a witness, not a library.

4. Automatic archiving

The only method that solves the actual problem, because the actual problem is not "how do I save a page" — it is that you do not know which pages matter until years later.

A tool that takes a copy of every page at the moment you save it removes the judgement call entirely. You are not deciding what is worth archiving; everything is archived, and the ones that turn out to matter are already there.

This is also what makes full-text search possible. Once the text of the page is stored, you can find the link you remember by a sentence inside it rather than by a title you never read. Anyone who has failed to find a link they knew they had saved has hit the limits of title-only search.

The tradeoff is honest: it requires trusting a provider, and the archive lives on their servers rather than yours. No export format carries archived copies, so the day you leave, the links come with you and the archive does not. Weigh that against how likely you are to press Cmd+S forty times.

5. Doing nothing, deliberately

Worth stating, because most advice on this topic assumes everything must be kept. Most of what you save is disposable. A recipe, a product page, an article you will read once. Archiving all of it is work you do not need to do.

The distinction that matters: is this page the record of something, or a pointer to something? A pointer can rot harmlessly. A record cannot.

What to actually do

  1. Find out how bad it already is

    Export your bookmarks and open the oldest twenty. If most still work you have less of a problem than you feared; if half are gone you now know what the next five years look like.

  2. Clear the dead weight first

    Run the export through the bookmark cleaner. It collapses duplicates and shows you everything saved more than five years ago — which is where the rot is concentrated, and much of it will be things you no longer want anyway.

  3. Separate records from pointers

    Whatever is a record of something — a spec, a brief, documentation for a version you run — gets a real copy. Everything else can stay a bookmark and rot in peace.

  4. Automate the ones you cannot predict

    For research and reference you save continuously, use something that archives at save time. Deciding what is important later is exactly what you cannot do.

The one-sentence version

A bookmark records where something was; an archive records what it said. That difference is the whole argument in choosing a Pocket alternative that archives. Almost everyone discovers which of those they needed several years after the moment they could have chosen.

Frequently asked questions

How do I save a web page permanently?

Save a copy rather than a link. Cmd/Ctrl+S with "Webpage, Single File" produces one self-contained file that opens with no network. For anything you must be able to prove later, submit the URL to the Wayback Machine as well, so the copy exists somewhere that is not your laptop.

Does bookmarking a page save it offline?

No, and this is the misunderstanding that costs people the most. A bookmark stores a URL and a title. If the page moves, changes or goes offline, the bookmark still opens — onto a 404, a redirect, or a rewritten version. Nothing about bookmarking preserves what you saw.

What percentage of saved links eventually break?

Studies of link rot consistently find meaningful decay within a few years, and the effect compounds: the older a link, the likelier it is dead. Rather than trust a number, check your own library — sort by oldest and open twenty. That gives you a rot rate for the kind of pages you actually save, which is the only figure that matters.

Can I search inside pages I have archived?

Only if whatever stored them indexed the text. A folder of saved .html files on your disk is searchable by your operating system’s file search; a PDF may or may not be, depending on whether the text layer survived. Tools that archive automatically usually index as they go, which is what makes "find the page with this sentence in it" possible.

Tools for this

Keep reading

Workflow

How to Import Your Pocket Export After the Shutdown

Pocket shut down in 2025 and took the app with it — but the export you downloaded is a standard file that any decent bookmark manager can still read. Here is how to open it, what survives, and what does not.

7 min read

Stop losing the good links

One branch per project, unlimited nesting, one-click capture from Chrome, a voice note for the why, and Cmd+K across all of it.

Open the dashboard