From Friday Deep-Dive · Friday, August 28, 2026 · 4 min read
Point it at the archive folder and ask for a search page
Claude Code is the version of Claude that works with the files on your own computer instead of only talking in a chat window, and it now comes as a normal desktop app: download it, sign in, click the Code tab, press Select folder. It does need a paid Claude plan, Pro or above, and on Windows there is one prerequisite the download page never puts in front of you, which is that Git must be installed or local sessions simply will not start. It is a free two-minute install from git-scm.com, Macs already have it, and skipping it produces an error that does not explain itself. Once you are in, the whole project is one request: point it at the folder of edition PDFs and ask for a single web page listing every edition by date with a search box at the top. I ran this on August 27 against a test archive built to be as ugly as a real one, fourteen files using five different filename date formats, a couple of ad proofs mixed in among the editions, and four editions that were scanned page images rather than digital exports. It worked. It sorted twelve editions by date, quietly left the ad proof and the rate card out on its own, and produced a page that opens by double-clicking, where typing one word finds it buried in the body text of a single edition from March. It needed exactly one correction, because the first version used a date format that works on a Mac and crashes on Windows, which is the whole rhythm of working this way: you read the first draft, you say what is wrong in plain English, it fixes it. And then it surfaced the thing I did not expect. One test file was named scan0042.pdf and was the May 2 edition, but because that date lived only inside the scanned image and nowhere in the filename, the finished page could not file it by date at all. It sat at the bottom of the list, undated, findable only by someone typing scan0042, which nobody will ever do. Your filenames are your index, and you do not find out otherwise until something finally reads the whole folder in order.

Trevor SletteCo-founder, Quadd.ai
Before you start
- Download the Claude desktop app at claude.com/download. There are builds for Mac and Windows.
- Install it, open it, and sign in.
- Click the Code tab at the top. If it asks you to upgrade, you need a paid plan. Pro is enough for this.
- Windows only, and this one is not on the download page: Git has to be installed or local sessions will not start. It is free at git-scm.com, and you click Next until it finishes. Macs already have it. If you skip this step, nothing works and the error message will not tell you why.
- Click Select folder and choose the folder your edition PDFs live in. Work on a copy the first time.
That is the whole setup. No terminal, no programming language to install.
The prompt
Paste this, adjusting the filename examples to match yours:
I have a folder of PDF files. Each one is a back issue of our newspaper. The filenames are inconsistent: some have dates like 2019-03-14, others like 3-21-19 or 03282019 or "april 11 2019".
Build me one HTML page, saved in that same folder, that:
- lists every edition sorted by date
- pulls the date out of the filename whatever format it is in
- reads the text inside each PDF and makes it searchable from a search box at the top
- links each entry to its PDF so it opens when I click it
- skips files that are obviously not editions, like ad proofs and rate cards
- clearly marks any edition that is a scanned image with no text inside, since those can only be found by date and filename
Some of my PDFs are scans with no text layer. Do not silently leave those out of the list. Show them and label them.
I am on Windows. Make sure everything works there.
That last line is there because I needed it. More on that below.
What happened when I ran it
I built a test archive on August 27 designed to be as ugly as a real one: 14 PDFs, five different filename date formats, two files that were not editions at all, and four editions that were scanned page images instead of digital exports.
| What went in |
What came out |
| 14 PDF files |
12 listed as editions |
| ad proof, rate card |
both left out automatically |
| 5 filename date formats |
all 5 parsed correctly |
| 8 born-digital editions |
full page text searchable |
| 4 scanned editions |
listed and labeled, date searchable only |
Then I tested the search the way a reader would. Typing referendum, a word that appears nowhere in any filename and only in the body text on one page, returned exactly that one edition. Typing april returned all four April issues despite three different filename formats. Typing flooding, a word that appears only inside a scanned edition, returned nothing, which is the honest and correct answer.
The correction round
The first version crashed on Windows. It used a date format that works on a Mac and throws an error on a PC. I said so in one sentence, it fixed it, and the second version ran clean.
This is the part people either love or hate about working this way. It is not a vending machine. It is a capable new hire handing you a first draft: you read it, you say what is wrong in plain English, and it corrects. If you have ever edited a stringer's copy, you already have the skill.
The thing I did not expect
One test file was named scan0042.pdf. It was the May 2 edition.
The finished page could not file it by date, because the date existed only inside the scanned image and nowhere in the filename. So it landed at the bottom of the list, marked undated, findable only by someone typing scan0042. Which nobody will ever do.
Your filenames are your index. Every archive has files like that, and you do not find out until something finally tries to read the whole folder in order. Building the index page is the cheapest audit of your own archive you will ever run, and that is true even if you throw the page away afterward.
If your archive turns out to be all scans
Then the index is not your project. OCR is. That is software that reads the picture of the page and writes the words back into the file, which turns a shelf you can browse into an archive you can search. Ask Claude Code what it would take for your specific files before you buy anything. The answer depends heavily on how the scans were made.
Two smaller projects if this one is too big
- The legals quoter. Paste in a legal notice, get the word count and the price at your rates. Genuinely one sitting, and you will use it every week.
- The subscriber list checker. Point it at a CSV export of your circulation list and ask it to flag duplicates, bad ZIP codes, and blank fields. Boring, and it will find things.