Point it at the archive folder and ask for a search page
Claude Code is the version of Claude that works with the files on your own computer instead of only talking in a chat window, and it now comes as a normal desktop app: download it, sign in, click the Code tab, press Select folder. It does need a paid Claude plan, Pro or above, and on Windows there is one prerequisite the download page never puts in front of you, which is that Git must be installed or local sessions simply will not start. It is a free two-minute install from git-scm.com, Macs already have it, and skipping it produces an error that does not explain itself. Once you are in, the whole project is one request: point it at the folder of edition PDFs and ask for a single web page listing every edition by date with a search box at the top. I ran this on August 27 against a test archive built to be as ugly as a real one, fourteen files using five different filename date formats, a couple of ad proofs mixed in among the editions, and four editions that were scanned page images rather than digital exports. It worked. It sorted twelve editions by date, quietly left the ad proof and the rate card out on its own, and produced a page that opens by double-clicking, where typing one word finds it buried in the body text of a single edition from March. It needed exactly one correction, because the first version used a date format that works on a Mac and crashes on Windows, which is the whole rhythm of working this way: you read the first draft, you say what is wrong in plain English, it fixes it. And then it surfaced the thing I did not expect. One test file was named scan0042.pdf and was the May 2 edition, but because that date lived only inside the scanned image and nowhere in the filename, the finished page could not file it by date at all. It sat at the bottom of the list, undated, findable only by someone typing scan0042, which nobody will ever do. Your filenames are your index, and you do not find out otherwise until something finally reads the whole folder in order.
The archive is the asset nearly every small paper is sitting on and nearly none can search. This is the cheapest possible way to learn what shape yours is actually in, and that answer is worth having whether or not you keep the page: if most of your editions turn out to be scans with no text inside them, then you have just learned that OCR is your real project and the index is the easy part. The wider point belongs in your next staff meeting. The barrier to building a small custom tool used to be hiring somebody who could build it. It is now an afternoon and a subscription you might already be paying for. The papers that work that out first will stop waiting on vendors to build the small things they need, and start building them.
Read the full post →
|
|
What is?
Text Layer
What it is: The invisible, selectable text stored inside a PDF. A PDF exported from your layout program carries one automatically. A PDF created by scanning paper does not: it is a photograph of a page, and to a computer it contains no words at all, only pixels.
Why publishers care: It is the difference between an archive you can search and a shelf you can only browse. Two PDFs can look identical on your screen while one is findable by any word printed on the page and the other is findable only by its filename. If your older editions were scanned from paper or microfilm, they almost certainly have no text layer, and the fix is OCR, software that reads the picture and writes the words back in.
|
|