Research Archive Matcher indexes thousands of PDFs on your own computer, searches inside every page, reads scanned documents with optional OCR, and matches publication lists against your library — without uploading a single file.
RAM reads your documents, understands their structure, and turns an unsorted directory into a searchable, reportable research library.
Search inside every page of every PDF. Exact phrases, or all your query words. Results show the article, page number, similarity score and a text snippet.
Double-click any result to open the document at the matching page, with your search terms highlighted in yellow and page-by-page navigation.
Optional Tesseract OCR reads image-only pages that contain no selectable text. Blank pages are skipped and your PDFs are never modified.
Identifies titles by font hierarchy, plus authors, DOIs, journals, years, abstracts and keywords — then classifies each document by type.
Finds exact duplicates by SHA-256 file hash and probable duplicates by fuzzy title comparison, so you can reclaim space with confidence.
Align a citation list from Excel, Word or plain text against your archive. See what you hold, what is missing, and the closest suggestions.
Exports formatted Excel matrices, a Word executive summary, and an interactive HTML dashboard you can open in any browser.
A SQLite database with FTS5 indexing keeps searches instant, even across thousands of documents and tens of thousands of pages.
A four-tab graphical interface for everyday work, plus a full CLI for scripting, automation and headless servers.
Every scan, search and match runs entirely on your own machine. Your research documents are never uploaded, never sent to a cloud service, and never shared with anyone.
These are live screenshots of the application, not mock-ups.
Point RAM at a folder of PDFs.
Metadata and page text are indexed locally.
Find words and phrases inside any page.
Compare a citation list against your library.
Export Excel, Word and HTML deliverables.
Locate a half-remembered line inside hundreds of downloaded papers.
Keep course reading organised and instantly searchable.
Manage a literature review without losing track of sources.
Audit a digital collection and identify duplicates and gaps.
Verify that a submitted reference list matches available material.
Reconcile a department's publication record against archived files.
Pre-built applications for all three desktop platforms. No Python installation required.
Browse every version on the releases page.
To read scanned documents, install Tesseract OCR, restart RAM, and tick Enable OCR for image-only pages. RAM detects it automatically. Everything else works without it.
Requires Python 3.10 or newer. Runs on Windows, macOS and Linux.
No. All scanning, indexing, searching and matching happens on your computer. The only optional network feature is Crossref metadata enrichment, which sends a DOI — never your document text — and can be left switched off.
Yes, when OCR is enabled. Install Tesseract, tick Enable OCR for image-only pages, and RAM will read pages that contain no selectable text. Only image-only pages are processed, blank pages are skipped, and your original files are never altered.
The index is a SQLite database with FTS5 full-text search, which comfortably handles thousands of documents and tens of thousands of pages. Searches remain fast because everything is indexed locally.
Publication lists can be supplied as Excel (.xlsx,
.xls), Word (.docx) or plain text
(.txt). RAM matches on DOI where available and falls back
to fuzzy title comparison.
Never. RAM only reads your documents. Extracted text and metadata are written to a separate local index file, and reports are written to an output folder you choose.
Yes. RAM is released under the MIT licence, free for personal, academic and commercial use. The full source code is on GitHub.
Your message opens in WhatsApp with the details filled in. You review it and press send there.