How to make WordPress search inside PDF and Word files Free & Pro
WordPress search skips the text inside uploaded files. To search PDFs and Word files, extract their text into an index, with OCR for scans. Here's how.
Updated September 26, 2026
WordPress’s built-in search only looks at text stored in its database — post titles, content and excerpts — so the words inside an uploaded PDF or Word file are invisible to it. To search file contents, a plugin has to extract each file’s text into a search index. Scanned PDFs need OCR first.
Why can’t WordPress search inside PDFs?
When you upload a file, WordPress stores it on disk and creates a Media Library entry holding a title, caption, alt text and description. Its search then runs a database query across the titles, content and excerpts of your posts and pages.
No part of that process reads the text inside the file. A 40-page policy PDF is found only by the words you typed into its title or description, not by anything on its pages.
Search engines such as Google can index the text of PDFs that are publicly reachable. That helps people find you, but it sends visitors away from your library, it works on Google’s schedule, and it cannot see files you have restricted.
What makes a file searchable?
Most PDFs exported from Word or another editor contain a text layer: the words are stored as text, so software can read them. A scanned or photographed PDF is a picture of a page, with no text layer at all.
A quick test: open the PDF and try to select a sentence with your cursor. If words highlight, there is a text layer. If the whole page selects as one image, the file needs OCR (optical character recognition) before any search can read it.
Modern Office files — .docx, .xlsx, .pptx — are zip archives of XML, so their text can be read directly. The older binary formats (.doc, .xls, .ppt) are much harder to read reliably.
What are the ways to add full-text search?
| Approach | Where document text goes | Scanned PDFs | Effort |
|---|---|---|---|
| Type key phrases into each file’s description | Your database | Yes, by hand | High, and it drifts as files change |
| A plugin that extracts text on your server into an index | Stays on your server | Only with OCR | Low once installed |
| A hosted search service | Copied to the provider | Varies | Moderate, often with a subscription |
| OCR | Wherever OCR runs — your server or a service | Yes | Processing-heavy; often a service |
| AI semantic search | Text chunks sent to an embedding provider | After OCR | Needs a provider account and key |
For most document libraries, server-side extraction does the heavy lifting. OCR covers scans, and semantic search is an optional layer for when visitors’ wording differs from your documents’.
How to search inside documents with FileDeck
Free: instant search across each document’s details
Free — included in every FileDeck install.
FileDeck’s library search filters as visitors type, matching each document’s title, excerpt, categories, tags, filename and reference number (a document with no excerpt is matched on the opening of its description instead). Libraries of up to about 500 documents search in the browser without a round trip; larger ones switch to server mode automatically. Add a standalone search box with [filedeck_search], or turn on the site-wide ⌘K / Ctrl-K quick search.
On the free version, the practical tip is to put the phrases people search for into each document’s title, excerpt or tags.
Pro: search inside PDF, Office and text files
FileDeck Pro — content search and OCR are included in all Pro plans. See pricing
- Activate Pro. Content search needs no separate switch. When you save a document, FileDeck extracts its text in the background, shortly afterwards, using WordPress’s scheduled tasks.
- Index your existing documents. Go to Documents → Settings → Search Index and click Rebuild search index. The rebuild runs in the background in batches, and the status line shows how many documents are indexed. Rebuild again after a bulk import.
- Search. Matches inside file contents appear alongside title matches in the same search box. For PDFs with a text layer, results show Found on page N; clicking it opens the document at that page. See jump to the page a match is on.
What content search reads:
- PDF — the text layer.
- Microsoft Office —
.docx,.xlsx,.pptx. - OpenDocument —
.odt,.ods,.odp. - Text —
.txt,.md,.csv,.log,.html,.xml.
Extraction runs in PHP on your own server, and the index is stored in your WordPress database; content search sends no document text anywhere. Reading PDFs uses PHP’s zlib, iconv and mbstring extensions, and reading Office and OpenDocument files uses ZipArchive — standard on most hosts.
Worth knowing:
- Older
.doc,.xlsand.pptfiles are not read. Re-save them in the modern format to make them searchable. - Documents that link to an external file (Google Drive, Dropbox, a URL) are searchable by their title and details, not their contents, because there is no local file to read.
- Very long files are capped at about 500 KB of extracted text each — roughly 80,000 words. PDFs larger than 40 MB are not read.
- Results respect access rules. Snippets and page numbers are only calculated for documents the visitor is allowed to view.
To tune matching, Pro also offers search synonyms and spelling tolerance.
Pro: make scanned PDFs searchable with OCR
- Go to Documents → Settings → Scanned OCR and tick Enable OCR. It is off by default, and there is no key to enter — your Pro licence connects to the service.
- From then on, each scanned PDF you publish is sent automatically. For the scans already in your library, and any an import brings in, go to Documents → Settings → Search Index and click Send N scanned PDF(s) for OCR. FileDeck processes them one at a time in the background.
Unlike content search, OCR sends document contents off your server. FileDeck sends only PDFs it could not read a text layer from on your server (normally scans) to the FileDeck OCR service, which extracts the text and returns it. Files are processed and immediately discarded, not stored. Each file is sent with its file name and your licence’s install ID, and, like any request WordPress makes, your site’s address. Access rules are not a filter: a scanned PDF in a restricted category is sent like a public one.
The managed service has fair-use limits: up to 100 scanned documents a day per site, and 2,000 pages a day in total. Only the first 40 pages of each scan are read. A larger backlog continues over the following days. You can point FileDeck at your own OCR service instead. OCR returns text without page breaks, so scanned documents are searchable but do not offer a page jump. See make scanned PDFs searchable with OCR.
AI tier: search by meaning
FileDeck AI — semantic search and "Ask this library" are part of the AI tier (from $119/year). They use your own AI provider API key and are off until you enable them. See pricing
Keyword search finds the words a visitor typed. AI semantic search finds documents that mean the same thing, so "time off after having a baby" can match a policy that only says "parental leave". Turn it on under Documents → Settings → AI search, then build the index under Documents → Settings → AI Index. Chunks of document text are sent to the embedding provider you choose, and the resulting vectors are stored in your own database. Ask this library builds on the same index to answer questions from your documents, with links to the sources.
FAQ
Does FileDeck send my documents anywhere to search them?
Content search does not: extraction and the index both stay on your server. Some optional features do send text elsewhere, and they are off until you enable them — OCR sends scanned PDFs to the FileDeck OCR service, and the AI-tier features send text to the AI provider you connect with your own key.
Can visitors search inside Word and Excel files?
Yes, with FileDeck Pro, for .docx and .xlsx files (and .pptx, OpenDocument and text files). Older .doc and .xls files need re-saving in the modern format first.
Why doesn’t a scanned PDF show up in search results?
It has no text layer to read. Enable OCR under Documents → Settings → Scanned OCR. Scans published after that are sent automatically; for one that was already in the library or came in with an import, open the Search Index tab and click Send scanned PDF(s) for OCR. If it still doesn’t match, the tab shows whether it is queued or why OCR stopped on it. See troubleshooting search.
Do I need AI search if I have full-text search?
Often not. If your visitors search with the same words your documents use, content search covers it. Semantic search earns its place in large or formal libraries where everyday wording and document wording differ.
Is searching inside documents free?
Searching titles, excerpts, categories, tags, filenames and reference numbers is free. Searching inside file contents, OCR, synonyms and spelling tolerance are Pro; semantic search and "Ask this library" are in the AI tier.
Next steps
Try the search box on the live demo, then read content search and scanned-PDF OCR for the settings in full. For the difference between keyword and meaning search, see can AI search inside your WordPress documents?, and see what each plan includes on the features page.
Still stuck? Email support@getfiledeck.com.