How to Remove Metadata from Word and PDF Documents Online
Word documents store author name, company, revision history, and deleted comments in hidden metadata. PDF files do too. Learn how to clean documents before sharing.
In 2003, a Microsoft Word document that the UK government published online contained hidden metadata revealing the document had been revised by an intelligence official. The metadata told a story the document itself didn't — that it had been modified at the last minute. This became part of the "dodgy dossier" controversy.
Document metadata has caused real problems in legal cases, business negotiations, and political scandals. Understanding what your documents reveal — and how to clean them — is practical professional knowledge.
What Word documents store in metadata
A Word document (.docx) is actually a zip archive containing multiple XML files. Inside that archive, metadata is stored in several places:
Core properties
- Author — who created the document (from their Office profile)
- Last modified by — who made the last change
- Company — from the Office account settings
- Created date — when the file was first created
- Modified date — when it was last saved
- Total editing time — cumulative hours the file was open
Application properties
- Template used — which Word template the document is based on
- Number of revisions — how many times it was saved
- Word count history
- Application version — which version of Word created it
Track changes and comments
- All tracked changes, including deleted text (if track changes was enabled)
- All comments, including deleted ones — this is frequently overlooked
- Who made each change and when
- Original text before modifications
File path information
- The full file path on the original computer: C:Usersjohn.doeDocumentsClientsAcme CorpProposal v3.docx
- This reveals the username, computer name, and folder structure
What PDF documents store in metadata
PDF files contain a metadata section (XMP or DocInfo) with:
- Title and Subject
- Author — from the creating application profile
- Creator — the application that created the PDF (Word, InDesign, etc.)
- Producer — the PDF converter/printer driver used
- Creation date and Modification date
- Keywords
PDFs created from Word documents often inherit Word metadata. A "clean" PDF exported from Word may still have the original Word author's name embedded.
High-risk scenarios for document metadata
Business proposals
A proposal sent to a client may contain internal comments like "accept 15% discount if they push back" or tracked changes showing price adjustments. The client sees the final numbers — but the metadata shows the negotiation history.
Legal contracts
Contract drafts reviewed by multiple parties retain every change and comment. Sending a "final" version without cleaning exposes the entire drafting process.
Résumés
Résumés edited from a template or someone else's CV may show the original person's name as the document author.
Academic submissions
Students who base work on classmates' documents may inadvertently submit files with the original author's metadata.
RFP and tender responses
In competitive bidding, document metadata can reveal which other companies provided templates, the internal pricing structure, and review comments.
How to remove metadata from Word documents
Method 1: Word's Document Inspector
1. Open the document in Microsoft Word
2. Go to File → Info
3. Click Check for Issues → Inspect Document
4. Select all inspection categories
5. Click Inspect
6. Click Remove All for each category found
7. Save the document
Covers: comments, revision history, personal information, hidden text, embedded documents.
Doesn't always catch: custom properties, some embedded object metadata.
Method 2: MetaLimpo (most thorough)
1. Go to limparmetadados.com.br/en
2. Upload the .docx file (up to 20MB)
3. Download the clean version
Removes all detectable metadata fields while preserving document content and formatting.
Method 3: Save as new document (manual)
1. Select all content (Ctrl+A)
2. Copy (Ctrl+C)
3. Create a new blank document
4. Paste as plain text (Ctrl+Shift+V or Paste Special → Unformatted Text)
5. Rebuild formatting as needed
6. Save with a new filename
Best for: when you need absolute certainty and are willing to reformat. Downside: formatting must be manually rebuilt.
How to remove metadata from PDFs
Using MetaLimpo
Same process as Word: upload the PDF at limparmetadados.com.br/en, download the clean version.
Using Adobe Acrobat (if available)
1. Open PDF in Acrobat
2. File → Properties → Description tab — manually clear fields
3. Tools → Redact → Sanitize Document (removes more thoroughly)
Free alternative: print to PDF
1. Open the PDF
2. Print → Select "Microsoft Print to PDF" or "Save as PDF"
3. The resulting file is a "fresh" PDF with minimal metadata
Verification checklist before sending any document
For Word (.docx):
- [ ] File → Info → Properties: Author field is blank or correct
- [ ] Document Inspector shows no personal information found
- [ ] No tracked changes visible (Review tab)
- [ ] No comments in the document
- [ ] Filename doesn't reveal the file path
For PDF:
- [ ] File → Properties: Author/Creator fields are blank or correct
- [ ] No form fields contain sensitive data
- [ ] No comments or annotations with personal information
Conclusion
Document metadata is invisible but significant. In professional contexts, it can reveal your negotiation strategy, your organization's internal structure, and your editing process. In legal contexts, it can contradict your claims about when or how a document was created.
The fix is simple: run your documents through MetaLimpo or Word's Document Inspector before sending. Takes under a minute and removes a meaningful privacy and professional risk.
Clean your documents at MetaLimpo — supports .docx, .pdf, and other common formats.