Turn a PDF into an editable Word document
- Text, faces, bold, italic, sizes, alignment and indents carry over.
- Paragraphs, page breaks, headings and tables whose columns line up carry over.
- Pictures, text colour and list numbering do not. A bullet stays a character.
- A scanned PDF has no text to convert, so it is refused instead of emptied.
Converting your PDF
Word document ready
Same page, two views
The Word preview is plain text. It does not load links, images or HTML.
Details
Convert a native-text PDF into an editable Word .docx without uploading it. This tab reads every drawn run with its box, size and face, builds paragraphs and headings, and recovers tables whose columns line up. Pictures, text colour, hyperlinks, running heads, list numbering and still-fillable form fields come across too. The .docx is reopened with a separate reader and 98% of the text must come back.
How to turn a PDF into a Word document
-
Choose a PDF
Press Choose a PDF, or drop one on the page. The file is opened in this browser tab.
-
The tool reads the text with its geometry
Every drawn run comes back with its box, its size and the face it used, so nothing depends on guessing from pixels.
-
It infers the structure
Runs that share a baseline become a line, lines that flow together become a paragraph, a larger face at the top of a block becomes a heading, and columns that line up across several lines become a table.
-
Check and download
The new .docx is reopened with the Word reader the other direction uses and its text compared with the text the PDF held. A failed check does not produce a download.
What structure can be recovered, and what cannot
A PDF records where each run of text was drawn. It does not record that two runs belong to the same sentence, that a line is a heading, or that four runs are a row. Those are inferences, and this tool makes them from geometry alone: a shared baseline makes a line, consistent left edges and line spacing make a paragraph, a larger face at the top of a block makes a heading, and a repeated column pattern across consecutive lines makes a table.
Where the evidence runs out the tool keeps the text and drops the claim. A table whose columns do not line up stays as paragraphs. A run of markers becomes real Word numbering only when a numbering definition built from them reproduces every marker exactly; one that does not stays the text it was. Rotated text is kept upright at the end of its page rather than woven into a paragraph it does not belong to.
A picture is copied out of the PDF rather than redrawn, so a stored JPEG arrives byte for byte. A picture the tool cannot copy faithfully is left behind and counted in the result rather than re-encoded into something worse. A line that repeats in the same place on every page becomes a real header or footer part, with its page number as a field Word renumbers. Text colour comes across as run colour.
A hyperlink is the one thing a PDF states outright: a rectangle on the page with a web address attached, held apart from the text it covers. The tool reads those rectangles and makes the runs under them clickable. A run is cut where the link changes between the pieces the PDF drew it in, and a rectangle that covers less than most of a piece names nothing, so a link over three words of a sentence stays over those three words and never spills onto the words beside them.
A link can also point at a page of the same PDF, which is what a contents page is made of. That is not an address, so Word cannot hold it as one: it needs a bookmark on the page being named, and an anchor that quotes the bookmark. The tool works the destination back to a page and a point on it, puts a bookmark in front of the nearest paragraph, and points the contents line at it. A destination that names a spot no paragraph came near keeps its text and loses its link, because an anchor that goes nowhere is worse than none. A web address on a running head is dropped as well: a header is a separate part with its own address list, and this writer keeps one. An anchor needs no address list, so a contents line in a header survives. Drawn vector art does not come across at all.
A form field keeps its answer nowhere near the page content, so reading the page alone finds nothing where the answers are. The tool reads the fields themselves and writes each one as a Word content control: a text box, a check box or a dropdown with its list. The answer is in the control, so the reader sees it and can still change it. An empty field gets a rule of underscores cut to its own width, and a check box gets a box glyph. Neither the rule nor the glyph is text the PDF held, so both are taken out of the check before it measures anything. A push button and a signature field are counted and left behind, because neither holds an answer Word could carry.
The check is the reason there is a download. The .docx this tool wrote is reopened with the same OOXML reader the Word to PDF direction uses, and its non-space characters are counted against the characters the PDF held. Below 98% nothing is offered.
Questions
- Does PDF to Word upload my file?
- No. The PDF is opened, converted, reopened and checked in this browser tab. The bytes never leave it.
- Is the result editable?
- Yes. Every line becomes a real Word paragraph, heading or table cell, with the face, weight, slant, size, alignment and indent the PDF used. It is not a picture of a page in a .docx wrapper.
- Will it look exactly like the PDF?
- No. A PDF stores finished pages; a .docx stores flowing content that Word lays out again. Page size, page breaks, paragraph shape and table columns are kept, but Word rebreaks the lines.
- Do hyperlinks still work in the Word file?
- Yes. A PDF keeps a link as a rectangle with an address, apart from the text under it. The tool reads those rectangles and writes real Word hyperlinks. A rectangle has to cover most of a piece of drawn text to name it, so a link over part of a sentence covers only that part and never spills onto the words beside it. A link that points at a page of the same PDF becomes a Word anchor into a bookmark the tool writes on that page. A web address on a running head is counted in the result instead of written.
- What does not carry over, and does it work on a scan?
- Drawn vector art does not come across. A picture stored in a form this tool cannot copy faithfully is left behind and counted rather than re-encoded into something worse. A scan does not work at all: this version needs a native text layer, and an image-only scan is refused instead of producing an empty document, because guessed text needs OCR and its own accuracy measure.
- Do form fields stay fillable?
- Yes. A text field, a check box, a radio button and a dropdown become Word content controls, with the answer already in them and the dropdown list intact. A push button and a signature field hold no answer, so they are counted instead. A password field keeps its control and drops the characters behind the dots.
- How are tables found?
- By alignment. Two or more consecutive lines that split into the same number of gap-separated cells on a shared set of column positions become a table. A single row, or a table whose columns do not line up, stays as paragraphs.
- What does the Word check prove?
- The tool reopens the .docx it wrote with its own OOXML reader and counts the non-space characters it finds. At least 98% of the text the PDF held must come back, or no download is offered.
The other pdfpix tools
All 33 of them run in a browser tab like this one. Each reads its own output back and checks it before offering the download, and none of them needs an account.
- Merge PDFs Several PDFs into one
- Organize pages Split, reorder, delete
- Rotate & crop Pages upright, margins cut
- Images to PDF JPG and PNG
- HEIC to PDF Open iPhone photos
- TIFF to PDF Open multi-page scans
- WebP to PDF Open images saved from the web
- BMP to PDF Open bitmap images
- Privacy scan Remove private metadata
- Protect & unlock Add or remove a password
- Repair PDF Rebuild a broken file
- Compress PDF Hit a size target
- PDF size breakdown See where every byte goes
- Print preflight Check a PDF before press
- PDF to images Export JPG or PNG
- PDF to TIFF One multi-page image file
- Extract images Save original pictures
- PDF to Markdown Keep the document structure
- Markdown to PDF Make a verified document
- PDF to CSV Extract a table
- PDF to Excel Every table on one sheet
- Excel to PDF Sheets on pages that fit
- Word to PDF Keep the layout
- Sign PDF Draw, type or upload
- Stamp PDF Number, label, watermark
- Compare PDFs Words and pages that changed
- Flatten PDF Make fields and marks permanent
- Fill PDF form Complete existing fields
- Create PDF form Add fillable fields
- N-up & booklet Print several pages per sheet
- Comic archive CBZ to PDF and back
- Redact PDF Destroy marked content
- Redaction checker Find text under black boxes
Where to go next
- Word to PDF Turn a Word document into a PDF that opens the same on every device.
- PDF to Markdown Turn selectable PDF text into a checked Markdown file.
- PDF to Excel Turn selected native-text PDF tables into a checked Excel workbook.
- Convert PDF to Word Turn a PDF back into an editable Word document, and know what the conversion cannot recover.
- Insert a PDF into Word Put a PDF inside a Word document.
- Add a PDF to google docs Get a PDF into a Google Doc.
- Privacy Know exactly what leaves the browser and what never does.