Turn a PDF into an editable Word document

or drop one here
  • Text, faces, bold, italic, sizes, alignment and indents carry over.
  • Paragraphs, page breaks, headings and tables whose columns line up carry over.
  • Pictures, text colour and list numbering do not. A bullet stays a character.
  • A scanned PDF has no text to convert, so it is refused instead of emptied.
Choose a PDF first.

Convert a native-text PDF into an editable Word .docx without uploading it. This tab reads every drawn run with its box, size and face, builds paragraphs and headings, and recovers tables whose columns line up. Pictures, text colour, hyperlinks, running heads, list numbering and still-fillable form fields come across too. The .docx is reopened with a separate reader and 98% of the text must come back.

How to turn a PDF into a Word document

  1. Choose a PDF

    Press Choose a PDF, or drop one on the page. The file is opened in this browser tab.

  2. The tool reads the text with its geometry

    Every drawn run comes back with its box, its size and the face it used, so nothing depends on guessing from pixels.

  3. It infers the structure

    Runs that share a baseline become a line, lines that flow together become a paragraph, a larger face at the top of a block becomes a heading, and columns that line up across several lines become a table.

  4. Check and download

    The new .docx is reopened with the Word reader the other direction uses and its text compared with the text the PDF held. A failed check does not produce a download.

What structure can be recovered, and what cannot

A PDF records where each run of text was drawn. It does not record that two runs belong to the same sentence, that a line is a heading, or that four runs are a row. Those are inferences, and this tool makes them from geometry alone: a shared baseline makes a line, consistent left edges and line spacing make a paragraph, a larger face at the top of a block makes a heading, and a repeated column pattern across consecutive lines makes a table.

Where the evidence runs out the tool keeps the text and drops the claim. A table whose columns do not line up stays as paragraphs. A run of markers becomes real Word numbering only when a numbering definition built from them reproduces every marker exactly; one that does not stays the text it was. Rotated text is kept upright at the end of its page rather than woven into a paragraph it does not belong to.

A picture is copied out of the PDF rather than redrawn, so a stored JPEG arrives byte for byte. A picture the tool cannot copy faithfully is left behind and counted in the result rather than re-encoded into something worse. A line that repeats in the same place on every page becomes a real header or footer part, with its page number as a field Word renumbers. Text colour comes across as run colour.

A hyperlink is the one thing a PDF states outright: a rectangle on the page with a web address attached, held apart from the text it covers. The tool reads those rectangles and makes the runs under them clickable. A run is cut where the link changes between the pieces the PDF drew it in, and a rectangle that covers less than most of a piece names nothing, so a link over three words of a sentence stays over those three words and never spills onto the words beside them.

A link can also point at a page of the same PDF, which is what a contents page is made of. That is not an address, so Word cannot hold it as one: it needs a bookmark on the page being named, and an anchor that quotes the bookmark. The tool works the destination back to a page and a point on it, puts a bookmark in front of the nearest paragraph, and points the contents line at it. A destination that names a spot no paragraph came near keeps its text and loses its link, because an anchor that goes nowhere is worse than none. A web address on a running head is dropped as well: a header is a separate part with its own address list, and this writer keeps one. An anchor needs no address list, so a contents line in a header survives. Drawn vector art does not come across at all.

A form field keeps its answer nowhere near the page content, so reading the page alone finds nothing where the answers are. The tool reads the fields themselves and writes each one as a Word content control: a text box, a check box or a dropdown with its list. The answer is in the control, so the reader sees it and can still change it. An empty field gets a rule of underscores cut to its own width, and a check box gets a box glyph. Neither the rule nor the glyph is text the PDF held, so both are taken out of the check before it measures anything. A push button and a signature field are counted and left behind, because neither holds an answer Word could carry.

The check is the reason there is a download. The .docx this tool wrote is reopened with the same OOXML reader the Word to PDF direction uses, and its non-space characters are counted against the characters the PDF held. Below 98% nothing is offered.

Questions

Does PDF to Word upload my file?
No. The PDF is opened, converted, reopened and checked in this browser tab. The bytes never leave it.
Is the result editable?
Yes. Every line becomes a real Word paragraph, heading or table cell, with the face, weight, slant, size, alignment and indent the PDF used. It is not a picture of a page in a .docx wrapper.
Will it look exactly like the PDF?
No. A PDF stores finished pages; a .docx stores flowing content that Word lays out again. Page size, page breaks, paragraph shape and table columns are kept, but Word rebreaks the lines.
Do hyperlinks still work in the Word file?
Yes. A PDF keeps a link as a rectangle with an address, apart from the text under it. The tool reads those rectangles and writes real Word hyperlinks. A rectangle has to cover most of a piece of drawn text to name it, so a link over part of a sentence covers only that part and never spills onto the words beside it. A link that points at a page of the same PDF becomes a Word anchor into a bookmark the tool writes on that page. A web address on a running head is counted in the result instead of written.
What does not carry over, and does it work on a scan?
Drawn vector art does not come across. A picture stored in a form this tool cannot copy faithfully is left behind and counted rather than re-encoded into something worse. A scan does not work at all: this version needs a native text layer, and an image-only scan is refused instead of producing an empty document, because guessed text needs OCR and its own accuracy measure.
Do form fields stay fillable?
Yes. A text field, a check box, a radio button and a dropdown become Word content controls, with the answer already in them and the dropdown list intact. A push button and a signature field hold no answer, so they are counted instead. A password field keeps its control and drops the characters behind the dots.
How are tables found?
By alignment. Two or more consecutive lines that split into the same number of gap-separated cells on a shared set of column positions become a table. A single row, or a table whose columns do not line up, stays as paragraphs.
What does the Word check prove?
The tool reopens the .docx it wrote with its own OOXML reader and counts the non-space characters it finds. At least 98% of the text the PDF held must come back, or no download is offered.

The other pdfpix tools

All 33 of them run in a browser tab like this one. Each reads its own output back and checks it before offering the download, and none of them needs an account.

Where to go next