← Home

REVIEWED & READY TO INSTALL

MarkItDown

microsoft/markitdown

Microsoft's official Python tool for converting PDF, Word, Excel, PowerPoint, Outlook MSG, images, audio and YouTube transcripts into tidy Markdown for prompts, RAG pipelines and agents. One command for every format, install only the extras you need, plus an OCR plugin and optional Azure Document Intelligence.

STEP BY STEP

How to install MarkItDown

The step-by-step guide is right on this page, written for people who have never installed anything from GitHub.

Updated 9/30/2026
GitHub preview image of MarkItDown
GitHub stars
187,793
Language
Python
Added
9/27/2026

How to install MarkItDown, step by step

Work from top to bottom. Each grey box is one command: copy the whole line, paste it into the terminal window and press Enter.

Never used a terminal? Read this first (1 minute)

A terminal is a window where you type commands for your computer. You do not need to understand them, just copy and paste them exactly.

  1. Windows: press the Windows key, type "PowerShell", then press Enter.
  2. macOS: press Command + Space, type "Terminal", then press Enter.
  3. Linux: press Ctrl + Alt + T.
  4. Copy a command from the guide, right-click inside the window you just opened to paste it (macOS: Command + V), then press Enter.
  5. Wait for the command to finish (the cursor blinks again on a new line) before running the next one.

Lines starting with # are notes; you do not need to copy them. If red error text appears, check the common errors section at the end of the guide.

1. What is MarkItDown?

MarkItDown is a free tool from Microsoft that turns files (PDF, Word, Excel, PowerPoint, images, audio, web pages and more) into Markdown text. Markdown is plain text with a handful of simple symbols for headings, lists and tables. AI assistants such as ChatGPT, Claude and Gemini read it very well, which is why MarkItDown is commonly used to tidy a document before pasting it into an AI.

To be direct: this is a command-line tool written in Python. It has no window and no buttons. Even so, basic use is a single command, so a non-technical person can follow along. The project also states that the output is meant for machines to read; if you need a faithful, good-looking conversion for human readers, it may not be the best choice.

It is released under the MIT licence and is completely free when it runs on your own machine. Only the optional features that call cloud services (Azure, or an AI model that describes images) cost money, billed by those providers.

A good fit if:

  • You often hand PDF, Word, Excel or PowerPoint files to an AI and want clean content.
  • You have many documents to extract text from in one go.
  • You would rather not upload your documents to an online converter.

Not a good fit if:

  • You need a conversion that keeps the original look, fonts and layout.
  • You only use a phone.
  • You need to read text in scanned images without any extra AI service (that feature needs a plugin plus an AI model).

2. What you need

  • A computer running Windows, macOS or Linux. No graphics card required.
  • Python 3.10 through 3.14: the programming language MarkItDown runs on. Download it from https://www.python.org/downloads/ . On Windows, tick "Add python.exe to PATH" during setup.
  • Internet to download the tool when installing. Converting files already on your computer with the built-in converters happens locally.
  • No account and no API key for basic use.

Things needed only for advanced features:

FeatureYou also needDoes it cost money?
Image descriptions, reading text in images (OCR)An API key for an OpenAI-compatible AI serviceYes, at the provider's rates
Azure Document IntelligenceAn Azure account and an endpoint addressYes, billed per call
Azure Content UnderstandingAn Azure account and an endpoint addressYes, every conversion is a billable call
Running in DockerDocker Desktop from https://www.docker.com/products/docker-desktop/No

3. Step-by-step installation

Step 1: Check Python

python --version

You should see something between Python 3.10 and Python 3.14. On macOS and Linux, if the command fails, try python3 --version.

Step 2: Create a separate environment (recommended)

A virtual environment is a private folder that holds MarkItDown's libraries so they do not interfere with other Python software. The project recommends using one.

Windows:

python -m venv .venv
.venv\Scripts\activate

macOS and Linux:

python -m venv .venv
source .venv/bin/activate

It worked if (.venv) appears at the start of your command line. Each time you open a new command-line window, run the second command (activate) again before using MarkItDown.

Step 3: Install MarkItDown

Windows:

pip install "markitdown[all]"

macOS and Linux:

pip install 'markitdown[all]'

The [all] part installs the readers for every supported format. It worked if the last line says Successfully installed.

Step 4: Check it

markitdown --list-plugins

If the command runs and prints plugin information (the list may be empty) rather than "command not found", MarkItDown is ready.

Alternative: install only the formats you need

For a leaner install, replace [all] with a list of formats. For example, PDF, Word and PowerPoint only:

pip install "markitdown[pdf, docx, pptx]"

Available groups: pptx, docx, xlsx, xls, pdf, outlook, az-doc-intel, az-content-understanding, audio-transcription, youtube-transcription.

Alternative: run it in Docker

For people already comfortable with Docker (software that runs apps in sealed containers). Download the repo's source code, then run this inside that folder:

docker build -t markitdown:latest .

Then convert a file (written for macOS and Linux):

docker run --rm -i markitdown:latest < ~/your-file.pdf > output.md

4. First use

Step 1: Convert a PDF

Put a PDF, say report.pdf, in the folder you are in and type:

markitdown report.pdf -o report.md

The -o option names the output file. If the command finishes without printing an error, it succeeded.

Step 2: Open the result

A file called report.md appears in the folder. Open it with Notepad or any text editor. You will see the document as text: headings start with #, tables are drawn with |. Copy the whole thing and paste it into your AI.

Step 3: Try another file type

The same command works for Word, Excel and PowerPoint:

markitdown sheet.xlsx -o sheet.md

Besides PDF and Office files, MarkItDown accepts images, audio, HTML, CSV, JSON, XML, ZIP files (it walks through the contents), EPUB and YouTube links.

Note: The official documentation also shows markitdown file.pdf > document.md. On Windows, prefer -o as above so the output file does not end up with a broken text encoding.

For people who write Python

from markitdown import MarkItDown

md = MarkItDown(enable_plugins=False)
result = md.convert("test.xlsx")
print(result.markdown)

Optional: read text inside images (OCR)

The markitdown-ocr plugin extracts text from images embedded in PDF, Word, PowerPoint and Excel files using an AI model that can see images. Install it with:

pip install markitdown-ocr
pip install openai

The plugin only does its job when you write Python code and pass in an llm_client and an llm_model. The details are at https://github.com/microsoft/markitdown/tree/main/packages/markitdown-ocr

5. Common errors and fixes

SymptomCauseFix
markitdown "is not recognized" or "command not found"The virtual environment is not active, or the install did not finish.Run the activate command from Step 2 and try again. If it still fails, repeat the install in Step 3.
The install complains about the Python versionMarkItDown supports Python 3.10 through 3.14 only.Check with python --version and install a Python release in that range.
Converting a PDF or Word file reports a missing libraryYou did a lean install without the matching format group.Reinstall with the group you need, for example pip install "markitdown[pdf]", or use [all].
On macOS/Linux the install says no matches found: markitdown[all]The quotes around the package name are missing, so the shell misreads the square brackets.Type pip install 'markitdown[all]' with the quotes.
A plugin is installed but has no effectPlugins are off by default.Add --use-plugins to the command, or set enable_plugins=True in Python.
markitdown-ocr is installed but text in images is still not readNo llm_client was provided; the plugin loads but silently skips OCR.Pass both llm_client and llm_model when creating MarkItDown(...).
You used uv for the environment and packages land in the wrong placeIn a uv environment you must use uv pip install.Replace pip install with uv pip install.
FileConversionExceptionNo converter could handle the file. For images, this only appears after the AI service has failed and no other converter could read the file either.Check that the file is not corrupted. If you use image descriptions, raise max_retries on the OpenAI client (the default is 2).

6. Uninstalling and updating

Update (activate the virtual environment first):

pip install -U "markitdown[all]"

On macOS and Linux use single quotes instead of double quotes. The latest release at the time of writing is 0.1.8, published on 21 September 2026.

Uninstall:

pip uninstall markitdown

If you created a virtual environment in Step 2, the cleanest removal is to delete the .venv folder: everything MarkItDown installed lives inside it. The .md files you produced are not touched.

7. Frequently asked questions

Does MarkItDown cost anything? No. The tool and its built-in converters run on your computer for free. Charges only arise if you switch on Azure Document Intelligence, Azure Content Understanding, or an AI model for image descriptions, and those come from the services themselves.

Are my documents sent anywhere? According to the project's comparison table, the built-in converters work offline and use only your own computer. A file leaves your machine only when you deliberately enable a cloud service or pass in an AI model. YouTube links and web addresses naturally need the internet to fetch their content.

Do I need an internet connection? For installing, yes. After that, converting local files does not need one.

What about a scanned PDF? The built-in converter extracts text that already exists in the file. For scanned pages you need the markitdown-ocr plugin with an AI model, or an Azure service.

Is there a version with a graphical interface? This project provides a Python library and a command-line tool only. The maintainers state that they do not accept web apps or desktop apps into the repository; any you find are separate products made by others.

Written by GitHot and checked against the project's official documentation. If a step is wrong or unclear, message GitHot through the channels at the bottom of the page.

SEE WHY IT'S HOT

Watch the original review

Go back to the video that brought you here.

Watch on TikTok ↗

Don't miss the next repo

Every new repo gets a video review, with its install guide right here.

Follow on TikTok ↗