REVIEWED & READY TO INSTALL
MarkItDown
microsoft/markitdown
Microsoft's official Python tool for converting PDF, Word, Excel, PowerPoint, Outlook MSG, images, audio and YouTube transcripts into tidy Markdown for prompts, RAG pipelines and agents. One command for every format, install only the extras you need, plus an OCR plugin and optional Azure Document Intelligence.
STEP BY STEP
How to install MarkItDown
The step-by-step guide is right on this page, written for people who have never installed anything from GitHub.
Updated 9/30/2026SIGNUP BONUS
Sign up and get 5 xu to try the AI tools
Every new GitHot account gets 5 xu — enough to run the AI Repo Explainer, MCP Config Generator or Repo Health Check. Free, just an email or Google.
- GitHub stars
- 187,793
- Language
- Python
- Added
- 9/27/2026
How to install MarkItDown, step by step
Work from top to bottom. Each grey box is one command: copy the whole line, paste it into the terminal window and press Enter.
Never used a terminal? Read this first (1 minute)
A terminal is a window where you type commands for your computer. You do not need to understand them, just copy and paste them exactly.
- Windows: press the Windows key, type "PowerShell", then press Enter.
- macOS: press Command + Space, type "Terminal", then press Enter.
- Linux: press Ctrl + Alt + T.
- Copy a command from the guide, right-click inside the window you just opened to paste it (macOS: Command + V), then press Enter.
- Wait for the command to finish (the cursor blinks again on a new line) before running the next one.
Lines starting with # are notes; you do not need to copy them. If red error text appears, check the common errors section at the end of the guide.
1. What is MarkItDown?
MarkItDown is a free tool from Microsoft that turns files (PDF, Word, Excel, PowerPoint, images, audio, web pages and more) into Markdown text. Markdown is plain text with a handful of simple symbols for headings, lists and tables. AI assistants such as ChatGPT, Claude and Gemini read it very well, which is why MarkItDown is commonly used to tidy a document before pasting it into an AI.
To be direct: this is a command-line tool written in Python. It has no window and no buttons. Even so, basic use is a single command, so a non-technical person can follow along. The project also states that the output is meant for machines to read; if you need a faithful, good-looking conversion for human readers, it may not be the best choice.
It is released under the MIT licence and is completely free when it runs on your own machine. Only the optional features that call cloud services (Azure, or an AI model that describes images) cost money, billed by those providers.
A good fit if:
- You often hand PDF, Word, Excel or PowerPoint files to an AI and want clean content.
- You have many documents to extract text from in one go.
- You would rather not upload your documents to an online converter.
Not a good fit if:
- You need a conversion that keeps the original look, fonts and layout.
- You only use a phone.
- You need to read text in scanned images without any extra AI service (that feature needs a plugin plus an AI model).
2. What you need
- A computer running Windows, macOS or Linux. No graphics card required.
- Python 3.10 through 3.14: the programming language MarkItDown runs on. Download it from https://www.python.org/downloads/ . On Windows, tick "Add python.exe to PATH" during setup.
- Internet to download the tool when installing. Converting files already on your computer with the built-in converters happens locally.
- No account and no API key for basic use.
Things needed only for advanced features:
| Feature | You also need | Does it cost money? |
|---|---|---|
| Image descriptions, reading text in images (OCR) | An API key for an OpenAI-compatible AI service | Yes, at the provider's rates |
| Azure Document Intelligence | An Azure account and an endpoint address | Yes, billed per call |
| Azure Content Understanding | An Azure account and an endpoint address | Yes, every conversion is a billable call |
| Running in Docker | Docker Desktop from https://www.docker.com/products/docker-desktop/ | No |
3. Step-by-step installation
Step 1: Check Python
python --version
You should see something between Python 3.10 and Python 3.14. On macOS and Linux, if the command fails, try python3 --version.
Step 2: Create a separate environment (recommended)
A virtual environment is a private folder that holds MarkItDown's libraries so they do not interfere with other Python software. The project recommends using one.
Windows:
python -m venv .venv
.venv\Scripts\activate
macOS and Linux:
python -m venv .venv
source .venv/bin/activate
It worked if (.venv) appears at the start of your command line. Each time you open a new command-line window, run the second command (activate) again before using MarkItDown.
Step 3: Install MarkItDown
Windows:
pip install "markitdown[all]"
macOS and Linux:
pip install 'markitdown[all]'
The [all] part installs the readers for every supported format. It worked if the last line says Successfully installed.
Step 4: Check it
markitdown --list-plugins
If the command runs and prints plugin information (the list may be empty) rather than "command not found", MarkItDown is ready.
Alternative: install only the formats you need
For a leaner install, replace [all] with a list of formats. For example, PDF, Word and PowerPoint only:
pip install "markitdown[pdf, docx, pptx]"
Available groups: pptx, docx, xlsx, xls, pdf, outlook, az-doc-intel, az-content-understanding, audio-transcription, youtube-transcription.
Alternative: run it in Docker
For people already comfortable with Docker (software that runs apps in sealed containers). Download the repo's source code, then run this inside that folder:
docker build -t markitdown:latest .
Then convert a file (written for macOS and Linux):
docker run --rm -i markitdown:latest < ~/your-file.pdf > output.md
4. First use
Step 1: Convert a PDF
Put a PDF, say report.pdf, in the folder you are in and type:
markitdown report.pdf -o report.md
The -o option names the output file. If the command finishes without printing an error, it succeeded.
Step 2: Open the result
A file called report.md appears in the folder. Open it with Notepad or any text editor. You will see the document as text: headings start with #, tables are drawn with |. Copy the whole thing and paste it into your AI.
Step 3: Try another file type
The same command works for Word, Excel and PowerPoint:
markitdown sheet.xlsx -o sheet.md
Besides PDF and Office files, MarkItDown accepts images, audio, HTML, CSV, JSON, XML, ZIP files (it walks through the contents), EPUB and YouTube links.
Note: The official documentation also shows
markitdown file.pdf > document.md. On Windows, prefer-oas above so the output file does not end up with a broken text encoding.
For people who write Python
from markitdown import MarkItDown
md = MarkItDown(enable_plugins=False)
result = md.convert("test.xlsx")
print(result.markdown)
Optional: read text inside images (OCR)
The markitdown-ocr plugin extracts text from images embedded in PDF, Word, PowerPoint and Excel files using an AI model that can see images. Install it with:
pip install markitdown-ocr
pip install openai
The plugin only does its job when you write Python code and pass in an llm_client and an llm_model. The details are at https://github.com/microsoft/markitdown/tree/main/packages/markitdown-ocr
5. Common errors and fixes
| Symptom | Cause | Fix |
|---|---|---|
markitdown "is not recognized" or "command not found" | The virtual environment is not active, or the install did not finish. | Run the activate command from Step 2 and try again. If it still fails, repeat the install in Step 3. |
| The install complains about the Python version | MarkItDown supports Python 3.10 through 3.14 only. | Check with python --version and install a Python release in that range. |
| Converting a PDF or Word file reports a missing library | You did a lean install without the matching format group. | Reinstall with the group you need, for example pip install "markitdown[pdf]", or use [all]. |
On macOS/Linux the install says no matches found: markitdown[all] | The quotes around the package name are missing, so the shell misreads the square brackets. | Type pip install 'markitdown[all]' with the quotes. |
| A plugin is installed but has no effect | Plugins are off by default. | Add --use-plugins to the command, or set enable_plugins=True in Python. |
markitdown-ocr is installed but text in images is still not read | No llm_client was provided; the plugin loads but silently skips OCR. | Pass both llm_client and llm_model when creating MarkItDown(...). |
You used uv for the environment and packages land in the wrong place | In a uv environment you must use uv pip install. | Replace pip install with uv pip install. |
FileConversionException | No converter could handle the file. For images, this only appears after the AI service has failed and no other converter could read the file either. | Check that the file is not corrupted. If you use image descriptions, raise max_retries on the OpenAI client (the default is 2). |
6. Uninstalling and updating
Update (activate the virtual environment first):
pip install -U "markitdown[all]"
On macOS and Linux use single quotes instead of double quotes. The latest release at the time of writing is 0.1.8, published on 21 September 2026.
Uninstall:
pip uninstall markitdown
If you created a virtual environment in Step 2, the cleanest removal is to delete the .venv folder: everything MarkItDown installed lives inside it. The .md files you produced are not touched.
7. Frequently asked questions
Does MarkItDown cost anything? No. The tool and its built-in converters run on your computer for free. Charges only arise if you switch on Azure Document Intelligence, Azure Content Understanding, or an AI model for image descriptions, and those come from the services themselves.
Are my documents sent anywhere? According to the project's comparison table, the built-in converters work offline and use only your own computer. A file leaves your machine only when you deliberately enable a cloud service or pass in an AI model. YouTube links and web addresses naturally need the internet to fetch their content.
Do I need an internet connection? For installing, yes. After that, converting local files does not need one.
What about a scanned PDF?
The built-in converter extracts text that already exists in the file. For scanned pages you need the markitdown-ocr plugin with an AI model, or an Azure service.
Is there a version with a graphical interface? This project provides a Python library and a command-line tool only. The maintainers state that they do not accept web apps or desktop apps into the repository; any you find are separate products made by others.
Written by GitHot and checked against the project's official documentation. If a step is wrong or unclear, message GitHot through the channels at the bottom of the page.
SEE WHY IT'S HOT
Watch the original review
Go back to the video that brought you here.
Don't miss the next repo
Every new repo gets a video review, with its install guide right here.