DOCUMENT CONVERTER

PDF → Markdown

Back to tool
Guide

How to Extract PDF Headings to a Markdown Outline

Learn how PDF bookmarks and TOC entries become Markdown headings, and how to build a clean outline from a 100+ page manual for Obsidian and LLM tools.

2026-10-05 · 7 min read

A 200-page service manual is a well-structured document in PDF form: chapters, sections, subsections, and a clickable table of contents. Export it to Markdown without that structure and you get a wall of text that is hard to navigate in Obsidian, awkward as documentation source, and worse as input for an LLM pipeline. The good news is that most PDF headings do not have to be guessed — they are already encoded in the file. This article explains where that heading data lives, how it maps to Markdown heading levels, and how to clean up the result on a manual that is far too long to fix by hand.

Where PDF heading information actually lives

PDF has no native heading concept the way HTML does. Heading structure comes from one of three sources, and they differ a lot in reliability.

  • Bookmarks (the document outline). A nested tree of entries, each pointing to a destination on a page. Generated by Word, LaTeX, InDesign, and most authoring tools when the export keeps navigation enabled. This is the most reliable source.
  • Tagged structure. PDFs built with accessibility support contain a structure tree with elements such as H1–H6, paragraphs, and tables. When it exists it is more precise than bookmarks, because it reflects author intent rather than just navigation.
  • Visual styling only. Larger, bolder, or differently colored text with extra space above it. There is no semantic data here, so a converter has to infer headings from font size, weight, and position. Inference works reasonably well but misfires on captions, table cells, and code listings.

Many real-world PDFs have none of the first two. Scanned documents have no text layer at all until OCR runs, and files exported from a word processor with bookmark creation turned off land in the same bucket.

Quick check before you convert

Open the PDF in any reader that shows a bookmarks or outline panel. If the panel lists chapters, the structure is there and your conversion will be mostly mechanical. If the panel is empty, expect to rely on font-based detection and plan for a review pass. Some readers also expose a "Tagged PDF" flag in the document properties, which tells you whether a structure tree exists.

How bookmarks and TOC entries become Markdown headings

The mapping rule used by most converters is simple: bookmark nesting depth becomes heading level. Top-level entries get the highest level the converter uses, and each nested level increments by one.

Bookmark depthTypical role in a manualMarkdown level
1ChapterH2 (##)
2SectionH3 (###)
3SubsectionH4 (####)
4Sub-subsectionH5 (#####)

Whether the top level becomes H1 or H2 matters more than it sounds. In note tools such as Obsidian, H1 is usually reserved for the note title, and in static site generators the page title comes from front matter. If a manual's eight chapters all arrive as H1, the outline pane shows eight competing titles instead of one document with eight sections. A common convention is to map chapter-level bookmarks to H2 and shift everything below accordingly. The same logic applies whether you convert PDF to Markdown with a desktop tool or an online service.

Headings inferred from font size

With no bookmarks, converters cluster text by font size and treat the largest recurring sizes as heading candidates. The word recurring matters: a size used once is more likely a cover title or a pull quote than a section heading. Two failure modes show up repeatedly — large bold text inside table headers becoming headings, and monospaced code listings being caught by weight or size thresholds.

Bookmarks that carry no useful text

Some outlines are generated from page labels: "1", "2", "3", or "Page 14". Others use generic labels such as "Untitled" or "Section 1" for every entry. In those cases the bookmark tree tells you where headings are but not what they say, so the heading text has to come from visual classification on the page itself. Check the first few entries before trusting the whole tree.

Building a clean outline from a 100+ page manual

For a long manual, the goal is not a perfect first pass — it is a predictable sequence that leaves a small, checkable amount of manual work. Run it in this order.

  1. Convert once and inspect only the outline. Do not read the body text yet. Open the Markdown with an outline pane expanded to every level and look at the shape. A free PDF to Markdown converter is enough for this first look; if the source is a scan, run OCR first so a text layer exists.
  2. Decide what H1 means for this document. Single note or docs page: H1 is the title, body headings start at H2. One file per chapter: each file's chapter title is H1 and its sections start at H2.
  3. Delete front matter and running elements. Cover pages, copyright blocks, the printed table of contents, and repeating headers or footers often arrive as headings because they are styled and isolated. Remove them at heading level.
  4. Merge headings split by line breaks. PDF line breaks become paragraph breaks, so "Chapter 4" and "Hydraulics" can land as two separate headings. Searching for consecutive heading lines catches most of these quickly.
  5. Fix hierarchy jumps. An H2 followed directly by an H4 means a level was lost. Promote the lower heading or insert the missing one — TOC generators and outline panes assume the sequence is continuous.
  6. Regenerate a Markdown table of contents and compare it to the printed one. If the PDF has a printed TOC, it is a free checklist. Section names present in the printed TOC but missing from your generated one are the headings the converter dropped.
  7. Work chapter by chapter. Split the file at chapter boundaries and fix one at a time. It keeps each review session short and avoids reprocessing the entire document after every correction.

Rough scale for planning: a manual with eight chapters typically produces a handful of chapter-level headings and several dozen section-level headings. Reviewing that many headings is a manageable pass; reviewing every paragraph is not. If you want the fuller end-to-end process, including splitting and renaming, see this PDF to Markdown workflow.

Why heading hierarchy matters downstream

  • Notes and knowledge bases. Outline panes, folding, and breadcrumbs are all driven by heading levels, and in Obsidian headings also determine how a long file gets split into separate notes. A flat or broken hierarchy means no usable navigation. The details are covered in this guide to PDF to Obsidian Markdown.
  • Docs sites and static generators. Anchors, sidebar navigation, and right-hand TOCs are generated from headings. A skipped level produces either a missing entry or one indented in the wrong place.
  • LLM and RAG pipelines. Most chunkers split on headings so each chunk carries a heading path as context, such as "Chapter 4 > Hydraulics > Pressure relief valve". If headings are missing, chunks are cut on fixed token counts and lose the context that makes retrieval accurate.
  • Re-export and accessibility. Markdown that maps cleanly back to HTML or DOCX keeps its structure, and consistent levels make it far easier to add captions and accessible headings later.

FAQ

My PDF has no bookmarks. Can I still extract headings?

Yes, but you are relying on inference rather than extraction. Run OCR first if the document is a scan, then let the converter classify headings by font size and weight. Expect lower accuracy around tables, captions, and code blocks, and budget a review pass before you rely on the outline.

Should the first heading in my Markdown file be H1 or H2?

Reserve H1 for the document title and start body headings at H2 if the file is a single note or a page on a docs site. If you split a manual into one file per chapter, give each file its chapter title as H1 and start the sections inside it at H2. Consistency matters more than the specific choice — downstream tools read the relative levels.

Why does my converted file have headings inside tables and code blocks?

Font-based detection sees large or bold text and assumes it is a heading, and table headers and monospaced listings often match those rules. Demote those headings manually after conversion, and consider converting table-heavy sections as a separate pass so their text never enters the heading classifier.

Ready to try it? Convert your first PDF for free — 200 pages per month, no credit card required. More questions? Check the FAQ or the blog.