markdowned-datasheets/PROCESS.md

568 lines
25 KiB
Markdown
Raw Normal View History

2026-08-22 15:33:55 +02:00
# Datasheet → Markdown: process
A repeatable runbook for turning a large vendor datasheet PDF into per-chapter
Markdown with extracted figures, a table of contents, and an accuracy audit.
Worked examples in this repo:
- Renesas RAA489400, 80 pages → 9 chapters, ~200 KB of Markdown, 25 extracted
figures.
- USB Power Delivery Specification r3.2 v1.2, 416 pages → 10 chapters + 2
appendices, ~1.1 MB of Markdown, 178 extracted figures. This one stress-tests
the process at 5x the page count and surfaced the gaps folded into this
document — see the caption-position note in step 4, the seam-defect table in
3c, and the audit-scaling notes in step 7.
The approach is: **split the PDF, transcribe chapters in parallel with agents
that read the rendered pages, extract figures mechanically, then audit each
chapter against the PDF with a second, independent set of agents.**
---
## 0. Why it is shaped this way
Two constraints drive the whole design.
**Text extraction is not enough.** `pdftotext` loses table structure, and
`pdfimages` finds nothing at all in a FrameMaker-authored datasheet because the
figures are vector line art, not embedded rasters. So transcription agents must
read the *rendered pages* as images, and figures must be rendered-and-cropped
rather than extracted.
**An agent cannot hold a 45-page chapter.** Context runs out silently — you get
a completed task with no file written. Large chapters must be split into
~8-page chunks processed in parallel, which then introduces seam defects that
have to be cleaned up afterwards. Budget for this.
---
## 1. Prerequisites
Everything lives in the flake devshell:
```bash
nix develop
```
Provides `qpdf`, `poppler-utils` (`pdftoppm`, `pdftotext`, `pdfinfo`,
`pdfimages`), `mupdf` (`mutool`), `imagemagick` (`magick`), and `python3`.
> If `nix develop` errors with *"Path 'flake.nix' … is not tracked by Git"*,
> either `git add flake.nix flake.lock`, or use `nix develop "path:$PWD"` to
> bypass the git-tracking requirement.
---
## 2. Split the source PDF into chapters
Find the chapter boundaries from the PDF bookmarks or the printed contents
page, then cut losslessly with `qpdf`:
```bash
pdfinfo REN_RAA489400_DST_20240827.pdf # page count
pdftotext -q REN_RAA489400_DST_20240827.pdf - | grep -n "^[0-9]\+\. "
mkdir -p chapters
qpdf source.pdf --pages . 5-10 -- chapters/01-overview.pdf
qpdf source.pdf --pages . 11-13 -- chapters/02-pin-information.pdf
# ...
```
Name files `NN-slug.pdf`. Everything downstream keys off that stem: the
transcription becomes `NN-slug.md` and figures go in `NN-slug/`.
Record the **printed** page range of each chapter — the number in the page
footer, not the position in the file. You need it in step 5.
```bash
for f in chapters/*.pdf; do
t=$(pdftotext -q "$f" - | grep -oE "Page [0-9]+" | sed "s/Page //")
echo "$(basename "$f") : $(echo $t | tr ' ' '\n' | head -1)-$(echo $t | tr ' ' '\n' | tail -1)"
done
```
---
## 3. Transcribe each chapter with a parallel agent
One agent per chapter PDF, all launched **in a single message** so they run
concurrently. Chapters over ~10 pages need chunking first — see step 3b.
### The prompt template
```
Convert a datasheet chapter PDF into Markdown.
Source: /abs/path/chapters/NN-slug.pdf
Output: /abs/path/chapters/NN-slug.md
Steps:
1. Use the native `read` tool with filePath=<source>. The read tool returns
PDFs as attachments so you can see the rendered pages directly. DO NOT use
pdftotext, python, or any bash-based extraction - use the read tool.
2. Transcribe the full contents of every page into GitHub-flavored Markdown,
then write it with the `write` tool to the output path.
Transcription rules:
- Transcribe all body text VERBATIM. Do NOT summarize, paraphrase, or omit.
- Preserve the heading hierarchy with #/##/###, keeping printed section numbers.
- Reproduce ALL tables as GFM tables. Every row and column, including units and
conditions. Do not truncate. GFM has no rowspan: where the source merges cells
vertically, repeat the value on each row.
- Preserve footnotes as a numbered list, keeping the markers in the cells.
- Transcribe printed errors, typos and inconsistencies EXACTLY as printed. Do
not silently correct them. They are load-bearing for an audit.
- Equations: use LaTeX in $...$, keeping printed equation numbers.
- Use inline code for register names, pin names, bit fields and hex values.
- Add a `<!-- page N -->` comment at each page boundary, placed BEFORE that
page's content, using the PRINTED page number. This chapter covers pages A-B.
FIGURES - IMPORTANT:
- For each figure, emit ONLY the printed caption, as bold text, exactly as it
appears (e.g. `**Figure 12. Startup from VBUS**`), plus an image link of the
form `![<printed caption>](NN-slug/figure-<printed number>.png)`.
- Do NOT write any description of what the figure shows. Do not describe axes,
curve labels, block names, signal paths, colours or annotations. The images
are extracted separately and the description is not in the datasheet.
- If a figure has no printed number or caption, say so plainly and do NOT
invent one.
Reply with: output path, pages transcribed, number of tables, number of figures.
```
### Why the figure rule matters
The first run of this project asked agents to describe each figure "for
accessibility". Every such description was invented text indistinguishable from
datasheet prose, and three of them were **factually wrong**:
- a waveform described as settling high when it actually falls back to baseline
(the exact behaviour the plot's measurement arrow was annotating);
- a 32-pin QFN described as "8 pins per side" when it is 11/5/11/5;
- an invented sentence about "RC filter components near the controller" for a
drawing that contains no such components.
In one chapter the invented prose was ~90% of the file. Removing it later cost
a full audit pass. **Do not ask for figure descriptions.**
### 3b. Chunking large chapters
An agent will silently fail on a big chapter — it reports success and writes
nothing. Pre-split anything over ~10 pages:
```bash
qpdf chapters/07-registers.pdf --pages . 1-8 -- /tmp/reg/07-1-8.pdf
qpdf chapters/07-registers.pdf --pages . 9-16 -- /tmp/reg/07-9-16.pdf
# ...
```
Give each chunk agent the same prompt plus:
```
IMPORTANT: this is chunk N of M and will be concatenated. Do NOT add a document
title, intro, or closing summary. Start directly with the content of page 1 of
this chunk. Do not write commentary into the file. If the chunk starts or ends
mid-table or mid-register, transcribe the partial content; do not invent the
missing part. This chunk covers printed pages A-B.
```
Then concatenate in order:
```bash
cat /tmp/reg/07-part{1,2,3,4,5,6}.md > chapters/07-registers.md
```
### 3c. Clean up the chunk seams — always required
Parallel chunks drift. Every one of these occurred in this project:
| Seam defect | Detection | Fix |
|---|---|---|
| Duplicate headings (source heading *and* the agent's own) | `rg "^#{3} " file.md` and look for adjacent pairs naming the same thing | Collapse to one; demote genuine sub-headings to `####` |
| Heading levels differ between chunks | compare `^#+ ` counts per chunk range | Normalize with a scripted pass — see below |
| A chunk emits valid heading text with no leading `#` at all | `rg '^[0-9]+\.[0-9]+(\.[0-9]+){0,4}\. [A-Z]' file.md \| rg -v '^#'` — matches printed-numbered lines that aren't markdown headings | Prepend the correct number of `#` (see the level formula below) |
| A table caption is written as a heading instead of bold text (`##### Table 3.7. ...` instead of `**Table 3.7. ...**`) | `rg -n '^#+ Table ' file.md` | `sed -E 's/^#+ (Table [0-9]+\.[0-9]+\. .+)$/**\1**/'`, then re-check table row/figure counts are unchanged |
| Derived metadata lines in 3 different formats (`\|` vs `·` vs `—`) | `rg -o "^Address:.*" \| sort \| uniq -c` | Pick one canonical form, rewrite |
| Some chunks bold table captions, others don't | `rg -c '^\*\*Table' ; rg -c '^Table '` | Bold them all |
| Address style `0x50` in one chunk, `50h` in another | `rg "0x[0-9A-F]" file.md` | Normalize to the printed style |
| RFC2119 keywords (Shall/Should/May/Shall Not/...) bold in some chunks, plain in others | count `\bShall\b` vs `\*\*Shall\*\*` per file | See "Fixing missing bold emphasis" below — this one has sharp edges |
### Normalizing heading levels from printed section numbers
Given a heading whose text starts with a printed number (`9.2.5. Title`), the
correct level is `(number of dots) + 1`: `9.1``##`, `9.2.5``###`,
`9.2.5.1``####`. `Chapter N.` / `Appendix X.` headings are always `#`.
```python
def desired_level(text):
if re.match(r'^(Chapter \d+|Appendix [A-Z])\.', text):
return 1
m = re.match(r'^(\d+(?:\.\d+)*)\.?\s', text)
return m and m.group(1).count('.') + 1
```
**Sharp edge:** if you rewrite a heading line with a regex substitution, make
sure the replacement keeps the line's trailing newline. A pattern like
`^(#{1,6})\s+(.*)$` followed by `f"{new_hashes} {text}"` silently drops the
`\n` for every line that actually changes, which merges that heading into the
following paragraph on the same line — invisible in a diff of line *counts*,
very visible once you view the file. Capture the newline explicitly and put it
back:
```python
heading_re = re.compile(r'^(#{1,6})[ \t]+(.*?)([\r\n]*)$')
# ... hashes, text, eol = heading_re.match(line).groups()
# ... out.append(f"{new_hashes} {text}{eol}")
```
Run the script twice — a correct normalizer is idempotent (reports 0 changes
on the second pass). If it isn't, something upstream (the level formula, the
newline handling) is still wrong.
### Fixing missing bold emphasis
Before trusting the text layer to tell you whether a word is bold, check a
rendered crop. `mutool draw -F stext`'s per-line `<font>` attribute is not
reliable for detecting an inline bold sub-run in the middle of a line — a
sentence like "...capacitance change **Shall** occur..." can report a single
non-bold font for the whole line even when "Shall" is visibly bold on the
page. Render the page and crop the relevant lines to confirm before writing a
fix:
```bash
pdftoppm -png -r 200 -f <page> -l <page> -singlefile source.pdf /tmp/page
magick /tmp/page.png -crop <w>x<h>+<x>+<y> /tmp/crop.png # then view it
```
If it's genuinely missing, a scripted fix needs two more guards or it
corrupts the file:
1. **Don't regex-exclude only the exact wrapped form.** Excluding `**Shall**`
but not `**Shall Discard**` means the bare `Shall` inside an
already-bold multi-word phrase gets wrapped again — `**Shall Discard**`
`****Shall**** Discard**`. Split each line on `**` first and only touch the
even-indexed (outside-any-bold-span) segments:
```python
parts = line.split('**')
for i in range(0, len(parts), 2): # even = outside bold, odd = inside
parts[i] = pattern.sub(lambda m: f'**{m.group(0)}**', parts[i])
line = '**'.join(parts)
```
2. **Watch for bold spans that cross a line break** (`**Shall` at the end of
one line, `Not**` at the start of the next — valid Markdown, renders fine).
The even/odd split above is computed fresh per line, so a line that
*starts* mid-bold-span (because the previous line left it open) gets the
parity backwards. This is rare — find candidates first with
`awk '{n=gsub(/\*\*/,"**"); if (n%2!=0) print FNR}' file.md` (an odd count
on a line is expected *only* for these) — and handle them by hand rather
than generalizing the script further for a handful of cases.
3. **Don't apply it to keyword/glossary-definition tables.** If a chapter
defines what "Shall" *means* (a keywords/conventions table), the term is
usually printed in plain text even though the same word is bold everywhere
else as normative emphasis — check the specific table's rendering before
assuming the chapter-wide convention applies inside it too.
### After any scripted rewrite, prove you changed nothing real
```bash
rg '^\|' before.md > /tmp/a; rg '^\|' after.md > /tmp/b; diff /tmp/a /tmp/b
```
An empty diff means no table cell moved. Also compare counts of headings and
`**Figure` captions before/after. This applies to every fix in this section —
heading-level normalization, missing-bold fixes, and any other scripted pass.
---
## 4. Extract the figures
`pdfimages` will not work on vector art. Use the caption-anchored cropper:
```bash
python3 tools/extract-figures.py --dry-run chapters/*.pdf # inspect first
python3 tools/extract-figures.py chapters/*.pdf # write PNGs
```
**Before running at scale, check the template's caption convention.** Render
one page that has a figure on it and look at where the caption sits relative
to the drawing:
```bash
pdftoppm -png -r 150 -f <page> -l <page> -singlefile source.pdf /tmp/page
```
The tool defaults to the Renesas-style convention — caption **below** the
figure, sequential integer numbers (`Figure 12.`). If the datasheet instead
prints the caption **above** the figure (common in USB-IF/standards-body
documents) and/or numbers figures per-chapter (`Figure 4.1`, `4.2`, ...), pass:
```bash
python3 tools/extract-figures.py --caption-position above chapters/*.pdf
```
`--caption-position above` flips the band direction (bounded above by the
caption, below by the next body text or caption) and the caption regex
accepts `Figure \d+(\.\d+)?\.`, writing `figure-4.1.png` instead of
zero-padded integers. It also requires the caption line's font to be bold —
without that check, a plain prose sentence like "Figure 4.2 illustrates the
application of..." false-matches as a caption and produces a `SKIP: no ink
found in band` entry (or worse, a bogus crop) once you get past the first one.
It locates `Figure N.` captions in the text layer via `mutool draw -F stext`,
derives a band bounded by the caption and the nearest body-text line, renders
the page, tightens the band to the actual ink, and crops at 300 dpi into
`chapters/<stem>/figure-NN.png` (or `figure-N.M.png`) using the datasheet's
own figure numbers.
Always `--dry-run` first and sanity-check the reported band sizes, then eyeball
the output as a contact sheet:
```bash
magick montage $(find chapters -name 'figure-*.png' | sort -t- -k2 -n) \
-tile 5x -geometry 300x300+4+4 -background gray70 /tmp/contact.png
```
Three tuning constants in the script are page-layout specific and are the first
things to adjust for a different vendor's template:
- `HEADER_BOTTOM` (62 pt) — must clear the running head **and the horizontal
rule beneath it**, which otherwise gets caught as figure ink.
- `BODY_GAP` (5 pt) — `stext` line boxes clip descenders, so without clearance
the tail of the heading above bleeds into the crop.
- `PAD` (6 pt) — clamped to the band so it cannot reach into the caption.
Also note captions **wrap**: bound the next figure with the caption's *last*
line, not its first, or the wrapped remainder lands at the top of the following
crop.
For a figure the datasheet never captioned, crop it explicitly rather than
inventing a number — and record the invocation:
```bash
python3 tools/extract-figures.py \
--region '1:195,550,410,640:current-sensing-trace-routing.png' \
chapters/08-layout.pdf
```
### Delegate this whole step to a subagent for big documents
For a document with 50+ figures, run the dry-run → adjust → contact-sheet →
re-dry-run loop in a subagent rather than the main thread. Each contact sheet
is a sizeable image, and getting the band-detection constants (or the
caption-position convention) right typically takes several rounds — none of
that iteration needs to live in the orchestrating conversation, only the
outcome does.
Give the subagent:
- the list of chapter PDFs and the expected figure count per chapter (count
`^\*\*Figure` in each transcribed `.md` — the subagent should treat a
mismatch against this count as a bug to chase down, not a fact to report);
- explicit instructions to render one figure page first and report which
caption convention applies before running the tool at scale;
- permission to extend `tools/extract-figures.py` if the template doesn't fit
(new flag, not a fork), **with an explicit instruction to re-run the tool
against the existing worked-example chapter(s) afterwards** to confirm the
change didn't break the default convention;
- a requirement to flag anything that looks visually wrong in a contact sheet
rather than silently accept it, and to use `--region` (recording the exact
invocation) for any figure the automatic detection can't handle.
Have it report back: final figure counts per chapter (matching the expected
counts above), which figures (if any) needed `--region`, any tool changes
made and the backward-compatibility check result, and one or two contact
sheets for you to spot-check before moving on to the TOC/audit phases.
---
## 5. Normalize the page markers
Chunk agents tend to number pages from 1 locally instead of using the printed
number. Check, then offset each file by the difference:
```bash
for f in chapters/*.md; do
echo "$f $(rg -o '<!-- page ([0-9]+) -->' -r '$1' "$f" | head -1)-$(rg -o '<!-- page ([0-9]+) -->' -r '$1' "$f" | tail -1)"
done
# e.g. chapter 1 printed pages start at 5 but markers start at 1 -> add 4
perl -i -pe "s/<!-- page (\d+) -->/'<!-- page '.(\$1+4).' -->'/ge" chapters/01-overview.md
```
Markers must sit **before** the content of the page they name. Anything that
later slices the file by page depends on this — see the pitfall in step 7.
---
## 6. Build the table of contents
Generate anchors mechanically; do not hand-write them. GitHub's slug rule is:
lowercase, strip punctuation except hyphens/underscores, spaces to hyphens.
```python
def anchor(h):
s = h.strip().lower().replace("`", "")
s = re.sub(r"[^\w\s-]", "", s, flags=re.UNICODE)
return re.sub(r"\s+", "-", s.strip())
```
A useful TOC has: a chapter table (file, printed page range, source PDF link,
figure count), the numbered section tree, a figure index, and a section listing
**known defects in the source datasheet** so nobody later "fixes" a faithful
transcription. Cap the depth — a 63-register chapter should link to its section,
not enumerate every register.
Verify:
```bash
python3 tools/check-links.py chapters
```
This checks every relative link, every anchor against real headings, and flags
orphaned PNGs. Exit code is non-zero on failure, so it can gate a commit.
---
## 7. Audit: transcription vs PDF
**Do not skip this, and use fresh agents.** The agent that wrote a file is not
a reliable auditor of it.
One agent per chapter, launched in parallel, large chapters split the same way
as in step 3b — but the chunk size doesn't have to match. Transcription chunks
are sized for *writing* a large amount of new content (~8 pages); audit
chunks only need to *read and compare*, so they tolerate roughly 2x the page
count before running into the same silent-context-loss failure mode. On a
400+ page document, auditing in ~15-page chunks instead of ~8-page chunks
roughly halves the number of audit agents needed with no loss of quality.
```
You are auditing a Markdown transcription of a datasheet chapter against the
original PDF. This is a READ-ONLY audit. Do NOT edit, write, or create files.
PDF (ground truth): <path>.pdf
Markdown to audit: <path>.md
Method:
1. Use the native `read` tool on the PDF. It returns rendered pages as an
attachment. DO NOT use pdftotext/python/bash extraction - compare against
what is actually printed.
2. Use the `read` tool on the .md.
3. Compare page by page.
Check:
- Missing or extra sections; wrong heading text or numbers.
- Body text summarized or paraphrased rather than transcribed verbatim.
- Every table: all rows and columns present, every value correct digit by
digit including sign, decimal point and unit.
- Figure captions: correct number, exact caption text.
- Any hallucinated content in the MD that is not in the PDF. THIS IS THE MOST
IMPORTANT THING TO CATCH.
For each finding give: severity (CRITICAL = wrong/invented, MAJOR = missing,
MINOR = formatting), printed page number, what the PDF says, what the MD says.
If a category is clean, say so in one line. End with a one-line verdict:
ACCURATE / MINOR ISSUES / SIGNIFICANT ISSUES. Do not fix anything.
```
Add per-chapter emphasis: numeric accuracy for a specifications chapter, every
pin row for a pinout chapter, bit-level detail for a register chapter, field
order and ACK/NACK bit values for protocol frame diagrams.
Feed forward any suspected source defects and ask the auditor to confirm what
the PDF *actually* prints. This is how you separate a transcription bug from a
vendor typo. In this project every such flag turned out to be a genuine defect
in the datasheet: a threshold printed in mV with volt-magnitude values, a pin
listed as both `NC` and `GPIO2`, `TYPE_C_PRAMETER` vs `TYPE_C_PARAMETER`
between a heading and its own table caption.
### Prefer letting the audit agent navigate by marker, not pre-slicing
The simplest way to avoid the slicing pitfall below is to not slice at all:
point the audit agent at the *full* chapter `.md` and tell it which
`<!-- page N -->` markers bound its assigned range ("only inspect the portion
between `<!-- page 254 -->` and `<!-- page 269 -->`"). The agent can `read`
the whole file and navigate to the right markers itself. This trades a larger
`read` for eliminating an entire class of off-by-one bugs in a slicing script,
and for the sizes involved (a chapter `.md` is at most a few thousand lines)
the read is cheap.
Reach for scripted pre-slicing only when a downstream tool (not an agent)
needs an exact byte range — and if you do, remember markers precede their
page:
```python
start = marks[a] # NOT marks[a-1] + 1
end = marks[b + 1]
```
Getting this wrong shifts every slice one page and produces six confident,
identical, **false** "MAJOR: last page is missing" reports. Before believing any
missing-content finding, grep the real file:
```bash
rg -o "Table ([0-9]+)\." -r '$1' chapters/07-registers.md | sort -nu | tr '\n' ' '
rg -o "^### 7\.2\.([0-9]+) " -r '$1' chapters/07-registers.md | sort -n | tr '\n' ' '
```
Contiguous sequences mean nothing is missing.
---
## 8. Act on the findings
Triage before editing. In this project, **all datasheet data — every table,
value, register and pin — was accurate**; every real defect was in prose the
agents added around the figures. Expect the same shape.
Distinguish:
- **Invented content** → delete.
- **Derived content** (e.g. an address/reset summary computed from the register
map) → keep if useful, but normalize it and add a note saying it is derived,
not printed. Do not let it masquerade as transcription.
- **Faithful transcription of a source defect** → leave alone, record in the
TOC's known-issues section.
When stripping generated text, anchor the pattern tightly. Figure descriptions
were italic paragraphs *directly following an image line*; that shape can be
matched without touching legitimate printed `*Note:*` paragraphs elsewhere.
Verify afterwards that the printed ones survived:
```bash
echo "before: $(rg -c '^\*Note' backup.md) after: $(rg -c '^\*Note' file.md)"
```
---
## Checklist
```
[ ] Split PDF into chapters/NN-slug.pdf; record printed page ranges
[ ] Transcribe in parallel, one agent per chapter (NO figure descriptions)
[ ] Chunk chapters over ~10 pages; concatenate; clean the seams
[ ] If normalizing heading levels with a script, verify it's idempotent
(0 changes on a second run) and preserves trailing newlines
[ ] Verify table-row diff is empty after every scripted rewrite
[ ] Check the figure caption convention (above/below the figure) on one
rendered page before extracting at scale; consider a subagent for the
extract/tune/contact-sheet loop on documents with 50+ figures
[ ] Extract figures; --dry-run, then contact-sheet review
[ ] Normalize <!-- page N --> to printed numbers
[ ] Build TOC with generated anchors; record known source defects
[ ] tools/check-links.py passes
[ ] Audit every chapter with fresh agents (chunks can be ~2x transcription
chunk size, since audits only read+compare)
[ ] Verify "missing content" findings against the real file before acting
[ ] Delete invented content; normalize derived content; keep source defects
```
## Tools in this repo
| Tool | Purpose |
|---|---|
| [`tools/extract-figures.py`](tools/extract-figures.py) | Caption-anchored figure cropping from vector PDFs. `--caption-position {below,above}` (default `below`) and decimal figure numbers (`4.1`) support both the Renesas-style and USB-IF-style caption conventions |
| [`tools/check-links.py`](tools/check-links.py) | Link, anchor and orphan-image validation. Anchor generation replicates GitHub's `-1`/`-2` disambiguation for duplicate headings (e.g. two companies with near-identical names in a contributors list) |
| [`chapters/TOC.md`](chapters/TOC.md) | Worked example of the generated TOC |