init
This commit is contained in:
commit
74718f7fb5
7 changed files with 1176 additions and 0 deletions
567
PROCESS.md
Normal file
567
PROCESS.md
Normal file
|
|
@ -0,0 +1,567 @@
|
|||
# Datasheet → Markdown: process
|
||||
|
||||
A repeatable runbook for turning a large vendor datasheet PDF into per-chapter
|
||||
Markdown with extracted figures, a table of contents, and an accuracy audit.
|
||||
|
||||
Worked examples in this repo:
|
||||
- Renesas RAA489400, 80 pages → 9 chapters, ~200 KB of Markdown, 25 extracted
|
||||
figures.
|
||||
- USB Power Delivery Specification r3.2 v1.2, 416 pages → 10 chapters + 2
|
||||
appendices, ~1.1 MB of Markdown, 178 extracted figures. This one stress-tests
|
||||
the process at 5x the page count and surfaced the gaps folded into this
|
||||
document — see the caption-position note in step 4, the seam-defect table in
|
||||
3c, and the audit-scaling notes in step 7.
|
||||
|
||||
The approach is: **split the PDF, transcribe chapters in parallel with agents
|
||||
that read the rendered pages, extract figures mechanically, then audit each
|
||||
chapter against the PDF with a second, independent set of agents.**
|
||||
|
||||
---
|
||||
|
||||
## 0. Why it is shaped this way
|
||||
|
||||
Two constraints drive the whole design.
|
||||
|
||||
**Text extraction is not enough.** `pdftotext` loses table structure, and
|
||||
`pdfimages` finds nothing at all in a FrameMaker-authored datasheet because the
|
||||
figures are vector line art, not embedded rasters. So transcription agents must
|
||||
read the *rendered pages* as images, and figures must be rendered-and-cropped
|
||||
rather than extracted.
|
||||
|
||||
**An agent cannot hold a 45-page chapter.** Context runs out silently — you get
|
||||
a completed task with no file written. Large chapters must be split into
|
||||
~8-page chunks processed in parallel, which then introduces seam defects that
|
||||
have to be cleaned up afterwards. Budget for this.
|
||||
|
||||
---
|
||||
|
||||
## 1. Prerequisites
|
||||
|
||||
Everything lives in the flake devshell:
|
||||
|
||||
```bash
|
||||
nix develop
|
||||
```
|
||||
|
||||
Provides `qpdf`, `poppler-utils` (`pdftoppm`, `pdftotext`, `pdfinfo`,
|
||||
`pdfimages`), `mupdf` (`mutool`), `imagemagick` (`magick`), and `python3`.
|
||||
|
||||
> If `nix develop` errors with *"Path 'flake.nix' … is not tracked by Git"*,
|
||||
> either `git add flake.nix flake.lock`, or use `nix develop "path:$PWD"` to
|
||||
> bypass the git-tracking requirement.
|
||||
|
||||
---
|
||||
|
||||
## 2. Split the source PDF into chapters
|
||||
|
||||
Find the chapter boundaries from the PDF bookmarks or the printed contents
|
||||
page, then cut losslessly with `qpdf`:
|
||||
|
||||
```bash
|
||||
pdfinfo REN_RAA489400_DST_20240827.pdf # page count
|
||||
pdftotext -q REN_RAA489400_DST_20240827.pdf - | grep -n "^[0-9]\+\. "
|
||||
|
||||
mkdir -p chapters
|
||||
qpdf source.pdf --pages . 5-10 -- chapters/01-overview.pdf
|
||||
qpdf source.pdf --pages . 11-13 -- chapters/02-pin-information.pdf
|
||||
# ...
|
||||
```
|
||||
|
||||
Name files `NN-slug.pdf`. Everything downstream keys off that stem: the
|
||||
transcription becomes `NN-slug.md` and figures go in `NN-slug/`.
|
||||
|
||||
Record the **printed** page range of each chapter — the number in the page
|
||||
footer, not the position in the file. You need it in step 5.
|
||||
|
||||
```bash
|
||||
for f in chapters/*.pdf; do
|
||||
t=$(pdftotext -q "$f" - | grep -oE "Page [0-9]+" | sed "s/Page //")
|
||||
echo "$(basename "$f") : $(echo $t | tr ' ' '\n' | head -1)-$(echo $t | tr ' ' '\n' | tail -1)"
|
||||
done
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Transcribe each chapter with a parallel agent
|
||||
|
||||
One agent per chapter PDF, all launched **in a single message** so they run
|
||||
concurrently. Chapters over ~10 pages need chunking first — see step 3b.
|
||||
|
||||
### The prompt template
|
||||
|
||||
```
|
||||
Convert a datasheet chapter PDF into Markdown.
|
||||
|
||||
Source: /abs/path/chapters/NN-slug.pdf
|
||||
Output: /abs/path/chapters/NN-slug.md
|
||||
|
||||
Steps:
|
||||
1. Use the native `read` tool with filePath=<source>. The read tool returns
|
||||
PDFs as attachments so you can see the rendered pages directly. DO NOT use
|
||||
pdftotext, python, or any bash-based extraction - use the read tool.
|
||||
2. Transcribe the full contents of every page into GitHub-flavored Markdown,
|
||||
then write it with the `write` tool to the output path.
|
||||
|
||||
Transcription rules:
|
||||
- Transcribe all body text VERBATIM. Do NOT summarize, paraphrase, or omit.
|
||||
- Preserve the heading hierarchy with #/##/###, keeping printed section numbers.
|
||||
- Reproduce ALL tables as GFM tables. Every row and column, including units and
|
||||
conditions. Do not truncate. GFM has no rowspan: where the source merges cells
|
||||
vertically, repeat the value on each row.
|
||||
- Preserve footnotes as a numbered list, keeping the markers in the cells.
|
||||
- Transcribe printed errors, typos and inconsistencies EXACTLY as printed. Do
|
||||
not silently correct them. They are load-bearing for an audit.
|
||||
- Equations: use LaTeX in $...$, keeping printed equation numbers.
|
||||
- Use inline code for register names, pin names, bit fields and hex values.
|
||||
- Add a `<!-- page N -->` comment at each page boundary, placed BEFORE that
|
||||
page's content, using the PRINTED page number. This chapter covers pages A-B.
|
||||
|
||||
FIGURES - IMPORTANT:
|
||||
- For each figure, emit ONLY the printed caption, as bold text, exactly as it
|
||||
appears (e.g. `**Figure 12. Startup from VBUS**`), plus an image link of the
|
||||
form ``.
|
||||
- Do NOT write any description of what the figure shows. Do not describe axes,
|
||||
curve labels, block names, signal paths, colours or annotations. The images
|
||||
are extracted separately and the description is not in the datasheet.
|
||||
- If a figure has no printed number or caption, say so plainly and do NOT
|
||||
invent one.
|
||||
|
||||
Reply with: output path, pages transcribed, number of tables, number of figures.
|
||||
```
|
||||
|
||||
### Why the figure rule matters
|
||||
|
||||
The first run of this project asked agents to describe each figure "for
|
||||
accessibility". Every such description was invented text indistinguishable from
|
||||
datasheet prose, and three of them were **factually wrong**:
|
||||
|
||||
- a waveform described as settling high when it actually falls back to baseline
|
||||
(the exact behaviour the plot's measurement arrow was annotating);
|
||||
- a 32-pin QFN described as "8 pins per side" when it is 11/5/11/5;
|
||||
- an invented sentence about "RC filter components near the controller" for a
|
||||
drawing that contains no such components.
|
||||
|
||||
In one chapter the invented prose was ~90% of the file. Removing it later cost
|
||||
a full audit pass. **Do not ask for figure descriptions.**
|
||||
|
||||
### 3b. Chunking large chapters
|
||||
|
||||
An agent will silently fail on a big chapter — it reports success and writes
|
||||
nothing. Pre-split anything over ~10 pages:
|
||||
|
||||
```bash
|
||||
qpdf chapters/07-registers.pdf --pages . 1-8 -- /tmp/reg/07-1-8.pdf
|
||||
qpdf chapters/07-registers.pdf --pages . 9-16 -- /tmp/reg/07-9-16.pdf
|
||||
# ...
|
||||
```
|
||||
|
||||
Give each chunk agent the same prompt plus:
|
||||
|
||||
```
|
||||
IMPORTANT: this is chunk N of M and will be concatenated. Do NOT add a document
|
||||
title, intro, or closing summary. Start directly with the content of page 1 of
|
||||
this chunk. Do not write commentary into the file. If the chunk starts or ends
|
||||
mid-table or mid-register, transcribe the partial content; do not invent the
|
||||
missing part. This chunk covers printed pages A-B.
|
||||
```
|
||||
|
||||
Then concatenate in order:
|
||||
|
||||
```bash
|
||||
cat /tmp/reg/07-part{1,2,3,4,5,6}.md > chapters/07-registers.md
|
||||
```
|
||||
|
||||
### 3c. Clean up the chunk seams — always required
|
||||
|
||||
Parallel chunks drift. Every one of these occurred in this project:
|
||||
|
||||
| Seam defect | Detection | Fix |
|
||||
|---|---|---|
|
||||
| Duplicate headings (source heading *and* the agent's own) | `rg "^#{3} " file.md` and look for adjacent pairs naming the same thing | Collapse to one; demote genuine sub-headings to `####` |
|
||||
| Heading levels differ between chunks | compare `^#+ ` counts per chunk range | Normalize with a scripted pass — see below |
|
||||
| A chunk emits valid heading text with no leading `#` at all | `rg '^[0-9]+\.[0-9]+(\.[0-9]+){0,4}\. [A-Z]' file.md \| rg -v '^#'` — matches printed-numbered lines that aren't markdown headings | Prepend the correct number of `#` (see the level formula below) |
|
||||
| A table caption is written as a heading instead of bold text (`##### Table 3.7. ...` instead of `**Table 3.7. ...**`) | `rg -n '^#+ Table ' file.md` | `sed -E 's/^#+ (Table [0-9]+\.[0-9]+\. .+)$/**\1**/'`, then re-check table row/figure counts are unchanged |
|
||||
| Derived metadata lines in 3 different formats (`\|` vs `·` vs `—`) | `rg -o "^Address:.*" \| sort \| uniq -c` | Pick one canonical form, rewrite |
|
||||
| Some chunks bold table captions, others don't | `rg -c '^\*\*Table' ; rg -c '^Table '` | Bold them all |
|
||||
| Address style `0x50` in one chunk, `50h` in another | `rg "0x[0-9A-F]" file.md` | Normalize to the printed style |
|
||||
| RFC2119 keywords (Shall/Should/May/Shall Not/...) bold in some chunks, plain in others | count `\bShall\b` vs `\*\*Shall\*\*` per file | See "Fixing missing bold emphasis" below — this one has sharp edges |
|
||||
|
||||
### Normalizing heading levels from printed section numbers
|
||||
|
||||
Given a heading whose text starts with a printed number (`9.2.5. Title`), the
|
||||
correct level is `(number of dots) + 1`: `9.1` → `##`, `9.2.5` → `###`,
|
||||
`9.2.5.1` → `####`. `Chapter N.` / `Appendix X.` headings are always `#`.
|
||||
|
||||
```python
|
||||
def desired_level(text):
|
||||
if re.match(r'^(Chapter \d+|Appendix [A-Z])\.', text):
|
||||
return 1
|
||||
m = re.match(r'^(\d+(?:\.\d+)*)\.?\s', text)
|
||||
return m and m.group(1).count('.') + 1
|
||||
```
|
||||
|
||||
**Sharp edge:** if you rewrite a heading line with a regex substitution, make
|
||||
sure the replacement keeps the line's trailing newline. A pattern like
|
||||
`^(#{1,6})\s+(.*)$` followed by `f"{new_hashes} {text}"` silently drops the
|
||||
`\n` for every line that actually changes, which merges that heading into the
|
||||
following paragraph on the same line — invisible in a diff of line *counts*,
|
||||
very visible once you view the file. Capture the newline explicitly and put it
|
||||
back:
|
||||
|
||||
```python
|
||||
heading_re = re.compile(r'^(#{1,6})[ \t]+(.*?)([\r\n]*)$')
|
||||
# ... hashes, text, eol = heading_re.match(line).groups()
|
||||
# ... out.append(f"{new_hashes} {text}{eol}")
|
||||
```
|
||||
|
||||
Run the script twice — a correct normalizer is idempotent (reports 0 changes
|
||||
on the second pass). If it isn't, something upstream (the level formula, the
|
||||
newline handling) is still wrong.
|
||||
|
||||
### Fixing missing bold emphasis
|
||||
|
||||
Before trusting the text layer to tell you whether a word is bold, check a
|
||||
rendered crop. `mutool draw -F stext`'s per-line `<font>` attribute is not
|
||||
reliable for detecting an inline bold sub-run in the middle of a line — a
|
||||
sentence like "...capacitance change **Shall** occur..." can report a single
|
||||
non-bold font for the whole line even when "Shall" is visibly bold on the
|
||||
page. Render the page and crop the relevant lines to confirm before writing a
|
||||
fix:
|
||||
|
||||
```bash
|
||||
pdftoppm -png -r 200 -f <page> -l <page> -singlefile source.pdf /tmp/page
|
||||
magick /tmp/page.png -crop <w>x<h>+<x>+<y> /tmp/crop.png # then view it
|
||||
```
|
||||
|
||||
If it's genuinely missing, a scripted fix needs two more guards or it
|
||||
corrupts the file:
|
||||
|
||||
1. **Don't regex-exclude only the exact wrapped form.** Excluding `**Shall**`
|
||||
but not `**Shall Discard**` means the bare `Shall` inside an
|
||||
already-bold multi-word phrase gets wrapped again — `**Shall Discard**` →
|
||||
`****Shall**** Discard**`. Split each line on `**` first and only touch the
|
||||
even-indexed (outside-any-bold-span) segments:
|
||||
|
||||
```python
|
||||
parts = line.split('**')
|
||||
for i in range(0, len(parts), 2): # even = outside bold, odd = inside
|
||||
parts[i] = pattern.sub(lambda m: f'**{m.group(0)}**', parts[i])
|
||||
line = '**'.join(parts)
|
||||
```
|
||||
|
||||
2. **Watch for bold spans that cross a line break** (`**Shall` at the end of
|
||||
one line, `Not**` at the start of the next — valid Markdown, renders fine).
|
||||
The even/odd split above is computed fresh per line, so a line that
|
||||
*starts* mid-bold-span (because the previous line left it open) gets the
|
||||
parity backwards. This is rare — find candidates first with
|
||||
`awk '{n=gsub(/\*\*/,"**"); if (n%2!=0) print FNR}' file.md` (an odd count
|
||||
on a line is expected *only* for these) — and handle them by hand rather
|
||||
than generalizing the script further for a handful of cases.
|
||||
|
||||
3. **Don't apply it to keyword/glossary-definition tables.** If a chapter
|
||||
defines what "Shall" *means* (a keywords/conventions table), the term is
|
||||
usually printed in plain text even though the same word is bold everywhere
|
||||
else as normative emphasis — check the specific table's rendering before
|
||||
assuming the chapter-wide convention applies inside it too.
|
||||
|
||||
### After any scripted rewrite, prove you changed nothing real
|
||||
|
||||
```bash
|
||||
rg '^\|' before.md > /tmp/a; rg '^\|' after.md > /tmp/b; diff /tmp/a /tmp/b
|
||||
```
|
||||
|
||||
An empty diff means no table cell moved. Also compare counts of headings and
|
||||
`**Figure` captions before/after. This applies to every fix in this section —
|
||||
heading-level normalization, missing-bold fixes, and any other scripted pass.
|
||||
|
||||
---
|
||||
|
||||
## 4. Extract the figures
|
||||
|
||||
`pdfimages` will not work on vector art. Use the caption-anchored cropper:
|
||||
|
||||
```bash
|
||||
python3 tools/extract-figures.py --dry-run chapters/*.pdf # inspect first
|
||||
python3 tools/extract-figures.py chapters/*.pdf # write PNGs
|
||||
```
|
||||
|
||||
**Before running at scale, check the template's caption convention.** Render
|
||||
one page that has a figure on it and look at where the caption sits relative
|
||||
to the drawing:
|
||||
|
||||
```bash
|
||||
pdftoppm -png -r 150 -f <page> -l <page> -singlefile source.pdf /tmp/page
|
||||
```
|
||||
|
||||
The tool defaults to the Renesas-style convention — caption **below** the
|
||||
figure, sequential integer numbers (`Figure 12.`). If the datasheet instead
|
||||
prints the caption **above** the figure (common in USB-IF/standards-body
|
||||
documents) and/or numbers figures per-chapter (`Figure 4.1`, `4.2`, ...), pass:
|
||||
|
||||
```bash
|
||||
python3 tools/extract-figures.py --caption-position above chapters/*.pdf
|
||||
```
|
||||
|
||||
`--caption-position above` flips the band direction (bounded above by the
|
||||
caption, below by the next body text or caption) and the caption regex
|
||||
accepts `Figure \d+(\.\d+)?\.`, writing `figure-4.1.png` instead of
|
||||
zero-padded integers. It also requires the caption line's font to be bold —
|
||||
without that check, a plain prose sentence like "Figure 4.2 illustrates the
|
||||
application of..." false-matches as a caption and produces a `SKIP: no ink
|
||||
found in band` entry (or worse, a bogus crop) once you get past the first one.
|
||||
|
||||
It locates `Figure N.` captions in the text layer via `mutool draw -F stext`,
|
||||
derives a band bounded by the caption and the nearest body-text line, renders
|
||||
the page, tightens the band to the actual ink, and crops at 300 dpi into
|
||||
`chapters/<stem>/figure-NN.png` (or `figure-N.M.png`) using the datasheet's
|
||||
own figure numbers.
|
||||
|
||||
Always `--dry-run` first and sanity-check the reported band sizes, then eyeball
|
||||
the output as a contact sheet:
|
||||
|
||||
```bash
|
||||
magick montage $(find chapters -name 'figure-*.png' | sort -t- -k2 -n) \
|
||||
-tile 5x -geometry 300x300+4+4 -background gray70 /tmp/contact.png
|
||||
```
|
||||
|
||||
Three tuning constants in the script are page-layout specific and are the first
|
||||
things to adjust for a different vendor's template:
|
||||
|
||||
- `HEADER_BOTTOM` (62 pt) — must clear the running head **and the horizontal
|
||||
rule beneath it**, which otherwise gets caught as figure ink.
|
||||
- `BODY_GAP` (5 pt) — `stext` line boxes clip descenders, so without clearance
|
||||
the tail of the heading above bleeds into the crop.
|
||||
- `PAD` (6 pt) — clamped to the band so it cannot reach into the caption.
|
||||
|
||||
Also note captions **wrap**: bound the next figure with the caption's *last*
|
||||
line, not its first, or the wrapped remainder lands at the top of the following
|
||||
crop.
|
||||
|
||||
For a figure the datasheet never captioned, crop it explicitly rather than
|
||||
inventing a number — and record the invocation:
|
||||
|
||||
```bash
|
||||
python3 tools/extract-figures.py \
|
||||
--region '1:195,550,410,640:current-sensing-trace-routing.png' \
|
||||
chapters/08-layout.pdf
|
||||
```
|
||||
|
||||
### Delegate this whole step to a subagent for big documents
|
||||
|
||||
For a document with 50+ figures, run the dry-run → adjust → contact-sheet →
|
||||
re-dry-run loop in a subagent rather than the main thread. Each contact sheet
|
||||
is a sizeable image, and getting the band-detection constants (or the
|
||||
caption-position convention) right typically takes several rounds — none of
|
||||
that iteration needs to live in the orchestrating conversation, only the
|
||||
outcome does.
|
||||
|
||||
Give the subagent:
|
||||
- the list of chapter PDFs and the expected figure count per chapter (count
|
||||
`^\*\*Figure` in each transcribed `.md` — the subagent should treat a
|
||||
mismatch against this count as a bug to chase down, not a fact to report);
|
||||
- explicit instructions to render one figure page first and report which
|
||||
caption convention applies before running the tool at scale;
|
||||
- permission to extend `tools/extract-figures.py` if the template doesn't fit
|
||||
(new flag, not a fork), **with an explicit instruction to re-run the tool
|
||||
against the existing worked-example chapter(s) afterwards** to confirm the
|
||||
change didn't break the default convention;
|
||||
- a requirement to flag anything that looks visually wrong in a contact sheet
|
||||
rather than silently accept it, and to use `--region` (recording the exact
|
||||
invocation) for any figure the automatic detection can't handle.
|
||||
|
||||
Have it report back: final figure counts per chapter (matching the expected
|
||||
counts above), which figures (if any) needed `--region`, any tool changes
|
||||
made and the backward-compatibility check result, and one or two contact
|
||||
sheets for you to spot-check before moving on to the TOC/audit phases.
|
||||
|
||||
---
|
||||
|
||||
## 5. Normalize the page markers
|
||||
|
||||
Chunk agents tend to number pages from 1 locally instead of using the printed
|
||||
number. Check, then offset each file by the difference:
|
||||
|
||||
```bash
|
||||
for f in chapters/*.md; do
|
||||
echo "$f $(rg -o '<!-- page ([0-9]+) -->' -r '$1' "$f" | head -1)-$(rg -o '<!-- page ([0-9]+) -->' -r '$1' "$f" | tail -1)"
|
||||
done
|
||||
|
||||
# e.g. chapter 1 printed pages start at 5 but markers start at 1 -> add 4
|
||||
perl -i -pe "s/<!-- page (\d+) -->/'<!-- page '.(\$1+4).' -->'/ge" chapters/01-overview.md
|
||||
```
|
||||
|
||||
Markers must sit **before** the content of the page they name. Anything that
|
||||
later slices the file by page depends on this — see the pitfall in step 7.
|
||||
|
||||
---
|
||||
|
||||
## 6. Build the table of contents
|
||||
|
||||
Generate anchors mechanically; do not hand-write them. GitHub's slug rule is:
|
||||
lowercase, strip punctuation except hyphens/underscores, spaces to hyphens.
|
||||
|
||||
```python
|
||||
def anchor(h):
|
||||
s = h.strip().lower().replace("`", "")
|
||||
s = re.sub(r"[^\w\s-]", "", s, flags=re.UNICODE)
|
||||
return re.sub(r"\s+", "-", s.strip())
|
||||
```
|
||||
|
||||
A useful TOC has: a chapter table (file, printed page range, source PDF link,
|
||||
figure count), the numbered section tree, a figure index, and a section listing
|
||||
**known defects in the source datasheet** so nobody later "fixes" a faithful
|
||||
transcription. Cap the depth — a 63-register chapter should link to its section,
|
||||
not enumerate every register.
|
||||
|
||||
Verify:
|
||||
|
||||
```bash
|
||||
python3 tools/check-links.py chapters
|
||||
```
|
||||
|
||||
This checks every relative link, every anchor against real headings, and flags
|
||||
orphaned PNGs. Exit code is non-zero on failure, so it can gate a commit.
|
||||
|
||||
---
|
||||
|
||||
## 7. Audit: transcription vs PDF
|
||||
|
||||
**Do not skip this, and use fresh agents.** The agent that wrote a file is not
|
||||
a reliable auditor of it.
|
||||
|
||||
One agent per chapter, launched in parallel, large chapters split the same way
|
||||
as in step 3b — but the chunk size doesn't have to match. Transcription chunks
|
||||
are sized for *writing* a large amount of new content (~8 pages); audit
|
||||
chunks only need to *read and compare*, so they tolerate roughly 2x the page
|
||||
count before running into the same silent-context-loss failure mode. On a
|
||||
400+ page document, auditing in ~15-page chunks instead of ~8-page chunks
|
||||
roughly halves the number of audit agents needed with no loss of quality.
|
||||
|
||||
```
|
||||
You are auditing a Markdown transcription of a datasheet chapter against the
|
||||
original PDF. This is a READ-ONLY audit. Do NOT edit, write, or create files.
|
||||
|
||||
PDF (ground truth): <path>.pdf
|
||||
Markdown to audit: <path>.md
|
||||
|
||||
Method:
|
||||
1. Use the native `read` tool on the PDF. It returns rendered pages as an
|
||||
attachment. DO NOT use pdftotext/python/bash extraction - compare against
|
||||
what is actually printed.
|
||||
2. Use the `read` tool on the .md.
|
||||
3. Compare page by page.
|
||||
|
||||
Check:
|
||||
- Missing or extra sections; wrong heading text or numbers.
|
||||
- Body text summarized or paraphrased rather than transcribed verbatim.
|
||||
- Every table: all rows and columns present, every value correct digit by
|
||||
digit including sign, decimal point and unit.
|
||||
- Figure captions: correct number, exact caption text.
|
||||
- Any hallucinated content in the MD that is not in the PDF. THIS IS THE MOST
|
||||
IMPORTANT THING TO CATCH.
|
||||
|
||||
For each finding give: severity (CRITICAL = wrong/invented, MAJOR = missing,
|
||||
MINOR = formatting), printed page number, what the PDF says, what the MD says.
|
||||
If a category is clean, say so in one line. End with a one-line verdict:
|
||||
ACCURATE / MINOR ISSUES / SIGNIFICANT ISSUES. Do not fix anything.
|
||||
```
|
||||
|
||||
Add per-chapter emphasis: numeric accuracy for a specifications chapter, every
|
||||
pin row for a pinout chapter, bit-level detail for a register chapter, field
|
||||
order and ACK/NACK bit values for protocol frame diagrams.
|
||||
|
||||
Feed forward any suspected source defects and ask the auditor to confirm what
|
||||
the PDF *actually* prints. This is how you separate a transcription bug from a
|
||||
vendor typo. In this project every such flag turned out to be a genuine defect
|
||||
in the datasheet: a threshold printed in mV with volt-magnitude values, a pin
|
||||
listed as both `NC` and `GPIO2`, `TYPE_C_PRAMETER` vs `TYPE_C_PARAMETER`
|
||||
between a heading and its own table caption.
|
||||
|
||||
### Prefer letting the audit agent navigate by marker, not pre-slicing
|
||||
|
||||
The simplest way to avoid the slicing pitfall below is to not slice at all:
|
||||
point the audit agent at the *full* chapter `.md` and tell it which
|
||||
`<!-- page N -->` markers bound its assigned range ("only inspect the portion
|
||||
between `<!-- page 254 -->` and `<!-- page 269 -->`"). The agent can `read`
|
||||
the whole file and navigate to the right markers itself. This trades a larger
|
||||
`read` for eliminating an entire class of off-by-one bugs in a slicing script,
|
||||
and for the sizes involved (a chapter `.md` is at most a few thousand lines)
|
||||
the read is cheap.
|
||||
|
||||
Reach for scripted pre-slicing only when a downstream tool (not an agent)
|
||||
needs an exact byte range — and if you do, remember markers precede their
|
||||
page:
|
||||
|
||||
```python
|
||||
start = marks[a] # NOT marks[a-1] + 1
|
||||
end = marks[b + 1]
|
||||
```
|
||||
|
||||
Getting this wrong shifts every slice one page and produces six confident,
|
||||
identical, **false** "MAJOR: last page is missing" reports. Before believing any
|
||||
missing-content finding, grep the real file:
|
||||
|
||||
```bash
|
||||
rg -o "Table ([0-9]+)\." -r '$1' chapters/07-registers.md | sort -nu | tr '\n' ' '
|
||||
rg -o "^### 7\.2\.([0-9]+) " -r '$1' chapters/07-registers.md | sort -n | tr '\n' ' '
|
||||
```
|
||||
|
||||
Contiguous sequences mean nothing is missing.
|
||||
|
||||
---
|
||||
|
||||
## 8. Act on the findings
|
||||
|
||||
Triage before editing. In this project, **all datasheet data — every table,
|
||||
value, register and pin — was accurate**; every real defect was in prose the
|
||||
agents added around the figures. Expect the same shape.
|
||||
|
||||
Distinguish:
|
||||
|
||||
- **Invented content** → delete.
|
||||
- **Derived content** (e.g. an address/reset summary computed from the register
|
||||
map) → keep if useful, but normalize it and add a note saying it is derived,
|
||||
not printed. Do not let it masquerade as transcription.
|
||||
- **Faithful transcription of a source defect** → leave alone, record in the
|
||||
TOC's known-issues section.
|
||||
|
||||
When stripping generated text, anchor the pattern tightly. Figure descriptions
|
||||
were italic paragraphs *directly following an image line*; that shape can be
|
||||
matched without touching legitimate printed `*Note:*` paragraphs elsewhere.
|
||||
Verify afterwards that the printed ones survived:
|
||||
|
||||
```bash
|
||||
echo "before: $(rg -c '^\*Note' backup.md) after: $(rg -c '^\*Note' file.md)"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Checklist
|
||||
|
||||
```
|
||||
[ ] Split PDF into chapters/NN-slug.pdf; record printed page ranges
|
||||
[ ] Transcribe in parallel, one agent per chapter (NO figure descriptions)
|
||||
[ ] Chunk chapters over ~10 pages; concatenate; clean the seams
|
||||
[ ] If normalizing heading levels with a script, verify it's idempotent
|
||||
(0 changes on a second run) and preserves trailing newlines
|
||||
[ ] Verify table-row diff is empty after every scripted rewrite
|
||||
[ ] Check the figure caption convention (above/below the figure) on one
|
||||
rendered page before extracting at scale; consider a subagent for the
|
||||
extract/tune/contact-sheet loop on documents with 50+ figures
|
||||
[ ] Extract figures; --dry-run, then contact-sheet review
|
||||
[ ] Normalize <!-- page N --> to printed numbers
|
||||
[ ] Build TOC with generated anchors; record known source defects
|
||||
[ ] tools/check-links.py passes
|
||||
[ ] Audit every chapter with fresh agents (chunks can be ~2x transcription
|
||||
chunk size, since audits only read+compare)
|
||||
[ ] Verify "missing content" findings against the real file before acting
|
||||
[ ] Delete invented content; normalize derived content; keep source defects
|
||||
```
|
||||
|
||||
## Tools in this repo
|
||||
|
||||
| Tool | Purpose |
|
||||
|---|---|
|
||||
| [`tools/extract-figures.py`](tools/extract-figures.py) | Caption-anchored figure cropping from vector PDFs. `--caption-position {below,above}` (default `below`) and decimal figure numbers (`4.1`) support both the Renesas-style and USB-IF-style caption conventions |
|
||||
| [`tools/check-links.py`](tools/check-links.py) | Link, anchor and orphan-image validation. Anchor generation replicates GitHub's `-1`/`-2` disambiguation for duplicate headings (e.g. two companies with near-identical names in a contributors list) |
|
||||
| [`chapters/TOC.md`](chapters/TOC.md) | Worked example of the generated TOC |
|
||||
Loading…
Add table
Add a link
Reference in a new issue