Make a PDF document
Activated Cloud✓ Officialactivated/make-pdf-document
Free · MIT
About
Makes finished PDF files on your own computer: reports and letters from HTML and CSS printed with the installed Chromium (or WeasyPrint), data-heavy layouts such as invoices and stock lists with reportlab, and merging, splitting, rotating or filling existing PDFs with pypdf, then renders pages to images to check them. Use when the owner wants a PDF to send, print or archive. Not for an editable Word file (see make-word-document) or for reading a PDF someone sent (see summarise-long-documents).
Documentation
Make a PDF document
A PDF is the final, fixed form of a document: what the owner sends, prints or files. This skill picks the right way to make one, builds it with Python on your computer, and proves it is right by rendering pages and looking at them. The standard: text is real and searchable (not a picture of text), fonts are embedded, tables do not split badly, every page is numbered, and nothing overflows the margins.
When to use
- "Send me the report as a PDF."
- "Make an invoice / quote / certificate / price list PDF."
- "Combine these five PDFs into one and put the summary first."
- "Rotate page 3 and take out the blank pages."
- "Print this web page to PDF."
What you need
- The content, already written (see
write-structured-report), or the data for a generated document. - Page size (A4 or US Letter), brand colours and logo if the owner has them (check
memory). - For legal, tax or financial documents (invoices with tax, contracts, statements): the figures and wording must come from the owner or their accounting software. You format; you do not invent tax rates or terms. Check rates for the country and date with
web_searchand name the source.
Choose the route
| Need | Route | Install |
|---|---|---|
| Report, letter, proposal, anything mostly text | HTML + CSS, printed by headless Chromium | nothing: Chromium is on your computer |
| Same, plus running headers, "Page X of Y" in older Chromium, contents page numbers | HTML + CSS with WeasyPrint | sudo apt-get install -y weasyprint |
| Invoices, labels, long data tables, exact positions | reportlab | pip (below) |
| The owner also wants Word | build the .docx (make-word-document), convert with LibreOffice |
libreoffice-writer-nogui |
| Merge, split, rotate, reorder, watermark, fill form fields | pypdf | pip (below) |
Install Python libraries only if missing (Debian needs --break-system-packages with --user; it writes only to your home folder):
python3 -c "import reportlab, pypdf, pypdfium2" 2>/dev/null || python3 -m pip install --user --break-system-packages reportlab pypdf pypdfium2
Method
Make a job folder in your work space, for example
/home/user/Desktop/Ada - Work space/supplier-review/.execute_coderuns in a temporary folder, so use absolute paths everywhere.HTML route (default for documents). Write the HTML with print CSS (template in
references/pdf-recipes.md:@pagesize and margins, page-number margin boxes, a table header that repeats, rules that stop headings being stranded). Then print:
cd "/home/user/Desktop/Ada - Work space/supplier-review"
chromium --headless --no-sandbox --disable-gpu --no-pdf-header-footer \
--user-data-dir="$(mktemp -d)" \
--print-to-pdf="$PWD/2026-10-05_supplier-review_v01.pdf" "file://$PWD/report.html"
--user-data-dir keeps this print run apart from your everyday browser profile. If an older Chromium rejects --no-pdf-header-footer, use --print-to-pdf-no-header. Images and fonts must be local files or reachable URLs; add --virtual-time-budget=5000 if the page loads web fonts.
reportlab route (data and exact layout). Use Platypus (
SimpleDocTemplate,Paragraph,Table) so text wraps and tables flow across pages, withrepeatRows=1for the header. Compute every number in the text from the data you are printing. The tested recipe with "Page X of Y" footers is inreferences/pdf-recipes.md. The built-in Helvetica covers Western European text; for other scripts register a TTF font such as/usr/share/fonts/truetype/noto/NotoSans-Regular.ttf(find fonts withfc-list).pypdf route (existing PDFs).
from pypdf import PdfReader, PdfWriter
JOB = "/home/user/Desktop/Ada - Work space/board-pack"
for name in ("summary", "accounts", "appendix"):
print(name, len(PdfReader(f"{JOB}/inputs/{name}.pdf").pages), "pages") # know the counts first
w = PdfWriter()
w.append(f"{JOB}/inputs/summary.pdf")
start = len(w.pages) # page indexes start at 0
w.append(f"{JOB}/inputs/accounts.pdf", pages=(0, 2)) # pages 1 and 2 only; the range must exist
w.append(f"{JOB}/inputs/appendix.pdf")
w.pages[start + 1].rotate(90) # accounts page 2 came in sideways
with open(f"{JOB}/2026-10-05_board-pack_v01.pdf", "wb") as f:
w.write(f)
print(len(PdfReader(f"{JOB}/2026-10-05_board-pack_v01.pdf").pages), "pages in the pack")
Splitting, deleting pages, watermarking, filling form fields and adding a password are in references/pdf-recipes.md. Never edit the owner's original; write a new file.
- Check it (below), fix, re-render, then tell the owner where it is.
Checking the output
Render the first page, the last page and any page with a big table to PNG, then look at each with vision_analyze:
import pypdfium2 as pdfium
from pypdf import PdfReader
PDF = "/home/user/Desktop/Ada - Work space/supplier-review/2026-10-05_supplier-review_v01.pdf"
r = PdfReader(PDF)
print(len(r.pages), "pages; page 1 text:", (r.pages[0].extract_text() or "")[:120])
doc = pdfium.PdfDocument(PDF)
for i in sorted({0, len(doc) - 1}):
doc[i].render(scale=1.5).to_pil().save(PDF.replace(".pdf", f"_check_p{i + 1}.png"))
Ask vision_analyze specific questions: "Is any text cut off or overlapping? Does the table header repeat? Is the page number present? Is anything misaligned?" Delete the _check_ images afterwards.
- Empty text from
extract_text()means the page is an image: fix that unless the owner wanted a scan. - File size: aim under 5 MB for email. Large files usually mean oversized images; resize them before building.
Output
- The PDF in the job folder, named
YYYY-MM-DD_topic_type_v01.pdf, plus its source (HTML or script) beside it so it can be rebuilt. - One message with the full path, page count and anything the owner should check (for example "the bank details on page 1 came from your last invoice; please confirm").
Checks before you finish
- Text is selectable (extract_text returns words) and fonts render correctly (no empty boxes for £, €, accents).
- Page size and margins are right; page numbers on every page of multi-page documents.
- Table headers repeat after page breaks; no heading sits alone at the foot of a page.
- Every figure in the text matches the data; totals recalculated.
- You looked at rendered pages, not just the code.
Pitfalls
- Hard-coded numbers in the narrative that disagree with the table. Compute them from the same data.
- Paragraphs inside reportlab tables keep their own style, so the table's FONTSIZE does not apply to them. Give cell Paragraphs a smaller style.
- Printing a live web page you do not control may pick up cookie banners and ads. Save the content into your own HTML first, or use the HTML route.
- Screenshots as pages. A PDF of images is unsearchable and inaccessible. Use real text.
- Missing glyphs for non-Latin text in reportlab's built-in fonts. Register a Unicode TTF.
- Invoices and legal documents with made-up details. Tax numbers, rates, payment terms and legal wording come from the owner or a checked source; a qualified person signs off contracts and tax documents.
See also: make-word-document (editable version), make-clear-charts (charts to embed), summarise-long-documents (reading PDFs).
Versions
Listed from the source repository.
Reviews
No reviews yet. Be the first.
