Skip to content

Read PDF text and tables

POST /pdf reads the text and tables that are already inside a digital PDF. No model is involved, so the result is exact and repeatable. For scanned pages, which are images of text, use OCR instead.

import base64, os, requests
from pathlib import Path
pdf = Path("statement.pdf")
body = {
"attachments": [{"name": pdf.name, "data": base64.b64encode(pdf.read_bytes()).decode()}],
"mode": "all", # all | text | tables
"table_format": "both", # rows | records | both
"pages": "all",
}
r = requests.post("https://dummydomain/pdf", json=body,
headers={"Authorization": "Bearer " + os.environ["API_KEY"]})
r.raise_for_status()
for f in r.json()["files"]:
print(f["pages_processed"], "pages,", f["table_count"], "tables, truncated:", f["truncated"])
Field Values
attachments List of {"name", "data"}, where data is the file as base64
mode all, text or tables
table_format rows, records or both
pages A maximum number of pages to read, or "all". It is not a page range.

The body is JSON, not a multipart upload.

Each item in files[] has text, page_texts, tables, table_count, pages_processed and truncated. Check truncated before treating the result as the whole document.

Up to 20 files per request, 50 MiB per PDF and 200 pages per file. Base64 makes the request about a third larger than the file, and your server’s request size limit still applies.