Read PDF text and tables
POST /pdf reads the text and tables that are already inside a digital PDF. No model is involved, so the
result is exact and repeatable. For scanned pages, which are images of text, use OCR instead.
import base64, os, requestsfrom pathlib import Path
pdf = Path("statement.pdf")body = { "attachments": [{"name": pdf.name, "data": base64.b64encode(pdf.read_bytes()).decode()}], "mode": "all", # all | text | tables "table_format": "both", # rows | records | both "pages": "all",}r = requests.post("https://dummydomain/pdf", json=body, headers={"Authorization": "Bearer " + os.environ["API_KEY"]})r.raise_for_status()for f in r.json()["files"]: print(f["pages_processed"], "pages,", f["table_count"], "tables, truncated:", f["truncated"])# Build the JSON body with the file base64-encoded (see the Python tab), then:curl --fail-with-body "https://dummydomain/pdf" \ -H "Authorization: Bearer $API_KEY" -H 'Content-Type: application/json' \ --data-binary @pdf-request.json --output pdf-result.jsonRequest
Section titled “Request”| Field | Values |
|---|---|
attachments |
List of {"name", "data"}, where data is the file as base64 |
mode |
all, text or tables |
table_format |
rows, records or both |
pages |
A maximum number of pages to read, or "all". It is not a page range. |
The body is JSON, not a multipart upload.
Response
Section titled “Response”Each item in files[] has text, page_texts, tables, table_count, pages_processed and
truncated. Check truncated before treating the result as the whole document.
Limits
Section titled “Limits”Up to 20 files per request, 50 MiB per PDF and 200 pages per file. Base64 makes the request about a third larger than the file, and your server’s request size limit still applies.