Skip to main content

Node Operations

Parameters​

ParameterDefaultDescription
Input Binary FielddataBinary property containing the PDF or image
PDF Extraction ModeAutomaticSelects embedded text, OCR, or per-page automatic fallback
OCR LanguageengTesseract language code; combine languages with +
Page SegmentationAutomaticTesseract's assumption about the page layout
Maximum PDF Pages20Limits processing; accepted range is 1 through 100
PDF Render Scale2Raster scale for PDF OCR; accepted range is 1 through 4
Minimum Embedded Text Length20Non-whitespace characters needed to avoid OCR in Automatic mode
PDF PasswordEmptyPassword used to open an encrypted PDF
Page SeparatorTwo newlinesJoins page text in the combined text field
Maximum Output Characters1000000Truncates extracted output beyond the configured limit
Document Timeout120000Limits total PDF and OCR processing time for each document
Destination FielddocumentTextContains all extraction output without overwriting common input fields
Keep Input BinaryfalsePreserves the source binary on the output item

Runtime-enforced limits​

These bounds are inclusive and enforced before binary retrieval and provider creation, including expression results and imported workflow settings. Values are rejected, never clamped or rounded; errors identify the serialized parameter name.

ParameterAccepted value
maxPagesInteger 1–100
renderScaleFinite number 1–4; fractions are supported
minimumTextLengthInteger 0–10,000
maxCharactersInteger 1,000–10,000,000
documentTimeoutMsInteger 1,000–300,000 milliseconds
modeauto, text, or ocr
pageSegmentationModeauto, singleBlock, singleColumn, or sparseText

Numeric strings, nonnumeric values, NaN and Infinity are rejected. All limits are validated even when the selected mode or file type does not use a setting. Mode, OCR language, segmentation and timeout remain execution-wide settings; expression-enabled limits are evaluated for each item. The page separator remains limited to 1,000 characters.

The processing timeout includes worker initialization. Worker shutdown is awaited afterward, so total wall-clock time can exceed the processing timeout while the worker exits.

PDF Modes​

Automatic​

Extracts embedded text from each page. Pages below Minimum Embedded Text Length are rendered as PNG images and passed through OCR. This is the recommended mode for unknown or mixed PDFs.

Embedded Text Only​

Extracts the PDF text layer without initializing OCR. Scanned pages generally return empty text.

OCR All Pages​

Renders every selected page and processes it with OCR. Use this when the PDF text layer is absent, corrupt, or has an unusable reading order.

Errors​

Invalid settings, unsupported file types, missing binary fields, oversized files or rendered pages, invalid PDFs, encrypted PDFs without the correct password, timeouts, PDF rendering failures, and OCR initialization failures produce errors. With n8n's Continue On Fail setting enabled, the failed input produces an error item paired to its input and subsequent inputs continue after required worker shutdown completes. An invalid per-item limit does not retrieve that item's binary and does not prevent a later item with valid limits from succeeding.

If worker shutdown itself rejects, execution stops even with Continue On Fail: worker exit is unconfirmed, so starting another document could overlap resource use. Cleanup errors are reported alongside the original extraction error when both occur. Successful results are returned only after final OCR cleanup succeeds.