Pipeline Automation (Automate)

The Automate tool builds, saves and reuses sequences of PDF operations. The backend pipeline API and folder scanner use the pipeline JSON format described below.


What is Pipeline Automation?#

Pipeline automation allows you to:

  • Chain operations - Combine multiple PDF tools in sequence
  • Save workflows - Reuse common operation sequences
  • Automate processing - Process files automatically with folder scanning
  • Standardize procedures - Ensure consistent processing across teams
  • Batch process - Apply same workflow to multiple files

For example, a saved Split → Watermark → Compress workflow runs those operations in sequence without manually transferring the output between tools.


Key Concepts#

Operations#

Individual PDF tools that perform specific tasks:

  • Split, Merge, Compress, Watermark, etc.
  • Each operation has configurable parameters
  • Operations execute in the order you define

Pipeline#

A sequence of operations with saved configurations:

  • Named workflow (e.g., "Invoice Processing")
  • Ordered list of operations
  • Pre-configured settings for each operation
  • Reusable across multiple files

Pipeline Configuration (JSON)#

Text file that defines your pipeline:

  • Lists operations in order
  • Specifies parameters for each operation
  • Can be shared, versioned, and backed up
  • Human-readable and editable

Folder Scanning#

Automated processing mode:

  • Watch a folder for new files
  • Automatically apply pipeline to new files
  • Move processed files to output folder
  • Unattended batch processing

Getting Started with Automate#

Accessing the Automate Tool#

  1. From Home Page

    • Click "Automate" in the Advanced Tools section
    • Or search for "automate" or "pipeline"
  2. Open the Automation Builder

    • Click "Create New Automation"
    • The automation builder opens

Building Your First Pipeline#

Steps to Configure and Use Your Pipeline#

  1. Start a New Automation

    • On the Automate screen, click Create New Automation.
  2. Enter Automation Name

    • Provide a name for your automation in the designated field (you can also add an optional description and icon).
  3. Add Tools

    • Choose the tools for your automation (e.g., Split Pages) with Add Tool. Tools run in the order you add them.
  4. Configure Tool Settings

    • Configure each added tool. A tool shows a ! Not Configured marker until its settings are saved.
  5. Add More Tools

    • You can add and reorder multiple tools. Make sure each tool is configured.
  6. Save Tool Settings

    • Click Save Configuration in each tool's settings dialog after customizing it.
  7. Save the Automation

    • Click Save Automation once all tools are configured. The Save button stays disabled until the automation is complete.
  8. Download Pipeline Configuration

    • The Automate UI offers two download buttons:
      • Export - downloads <name>.automate.json in the native Automate format with frontend tool IDs. Use this for re-importing into another Stirling PDF instance via the UI.
      • Export for Folder Scanning - downloads <name>.folder-scan.json in the backend format with full endpoint paths. Use this for the REST API and for folder scanning.
    • To pre-load a pipeline for all users, place a folder-scanning-format JSON file in /pipeline/defaultWebUIConfigs/ - it will appear in the dropdown.
  9. Run the Automation

    • Once saved, select the automation from the list, add your files, and run it.
  10. Note on Web UI Limitations

    • The current web UI version does not support operations that require multiple different types of inputs, such as adding a separate image to a PDF.

Current Limitations#

  • The same operation can appear more than once, with separate parameters for each step.
  • Cannot input additional files via UI.
  • All files and operations run in serial mode.

Example Pipelines#

Example 1: Invoice Processing#

Goal: Process scanned invoices for archival

Pipeline Steps:

  1. OCR - Make invoices searchable
    • Language: English
    • Preserve formatting: Yes
  2. Crop - Remove scanner edges
    • Margins: 0.5 inches all sides
  3. Add Watermark - Mark as processed
    • Text: "PROCESSED [DATE]"
    • Position: Bottom right
    • Opacity: 50%
  4. Compress - Reduce file size
    • Level: Medium
  5. Add Password - Secure documents
    • Password: [configured per run]

Use Case: Accounting department processing hundreds of invoices monthly


Example 2: Report Distribution#

Goal: Prepare reports for external sharing

Pipeline Steps:

  1. Remove Pages - Remove internal pages
    • Pages: 2,3 (remove cover sheets)
  2. Add Page Numbers - Number all pages
    • Position: Bottom center
    • Format: "Page X of Y"
  3. Add Stamp - Add "CONFIDENTIAL" stamp
    • Position: Top right
    • Color: Red
  4. Change Permissions - Restrict editing
    • Allow printing: Yes
    • Allow editing: No
  5. Compress - Optimize for email
    • Level: High

Use Case: Monthly reports sent to clients


Example 3: Document Standardization#

Goal: Standardize format of received documents

Pipeline Steps:

  1. Rotate - Fix orientation
    • Mode: Auto-detect
  2. Scale Pages - Standardize to Letter size
    • Target: 8.5 x 11 inches
  3. Add Metadata - Tag documents
    • Title: [Auto-extracted]
    • Author: "Company Name"
    • Keywords: "Standardized, Processed"
  4. Remove Annotations - Clean markup
  5. Flatten - Remove form fields

Use Case: HR department standardizing employee submissions


Example 4: Batch Conversion#

Goal: Convert and optimize image scans

Pipeline Steps:

  1. Convert - Images to PDF
    • Source: JPG, PNG
  2. OCR - Add text layer
    • Language: Multiple
  3. Remove Blanks - Delete empty pages
    • Threshold: 95%
  4. Compress - Optimize size
    • Level: Medium
  5. PDF/A - Convert for archival
    • Version: PDF/A-2b

Use Case: Digitization project for paper archives


Common Pipeline Patterns#

Quality Enhancement Pipeline#

Pattern: Improve scanned document quality

OCR → Remove Blanks → Adjust Contrast → Compress → Add Metadata

Security Pipeline#

Pattern: Secure documents for distribution

Remove Metadata → Add Watermark → Add Password → Change Permissions

Compression Pipeline#

Pattern: Reduce file sizes for storage/email

Remove Annotations → Remove Images (optional) → Compress → Validate

Branding Pipeline#

Pattern: Add company branding to documents

Add Watermark → Add Stamp → Add Page Numbers → Add Metadata

Preparation Pipeline#

Pattern: Prepare documents for printing

Rotate → Scale Pages → Booklet Imposition → Remove Annotations

JSON Configuration#

Basic Structure#

A pipeline JSON file has a name and a pipeline array. Each entry has an operation (the full API endpoint path) and a parameters object:

json
{
  "name": "My Pipeline",
  "pipeline": [
    {
      "operation": "/api/v1/general/split-pages",
      "parameters": {
        "pageNumbers": "5"
      }
    },
    {
      "operation": "/api/v1/misc/compress-pdf",
      "parameters": {
        "optimizeLevel": 5,
        "expectedOutputSize": ""
      }
    }
  ]
}
Operation names are full endpoint paths

Pipeline operation names use the full REST API path, not short names. For example, use /api/v1/general/split-pages (not just split-pages). The Folder-Scanning export from the Automate UI produces these paths automatically.

If you have older pipeline JSONs that used short names, regenerate them from the Automate UI using Export for Folder Scanning.

Optional fields#

For folder scanning (not used by the REST API):

json
{
  "name": "...",
  "pipeline": [ ... ],
  "outputDir": "{outputFolder}/{folderName}",
  "outputFileName": "{filename}-{pipelineName}-{date}-{time}"
}

outputDir supports {outputFolder} and {folderName}. outputFileName supports {filename}, {pipelineName}, {date} and {time}; the output extension is appended by the processor.


Operation and parameter reference#

Pipeline operations use the full endpoint paths of Stirling PDF's REST API, with the same field names. So once you know the underlying endpoint, you know the pipeline operation - no separate vocabulary to learn.

For the canonical list of operations and the full parameter schema for each, see:

See API Documentation for authentication and general API usage.

Pipelines support operations under /api/v1/general/, /api/v1/misc/, /api/v1/security/, /api/v1/convert/, /api/v1/filter/, /api/v1/integration/, /api/v1/docparse/, and /api/v1/ai/tools/. This also applies to Processor pipelines.

For external-api-call, create an integration first and use its connection ID in the step.

The /api/v1/docparse/rag-ingest operation prepares document chunks for knowledge search or export. For source, trigger, and database destination setup, use the Processor's Ingestion policy.

AI steps such as math-auditor-agent and pdf-comment-agent require an enabled AI engine and the corresponding AI tool.

Build it in the UI, export it as JSON

The fastest way to get a correct pipeline JSON for any combination of operations is to build it visually in the Automate tool and click Export for Folder Scanning. The exported file uses exactly the format the API expects, with the right operation paths and parameters already filled in for you.


Filter / conditional operations#

Filter operations keep or drop files based on a condition. Matching files continue to later steps; non-matching files are removed from that pipeline run. For example, a page-count filter can select longer documents for watermarking.

A file that does not match a filter is simply removed from the rest of the pipeline. It is not treated as an error.

These operation names go in your pipeline configuration:

Operation name Keeps the file when...
filter-contains-text the PDF contains a given piece of text (you can limit the check to specific pages)
filter-contains-image the PDF contains an image (you can limit the check to specific pages)
filter-page-count the page count is greater than, equal to, or less than a value you set
filter-page-size the first page's size compares to a standard page size you choose
filter-file-size the file size compares to a value you set
filter-page-rotation the first page's rotation compares to a value you set

The four comparison filters (filter-page-count, filter-page-size, filter-file-size, filter-page-rotation) take a comparator of Greater, Equal, or Less.

Example - OCR PDFs containing images: filter-contains-image keeps PDFs with an image, including PDFs that also contain text. The OCR step uses skip-text to skip pages that already have text; the filter itself does not detect an absent text layer.

json
{
  "name": "OCR PDFs containing images",
  "pipeline": [
    {"operation": "/api/v1/filter/filter-contains-image", "parameters": {"pageNumbers": "all"}},
    {"operation": "/api/v1/misc/ocr-pdf", "parameters": {"languages": ["eng"], "ocrType": "skip-text"}}
  ]
}

These filter operations work both in the REST API and in folder scanning.


REST API: POST /api/v1/pipeline/handleData#

Trigger a pipeline programmatically via the REST API. Use this from scripts, automation platforms (n8n, Zapier, Make, Power Automate), or your own integrations.

Request#

  • Method: POST
  • URL: /api/v1/pipeline/handleData
  • Content-Type: multipart/form-data
  • Authentication: When security is enabled, set the X-API-KEY header. See API Documentation for details.

Multipart fields#

Field Type Required Purpose
fileInput file yes One or more PDF files. Repeat the field for multiple files.
json string yes The pipeline configuration JSON.

You don't need to include fileInput inside the parameters object - the pipeline processor injects each uploaded file automatically. The Automate UI's "Export for Folder Scanning" includes "fileInput": "automated" as a marker in every step, which the backend ignores; you can leave it in or strip it out, both work.

Optional query parameters#

  • ?async=true - run the pipeline asynchronously and return a job ID instead of the file. Poll GET /api/v1/general/job/{id} for progress.

Response#

  • Single output file: returned directly as application/octet-stream with Content-Disposition: attachment; filename=....
  • Multiple output files: returned as output.zip.
  • Async mode: returns a JSON body with the job ID.

Working curl example#

bash
curl -X POST "http://localhost:8080/api/v1/pipeline/handleData" \
  -H "X-API-KEY: $STIRLING_API_KEY" \
  -F "fileInput=@/path/to/input.pdf" \
  -F 'json={
    "name": "Repair-then-compress",
    "pipeline": [
      {"operation": "/api/v1/misc/repair", "parameters": {}},
      {"operation": "/api/v1/misc/compress-pdf", "parameters": {"optimizeLevel": 2}}
    ]
  }' \
  --output result.pdf

For multiple files use repeated -F "fileInput=@..." flags; the response will be output.zip. For the full parameter list for each operation, see the API docs linked above.

Error responses#

Situation HTTP status Body
Auth required and no key supplied 401 {"error":"Unauthorized","message":"Authentication required...","status":401}
Multipart parsing failed (missing field, bad JSON) 400 Spring's standard error JSON
Invalid operation name, disallowed endpoint, or missing required parameter 200 with empty body The server logs an IllegalArgumentException but returns an empty response.
Downstream endpoint returned non-2xx 200 with partial/empty body The error is logged but does not surface in the HTTP response.

Tips#

  • Build in the UI, export the JSON. The fastest way to get a correct JSON is to build the pipeline in the Automate UI, then click Export for Folder Scanning. The exported file works directly with handleData. The other button, Export, produces a different "native Automate" format (uses an operations key with frontend tool IDs like "merge") that is only for re-importing into another Automate UI, not for the API.
  • No image / file parameters. Operations that take an additional file input (image watermarks, separate overlay PDFs, attaching files) cannot be expressed in pipeline JSON via the REST API. Call those endpoints directly instead.
  • List parameters become repeated form fields. Internally the processor expands ["eng","deu"] into two languages=eng and languages=deu form parts, which is what the underlying endpoints expect.
  • Filters drop files. A filter step that doesn't match keeps the file out of later steps. Useful for "process only PDFs that contain X".
  • Multi-input operations batch. Operations marked multi-input (e.g. merge-pdfs) receive every matching file in a single call. If no files in the working set match the operation's expected extension, the step logs No files with extension X found for operation Y... and continues with the other files.
  • Unknown JSON fields are ignored. The pipeline parser silently drops fields it doesn't recognise, so you can add description, icon, or other metadata at the top level without breaking anything.

Folder Scanning Setup#

Automate processing of files placed in watched folders.

How Folder Scanning Works#

  1. Watch Input Folder - Monitor for new files
  2. Detect New Files - Identify PDFs added to folder
  3. Apply Pipeline - Process with configured pipeline
  4. Output Results - Save to output folder
  5. Archive Originals - Move processed files (optional)

Directory Structure#

/pipeline/
  ├── watchedFolders/
  │   ├── invoice-processing/
  │   │   ├── my-pipeline.json   # any *.json file in the folder is the pipeline config
  │   │   ├── invoice-001.pdf    # drop PDFs directly into the folder root
  │   │   ├── invoice-002.pdf
  │   │   └── processing/        # auto-created by the scanner while a file is in flight
  │   └── report-prep/
  │       └── ...
  ├── finishedFolders/           # outputs appear here by default (per `outputDir` placeholder)
  └── defaultWebUIConfigs/       # pre-loaded pipelines exposed in the Automate UI dropdown
      ├── invoice.json
      └── reports.json

Configuration File#

Drop a .json file (any name) into each watched folder. The first .json the scanner finds is used as the pipeline:

json
{
  "name": "Invoice Processing",
  "pipeline": [
    {"operation": "/api/v1/misc/ocr-pdf", "parameters": {
       "languages": ["eng"], "ocrType": "skip-text",
       "ocrRenderType": "hocr", "deskew": true, "clean": false,
       "cleanFinal": false, "sidecar": false, "removeImagesAfter": false}}
  ],
  "outputDir": "{outputFolder}/{folderName}",
  "outputFileName": "{filename}-processed-{date}"
}

PDFs go directly in the watched folder root (NOT in an input/ subdirectory). The scanner auto-creates a processing/ subfolder while a file is being worked on, and writes outputs to wherever outputDir resolves to (typically /pipeline/finishedFolders/... via the {outputFolder} placeholder).

The watched-folder scanner runs every 60 seconds.

Learn more: Folder Scanning Guide


Testing and Maintaining Workflows#

  • Add one operation at a time and inspect its output before extending the pipeline.
  • Run OCR before operations that require searchable text; remove unwanted pages before processing the remaining pages.
  • Test locked, empty and representative large files, and check the server logs when output is missing or incomplete.
  • Keep exported JSON in version control and retest it after changes to the workflow or server.
  • Measure processing time and resource use with your own files before increasing batch size or concurrency.

Troubleshooting#

Pipeline Fails to Execute#

Symptoms: Pipeline starts but doesn't complete

Common Causes:

  • Invalid parameter values
  • Unsupported file format
  • Missing dependencies (OCR languages, fonts)
  • File permissions issues

Solutions:

  1. Validate JSON configuration
  2. Test each operation individually
  3. Check server logs for errors
  4. Verify required dependencies installed

handleData Returns Empty Response#

Symptoms: REST API call returns HTTP 200 with an empty body.

Cause: Errors after multipart parsing (invalid operation name, missing required parameter, downstream endpoint failure) currently collapse to 200 OK with no body. Check the server logs for the actual error.

Common reasons:

  • Operation name used short form (e.g. compress-pdf) instead of full path (/api/v1/misc/compress-pdf)
  • Operation references an endpoint outside the allowed namespaces (only general, misc, security, convert, filter, integration, docparse, ai/tools are permitted)
  • A required parameter was omitted (check the schema for the underlying endpoint in the Swagger UI / API reference)
  • The pipeline tries to call /api/v1/pipeline/handleData recursively

Folder Scanning Not Working#

Symptoms: Files not processed automatically

Possible Issues:

  • Folder permissions incorrect
  • Pipeline configuration invalid
  • Folder scanning not enabled

Solutions:

  1. Check folder permissions (read/write access)
  2. Test pipeline manually first
  3. Check docker logs for errors
  4. Ensure folder scanning feature enabled
  5. The scanner runs once every 60 seconds - allow that long after dropping a file

Operation Parameters Not Applying#

Symptoms: Pipeline runs but doesn't use specified settings

Causes:

  • Incorrect parameter names
  • Wrong parameter data types
  • Parameters not supported in operation

Solutions:

  1. Check parameter names against the endpoint's schema in the Swagger UI / API reference
  2. Verify parameter value types (string, number, boolean)
  3. Test the same parameters by calling the endpoint directly first

Results Not as Expected#

Symptoms: Pipeline completes but output incorrect

Debugging Steps:

  1. Test each operation individually
  2. Check intermediate outputs
  3. Verify operation order makes sense
  4. Review parameter values
  5. Test with simpler input files

Pipeline vs. Multi-Tool vs. Manual#

Use Pipeline/Automate When:#

  • Same workflow repeated frequently
  • Predictable, consistent operations
  • Automated folder processing needed
  • No manual intervention required
  • Standardizing team processes
  • Large batch processing
  • Scheduled/unattended processing

Use Multi-Tool When:#

  • Workflow varies per file
  • Need visual feedback at each step
  • Experimenting with different settings
  • Manual decision points in workflow
  • One-time complex tasks

Use Individual Tools When:#

  • Single, simple operation
  • Quick one-off task
  • Learning how operations work
  • No need for automation

Security Considerations#

Pipeline Files#

  • Protect JSON configs - May contain passwords or sensitive settings
  • Restrict folder access - Limit who can create/modify pipelines
  • Review before deploying - Audit pipelines for security issues

Folder Scanning#

  • Isolate watched folders - Don't expose to untrusted users
  • Monitor activity - Log all processing for audit trail
  • Secure output folders - Protect processed documents appropriately

Automated Processing#

  • Validate inputs - Ensure only expected files processed
  • Error handling - Don't expose sensitive error messages
  • Resource limits - Prevent resource exhaustion attacks