Stirling-PDF on a VPS: Self-Hosted PDF Tools for Merge, Split, Convert, and OCR

Stirling-PDF on a VPS: Self-Hosted PDF Tools for Merge, Split, Convert, and OCR

Stirling-PDF is a self-hosted web application providing 50+ PDF operations — merge, split, rotate, compress, convert (PDF to Word, Excel, images and back), OCR with Tesseract, watermarking, page ordering, redaction, and more. It replaces paid PDF services (Adobe Acrobat at $20/month, Smallpdf, ILovePDF) with a private, self-hosted tool where sensitive documents never leave your infrastructure.

What Stirling-PDF Provides

  • Merge & Split: Combine multiple PDFs, extract page ranges, split at specified pages
  • Convert: PDF→Word/Excel/HTML/images, Word/Excel/PowerPoint→PDF, images→PDF
  • OCR: Make scanned PDFs searchable using Tesseract (50+ languages)
  • Compress: Reduce file size while maintaining quality
  • Security: Add/remove passwords, permissions, digital signatures
  • Edit: Watermark, stamp, add page numbers, rotate, crop, redact
  • Compare: Visual diff between two PDFs
  • API: REST API for every operation — integrate with workflows

Step 1: Docker Compose Setup

<code">mkdir -p /opt/stirling-pdf && cd /opt/stirling-pdf
nano docker-compose.yml
<code">version: '3.8'

services:
  stirling-pdf:
    image: frooodle/s-pdf:latest
    container_name: stirling-pdf
    restart: always
    ports:
      - "127.0.0.1:8080:8080"
    environment:
      DOCKER_ENABLE_SECURITY: "false"   # Set to true for login/user management
      INSTALL_BOOK_AND_ADVANCED_HTML_OPS: "false"   # Skip heavy Calibre install
      LANGS: "en_GB"
    volumes:
      - ./trainingData:/usr/share/tesseract-ocr/5/tessdata  # OCR training data
      - ./extraConfigs:/configs
      - ./logs:/logs
<code">docker compose up -d
docker compose logs -f stirling-pdf   # Watch for startup completion

Step 2: Nginx Reverse Proxy

<code">sudo nano /etc/nginx/sites-available/stirling-pdf
<code">server {
    listen 80;
    server_name pdf.yourdomain.com;
    return 301 https://$host$request_uri;
}

server {
    listen 443 ssl http2;
    server_name pdf.yourdomain.com;

    ssl_certificate /etc/letsencrypt/live/pdf.yourdomain.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/pdf.yourdomain.com/privkey.pem;

    # Large file uploads (PDFs can be large)
    client_max_body_size 500M;
    proxy_read_timeout 300s;   # OCR on large documents takes time

    location / {
        proxy_pass http://127.0.0.1:8080;
        proxy_http_version 1.1;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}
<code">sudo ln -s /etc/nginx/sites-available/stirling-pdf /etc/nginx/sites-enabled/
sudo certbot --nginx -d pdf.yourdomain.com
sudo systemctl reload nginx

Step 3: Enable Authentication

<code"># For shared/team use, enable login:
# docker-compose.yml:
environment:
  DOCKER_ENABLE_SECURITY: "true"
  SECURITY_ENABLE_LOGIN: "true"
  SECURITY_INITIALLOGIN_USERNAME: admin
  SECURITY_INITIALLOGIN_PASSWORD: ${ADMIN_PASSWORD}

# Change default password on first login

Step 4: Use the REST API

<code"># Every Stirling-PDF operation has an API endpoint
# API docs available at: https://pdf.yourdomain.com/swagger-ui/index.html

# Merge PDFs:
curl -X POST https://pdf.yourdomain.com/api/v1/general/merge-pdfs \
  -F "fileInput=@file1.pdf" \
  -F "fileInput=@file2.pdf" \
  -o merged.pdf

# Convert PDF to Word:
curl -X POST https://pdf.yourdomain.com/api/v1/convert/pdf/word \
  -F "fileInput=@document.pdf" \
  -o document.docx

# OCR a scanned PDF:
curl -X POST https://pdf.yourdomain.com/api/v1/misc/ocr-pdf \
  -F "fileInput=@scanned.pdf" \
  -F "languages=eng" \
  -o searchable.pdf

# Compress a PDF:
curl -X POST https://pdf.yourdomain.com/api/v1/general/compress-pdf \
  -F "fileInput=@large.pdf" \
  -F "optimizeLevel=3" \
  -o compressed.pdf

# Extract pages 3-7 from a PDF:
curl -X POST https://pdf.yourdomain.com/api/v1/general/extract-pages \
  -F "fileInput=@document.pdf" \
  -F "pageNumbers=3,4,5,6,7" \
  -o extracted.pdf

Step 5: Automate with n8n or Windmill

<code"># Example: n8n workflow for automatic PDF processing
# Trigger: Email attachment received → PDF detected
# Step 1: Download attachment
# Step 2: HTTP POST to Stirling-PDF /api/v1/misc/ocr-pdf
# Step 3: Save searchable PDF to storage
# Step 4: Send processed PDF via email

# Windmill script (Python):
import httpx

def process_pdf(pdf_path: str, operation: str = "ocr") -> bytes:
    with open(pdf_path, "rb") as f:
        response = httpx.post(
            "https://pdf.yourdomain.com/api/v1/misc/ocr-pdf",
            files={"fileInput": f},
            data={"languages": "eng"},
            timeout=300,
        )
    return response.content

Getting Started

Stirling-PDF uses 500 MB–1 GB RAM when processing documents (OCR and conversion are CPU-intensive). A 2 GB Ubuntu VPS at VPS.DO handles typical document workloads comfortably. For heavy batch processing, a 4 GB VPS provides headroom for concurrent operations. Sensitive documents — contracts, financial records, medical files — stay entirely on your VPS with no cloud processing.

Conclusion

Self-hosted Stirling-PDF provides 50+ PDF operations in one Docker container — replacing Adobe Acrobat, Smallpdf, and similar services for teams handling sensitive documents. The REST API enables automation pipelines for batch document processing, OCR workflows, and format conversion without manual steps. For legal, medical, or financial organizations where documents cannot be uploaded to cloud services, Stirling-PDF provides a GDPR-compliant, private PDF processing solution.

2 Comments

Post Your Comment

Fast • Reliable • Affordable VPS - DO It Now!

Get top VPS hosting with VPS.DO’s fast, low-cost plans. Try risk-free with our 7-day no-questions-asked refund and start today!