Scan documents without uploading them

Passports, contracts, payslips, medical letters. The documents most worth scanning are the ones least worth handing to a stranger's server. ScanPDF was built so that handing them over is not part of the deal.

What a typical online converter does

The usual "photo to PDF" site is a web form in front of a server. Your file is uploaded, converted somewhere you cannot see, and offered back as a download. Somewhere in the terms there is a promise to delete it after an hour or a day.

That promise may well be kept. But it is a promise, not a mechanism, and it does not cover the parts you were never told about: how many machines the file passed through, what was logged, which subprocessor stored the backup, or what a future breach exposes. For an ID document or a signed contract, the honest risk assessment is that once the upload completes, the outcome is no longer yours to control.

What ScanPDF does instead

There is no upload step, because there is nothing to upload to. ScanPDF is a static page. Your browser reads the photo from disk, and everything after that happens in the tab:

  • Edge detection runs on a WebAssembly build of OpenCV, executing on your CPU.
  • Perspective correction and filtering are applied to the full-resolution image in memory.
  • The PDF is assembled locally with pdf-lib and handed to the browser as a normal download.

No request carries your image data, because no such request exists in the code. The whole thing works with the network cable pulled out once the app has loaded. Photos you send from another app through Share and then ScanPDF are handed over on the device as well: the app's service worker answers that request itself, and nothing is transmitted. One edge case is worth knowing about: if you clear the site's data while the app is still installed on Android, there is nothing left on the device to receive the next share, so the browser sends those photos to the web host, which does not accept them. Open the app once after clearing its data and sharing stays on the device again.

Verifiable, not just stated

A privacy claim you cannot check is worth very little, so this one is arranged to be checkable three ways.

  • Watch it yourself. Open your browser's developer tools, switch to the Network tab, and scan a document. After the app's own files have loaded, nothing further is sent.
  • Read the source. The whole app is MIT-licensed and lives on GitHub - a few hundred lines of plain JavaScript modules with no build step to obscure them, and two dependencies pinned by URL and SHA-256.
  • Let the browser enforce it. The self-hosted container ships a strict Content-Security-Policy limiting connections to its own origin. Under it, an accidental request to any external host is blocked by the browser rather than merely absent from the code.

The hosted demo versus self-hosting

It is worth being precise about the difference, because most sites are not.

The demo at scanpdf.io is the same code, served from GitHub Pages. Processing is fully client-side there too, and your documents still never leave the machine. What differs is who controls the response headers: on GitHub Pages that is GitHub, so the strict CSP from the project's own nginx config is not applied. GitHub also sees the ordinary web-server request for the page itself, as any host would.

If you want the guarantee enforced by the browser rather than resting on the code being what it says it is, run it yourself. One command starts the container, and it needs no network access at run time at all.

No account, no telemetry, no cookies

There is nothing to sign up for, so there is no email address on file. There is no analytics script, no tracking pixel, no error reporting service and no A/B testing framework. The app sets no cookies. What it keeps on your device is limited to your interface preferences, such as the theme and the page size, a copy of the app's own files so that it opens offline, and, briefly, photos shared from another app, which are deleted as soon as the scanner opens. Your pages exist in memory only; reload the page and every trace of the session is gone.

This is less a policy than an architecture. A static page with no backend has nowhere to put the data it would need to misuse.

Try it

Open the scanner and drop in a document you would not have uploaded anywhere. That is rather the point. If you are new to photographing paper, the scanning guide covers getting a clean result on the first try.