Referral Program: Refer a company & earn 10% commission on the service fees they pay us. Learn more

Blog

OCR to XSS

OCR to XSS

Optical Character Recognition (OCR) is the process of electronically extracting text from images or documents like PDFs, and reusing it in a variety of ways such as full-text search, invoice processing, and document verification. This becomes a security issue when the extracted text is reflected back into an application without validation.

If an application parses an image and reflects the extracted text to a user, embedding an XSS vector as text inside that image can lead to XSS.

Proof of Concept

Take a simple JPG containing an XSS payload as text. You can create an image like that here.

This PoC uses tesseract for OCR, along with a simple Flask server that accepts an image as input, parses it, and reflects the extracted content back to an admin or another user. The server code is available here.

XSS payload rendered as text inside an image, ready for OCR extraction

Steps to reproduce:

  1. Start the server: python ocr.py
  2. Visit the local server at 127.0.0.1:5000
  3. Upload the image above
  4. Visit /admin/ocr/files
  5. The XSS payload fires

The XSS alert firing after the OCR server reflects the extracted payload

Similarly, you can create an image containing a blind XSS payload to confirm a pingback to a server you control.

Note: different OCR parsers handle certain characters differently. Tesseract, for example, treats a forward slash ("/") as the letter "L", so http:// becomes http:/l, which won't resolve correctly in a browser. Using backslashes works around this for Tesseract; other parsers may need their own workaround.

I used ngrok here just to confirm the ping. Burp Collaborator or any similar out-of-band tool works just as well. Create an image with this content, upload it, and check whether you get a hit.

Payload image crafted for a blind XSS pingback

Response

Out-of-band pingback confirming the blind XSS fired

Fix

If you're using OCR services, don't just sanitize the filename. Sanitize the extracted text pulled from the image or PDF before it's stored in the database or reflected back to a user.

After uploading an image, check whether its contents are reflected anywhere in the response. If so, and there's no validation on how that extracted text is displayed, it can lead to XSS, particularly in applications using OCR for KYC, document verification, or scanned-document uploads.

So next time an application asks you to upload a scanned document, a passport photo, or anything else for verification, it's worth testing how that OCR pipeline handles what it extracts.

Back to the Blog