Every Upload Is Untrusted Input Until You Prove Otherwise

An upload form accepts a file of the user's choosing, and your server then stores it, processes it and serves it back to other people. Each step can be abused. The worst cases are remote code execution (a script uploaded into a directory the web server executes) and stored XSS served from your own origin; the everyday ones are disk exhaustion, crashing image libraries and malware distribution. The fix is a pipeline, not a single check.

What Can Go Wrong

Abuse How it works Main control
Web shell shell.php lands in a folder the server executes Store outside the web root, never execute uploads
Stored XSS An HTML or SVG file is served inline from your domain Separate domain, attachment, CSP sandbox
Filename tricks ../../app/config.py, .htaccess, null bytes Generate the stored name yourself
Type confusion Extension and Content-Type say "image", bytes say otherwise Sniff, parse and re-encode
Resource exhaustion Huge files, decompression bombs, zip bombs Size, pixel and uncompressed-size caps
Parser exploits Malformed files, or XML inside SVG/DOCX (XML external entities) Patch, sandbox the worker, safe parsers
Malware hosting Your domain distributes someone else's payload Scanning, quarantine, access control

Magic Bytes Help, But Are Not Enough

In a multipart upload, the filename, the part's Content-Type and the body are all chosen by the client; with curl or Burp an attacker sends cat.png, image/png and a PHP web shell. An accept="image/*" attribute or a JavaScript check changes nothing. What the file is must be decided from the bytes, on the server.

Most formats start with a fixed signature, and checking it catches the lazy case:

SIGNATURES = {b"\x89PNG\r\n\x1a\n": "png", b"\xff\xd8\xff": "jpg",
              b"GIF89a": "gif", b"%PDF-": "pdf", b"PK\x03\x04": "zip"}

def sniff(head: bytes) -> str | None:
    return next((kind for magic, kind in SIGNATURES.items()
                 if head.startswith(magic)), None)

print(sniff(b"<?php echo 1; ?>"))                        # None
print(sniff(b"GIF89a<?php system($_GET['c']); ?>"))      # gif  <- still a webshell

The second line is the problem. Prefixing a script with GIF89a passes a signature check, and if that file lands somewhere PHP executes it, the interpreter prints the six junk bytes and runs the rest. Files valid in two formats at once (polyglots) are easy to build. Magic bytes tell you what a file claims to start with, not that it is safe.

For images, the stronger approach is to decode the whole file and re-encode it, which writes fresh bytes from pixels and drops metadata (including EXIF GPS coordinates), trailing data and anything hidden in comments.

A Safe Image Upload Handler

A Flask endpoint applying, in order: size cap, full decode, format allowlist, pixel cap, re-encode, random name, storage outside the web root.

import io
import secrets
from pathlib import Path

from flask import Flask, request, abort, jsonify
from PIL import Image, UnidentifiedImageError

STORAGE = Path("/srv/upload-store").resolve()      # NOT under the web root
STORAGE.mkdir(parents=True, exist_ok=True)
ALLOWED_FORMATS = {"JPEG": "jpg", "PNG": "png", "WEBP": "webp"}
Image.MAX_IMAGE_PIXELS = 40_000_000                 # refuse decompression bombs

app = Flask(__name__)
app.config["MAX_CONTENT_LENGTH"] = 5 * 1024 * 1024  # 5 MB, enforced before parsing

@app.post("/avatar")
def upload_avatar():
    f = request.files.get("file")
    if f is None:
        abort(400, "no file")
    data = f.read()
    try:
        with Image.open(io.BytesIO(data)) as probe:
            probe.verify()                           # structural check, no decode
        img = Image.open(io.BytesIO(data))           # verify() leaves it unusable
        fmt = img.format
        if fmt not in ALLOWED_FORMATS:
            abort(415, f"format {fmt} not allowed")
        img.load()                                   # full decode, catches truncation
    except (UnidentifiedImageError, Image.DecompressionBombError, OSError, SyntaxError):
        abort(415, "not a valid image")

    # Re-encode: throws away metadata, trailing bytes and anything smuggled inside
    img = img.convert("RGB") if fmt == "JPEG" else img
    ext = ALLOWED_FORMATS[fmt]
    stored_name = f"{secrets.token_hex(16)}.{ext}"
    out = io.BytesIO()
    img.save(out, format=fmt)
    (STORAGE / stored_name).write_bytes(out.getvalue())
    return jsonify(id=stored_name, format=fmt, display_name=f.filename[:100])

Non-obvious lines:

  • MAX_CONTENT_LENGTH rejects oversized requests with 413 before reading the body. Set the same limit at the proxy (client_max_body_size in nginx).
  • Image.MAX_IMAGE_PIXELS stops a tiny file that declares enormous dimensions from allocating gigabytes. Pillow warns above the limit and raises DecompressionBombError at twice it.
  • The extension comes from the detected format, never the uploaded name. The original name is only a display label, escaped wherever it is rendered (XSS).

Testing it with Flask's test client and five hostile inputs:

real PNG                 200  {'display_name': 'cat.png', 'format': 'PNG', 'id': '96ec1ed35f5a7a8cbec1bcd456ebf92f.png'}
PHP renamed to .png      415  415 UNSUPPORTED MEDIA TYPE
PNG with PHP appended    200  {'display_name': 'cat.php.png', 'format': 'PNG', 'id': '9e67c5fa33ad409e2b914d796fe106c4.png'}
SVG claiming image/png   415  415 UNSUPPORTED MEDIA TYPE
6 MB of zeros            413  413 REQUEST ENTITY TOO LARGE

The third case is a real image, so it is accepted, but the stored file is the re-encoded output and contains no <?php. Without re-encoding you would have stored the polyglot byte for byte. For files you cannot re-encode (PDFs, office documents), rely on format allowlists, scanning and safe serving.

Filenames: Never Use Them for Storage

The client-supplied filename is an attacker-controlled path. Werkzeug's secure_filename shows how much normalisation a name needs:

from werkzeug.utils import secure_filename

for name in ["holiday photo.jpg", "../../app/config.py", "shell.php.jpg",
             "report.pdf\x00.exe", "résumé.pdf", ".."]:
    print(repr(name), "->", repr(secure_filename(name)))
'holiday photo.jpg' -> 'holiday_photo.jpg'
'../../app/config.py' -> 'app_config.py'
'shell.php.jpg' -> 'shell.php.jpg'
'report.pdf\x00.exe' -> 'report.pdf.exe'
'résumé.pdf' -> 'resume.pdf'
'..' -> ''

Traversal is neutralised (see path traversal), but shell.php.jpg survives and an empty string is a possible result. Double extensions matter on Apache, where AddHandler matches an extension anywhere in the name, so shell.php.jpg can run as PHP. Keeping user names also enables uploading an .htaccess containing AddType application/x-httpd-php .png (every PNG in the folder becomes executable) and overwriting another user's file.

The robust rule: the stored name is random (secrets.token_hex(16)) plus an extension you derived from content. The user's name lives in the database as a label.

Storage and Execution

Code execution needs an attacker-controlled file and a server willing to execute it. Remove the second:

  • Store uploads outside the document root, or in object storage (S3, GCS, Azure Blob) that has no interpreter at all.
  • If files must sit under a web-served path, make that path static-only. In nginx, the ^~ modifier stops regex locations (such as your \.php$ handler) from matching anything under it:
location ^~ /media/ {
    alias /srv/upload-store/;
    default_type application/octet-stream;
    add_header X-Content-Type-Options "nosniff" always;
    add_header Content-Security-Policy "sandbox; default-src 'none'" always;
}
  • Mount the upload volume noexec.
  • For large files, let browsers upload directly to a quarantine bucket with a presigned POST whose policy has a content-length-range condition. A worker validates and scans, then copies clean files to the serving bucket.

Serving Without Becoming an XSS Host

SVG is XML that can contain <script>, and HTML uploads are HTML. If either is served from https://app.example/media/x.svg with its native content type, opening the link runs the script on your origin, with access to your users' sessions:

<svg xmlns="http://www.w3.org/2000/svg"><script>alert(document.domain)</script></svg>

Your options, from strongest to weakest:

  1. Do not accept SVG or HTML. For avatars and photos, raster formats are enough.
  2. Rasterise SVGs to PNG server-side (in a sandboxed worker, since the renderer parses untrusted XML).
  3. Serve user content from a separate registrable domain (for example app-usercontent.example, not a subdomain that shares cookies), so script there has no access to your application's origin.
  4. Send defensive headers on every user file:
Content-Type: image/png
X-Content-Type-Options: nosniff
Content-Disposition: attachment; filename="report.pdf"; filename*=UTF-8''r%C3%A9sum%C3%A9.pdf
Content-Security-Policy: sandbox; default-src 'none'

Content-Type comes from your mapping of the verified format, not the upload request. nosniff stops browsers guessing a scriptable type. attachment forces a download, and CSP sandbox puts anything that does render into an opaque origin with scripts disabled (details in security headers).

Archives and Documents

A zip entry named ../../etc/cron.d/job escapes the extraction folder (zip-slip), and a zip bomb expands a small archive into gigabytes. The sizes in zip headers are attacker-written, so count bytes as you decompress:

import io
import zipfile
from pathlib import Path

MAX_TOTAL = 200 * 1024 * 1024      # 200 MB uncompressed, counted as we read

def safe_extract(zip_bytes: bytes, dest: Path) -> int:
    dest = dest.resolve()
    total = 0
    with zipfile.ZipFile(io.BytesIO(zip_bytes)) as zf:
        for info in zf.infolist():
            target = (dest / info.filename).resolve()
            if not target.is_relative_to(dest):
                raise ValueError(f"zip-slip entry: {info.filename}")
            if info.is_dir():
                continue
            target.parent.mkdir(parents=True, exist_ok=True)
            with zf.open(info) as src, open(target, "wb") as out:
                while chunk := src.read(65536):
                    total += len(chunk)
                    if total > MAX_TOTAL:
                        raise ValueError("uncompressed size limit exceeded")
                    out.write(chunk)
    return total

Feeding it a zip containing 1 GB of zeros:

archive size: 1019 KiB
rejected: uncompressed size limit exceeded

Also cap the number of entries, and extract into a fresh temporary directory that you delete on any exception, since a partial file is left behind when the limit trips. Office formats (DOCX, XLSX) are zip archives of XML, so the same limits apply, plus a safe XML parser.

Malware Scanning

If users download each other's uploads, you are a distribution channel. Scan in quarantine with an engine such as ClamAV:

clamdscan --fdpass --no-summary /srv/quarantine/9e67c5fa.bin
echo $?     # 0 = clean, 1 = virus found, 2 = error (treat as "not clean")

Treat an error or timeout as a failure. Signature scanners catch known malware, not novel payloads, so scanning complements the other controls. Run all file processing (image decoding, PDF thumbnails, document conversion) in an isolated worker with no network access and tight CPU and memory limits, because those parsers are where memory-corruption bugs live.

Testing Your Own Upload Feature

Run these against a staging copy of your app, or practise on DVWA's upload module and PortSwigger's file upload labs:

  • A script renamed to .png with Content-Type: image/png; a real image with a script appended; a GIF89a-prefixed script.
  • Filenames with ../, a leading /, a null byte, .htaccess, and the name of an existing upload.
  • An SVG and an HTML file containing script, then open the returned URL directly.
  • A file one byte over the limit, a decompression-bomb image, and a zip bomb.
  • Finally, fetch an uploaded file and confirm: random name, your own Content-Type, nosniff, attachment where appropriate, and a separate domain.