Secure File Uploads
Every Upload Is Untrusted Input Until You Prove Otherwise
An upload form accepts a file of the user's choosing, and your server then stores it, processes it and serves it back to other people. Each step can be abused. The worst cases are remote code execution (a script uploaded into a directory the web server executes) and stored XSS served from your own origin; the everyday ones are disk exhaustion, crashing image libraries and malware distribution. The fix is a pipeline, not a single check.
What Can Go Wrong
| Abuse | How it works | Main control |
|---|---|---|
| Web shell | shell.php lands in a folder the server executes |
Store outside the web root, never execute uploads |
| Stored XSS | An HTML or SVG file is served inline from your domain | Separate domain, attachment, CSP sandbox |
| Filename tricks | ../../app/config.py, .htaccess, null bytes |
Generate the stored name yourself |
| Type confusion | Extension and Content-Type say "image", bytes say otherwise |
Sniff, parse and re-encode |
| Resource exhaustion | Huge files, decompression bombs, zip bombs | Size, pixel and uncompressed-size caps |
| Parser exploits | Malformed files, or XML inside SVG/DOCX (XML external entities) | Patch, sandbox the worker, safe parsers |
| Malware hosting | Your domain distributes someone else's payload | Scanning, quarantine, access control |
Magic Bytes Help, But Are Not Enough
In a multipart upload, the filename, the part's Content-Type and the body are all chosen by the client; with curl or Burp an attacker sends cat.png, image/png and a PHP web shell. An accept="image/*" attribute or a JavaScript check changes nothing. What the file is must be decided from the bytes, on the server.
Most formats start with a fixed signature, and checking it catches the lazy case:
SIGNATURES = {b"\x89PNG\r\n\x1a\n": "png", b"\xff\xd8\xff": "jpg",
b"GIF89a": "gif", b"%PDF-": "pdf", b"PK\x03\x04": "zip"}
def sniff(head: bytes) -> str | None:
return next((kind for magic, kind in SIGNATURES.items()
if head.startswith(magic)), None)
print(sniff(b"<?php echo 1; ?>")) # None
print(sniff(b"GIF89a<?php system($_GET['c']); ?>")) # gif <- still a webshell
The second line is the problem. Prefixing a script with GIF89a passes a signature check, and if that file lands somewhere PHP executes it, the interpreter prints the six junk bytes and runs the rest. Files valid in two formats at once (polyglots) are easy to build. Magic bytes tell you what a file claims to start with, not that it is safe.
For images, the stronger approach is to decode the whole file and re-encode it, which writes fresh bytes from pixels and drops metadata (including EXIF GPS coordinates), trailing data and anything hidden in comments.
A Safe Image Upload Handler
A Flask endpoint applying, in order: size cap, full decode, format allowlist, pixel cap, re-encode, random name, storage outside the web root.
import io
import secrets
from pathlib import Path
from flask import Flask, request, abort, jsonify
from PIL import Image, UnidentifiedImageError
STORAGE = Path("/srv/upload-store").resolve() # NOT under the web root
STORAGE.mkdir(parents=True, exist_ok=True)
ALLOWED_FORMATS = {"JPEG": "jpg", "PNG": "png", "WEBP": "webp"}
Image.MAX_IMAGE_PIXELS = 40_000_000 # refuse decompression bombs
app = Flask(__name__)
app.config["MAX_CONTENT_LENGTH"] = 5 * 1024 * 1024 # 5 MB, enforced before parsing
@app.post("/avatar")
def upload_avatar():
f = request.files.get("file")
if f is None:
abort(400, "no file")
data = f.read()
try:
with Image.open(io.BytesIO(data)) as probe:
probe.verify() # structural check, no decode
img = Image.open(io.BytesIO(data)) # verify() leaves it unusable
fmt = img.format
if fmt not in ALLOWED_FORMATS:
abort(415, f"format {fmt} not allowed")
img.load() # full decode, catches truncation
except (UnidentifiedImageError, Image.DecompressionBombError, OSError, SyntaxError):
abort(415, "not a valid image")
# Re-encode: throws away metadata, trailing bytes and anything smuggled inside
img = img.convert("RGB") if fmt == "JPEG" else img
ext = ALLOWED_FORMATS[fmt]
stored_name = f"{secrets.token_hex(16)}.{ext}"
out = io.BytesIO()
img.save(out, format=fmt)
(STORAGE / stored_name).write_bytes(out.getvalue())
return jsonify(id=stored_name, format=fmt, display_name=f.filename[:100])
Non-obvious lines:
MAX_CONTENT_LENGTHrejects oversized requests with 413 before reading the body. Set the same limit at the proxy (client_max_body_sizein nginx).Image.MAX_IMAGE_PIXELSstops a tiny file that declares enormous dimensions from allocating gigabytes. Pillow warns above the limit and raisesDecompressionBombErrorat twice it.- The extension comes from the detected format, never the uploaded name. The original name is only a display label, escaped wherever it is rendered (XSS).
Testing it with Flask's test client and five hostile inputs:
real PNG 200 {'display_name': 'cat.png', 'format': 'PNG', 'id': '96ec1ed35f5a7a8cbec1bcd456ebf92f.png'}
PHP renamed to .png 415 415 UNSUPPORTED MEDIA TYPE
PNG with PHP appended 200 {'display_name': 'cat.php.png', 'format': 'PNG', 'id': '9e67c5fa33ad409e2b914d796fe106c4.png'}
SVG claiming image/png 415 415 UNSUPPORTED MEDIA TYPE
6 MB of zeros 413 413 REQUEST ENTITY TOO LARGE
The third case is a real image, so it is accepted, but the stored file is the re-encoded output and contains no <?php. Without re-encoding you would have stored the polyglot byte for byte. For files you cannot re-encode (PDFs, office documents), rely on format allowlists, scanning and safe serving.
Filenames: Never Use Them for Storage
The client-supplied filename is an attacker-controlled path. Werkzeug's secure_filename shows how much normalisation a name needs:
from werkzeug.utils import secure_filename
for name in ["holiday photo.jpg", "../../app/config.py", "shell.php.jpg",
"report.pdf\x00.exe", "résumé.pdf", ".."]:
print(repr(name), "->", repr(secure_filename(name)))
'holiday photo.jpg' -> 'holiday_photo.jpg'
'../../app/config.py' -> 'app_config.py'
'shell.php.jpg' -> 'shell.php.jpg'
'report.pdf\x00.exe' -> 'report.pdf.exe'
'résumé.pdf' -> 'resume.pdf'
'..' -> ''
Traversal is neutralised (see path traversal), but shell.php.jpg survives and an empty string is a possible result. Double extensions matter on Apache, where AddHandler matches an extension anywhere in the name, so shell.php.jpg can run as PHP. Keeping user names also enables uploading an .htaccess containing AddType application/x-httpd-php .png (every PNG in the folder becomes executable) and overwriting another user's file.
The robust rule: the stored name is random (secrets.token_hex(16)) plus an extension you derived from content. The user's name lives in the database as a label.
Storage and Execution
Code execution needs an attacker-controlled file and a server willing to execute it. Remove the second:
- Store uploads outside the document root, or in object storage (S3, GCS, Azure Blob) that has no interpreter at all.
- If files must sit under a web-served path, make that path static-only. In nginx, the
^~modifier stops regex locations (such as your\.php$handler) from matching anything under it:
location ^~ /media/ {
alias /srv/upload-store/;
default_type application/octet-stream;
add_header X-Content-Type-Options "nosniff" always;
add_header Content-Security-Policy "sandbox; default-src 'none'" always;
}
- Mount the upload volume
noexec. - For large files, let browsers upload directly to a quarantine bucket with a presigned POST whose policy has a
content-length-rangecondition. A worker validates and scans, then copies clean files to the serving bucket.
Serving Without Becoming an XSS Host
SVG is XML that can contain <script>, and HTML uploads are HTML. If either is served from https://app.example/media/x.svg with its native content type, opening the link runs the script on your origin, with access to your users' sessions:
<svg xmlns="http://www.w3.org/2000/svg"><script>alert(document.domain)</script></svg>
Your options, from strongest to weakest:
- Do not accept SVG or HTML. For avatars and photos, raster formats are enough.
- Rasterise SVGs to PNG server-side (in a sandboxed worker, since the renderer parses untrusted XML).
- Serve user content from a separate registrable domain (for example
app-usercontent.example, not a subdomain that shares cookies), so script there has no access to your application's origin. - Send defensive headers on every user file:
Content-Type: image/png
X-Content-Type-Options: nosniff
Content-Disposition: attachment; filename="report.pdf"; filename*=UTF-8''r%C3%A9sum%C3%A9.pdf
Content-Security-Policy: sandbox; default-src 'none'
Content-Type comes from your mapping of the verified format, not the upload request. nosniff stops browsers guessing a scriptable type. attachment forces a download, and CSP sandbox puts anything that does render into an opaque origin with scripts disabled (details in security headers).
Archives and Documents
A zip entry named ../../etc/cron.d/job escapes the extraction folder (zip-slip), and a zip bomb expands a small archive into gigabytes. The sizes in zip headers are attacker-written, so count bytes as you decompress:
import io
import zipfile
from pathlib import Path
MAX_TOTAL = 200 * 1024 * 1024 # 200 MB uncompressed, counted as we read
def safe_extract(zip_bytes: bytes, dest: Path) -> int:
dest = dest.resolve()
total = 0
with zipfile.ZipFile(io.BytesIO(zip_bytes)) as zf:
for info in zf.infolist():
target = (dest / info.filename).resolve()
if not target.is_relative_to(dest):
raise ValueError(f"zip-slip entry: {info.filename}")
if info.is_dir():
continue
target.parent.mkdir(parents=True, exist_ok=True)
with zf.open(info) as src, open(target, "wb") as out:
while chunk := src.read(65536):
total += len(chunk)
if total > MAX_TOTAL:
raise ValueError("uncompressed size limit exceeded")
out.write(chunk)
return total
Feeding it a zip containing 1 GB of zeros:
archive size: 1019 KiB
rejected: uncompressed size limit exceeded
Also cap the number of entries, and extract into a fresh temporary directory that you delete on any exception, since a partial file is left behind when the limit trips. Office formats (DOCX, XLSX) are zip archives of XML, so the same limits apply, plus a safe XML parser.
Malware Scanning
If users download each other's uploads, you are a distribution channel. Scan in quarantine with an engine such as ClamAV:
clamdscan --fdpass --no-summary /srv/quarantine/9e67c5fa.bin
echo $? # 0 = clean, 1 = virus found, 2 = error (treat as "not clean")
Treat an error or timeout as a failure. Signature scanners catch known malware, not novel payloads, so scanning complements the other controls. Run all file processing (image decoding, PDF thumbnails, document conversion) in an isolated worker with no network access and tight CPU and memory limits, because those parsers are where memory-corruption bugs live.
Testing Your Own Upload Feature
Run these against a staging copy of your app, or practise on DVWA's upload module and PortSwigger's file upload labs:
- A script renamed to
.pngwithContent-Type: image/png; a real image with a script appended; aGIF89a-prefixed script. - Filenames with
../, a leading/, a null byte,.htaccess, and the name of an existing upload. - An SVG and an HTML file containing script, then open the returned URL directly.
- A file one byte over the limit, a decompression-bomb image, and a zip bomb.
- Finally, fetch an uploaded file and confirm: random name, your own
Content-Type,nosniff,attachmentwhere appropriate, and a separate domain.