What we defend against

The realistic threats are not server breaches, because there is no file server to breach. They are:

  • A malicious file — something crafted to crash the browser, exhaust memory, or exploit a parser.
  • A disguised file — a script or archive renamed to .pdf or .png so it slips past a naive check.
  • Cross-site scripting — pasted text, filenames and file contents all reach the DOM, and none of them are trusted.
  • Resource exhaustion — a decompression bomb, an enormous image, or a regex that backtracks catastrophically and freezes the tab.
  • Supply chain — a compromised dependency shipping malicious code to every visitor.

File validation

No file extension is trusted, because an extension is a claim rather than evidence. Every file is checked on three independent axes before any parser or codec sees it:

  1. Extension — cheapest, and the easiest to forge. Used to give a sensible error message.
  2. MIME type — reported by the operating system, and also forgeable.
  3. Magic bytes — the actual file signature. This is the check that matters, and it is the one that gates processing.

A file whose contents do not match its extension is rejected before processing, with an explanation, rather than being handed to a parser that might not survive it. Multi-format inputs such as DOCX are validated as ZIP containers first and then inspected internally.

Processing isolation

Processing happens in a Web Worker wherever the work is heavy enough to justify it, which keeps a runaway parse from freezing the interface and lets a job be cancelled. Image work uses createImageBitmap and canvas, which decode in the browser's sandboxed image pipeline rather than in a hand-written parser.

Object URLs, image bitmaps and canvases are released explicitly on both the success and the error path. A long session that leaked one decoded bitmap per file would eventually take the tab down.

PDF and OOXML parsing runs with isEvalSupported disabled, so content-stream expressions inside a document are not evaluated.

Output encoding

Nothing that came from a user or a file is inserted as raw HTML anywhere in Furtu. The Markdown preview tool is the sharpest edge here, so it builds React elements for a documented subset of Markdown and escapes everything else, rather than setting innerHTML. Link targets are filtered to http, https and mailto, so a pasted javascript: URL is dropped rather than rendered as a clickable link.

Structured data is serialised with < escaped, so a value inside a JSON-LD block cannot terminate the surrounding script tag.

Denial of service

A tool that runs on a visitor's machine can be turned into a way to freeze that machine, so several guards exist specifically for that:

  • Decompression bombs — ZIP-based documents are checked against their declared uncompressed size before extraction, and rejected above a ceiling.
  • Catastrophic regex backtracking — the regular expression tester detects zero-length matches, advances past them, and caps the number of matches returned, so a pattern like (a*)* cannot hang the tab.
  • File and total size ceilings — enforced before parsing, and again during batch processing.
  • Bounded concurrency — batch operations process a few files at a time rather than decoding twenty images simultaneously, which reliably crashes a tab.
  • Target-size search limits — the image target-size tool bounds its quality and dimension search so it always terminates with a result or an honest failure.

Supply chain

Furtu is deliberately built on a small number of well-known, permissively licensed dependencies. Heavy libraries are loaded lazily, so a dependency in the PDF path cannot affect a visitor who only ever uses a text tool.

The complete list of dependencies, their versions and their licences is in the third-party notices in the project repository. Every open-source component retains its own copyright notice, and none of it is presented as Furtu-original work.

Where a project had no licence at all, its code was not used. That is not a technical constraint we worked around; unlicensed code is not reusable by anyone, including us.

Security headers

Static hosting is expected to send a restrictive content security policy, frame restrictions, a strict referrer policy and a permissions policy. The application makes no third-party requests during tool use, so the policy can stay tight: no inline scripts, no remote script hosts, and no connections anywhere except the page's own origin and the font CDN.

The recommended header set is documented in the deployment guide in the repository, so a host configuration can reproduce it rather than approximate it.

Reporting a vulnerability

Report anything that looks like a real problem to security@furtu.xyz. Please include the tool, the steps, and a file if you can share one safely.

If your report turns out to be a privacy claim on a tool page that is not technically accurate, that is a security report too — a false claim is a defect we want to know about.