Inpainting.app

How Inpainting.app runs AI image models in your browser

Updated September 28, 2026 · Inpainting.app

Every tool on Inpainting.app runs its AI model inside your browser tab. There is no image upload, no server-side processing and no account. This article explains how that works, what it costs you as a visitor, and the practical browser limits we ran into while building it. It is written for curious users and for developers who want to do something similar.

The building blocks

The models are ordinary neural networks exported to the ONNX format, an open standard for describing a trained model as a graph of operations. We run them with ONNX Runtime Web, Microsoft's JavaScript and WebAssembly build of the ONNX Runtime engine. It can execute the graph in two ways:

All inference happens in a Web Worker, a background thread, so the page stays responsive while a model is working.

The models

Each model file is pinned to an exact published version and checked against a SHA-256 checksum when the site is built.

Downloading models only when you ask

A 90 MB model should not be downloaded just because someone opened a web page. The tool pages load nothing heavy: no model, no AI runtime, not even the worker script. When you start a task for the first time, the tool shows how many megabytes it needs and waits for you to click Download and continue.

Our host serves files of at most 25 MB, so every model and the WebGPU build of ONNX Runtime are split into chunks of up to 20 MB. The browser downloads up to three chunks at a time, verifies each against its own SHA-256 checksum, joins them in memory and hands the complete model to ONNX Runtime.

Verified chunks are stored in the browser's Cache Storage. On your next visit the tool finds them there, skips the question and starts almost instantly, even offline. Chunk file names contain a hash of their content, so when a model is updated the new version is fetched and the old one simply stops being used. The site also asks the browser to treat this storage as persistent, so it is less likely to be cleared when disk space runs low. You can remove everything at any time by clearing the site data for tool.inpainting.app in your browser settings.

Keeping images private by design

The tools run on a separate address, tool.inpainting.app, embedded in the pages you read. That page loads no advertising, analytics or third-party scripts, and its Content Security Policy forbids it from connecting to any server other than its own. Even a bug could not send your image anywhere. The only message the tool sends to the surrounding page is its own height, so the frame can resize.

What we learned about browser limits

Running large models in a browser tab is still an engineering exercise. Three findings from building this site may save other developers some time.

1. GPU limits decide which models you can use

For background removal we first wanted BiRefNet-lite, which produced the cleanest cut-outs in our offline tests. It does not run in the browser on Apple hardware today: one of its operations needs 11 storage buffers in a single GPU shader, while WebGPU on Apple GPUs allows 10. On the WebAssembly path the model needs more than the 4 GB of memory a WebAssembly module can address at its 1024 × 1024 working size. The same happened with BEN2. RMBG-1.4, a convolutional model, runs comfortably on both paths, so that is what the background remover uses.

2. A model can fail silently

LaMa ran on WebGPU without any error, but the filled area came back almost pure white: some of its Fourier-transform operations produce wrong numbers on the GPU path. The same model on WebAssembly is correct. We now run LaMa on the processor only, and our automated tests check the colours of the actual output, not just that a run finished. If you ship models to the web, test the pictures, not the absence of errors.

3. Half precision now arrives as Float16Array

Some models use 16-bit floating point numbers to halve their size and speed up GPU work. Older browsers had no native 16-bit float type, so ONNX Runtime Web returned such outputs as raw bits in a Uint16Array. Current Chrome supports Float16Array, and the runtime now returns real numbers instead. Code that assumed raw bits turned a perfectly good upscale into a black image. Handle both representations.

The trade-off

Running locally means the models have to be small enough to download and fast enough for a laptop or phone. Large image-generation models behind commercial tools can reconstruct bigger and more complex areas than a 28 to 93 MB model can. In exchange, local tools cost nothing to run, work offline after the first use, and never see your photos. For everyday edits such as removing a stranger from a holiday snap, cutting out a product or enlarging an old picture, that is a trade we think is worth making.

Try the tools →

Example photos: coffee cup (Rachel Michetti) and cat (Stefan van der Walt), CC0; Falcon 9 launch pad (SpaceX) and Eileen Collins (NASA), public domain. All results shown were produced with the tools on this site.

More guides