Comic text ONNX models, half precision

Float16 conversions of three ONNX models for manga and comic text processing, aimed at onnxruntime-web on the WebGPU execution provider. The weights are unchanged apart from the precision cast; no retraining, pruning, or architecture edit.

These exist because WebGPU rewards half precision and, in one case, punishes integer quantisation badly.

Files

file converted from licence
detector-fp16.onnx ogkalu/comic-text-and-bubble-detector detector.onnx Apache-2.0
comic-text-detector-fp16.onnx mayocream/comic-text-detector-onnx comic-text-detector.onnx GPL-3.0
aot-fp16.onnx lemondouble/lemon-manga-translator onnx/aot-inpainting/aot_folded.onnx GPL-3.0

Each file keeps the licence of the model it was converted from. They are distributed together as an aggregate, which GPL-3.0 section 5 permits; the GPL-3.0 files are not combined with the Apache-2.0 one.

Measured

Chrome, Apple M3, onnxruntime-web on WebGPU, median of repeated runs at the input size these are run at. "Shipping" is the artifact each was converted from, except for detection, where it is the int8 build that was in use before.

model shipping this file agreement with float32
detector 1610 ms (int8) 140 ms same 10 regions on a real page, worst box IoU 0.969
comic-text-detector 468 ms 290 ms 0.006% of mask pixels land on the other side of the threshold
aot 310 ms 237 ms mean absolute difference 0.6/255, worst 15/255

The detector number is not a half precision result so much as an integer one. WebGPU has no kernels for ConvInteger or MatMulInteger, so every quantised node in the int8 build crosses back to the CPU. Its float32 build already runs in 188 ms; half precision takes it to 140 ms.

A LaMa conversion is deliberately absent. It overflows in the fast Fourier convolution branch and returns NaN, for a 16% gain that would not have been worth it anyway.

How they were made

onnxconverter-common's convert_float_to_float16 with keep_io_types=True, followed by a repair pass. The converter retypes tensors but leaves each Cast node's to attribute on the original dtype, so the graph declares float32 where it now produces float16 and the session refuses to build. Setting to from the inferred output type fixes 5 casts in the detector and 10 in AOT.

Inputs and outputs stay float32, so these are drop-in replacements: no caller change beyond the file name.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support