QuadTok: quadtree visual tokenizer cuts tokens ~10% for autoregressive image generation
The paper introduces QuadTok, a hierarchical quadtree framework for visual tokenization that bridges 2D spatial binding and 1D sequence flexibility, allocating representational capacity to visually intricate areas while leaving homogeneous regions at coarse resolution. Its ImageNet-trained tokenizer reportedly saves about 10% of tokens on ImageNet versus a fixed 256-token grid, and about 9% when transferred zero-shot to COCO, while maintaining comparable reconstruction fidelity. Conditioned on a quadtree topology supplied before generation, a 947M GPT-style generative model reaches 2.08 gFID on ImageNet 256×256, and the preserved spatial correlation also enables zero-shot spatially controlled image generation.