← Eduard Zamfir

Elastic Token Compression for
Pixel-Space Diffusion Transformers

1Computer Vision Lab, University of Würzburg    2Google

A pretrained diffusion transformer spends one token on every patch, in every layer and at every step. We keep the first and last blocks dense and run the middle on a few hundred content-sized region tokens: large over flat areas, single patches over detail.

We cut a Hilbert ordering of the patch tokens where the model's own features change most, so grouping in two dimensions becomes cutting a line in one. A learned Read pools each run into one token carrying its size and position, and a learned Write returns the core's update to its patch tokens. The backbone stays frozen apart from low-rank adapters, and the region count is drawn at random while fine-tuning, so one checkpoint runs at every budget, 1024 tokens down to 64. On half the tokens, RTI holds the frozen model's quality on GenEval and the three preference scores at 2.11× the speed. FID is the exception: it recovers most of the reduction's cost but not all of it, +2.4 at that setting.

Three samples with the uniform patch grid overlaid on top and our region partition below; arrows report how many region tokens cover the area where the grid spends about a hundred patches.
Content-adaptive tokenization. The same three images under a uniform grid (top) and our partition (bottom). Each arrow reports how many region tokens cover the area where the grid spends about a hundred patch tokens: 4 over a plain backdrop and 3 over empty sky, but 23 over the balloon, where the detail is. Size follows content in both directions. The overlays compare tokenizations, not outputs.
01

Samples

MiniT2I[1]-L/16 · 512² · 100 steps · cfg 5.0 · one seed

Select a prompt and use the slider to set the compression rate. The middle column is one RTI checkpoint queried at five budgets. The right column is feature-similarity merging, the strongest prior reduction, at the same budgets and the same speed. The left column never changes: it is the frozen backbone at all 1024 patch tokens.

Prompt
Showing

Dense1024 / 1024 · 1.00×
Sample from the frozen dense backbone
RTI (ours)
Sample from RTI at the selected budget
Feat Sim
Sample from feature-similarity merging at the selected budget
cropdense
crop
crop
Click any sample to move the crop ·

The dense backbone runs on all 1024 patch tokens; every other setting is that sequence cut to the length shown. Dropping a quarter of it, 728 of 1024, leaves the image essentially unchanged — and 728 is a length neither checkpoint ever trained on, since the budget is drawn from {64, 128, 256, 512}. Both reductions run inside the same interface, so the only difference between the two right-hand columns is the rule that forms the groups. Measured speedups are in the results section below.

02

The partition on the trajectory

Every video: 100 sampling steps

The partition is not fixed: we rebuild it from the model's own features at every step, so it tracks the image as the image appears. Early on there is nothing to track and the regions are large everywhere; as detail forms, the cuts crowd into it and the flat parts of the frame stay in a handful of tokens.

Prompt
Budget
Current image estimateThe same estimate, region boundaries drawn

Left: the model's running estimate of the clean image. Right: the same frame with the boundaries of the regions the core is running on. Same sampler, seed and prompt as the gallery above.

02b

Contiguous, or scattered

Same prompt · same seed · R = 512

Both methods turn the 1024 patch tokens into 512 tokens, two patches each on average; they differ in which patches each token is built from. We mark six grid positions — the same six on both sides — and paint the whole group that each one belongs to.

Ours are connected: a group covers one place, wide where the image is flat and down to a single patch where it is not. Similarity merging chooses members by appearance alone, so one group collects patches from opposite corners of the frame. Its token stands nowhere in particular, and the pretrained rotary attention is left with no position to read for it.

Prompt
Columns: the sample, then its six marked groupsRows: RTI (ours) on top, Feat Sim below

Measured over 64 prompts at R = 256: a patch sits 0.5 grid cells from its group's centre under our partition, 5.2 under similarity merging, against 12.2 for a random assignment.

02c

One checkpoint, every budget

Partition swept at a fixed step

We cut the Hilbert order at the largest feature gaps, so the cut sets are nested: lowering R only ever merges neighbouring regions, and one trained checkpoint serves every setting.

We hold the image fixed and sweep the budget from R = 1024 down to R = 32 and back. Carved stone and feather barbs hold their cuts longest; flat ground and blank wall merge first.

03

Method

Region Token Interface

We probe the frozen MiniT2I[1]-B/16 before any training. Three measurements each fix one design choice. Detail is concentrated in space, so regions must be sized by content (a). The representation grows more poolable as the image forms, so the partition is rebuilt at every step (b). The redundancy sits in the middle blocks, so only those are compressed (c).

a · Space

Detail is concentrated

The most-detailed 15% of patches hold half of the detail; half of them hold 88%.

b · Time

More poolable as the image forms

c · Depth

Poolable in the core, sensitive at the ends

Representation retained = explained variance of a middle block's tokens after pooling the N=1024 patch tokens to R=256 and scattering back. Adaptive is our partition, uniform is the equal-length cut of the same order, skip drops tokens instead of pooling them. Panel (c) is why the first three and last three blocks stay dense: the core is blocks 3–13 of the 17 on B/16 and 3–19 of the 23 on L/16, counted from zero and inclusive, so both leave a three-block prelude and a three-block coda at full resolution.

03b

Four properties of the grouping

And what each prior reduction gives up

We measure the effect of coarsening the token grid on the frozen model to investigate the redundancy, and the measurements fix four properties a reduction must keep. Redundant patches lie next to one another in flat stretches of the frame, so groups must be contiguous. In natural images detail concentrates in a small part of the frame, so group size must follow the content, large over a flat wall and small over foliage. A group covers a definite place that its token has to stand for, so each must carry a position. And dropping loses what pooling keeps, so the redundant content must be summarized rather than deleted.

Each prior reduction forfeits at least one of the four, and each is worst precisely where it does. Token Skip[3] routes the most important patch tokens through a block and lets the rest bypass it: its units are single patch tokens, connected and positioned by construction, but what bypasses is deleted, not summarized. Feature-similarity merging[4][5] pools patches that look alike regardless of where they are, so a group can gather patches from opposite corners of the frame; its token stands nowhere in particular, losing contiguity and position at once. A latent array[6][7] replaces grouping with R free latents that read the whole frame by cross-attention, giving up grouping and position together.

GenEval on MiniT2I-B/16 at matched budget, with the difference to RTI. (✓) holds only trivially: Token Skip never groups, so each unit is a single patch token, connected and positioned by itself, while the admitted set scatters across the frame and nothing stands for the patches that bypass. It is also scored at its own operating point, 576 mean tokens at 1.19×, against RTI at 256 and 2.15×, so its difference understates the gap.
MethodHow it groupsContiguousContent-sizedCarries positionSummarizesGenEval ↑Δ
RTI (ours)contiguous runs of the Hilbert order✓✓✓✓84.7—
Token Skip[3]top 12.5% of patch tokens enter the block, the rest bypass(✓)✗(✓)✗82.3−2.4
Feat Sim[4]nearest of R anchors by feature similarity✗✗✗✓80.5−4.2
Latent Array[6][7]R learned latents read every patch token by cross-attention✗✗✗✓45.7−39.0

Removing position costs an order of magnitude more than anything else.

03c

The interface

Cut, Read, core, Write
RTI architecture: prelude blocks at full resolution, region partition and Read, the DiT core on R region tokens, Write, coda blocks at full resolution.
An elastic region-token interface. The first and last blocks keep all N patch tokens. The core between them runs on R ≪ N region tokens. Left inset: cuts at the R−1 largest feature gaps along the Hilbert order give the regions; Read pools each into one token carrying its size and position. Right inset: Write returns each region token's change to its patch tokens, specialised per token by a learned map g. Grey: LoRA. Green: the interface.

From an order to regions

MiniT2I[1] takes a 512² image as a 32×32 grid of N = 1024 patch tokens and follows the JiT[2] convention of predicting the clean image instead of the noise or the velocity, so its features track image structure. A Hilbert curve[8] visits every patch of that grid exactly once, and consecutive positions along it are always neighbours in the image. Any contiguous run of the order is therefore a connected region, and grouping in two dimensions becomes cutting a line in one.

Each boundary along the curve is scored by the feature difference across it:

dj = ‖ hπ(j+1) − hπ(j) ‖2

This is a first difference along the curve, an edge detector at patch resolution on mid-trunk features rather than on pixels. We cut at the R−1 largest scores: homogeneous stretches stay merged, detailed areas split down to single patches. The rule has no parameters and costs one pass over the sequence, cheap enough to recompute at every sampling step, so the partition follows the image as it forms.

Read: patch tokens into a region token

A region token is a learned weighted average of the region's patch tokens, plus an embedding of the region's size:

αj = softmaxj∈c( w⊤hj )
zc = Σj∈c αj hj + esize(⌊log₂|c|⌋)

The softmax runs within the region, so Read can put its weight on a region's informative tokens instead of averaging blindly, and the size embedding tells the core whether a token stands for one patch or sixteen. The token's position is the average of its member tokens' rotary embeddings. Those are unit phasors, so the average survives at wavelengths longer than the region and shrinks below it: position is sharp for one patch and coarse for a large region, which is what a token covering an area needs.

Write: the update back to the patch tokens

After the core, a region's token has changed by z′c − zc. Write adds that change to each patch token as a residual on its pre-core features, with a small learned map g conditioning the update on the token itself:

hj ← hj + g( [ hj ‖ z′c(j) − zc(j) ] )

Tokens in one region therefore receive individual updates rather than one shared copy. And because every patch token keeps its pre-core features, the core's work arrives as a correction on top of them, which is how the coda still produces per-pixel output although the core never saw individual patch tokens.

Core span and fine-tuning

The core is the contiguous span of middle blocks where the depth probe puts the redundancy, 3–13 of 17 on B/16 and 3–19 of 23 on L/16 — counted from zero and inclusive, so three blocks of prelude and three of coda in both cases; those stay at full resolution. Its extent trades against its width: a longer core must run on more region tokens to hold the same quality, a shorter one needs fewer, and at matched compute the two are equivalent. The probe fixes where the span sits, not a unique length.

The trunk stays frozen. We train only LoRA[9] adapters and the Read/Write parameters, on MiniT2I's own fine-tuning data and settings. Both are initialised to exact mean-pool and broadcast, so training starts from the frozen model's own behaviour. We draw the budget R at each step from {64, 128, 256, 512}, so one checkpoint covers the range.

04

Results

Text-to-image · MiniT2I[1]-B/16 and L/16

Quality against speed

MiniT2I-L/16 · dashed rule is the frozen dense backbone

Data

MJHQ-30K FID

L/16 · 30 000 prompts · cfg 5.0 · clean-fid · lower is better

RTI at R=256 (14.71) beats Feat Sim at R=512 (19.99): better images at half the tokens and a higher speedup.

Text-to-image benchmarks. GenEval[11] measures compositional accuracy; CLIP[13], PickScore[14] and ImageReward[15] measure alignment and preference, over PartiPrompts[12]. Speedups are wall-clock on one RTX 4090. Token Skip sets a capacity rather than a budget, so 576 is its derived mean image-token count per core block. The subscript on GenEval is the 95% spread of the paired difference against the backbone row of its own block, which therefore carries none.
MethodRGenEval ↑CLIP ↑Pick ↑ImageReward ↑Speedup
MiniT2I-B/16
Dense backbone102487.228.2322.451.181.00×
Token Skip[3]576*82.3±1.728.0821.931.061.19×
Feat Sim51284.7±1.628.0422.041.071.84×
RTI (ours)51287.8±1.528.0922.301.111.84×
Feat Sim25680.5±1.927.7321.570.932.15×
Latent Array[6][7]25645.7±3.123.3019.78−0.652.15×
RTI (ours)25684.7±1.728.0422.061.062.15×
MiniT2I-L/16
Dense backbone102488.128.5022.751.231.00×
Feat Sim51283.1±1.928.8022.341.162.11×
RTI (ours)51287.5±1.328.5822.631.202.11×
Feat Sim25676.5±2.328.3121.761.012.59×
RTI (ours)25686.3±1.628.5222.371.172.59×

At B/16 and R=512, RTI scores 87.8 against its dense backbone's 87.2 while running 1.84× faster. That +0.6 sits inside the ±1.5 interval on the paired difference, so the row says dense quality at 1.84×, not better than dense. What the block does settle is the comparison between reductions: Feat Sim, trained on the same data with the same LoRA recipe and the same elastic budget protocol, reaches 84.7 at that budget, 3.1 points down and outside the interval. CLIPScore is the one metric where Feat Sim leads at R=512; it is also the least discriminative of the four, and RTI leads on GenEval, PickScore and ImageReward.

MJHQ-30K FID on MiniT2I-L/16. FID is the one metric where RTI does not hold dense quality: it recovers most of the reduction's cost. The gap to similarity merging is also widest here.
MethodRFID ↓vs. denseSpeedup
Dense backbone10249.07—1.00×
RTI (ours)51211.43+2.362.11×
Feat Sim51219.99+10.922.11×
RTI (ours)25614.71+5.642.59×
Feat Sim25628.02+18.952.59×
05

Ablations

Elasticity · the components · the order · a second task

Budgets it never trained on

The budget is drawn from {64, 128, 256, 512} during fine-tuning, so those four are the only lengths the checkpoint has ever run. Queried at three it has not — 96, 192, 384 — it lands between its trained neighbours every time, with no step at the trained values. The budget is a knob that can be set after training, not a set of four modes. The gallery at the top of this page is the same point: its richest stop is R=728, which neither checkpoint trained on.

MiniT2I-B/16, one elastic checkpoint, GenEval overall. The dense backbone scores 87.2. The same holds on the class-conditional JiT-B/16: the untrained R=130 gives 6.23 FID against 6.34 at the trained R=128.
RDrawn in trainingGenEval ↑
512✓87.8
384✗86.4
256✓84.7
192✗84.6
128✓81.3
96✗79.0
64✓71.8

What each part contributes

One design choice swapped at a time, everything else held, retrained under the same recipe. Cutting at the largest feature gaps rather than evenly is worth 2.7 points at R=128, and less as the budget grows, since at R=256 a region holds four patches and an even cut already lands close to the adaptive one. The learned Read/Write is worth 7.2 over a plain mean and broadcast. Anchoring is worth the most: free latents in place of region tokens fall to the Latent Array row.

MiniT2I-B/16, GenEval overall, each arm against RTI at its own budget. The interface arm swaps Read and Write together, so the attention pooling, the size embedding and the per-token g of Write are not separated from one another. Rebuilding the partition at every step is not ablated here either; it is motivated by panel (b) of the method section and costs under 2% of the forward pass.
SwapRGenEval ↑Δ
RTI, unchanged12881.3—
Even cuts, not feature gaps12878.6−2.7
RTI, unchanged25684.7—
Mean pool + broadcast25677.5−7.2
Free latents, no anchoring25645.7−39.0

The choice of order

Contiguity is a property of the order, so we swapped it and retrained. Raster order cuts row strips that wrap at row ends; Morton keeps dyadic cells together but jumps diagonally; only Hilbert steps to a neighbour every time. At R=128 the ranking is the one locality predicts, raster worst and Hilbert best, and it widens at R=64. At R=256 all three sit inside noise, because with that many regions even a bad order cuts short runs.

Budget against steps

A region budget and fewer sampling steps both buy speed, and they do not compose. Below about 5 seconds an image the budget is the better purchase: at 1.8 s, 82.7 GenEval against 65.9 for the dense model run with fewer steps. Above it the ordering reverses. The results table reports the quality-matched point, not the compute-optimal one.

05b

Class-conditional transfer

Frozen pixel-space JiT[2]-L/16 · ImageNet-256

The same interface on a different task and backbone family: no text stream, the partition read from the new model's features unchanged, everything else the same recipe. RTI leads feature-similarity merging at every budget on both FID and Inception Score. The gap is largest at the middle budget and narrows at both ends: at R=193 little is reduced, at R=64 both methods degrade.

ImageNet-256, 50K samples, cfg 3.0, 50 Heun steps, torch-fidelity — not comparable to the MJHQ numbers above. The grid is 16×16=256 patch tokens, so the three budgets are 75/50/25% of them; speedup is token compute, not wall-clock. The dense row is the published JiT-L/16 result under this protocol, not our own re-measurement.
MethodRFID ↓IS ↑Token compute
Dense JiT-L/16[2]2562.36298.51.00×
RTI (ours)1932.42313.51.14×
Feat Sim1932.73291.61.14×
RTI (ours)1302.69303.81.33×
Feat Sim1304.53252.81.33×
RTI (ours)644.71253.01.61×
Feat Sim646.28229.21.61×

At R=193, three quarters of the patch tokens and 88% of the dense token compute, RTI is 0.06 FID off the dense model and above it on Inception Score. A higher Inception Score is not by itself evidence of better images — the metric rewards confident, prototypical class predictions — so the FID column carries the claim.

06

Where it breaks

R = 128 and R = 64 · L/16

Below R=256 the budget costs quality: on L/16, GenEval falls from 86.3 at R=256 to 80.3 at 128 and 69.1 at 64. The layout survives; lettering goes first, then fine geometry. Feature-similarity merging degrades further at both settings.

Speed stops improving over the same range. Only the image tokens inside the core shrink, while the dense prelude, the coda and the full-length text stream run unchanged, so the speedup saturates near 3× and halving the budget from 128 to 64 buys only 0.15×. Applying the same interface to the text stream is future work.

Budget
Prompt
Dense backboneN = 1024
Dense sample
RTI (ours)
RTI at a low budget
Feat Sim
Feature-similarity merging at a low budget

06b

Style and identity preservation

One RTI checkpoint · same prompt · same seed

A shorter sequence does not blur the dense image, it re-runs the whole trajectory. The sampler can therefore land on a different sample rather than a coarser one, and identity, count and style can all move with the budget. What decides whether they do is the prompt.

Top row: the prompt names a robot and a lab but not a style. The model renders a photoreal robot at 1024 tokens, a flat cartoon at 728, and a photoreal robot again at 512 and below. The style is not a monotone function of the budget; it is a mode the sampler picks, and one budget in the range picks a different one. Middle and bottom: the same subject and seed with the style stated in the prompt, once as a photograph and once as a vector cartoon. Both hold their style at every budget down to 64. Detail still degrades — the lettering breaks first, as everywhere else — but the register does not move.

The same effect is milder elsewhere: the crowd in the Tokyo scene goes from three figures to five to two, and the cathedral's rose window fades at 256 while its floor pattern survives. Subjects with one strong reading, like the marble statue, keep their identity down to 128. Naming the style is the cheapest way to pin it.

✳

Acknowledgements

We thank the authors of MiniT2I[1] and JiT[2] for open-sourcing their code and pretrained checkpoints. Every result here starts from those models; we train only LoRA adapters and the Read/Write interface on top of them.

MiniT2I is pretrained on the recaptioned Conceptual Captions 12M[16] and fine-tuned on a 120K-image mix of BLIP3o-60k[17], OpenDatasets/dalle-3-dataset, and ShareGPT-4o-Image[18]; our text-to-image runs use that same fine-tuning mix. The class-conditional runs use ImageNet-1K[19].

The MiniT2I and JiT code releases are MIT-licensed, and so are our code and checkpoints. The datasets keep their own terms: BLIP3o-60k and ShareGPT-4o-Image are Apache-2.0 and the DALL·E 3 set is CC0-1.0, while the CC12M recaption release is CC BY-SA 4.0 and ImageNet is released under its own non-commercial research terms. Those last two are the ones to read before any commercial use.

Every prompt behind a figure in the paper or a visual on this page — 85 of them, grouped by figure, with the file each is read from — is listed in PROMPTS.md.

✳

References

  1. X. Wang, H. Zhao, Y. Lu, K. Zhou, L. Ma, K. He. MiniT2I: A Minimalist Baseline for Text-to-Image Generation. 2026. blog
  2. T. Li, K. He. Back to Basics: Let Denoising Generative Models Denoise. 2026. arXiv:2511.13720
  3. D. Raposo, S. Ritter, B. Richards, T. Lillicrap, P. C. Humphreys, A. Santoro. Mixture-of-Depths: Dynamically Allocating Compute in Transformer-based Language Models. 2024. arXiv:2404.02258  — Token Skip
  4. D. Bolya, C.-Y. Fu, X. Dai, P. Zhang, C. Feichtenhofer, J. Hoffman. Token Merging: Your ViT But Faster. ICLR 2023. arXiv:2210.09461  — Feat Sim
  5. D. Bolya, J. Hoffman. Token Merging for Fast Stable Diffusion. CVPR Workshops 2023.  — Feat Sim, diffusion
  6. A. Jaegle, S. Borgeaud, J.-B. Alayrac, C. Doersch, C. Ionescu, D. Ding et al. Perceiver IO: A General Architecture for Structured Inputs & Outputs. ICLR 2022. arXiv:2107.14795  — Latent Array
  7. A. Jabri, D. J. Fleet, T. Chen. Scalable Adaptive Computation for Iterative Generation. ICML 2023. arXiv:2212.11972  — Latent Array, RIN
  8. D. Hilbert. Über die stetige Abbildung einer Linie auf ein Flächenstück. Mathematische Annalen 38:459–460, 1891.
  9. E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen. LoRA: Low-Rank Adaptation of Large Language Models. ICLR 2022. arXiv:2106.09685
  10. J. Geiping, S. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson et al. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach. 2025. arXiv:2502.05171  — prelude / core / coda
  11. D. Ghosh, H. Hajishirzi, L. Schmidt. GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment. NeurIPS 2023. arXiv:2310.11513
  12. J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid et al. Scaling Autoregressive Models for Content-Rich Text-to-Image Generation. TMLR 2022.  — PartiPrompts
  13. J. Hessel, A. Holtzman, M. Forbes, R. Le Bras, Y. Choi. CLIPScore: A Reference-free Evaluation Metric for Image Captioning. EMNLP 2021.
  14. Y. Kirstain, A. Polyak, U. Singer, S. Matiana, J. Penna, O. Levy. Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation. NeurIPS 2023.  — PickScore
  15. J. Xu, X. Liu, Y. Wu, Y. Tong, Q. Li et al. ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation. NeurIPS 2023.
  16. S. Changpinyo, P. Sharma, N. Ding, R. Soricut. Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts. CVPR 2021. arXiv:2102.08981  — MiniT2I pretraining
  17. J. Chen, J. Xu, W. Peng, Y. Zhang et al. BLIP3-o: A Family of Fully Open Unified Multimodal Models. 2025. arXiv:2505.09568  — BLIP3o-60k
  18. J. Chen, P. Cai, Z. Chen, H. Chen, Y. Ji, X. Wang, S. Yang, B. Wang. ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation. 2025. arXiv:2506.18095
  19. J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei. ImageNet: A Large-Scale Hierarchical Image Database. CVPR 2009. image-net.org  — class-conditional runs
✳

BibTeX

@misc{zamfir2026elastic,
  title  = {Elastic Token Compression for Pixel-Space Diffusion Transformers},
  author = {Zamfir, Eduard and Reisswig, Christian and Wu, Zongwei
            and Xian, Yongqin and Timofte, Radu},
  year   = {2026},
  eprint = {2608.29281},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url    = {https://arxiv.org/abs/2608.29281}
}