Apple’s Machine Learning Research team published a paper detailing a new generative architecture called Normalising Trajectory Models (NTM). The approach achieves high-quality image generation with only four steps of sampling while preserving the ability to compute exact likelihoods across the generation path.
The architecture treats each reverse sampling step as a conditional normalising flow, using shallow invertible blocks per step combined with a deep parallel predictor spanning the trajectory. NTM can be built from scratch or started from pretrained flow-matching models. The exact likelihood over the trajectory enables self-distillation, where a lightweight denoiser learns from the model’s own score function.
On text-to-image benchmarks, NTM performed as well as or better than strong baselines with just four sampling steps, according to the paper.
Why it matters
In our view, NTM’s combination of fast sampling and exact likelihood is notable because generative models often trade sample quality for speed or probabilistic expressiveness. This research points towards efficient, principled generation that could benefit on-device applications such as image editing and content creation. The paper is a research publication, however; Apple has not indicated plans to productise the approach.
