GLM-Image is an image generation model adopts a hybrid autoregressive + diffusion decoder architecture. In general image generation quality, GLM‑Image aligns with mainstream latent diffusion approaches, but it shows significant advantages in text-rendering and knowledge‑intensive generation scenarios. It performs especially well in tasks requiring precise semantic understanding and complex information expression, while maintaining strong capabilities in high‑fidelity and fine‑grained detail generation. In addition to text‑to‑image generation, GLM‑Image also supports a rich set of image‑to‑image tasks including image editing, style transfer, identity‑preserving generation, and multi‑subject consistency.
Capability badges appear only where the provider catalog explicitly lists support. A blank cell means the source is silent, not that the feature is absent.
Specification
Capabilities
Recorded from the source catalog and provider listings.
Every result stays attached to its source, version, metric and harness. Scores from different versions are never merged, and the profile below plots each benchmark against its own population rather than on a shared scale.
ModelBench does not infer quality from price, context size, or model name. When a source publishes a comparable result, it appears here with its version and link.
Answered directly from the stored record — nothing here is generated beyond the catalog's own fields.
What is GLM-Image?
GLM-Image is an image generation model adopts a hybrid autoregressive + diffusion decoder architecture. In general image generation quality, GLM‑Image aligns with mainstream latent diffusion approaches, but it shows significant advantages in text-rendering and knowledge‑intensive generation scenarios. It performs especially well in tasks requiring precise semantic understanding and complex information expression, while maintaining strong capabilities in high‑fidelity and fine‑grained detail generation. In addition to text‑to‑image generation, GLM‑Image also supports a rich set of image‑to‑image tasks including image editing, style transfer, identity‑preserving generation, and multi‑subject consistency. It is published by Zhipu AI and catalogued here from Models.dev.
What is the context length of GLM-Image?
GLM-Image accepts up to 10.24K tokens of context.
Does GLM-Image support tool calling and structured output?
Provider catalogs list support for image input.
Which providers serve GLM-Image?
1 provider list this model: ZenMux.
Are the weights for GLM-Image open?
Yes. The weights are published and downloadable from Hugging Face under the MIT license.
When was GLM-Image released?
The catalog records a release date of 2026-01-19, last verified Sep 19, 2026.