Project MonetRequest demo
Home/Blog/GLM-5.3-Flash: Pricing, API, Benchmarks, Context & Open Weights

AI · Project Monet Briefing

GLM-5.3-Flash: What It Is, Pricing, API, Benchmarks and Open Weights

Z.ai’s native-multimodal 320B/18B-active model combines long context, open MIT-licensed weights and lower-cost agentic inference.

Published 2026-08-28 · Updated 2026-08-28 · By Project Monet

Project Monet editorial graphic for GLM-5.3-Flash: Pricing, API, Benchmarks, Context & Open Weights

01

In brief

GLM-5.3-Flash is Z.ai’s newly released native-multimodal model aimed at efficient coding, agentic workflows, visual understanding and professional work. Z.ai says it has 320 billion total parameters but activates about 18 billion per token, uses a new base model rather than being a simple GLM-5.3 post-train, and combines sparse and linear attention to reduce long-context inference cost. The weights are publicly available under the MIT License.

02

GLM-5.3-Flash in brief

GLM-5.3-Flash appeared across Z.ai product surfaces on August 26, 2026, while Z.ai and AutoClaw's detailed launch posts are dated August 27. Before the named release, the company says it evaluated the model anonymously as Ox Alpha under real-world traffic. That earlier identity matters because people who encountered Ox Alpha may now search for what model it became.

The model is the first natively multimodal member of the GLM-5 family. Z.ai says it supports text, images, video and files, allowing visual information to participate directly in multi-step reasoning and tool workflows rather than acting as a separate bolt-on capability.

03

Architecture: 320B total, 18B active

GLM-5.3-Flash has 320B total parameters and about 18B active parameters. Z.ai says it redesigned the architecture and training recipe around efficiency, reducing the active parameter count and number of layers relative to the similarly sized GLM-4.5 series.

Its hybrid attention design combines linear attention for local/state-based dependencies with sparse attention that retrieves selected global context. Z.ai also describes Manifold-Constrained Hyper-Connections and a 30T-token multimodal training corpus as parts of the model design.

For long-context workloads, Z.ai says IndexPool compresses cached key vectors and helps reduce indexer latency and memory overhead. In its own comparison with GLM-5.3, the company reports about 3× lower attention compute and a 4.4× smaller KV cache. These are vendor-reported figures, not independent measurements.

04

Context window

Z.ai’s launch material describes the architecture operating at context lengths up to one million tokens. Exact context and maximum-output limits should be checked against the current official model card/API documentation before production use because provider endpoints and serving configurations can differ.

05

Native multimodal support

GLM-5.3-Flash is trained to work across text and visual information. Z.ai specifically describes image, video and file support and highlights use cases involving screenshots, charts, documents, interfaces, presentations and spreadsheets.

For creators and business teams, the interesting part is not simply image recognition. A multimodal agent can inspect a rendered document, slide deck, chart or webpage, reason about the visual result and continue making changes. That makes the model relevant to content workflows, design QA, document automation and agentic web/product work.

06

GLM-5.3-Flash benchmarks

Z.ai reports that GLM-5.3-Flash outperforms GLM-5.2 across several coding, tool-use, automation and professional-work evaluations. Its published results include 84.3 on Terminal Bench 2.1, 63.4 on DeepSWE v1.1, 78.4 on Toolathlon Verified and 48.8 on AutomationBench v1.0.6.

The company also reports multimodal results including 89.4 on CharXiv Reasoning with Tools, 78.0 on Chartography with Tools and 80.5 on MMVU.

These numbers come from Z.ai’s own evaluation material. They should be treated as vendor-reported benchmarks; real performance can change with inference settings, tool frameworks, prompts and serving environments.

07

Open weights and license

Z.ai says the GLM-5.3-Flash weights are publicly available on Hugging Face under the MIT License. This is materially different from a model that is only accessible through a hosted API because teams can evaluate self-hosted deployment and community quantization paths.

The official launch material lists SGLang, vLLM and TokenSpeed among supported inference frameworks. Day-zero community reports already show active work on vLLM/SGLang deployment, but community bug reports and unofficial quantizations should not be confused with official support guarantees.

08

GLM-5.3-Flash API and pricing

Z.ai's current developer documentation lists the API model code as glm-5.3-flash and says the model is fully available through the GLM Coding Plan. Its pay-as-you-go pricing table lists standard rates of $0.15 per million input tokens, $0.03 per million cached-input tokens and $0.50 per million output tokens.

A 50% launch promotion currently reduces those rates to $0.075 input, $0.015 cached input and $0.25 output per million tokens. Z.ai says that promotion ends at 24:00 on September 9, 2026 in Singapore time (UTC+8). Cached-input storage is listed as limited-time free. These promotional terms are time-sensitive; production budgets should use the live pricing page.

09

What was Ox Alpha?

Z.ai says GLM-5.3-Flash was tested anonymously as Ox Alpha before the official release. The anonymous test let the company evaluate the model under real-world traffic before publicly attaching the GLM name.

So if you used or saw Ox Alpha shortly before this release, GLM-5.3-Flash is the model Z.ai identifies behind that preview.

10

GLM-5.3-Flash vs GLM-5.3

The clearest current distinction is positioning. GLM-5.3 is the heavier flagship model, while Flash is designed around lower inference cost, native multimodality and frequent agent/workflow use. A dedicated comparison can be useful because users choosing between the two care about cost, benchmark trade-offs, multimodality and deployment requirements rather than simply which model is newer.

A separate comparison article should use current first-party pricing and equivalent benchmark conditions before making a recommendation.

11

Who should consider GLM-5.3-Flash?

The model is especially relevant to developers building multimodal agents, coding assistants, document and presentation workflows, visual QA systems, long-context automation and high-volume applications where inference efficiency matters.

It may also be interesting for teams that want open weights but still need native visual understanding and modern agent/tool-use performance.

12

Frequently asked questions

When was GLM-5.3-Flash released?

GLM-5.3-Flash appeared across Z.ai product surfaces on August 26, 2026, with detailed official launch posts dated August 27.

How large is GLM-5.3-Flash?

Z.ai describes it as a 320B-total-parameter mixture-of-experts model with about 18B active parameters.

Is GLM-5.3-Flash multimodal?

Yes. Z.ai describes native support for text, images, video and files.

Is GLM-5.3-Flash open weight?

Yes. Z.ai says the model weights are available on Hugging Face under the MIT License.

Can GLM-5.3-Flash run with vLLM or SGLang?

Z.ai lists vLLM and SGLang among supported inference frameworks. Because this is a very new architecture, check current framework releases and known issues before deployment.

Is Ox Alpha GLM-5.3-Flash?

Z.ai says it anonymously evaluated GLM-5.3-Flash as Ox Alpha before the named release.

13

What to watch next

The highest-value updates are stable API pricing after the launch period, wider provider availability, independent benchmarks, mature quantizations and clearer hardware/deployment recipes. Those developments can justify supporting guides without fragmenting the main overview into thin pages.

Sources

Primary and supporting sources

Facts were rechecked against the linked sources immediately before publication. Pricing, product availability and rollout status can change.

Project Monet

Useful signals. Clear decisions. Better digital work.

Project Monet turns relevant shifts in AI, creator tools and the web into practical context—and builds focused websites for businesses ready to grow.

Request a free homepage concept