In the run-up to Z.ai launching GLM-5.3-Flash, an unbranded model calling itself ox-alpha showed up on OpenRouter and OpenCode with no declared owner — free to use, with a 1,048,576-token context window and text, image, and video input. Multiple independent usage-tracking sites place the stealth window at roughly August 20–26, 2026, about six days.
An anonymous model, then a reveal
Nobody using ox-alpha during that window was told who built it. That’s not unusual on its own — labs have quietly stealth-tested models under codenames on public routers before, using real traffic and real developer feedback before a formal launch. What made ox-alpha notable was the scale: OpenCode’s own usage dashboard, which still labels the current model “formerly ox-alpha,” confirms the identity switch happened, even though it doesn’t publish exact stealth-period dates itself.
After the free window closed, Z.ai confirmed what several trackers had already suspected: ox-alpha and GLM-5.3-Flash were the same model. The company’s official documentation describes GLM-5.3-Flash as the first native multimodal model in its GLM-5 series — 320B total parameters with 18B activated, built on a hybrid sparse-and-linear attention architecture with Manifold-Constrained Hyper-Connections and IndexPool compression to make the 1M-token context window workable at that scale.
Why run a stealth free window at all
An unbranded model removes brand expectations from the equation: developers rate what they experience, not what they assume a known lab’s next release should do. A large, free, high-limit window also generates exactly the kind of real-world, long-context agent traffic that’s hard to simulate in a benchmark suite. Whether that was Z.ai’s actual reasoning hasn’t been confirmed on the record — The Signal hasn’t seen an official statement from the company explaining the stealth period itself, only the eventual confirmation that ox-alpha and GLM-5.3-Flash were one and the same.
What the numbers show
As of this week, OpenCode’s live dashboard shows the model — now listed under its real name — has processed roughly 52 trillion tokens across about 776,000 unique users and 14.2 million sessions since mid-July, with current pricing around $0.07 per million input tokens and $0.25 per million output tokens. Those are OpenCode’s own cumulative figures as of publication, not a stealth-window-only count; separate community trackers reported narrower stealth-period numbers (on the order of tens of trillions of tokens and low hundreds of thousands of users) during the free window itself, which The Signal has not independently reconciled against OpenCode’s running total.
What’s confirmed and what isn’t
- Confirmed: an anonymous model called ox-alpha ran on OpenRouter and OpenCode; OpenCode’s own listing now reads “formerly ox-alpha” for GLM-5.3-Flash.
- Confirmed: GLM-5.3-Flash’s specs (320B/18B parameters, 1M context, hybrid attention architecture) per Z.ai’s own documentation.
- Reported, not independently verified by The Signal: the exact Aug 20–26 stealth-window dates and the stealth-period-only usage numbers, which come from third-party trackers rather than an official Z.ai statement.
- Not confirmed: Z.ai’s own stated reason for running the stealth period, since the company has not published an explanation on the record.
Sources
- OpenCode — GLM-5.3-Flash usage data (confirms the ox-alpha rename and current cumulative usage/pricing)
- Z.ai Developer Docs — GLM-5.3-Flash (official model specs)




