On Wednesday morning, an AI model called Ox Alpha appeared on OpenRouter. The provider name was "stealth." There was no company, no press release, no LinkedIn announcement, no founder doing a Twitter thread. Just a model, listed at $0 per token, with near-unlimited capacity.
By Thursday, the AI community had already done enough forensic work to suspect who built it.
By Friday, developers were switching their coding agents to it and sharing benchmarks.
If you have a small team that uses AI for any amount of code work — writing, reviewing, debugging, automating — you should know about this, and you should try it before the free window closes around August 27.
What it can actually do
The number that caught people's attention is the DeepSWE benchmark score. Ox Alpha hit 80% Pass@1 on DeepSWE, a real-world software engineering evaluation. For comparison: GPT-5.6 Sol sits at 52%. Claude Fable 5 is at 65%.
That's not a small gap. DeepSWE tests things that matter for actual software work — debugging across codebases, writing tests, implementing features from natural language specs, chaining multiple steps without losing context. It's not a trick benchmark built to be gamed.
The model also comes with a 1-million-token context window and accepts text, images, and video input. You can throw a full codebase at it and ask it to reason across the whole thing.
OpenCode is also hosting it with what they're calling near-unlimited capacity — they mentioned 100 trillion tokens per day. That's not a number you publish unless you're trying to prove something about your infrastructure at scale. Someone is spending real money on this preview.
How to use it right now
The access path is simple. OpenRouter uses an OpenAI-compatible API, so if you've already got code pointed at GPT or Claude, changing it to Ox Alpha takes about 30 seconds:
- Go to openrouter.ai and generate a free API key
- Set your base URL to
https://openrouter.ai/api/v1 - Set your model to
stealth/ox-alpha
That's it. Any tool that can hit an OpenAI-compatible endpoint — Cursor's API mode, your own scripts, n8n, LangChain, whatever your team is running — can use this model today with no configuration changes beyond those three lines.
If your team isn't writing code directly against an API, OpenCode also has a chat interface where you can use it without touching a terminal.
Who built this?
Nobody's officially saying. But the AI community didn't just shrug and move on.
By August 22 — two days after launch — researchers had matched Ox Alpha's serving layer to Zhipu AI's GLM series. The forensic work included: a Java stack trace in error responses, an error code (1214) that's a known GLM signature, and 30 out of 30 tokenizer matches to GLM-5.3. The Next Web covered the detective work.
Zhipu has not confirmed or denied any of it. But the pattern fits — the last several times an anonymous model appeared on OpenRouter with this approach, it was eventually claimed by a Chinese lab testing distribution before a formal launch.
The provider did state publicly that they won't train on prompts or completions submitted during the testing window, which matters if you're considering what to send it.
The actual opportunity here
Here's the frame for a 5-50 person org: right now, this week, you can have access to a coding AI that outperforms what OpenAI charges $20 per million output tokens for — and you can use it for nothing.
That's useful in ways that go beyond "neat demo." If you have a developer who's been using GPT-5.6 or Claude Fable 5 for code review, debugging, or boilerplate generation, swapping in Ox Alpha costs nothing to test. If the results hold up for your actual work, you've found a free upgrade that your team can use for any tasks that don't touch sensitive data.
This also illustrates something worth understanding: the top of the frontier model market is getting crowded fast, and the competitive pressure is driving prices down and free preview periods up. Three months ago, getting 1 million tokens of context on a frontier-level coding model cost real money. Now a lab is testing the waters by giving it away.
For a 12-person nonprofit doing data work, a 20-person government contractor with a small dev team, or any SMB with an IT function — these shifts are cumulative. The teams that know to look for them are running on better tools at lower costs than the ones waiting for their SaaS vendor to tell them what's available.
What not to do
This is an anonymous preview from a provider that hasn't publicly identified itself. The commitment to not training on your prompts is a public statement, not a signed contract.
Don't send: client data, internal credentials, code that contains API keys or environment variables, anything regulated by privacy law. Use it for work that would be fine as a public GitHub repo — generic utilities, open-source integrations, infrastructure scripts, anything where the code itself isn't sensitive.
If your team wants to use something like this for more sensitive code, that's a different conversation about self-hosted models — which is exactly the space that GLM-5.3's delayed open-weight release and the broader Qwen/Llama/GLM ecosystem is opening up.
The bottom line
A model appeared Wednesday with no name attached, beating the frontier on coding benchmarks, available free through next week. The community's best guess is it's Zhipu testing a new GLM before an official launch. The access path is three lines of config.
If your team writes any code, point your AI tools at stealth/ox-alpha this week and see how it performs on real tasks. Form your own opinion. The window is short.
We work with small teams across NGOs, government, and business to help them stay current on tools like this — figuring out what's actually worth adopting and what's noise. If you want help thinking through your AI tool stack, that's what we're here for.