All Insights

Frontier AI at $0.83 Per Million Tokens: What Tencent's Hy4 Means for Your Budget

CivSafe Team·August 31, 2026·6 min read

Tencent released Hy4 Preview on August 28. It's a 770-billion parameter model with a 1-million token context window, released under the Apache 2.0 license. The API is live at $0.83 per million input tokens.

For context: GPT-5 via the OpenAI API runs roughly $15 per million input tokens. Same frontier-quality reasoning tier. Hy4 is 18x cheaper.

That's not a benchmark footnote. For any org running real volume through an AI API, that number changes the math.

Why a 770B model costs like a 49B model

Hy4 uses a Mixture-of-Experts architecture. It has 770 billion total parameters, but only 49 billion activate per inference pass. The model routes each request through a relevant subset of its specialized sub-networks rather than the whole thing.

The upshot: 770B-quality reasoning at inference costs much closer to a 49B model. Tencent can price this at $0.83/M because the per-token compute burden is far lower than a traditional dense model at that scale.

Benchmark numbers: on engineering and office productivity tasks, Hy4 Preview edges out GLM 5.3 and Kimi K3, two other recent frontier-class Chinese open-weight releases. Community testing is still early — independent evaluations will land over the next few weeks — but the internal numbers match what comparable models in this architecture class have shown.

The math for a real small-org workflow

Here's what this looks like on an actual workflow.

A 15-person NGO processes grant applications, funder reports, and policy documents. Call it 400,000 input tokens per day — realistic for a team actively using AI for drafting, synthesis, and Q&A across long documents.

At $0.83/M: $12 per day. $360 per month.

At GPT-5 pricing ($15/M): $6,000 per day. $180,000 per month.

At lower volumes, the ratio is the same. At 50,000 tokens per day: Hy4 runs $1,245/month, GPT-5 runs $22,500/month. Both of those are real budget decisions for small orgs. The 18x difference is the difference between "this fits in our operating budget" and "this requires a major grant just to sustain."

For a lot of organizations that looked at AI API costs and concluded they couldn't justify the expense at scale, these numbers are worth revisiting.

Why the Apache 2.0 license matters more than the price

Most organizations don't track model licenses. They should.

Proprietary API (OpenAI, Google, most paid tools): You rent access. Terms change, pricing changes, model versions get deprecated. The vendor controls the relationship.

Gated open weights (some Meta Llama variants): Free to download but with restrictions. Commercial usage caps, required license agreements above certain revenue thresholds, no redistribution rights.

Apache 2.0 (Hy4): Commercial use, modification, and redistribution — all allowed, no restrictions. If you build a product on top of it, you owe nothing. If you fine-tune it on your data, you own the output.

Apache 2.0 also means self-hosting is a genuine option. If you have GPU infrastructure — cloud rental through RunPod or Lambda Labs, or on-prem hardware — you can run Hy4 yourself. At that point the per-token cost disappears and you're paying only for compute.

This matters for two types of organizations: those with compliance requirements that restrict external API calls (government-adjacent, healthcare, legal), and those building internal tools where per-token costs would erode margins over time.

What a 1-million token context window actually changes

The 1M context window sounds like a marketing number until you hit a real document workflow.

Previous workaround: chunk a 400-page report into sections, run each through the model, synthesize the outputs. It works, but cross-section reasoning breaks down and the inconsistencies compound.

With 1M tokens, a 400-page environmental assessment, a 300-page RFP, or a full grant application cycle fits in a single API call. The model reads the whole thing at once and reasons across it.

We've been running large-context models on policy document synthesis and grant review workflows with a public sector client. The quality gap between chunked and full-context analysis is real — mostly on questions that require understanding relationships across a document, not just what's on each page. For legal, compliance, regulatory, or grant work, this is the part worth testing.

One thing to factor in before you route data through it

Hy4 comes from Tencent, which operates under Chinese law. Most organizations won't have a compliance issue, but some will.

If you're at a government org, a defense-adjacent contractor, or an organization with data residency requirements, check with your legal team before routing data through the Tencent Cloud API endpoint. The model weights are Apache 2.0 and clean, but API traffic is a different question.

The path for compliance-sensitive orgs: self-host. The license allows it. You get the same model with no third-party API in the data path. The weights are available on Hugging Face. If you have a server with 4-8 A100s or H100s, this is feasible. If not, a cloud GPU rental for an internal deployment is often cheaper than the API cost difference alone.

What to do with this right now

If you're currently paying GPT-5 or similar prices for high-volume tasks:

  • Test Hy4 on OpenRouter at $0.83/M — it's live now, no setup required
  • Run your existing prompts against it and compare outputs
  • If quality holds for your use case, that's a direct budget reduction you can implement this week

If you're hitting chunking issues on long documents:

  • Hy4's 1M context window removes that constraint for most document types
  • Worth a direct test on legal, policy, grant, or procurement workflows where cross-document reasoning matters

If you're building an internal tool and concerned about vendor lock-in:

  • Apache 2.0 means no licensing risk for commercial products
  • Self-hosting on rented or owned GPU infrastructure is viable at this license tier

The model access question has shifted. Three months ago, frontier-quality AI required either a large API budget or serious infrastructure to run locally. The open-weight releases from Chinese labs — Kimi K3, GLM 5.3, DeepSeek V4, and now Hy4 — have collectively moved the line. Hy4 is the clearest example yet of frontier-class reasoning at a price point that doesn't require a budget negotiation.

The question now is whether your team knows how to wire it into your actual workflows.


That's the problem we solve. If you're running document-heavy workflows and want to know whether Hy4 is the right fit, get in touch.

CivSafe — Strategic Innovation. Community Impact.