Something landed on Hacker News this week that a lot of people scrolled past. A solo developer named Leo Nickson posted a Swift and Metal runtime called Swiftlet that runs an 80-billion-parameter Qwen model in 4.3 gigabytes of RAM on a Mac — and a 35B version on an iPhone 17.
Let's be clear about what those numbers mean. An 80B parameter model is roughly the size of what was considered frontier-class a year ago. The kind of model that costs serious money to run in the cloud, or requires a server rack of GPUs to self-host. It's now running at 4.5 to 5 tokens per second on a laptop, in less RAM than most browser sessions consume.
How it works
Most inference runtimes load the entire model into memory before you can ask it anything. Swiftlet exploits a property of Mixture-of-Experts (MoE) models — like the Qwen series — where only a small slice of the weights activates per token. The dense core (attention layers, routing logic, embeddings) stays in RAM. The expert blocks stream from your SSD on demand.
The result: 42 GB of model on disk, 4.3 GB in memory at runtime. At any moment, Swiftlet is only holding the parts of the model that are actually doing work.
This is not a hack. It's a correct application of how MoE models function, written in Swift with Apple's Metal GPU API. Everything is Apache 2.0 licensed. You can read the source, modify it, and deploy it however you want.
Why this matters for your org
Here's the conversation we keep having with NGOs and public-sector clients: "We'd love to use AI for [sensitive document review / case notes / intake forms / internal search], but our data can't leave our systems."
That conversation has historically ended with either expensive private cloud infrastructure or nothing at all.
Swiftlet changes the "nothing" option. If your team is running Macs — and most office workers are — you now have a path to near-frontier AI that:
- Never sends data to anyone's server
- Costs zero dollars in API fees
- Works with no internet connection
- Is fully auditable because the code is open
For a legal aid clinic, a health-services nonprofit, or a government department handling personal information, these aren't nice-to-haves. They're the difference between being able to use AI at all and being shut out entirely.
The honest limits
4.5 tokens per second is slow. You're not running batch processing pipelines on this. For one person doing a document summary, checking a draft, or extracting structured data from a form, it's workable. For parallel workloads or anything real-time, you need different infrastructure.
The iPhone version, running the 35B model at about 1 tok/s and 2.5 GB of RAM, is more proof-of-concept than production tool right now. The Mac version is the real story.
The 80B model also requires 42 GB of disk space. Modern Macs ship with 512 GB to 1 TB SSDs as standard, so this isn't a blocker for most teams, but it's worth planning around.
Who should move on this now
If data sensitivity has been your excuse for avoiding AI adoption, Swiftlet is worth testing this week. The GitHub repo has setup instructions and is actively maintained.
Specifically useful for:
- Healthcare NGOs and clinics operating under PIPEDA or provincial health privacy rules
- Legal services orgs with client confidentiality requirements
- Government departments with data residency or sovereignty mandates
- Any team where uploading documents to a cloud API feels legally or politically uncomfortable
You don't need a private cloud contract. You don't need a GPU server. You need a Mac with SSD space and an afternoon to get it running.
One thing worth flagging
The Qwen model family comes from Alibaba's research group. If your organization has policies around the provenance of AI models, particularly around Chinese-developed technology, that's worth evaluating. Qwen is open-weight and Apache 2.0 licensed, so the model itself is inspectable, but you can't audit its training data origins. For most orgs we work with, the concrete privacy benefits of fully on-device inference outweigh that consideration. But it's your call.
The short version
Private, offline, near-frontier AI just got accessible on hardware your team already owns. The open-source community keeps removing the technical barriers that used to justify "wait and see." If your org has been sitting out AI adoption because of data sensitivity requirements, the technology just caught up to your constraints.
This is exactly the kind of thing we help orgs evaluate, configure, and actually put to work. If you want to know whether Swiftlet fits your setup, reach out.