What a Mac unlocks
Same privacy as the free tools. Different ceiling.
Every tool on pocketweb.tools already runs on your device. What a browser cannot do is hold a large model: a tab tops out at about 4 GB of memory, which is enough for 2B to 9B-class models and 1-bit tricks, not for a dense 27B. The Mac app uses your Mac's actual memory, so the lineup goes where a tab cannot.
| In your browser (free) | PocketWebTools for Mac | |
|---|---|---|
| Model size | Under about 4 GB | Up to 22 GB on a 32 GB Mac |
| Largest model in the lineup | Qwen3.5 2B, Bonsai 27B 1-bit | Qwen3.8 27B dense, Qwen3.6 35B-A3B |
| Memory available | About 4 GB per tab | Your Mac's RAM, with tiers for 8, 16 and 32 GB |
| Engine | WebGPU or WebAssembly | llama.cpp and transcribe.cpp on Metal, every layer on the GPU |
| Audio input | Files up to 500 MB and the microphone | Files of any length, the microphone, and the audio the Mac is playing |
| Speaker labels | No | Yes, up to four speakers |
| Where it runs | On your device | On your device |
| Works offline | After the model is cached | Always, after the download |
| Where models live | Browser cache, can be evicted | Application Support, until you delete them |
| Price | Free, unlimited | $29 once, free lifetime updates |
| Account | None | None. A license key from your purchase page. |
What's inside
Version 0.1.0
Chat with a local model
Threads saved on your Mac, Markdown replies with highlighted code, copy and regenerate on every turn, and a context meter so a long thread never overflows unnoticed.
Models on demand
Download only the models you want. Downloads resume after an interruption, are checked against the publisher's checksum, and stay on disk until you delete them.
Every layer on the GPU
llama.cpp on Metal with the whole model resident on Apple Silicon, plus prompt caching so follow-up messages start streaming in well under a second.
Transcribe anything, including system audio
Files of any length, your microphone, or whatever the Mac is playing: a call, a video, any app. Word timestamps, speaker labels, subtitles export. The part a browser tab categorically cannot do.
Next in line: the rest of the local-AI tools from pocketweb.tools with bigger models and batch folders. Whatever ships is a free update for every buyer. No dates are promised; the changelog is the record.
The lineup
Commercially redistributable licenses only. Download only what you want.
32 GB MACS AND UP
Qwen3.8 27B
Flagship17.6 GB download · 32k context · Apache 2.0
The strongest model a Mac can run today. Dense 27B, far beyond what any browser tab can hold. Best for writing, code and long documents.
Qwen3.6 35B-A3B
22.4 GB download · 32k context · Apache 2.0
Mixture-of-experts: 35B of knowledge, only 3B active per token, so replies stream noticeably faster than the flagship.
16 GB MACS
Qwen3.5 9B
Best pick for 16 GB6.0 GB download · 16k context · Apache 2.0
The fastest replies in the lineup from a 6 GB download, and still a clear step up from anything that runs in a browser.
Gemma 4 12B
7.4 GB download · 16k context · Apache 2.0
Google's 12B instruct model, the alternative for 16 GB Macs. Slower to reply than Qwen3.5 9B; pick it when you want a second model to compare answers.
8 GB MACS
Bonsai 27B (1-bit)
3.8 GB download · 8k context · Apache 2.0
A 27B model squeezed to under 4 GB with 1-bit weights. Runs on 8 GB Macs; trades some accuracy for size and speed.
Qwen3.5 2B
Instant1.3 GB download · 8k context · Apache 2.0
Small and instant. Downloads in a couple of minutes and answers fast; the right first model while a bigger one downloads.
SPEECH, EVERY MAC
Parakeet TDT v3
FilesThe default for files: 25 European languages, word-level timestamps, and roughly ten times the speed of the browser tool on the same Mac.
NVIDIA · 0.7 GB · CC BY 4.0
Whisper Large v3 Turbo
Files99 languages with auto-detect, plus translation into English. The same model the web tool's high-quality mode uses, now on Metal with no tab memory limit.
OpenAI · 0.9 GB · MIT
Nemotron 3.5 Streaming
LiveBuilt for live transcription: words appear as they are spoken, in 32 languages, with word-level timestamps.
NVIDIA · 0.8 GB · OpenMDW 1.1
Sortformer speaker labels
SpeakersAdds who-spoke-when labels to any transcript, up to four speakers. Small and fast, runs alongside the transcription model.
NVIDIA · 0.1 GB · NVIDIA Open Model License
The app checks your Mac's memory and marks which tiers fit before you download anything. Every model here was chosen because its weights may be redistributed commercially, so there is no license surprise later; the full license texts ship inside the app.
Pricing
Pay once. That is the whole model.
$29
one payment, forever.Founding users: $14 during launch week, offered once.
- Every tool in the app, today and in every future version
- Every model in the lineup, downloaded on demand
- One license key for up to 3 Macs
- Free lifetime updates, no renewal, no subscription
- 30-day money-back guarantee, no questions asked
Payments are handled by Polar as merchant of record, so VAT and sales tax are sorted at checkout. Refunds within 30 days come from your purchase page, no email required.
The only network calls it makes
Three. Listed in full, so you can hold us to it.
- 1
License activation and checks
When you enter your key, and then at most once a day while the app is open, it asks Polar (our payment provider) whether the key is still valid. The request carries the key, a device label (your Mac's name and chip) and a hashed hardware id, never the raw one. Chatting with a model already on your Mac never checks the license.
- 2
Update check
On launch the app asks whether a newer version exists so you get every update. It sends the version you are running and nothing about how you use the app.
- 3
Model downloads
When you choose a model, the weights download to your Mac once, verified against a checksum. Only the file you asked for is fetched.
Nothing else. No analytics, no crash reports, no telemetry of any kind. Your chats are stored on your Mac and never sent anywhere. Turn Wi-Fi off after a model has downloaded and everything keeps working, which is a test no cloud AI can pass.
Requirements
- Chip
- Apple Silicon (M1 or later)
- macOS
- macOS 14 Sonoma or later
- Memory
- 8 GB minimum, 16 GB recommended, 32 GB for the 27B flagship
- Disk
- The app is small; each model is its download size, from 1.3 GB to 22.4 GB
Intel Macs are not supported: these models need unified memory and Metal to run at a usable speed.
Frequently asked questions
- Is there a free trial?
- No. The free tier is pocketweb.tools itself: the same tools, in your browser, with the models a browser can run. The Mac app is for the models it cannot. If it is not for you, ask for a refund within 30 days and you get every cent back, no questions asked.
- What does free lifetime updates mean?
- Every future version of the app, including new models and new tools, at no extra cost, with no renewal. It is a promise about price, not about timing: we do not commit to a release schedule. The changelog shows what has shipped and when.
- Which Macs does it run on?
- Any Apple Silicon Mac (M1 or later) on macOS 14 Sonoma or later. Intel Macs are not supported because these models do not run usefully on them. 8 GB of memory runs the smallest tier, 16 GB is the sweet spot, and the 27B flagship needs 32 GB.
- Does it need an internet connection?
- Only for three things: activating and re-checking your license, checking for updates, and downloading a model. Everything else, including every chat, runs on your Mac and works with Wi-Fi off.
- How many Macs can I use it on?
- One key activates up to 3 Macs. You can deactivate a Mac from inside the app or from your purchase page to free a slot for another one.
- Where do my chats go?
- Nowhere. Chats, audio and transcripts stay on your Mac and are never sent to us or anyone else. The app has no analytics, no crash reporting and no telemetry of any kind. The privacy section above lists the only network calls it ever makes.
- Is it a native Mac app?
- The engines are native: a Rust core running llama.cpp for chat and transcribe.cpp for speech, both on Metal with the whole model on the GPU, and Core Audio for microphone and system-audio capture. The interface is the same one you use on pocketweb.tools, shown in a system web view. That is what keeps the app small and lets web and Mac share one design.
- Why not the Mac App Store?
- The app is signed and notarized by Apple like any other download, but sold directly. That keeps the price at $29 instead of paying a store commission, and it means multi-gigabyte model files are not squeezed through the store's sandbox rules.
- Can I use the models commercially?
- Yes. The chat models are Apache 2.0; the speech models are CC BY 4.0 (Parakeet), MIT (Whisper), OpenMDW 1.1 (Nemotron) and the NVIDIA Open Model License (speaker labels). All of them allow commercial use, the app ships every license text, and we add no restrictions of our own to what you make with them.
Track record
Every release and model refresh, dated.
- Transcribe: files, live audio, system audio, speaker labels
- Qwen3.5 9B becomes the recommended model for 16 GB Macs
- Markdown replies, prompt caching, turn actions
