I'm a Data Analyst and Web Developer with a passion for turning complex data into actionable insights and building user-friendly, efficient web applications. With [X years] of experience in both fields, I bring a unique combination of analytical thinking and technical expertise. My goal is to create value by leveraging data and technology to solve real-world problems.
📱 POCKET — a 35-billion-parameter model that runs on your iPhone, and on your PC with no GPU
We're releasing POCKET, VIDRAFT's flagship Darwin-36B-Opus compressed for on-device use. No fork, no CUDA, no cloud — it runs on stock llama.cpp. It's a sparse Mixture-of-Experts model (256 experts, only 8 active per token), so the file can be large while the work per token stays small. That's what lets a 35B model run on a phone, and generate fast on a CPU with no graphics card.
Measured (POCKET-35B IQ1_M vs Bonsai-27B Q1_0): • CPU generate (Xeon, 16 threads): 27.0 vs 10.1 tok/s → 2.69× faster • GPU generate (H100): 197 vs 89 tok/s → 2.22× faster • GPU prompt processing (H100): 753 vs 1816 → 0.41× (Bonsai wins this one — MoE prefill wakes every expert, so sparsity stops helping there. We say so.) • Quality (HellaSwag, 400 q): 61.0% vs 60.0% → a tie (confidence intervals overlap)
On a real consumer laptop — MacBook M3 Pro (18 GB) — POCKET wins every axis, prompt processing included: • Metal generate: 25.4 vs 12.8 → 1.99× • CPU generate: 13.8 vs 4.4 → 3.13× • Metal prompt: 240.7 vs 73.4 → 3.28×
One more quiet fact: the same-size, quality-oriented rival Ternary-Bonsai-27B (7.2 GB) fails to load in upstream llama.cpp at all — it needs the PrismML fork. POCKET runs on the tools you already have: LM Studio, Ollama, PocketPal, MLX.
Nobody can be the originator of every great idea. And the rate of innovation in AI is difficult to track in detail.
But the next great idea for your codebase may have been described in open-science publications like arXiv. Who among us can keep up with reading, smoke testing, adapting every viral method in the feed? (not me)
And what about the long-tail of results that you won't discover for another 6 months? That's what we built Outrider for!
Now I get alerts when candidate methods look like a good fit for the engineering challenges I'm addressing in my application. I don't waste time running down dead end leads. Instead, I choose to invest my attention into ones which have earned it with evidence.