Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
๐
In a Training Loop
66657.2
TFLOPS
VIDRAFT_LAB
SeaWolf-AI
111
26
226
Follow
John6666's profile picture
mepunit's profile picture
johnnycakes112358's profile picture
245 followers
ยท
219 following
https://www.vidraft.net
AI & ML interests
Contact: arxivgpt@gmail.com
Recent Activity
upvoted
a
paper
about 12 hours ago
Quantum Cryptanalysis on IBM Quantum Hardware: Extending Even--Mansour Period Recovery from N=4 to N=10
upvoted
an
article
about 15 hours ago
POCKET: a 35-billion-parameter model that runs on your iPhone โ and on your PC with no GPU
reacted
to
their
post
with โค๏ธ
about 15 hours ago
๐ฑ POCKET โ a 35-billion-parameter model that runs on your iPhone, and on your PC with no GPU We're releasing POCKET, VIDRAFT's flagship Darwin-36B-Opus compressed for on-device use. No fork, no CUDA, no cloud โ it runs on stock llama.cpp. It's a sparse Mixture-of-Experts model (256 experts, only 8 active per token), so the file can be large while the work per token stays small. That's what lets a 35B model run on a phone, and generate fast on a CPU with no graphics card. Measured (POCKET-35B IQ1_M vs Bonsai-27B Q1_0): โข CPU generate (Xeon, 16 threads): 27.0 vs 10.1 tok/s โ 2.69ร faster โข GPU generate (H100): 197 vs 89 tok/s โ 2.22ร faster โข GPU prompt processing (H100): 753 vs 1816 โ 0.41ร (Bonsai wins this one โ MoE prefill wakes every expert, so sparsity stops helping there. We say so.) โข Quality (HellaSwag, 400 q): 61.0% vs 60.0% โ a tie (confidence intervals overlap) On a real consumer laptop โ MacBook M3 Pro (18 GB) โ POCKET wins every axis, prompt processing included: โข Metal generate: 25.4 vs 12.8 โ 1.99ร โข CPU generate: 13.8 vs 4.4 โ 3.13ร โข Metal prompt: 240.7 vs 73.4 โ 3.28ร One more quiet fact: the same-size, quality-oriented rival Ternary-Bonsai-27B (7.2 GB) fails to load in upstream llama.cpp at all โ it needs the PrismML fork. POCKET runs on the tools you already have: LM Studio, Ollama, PocketPal, MLX. ๐ Full story (tech, measurements, recipes): https://huggingface.co/blog/FINAL-Bench/pocket Models: ๐ฆ POCKET-35B-GGUF (PC / server, no GPU): https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF ๐ฐ๐ท POCKET-KR-GGUF (Android): https://huggingface.co/FINAL-Bench/POCKET-KR-GGUF ๐ POCKET-KR-MLX (iPhone / Mac): https://huggingface.co/FINAL-Bench/POCKET-KR-MLX ๐ POCKET-EN-GGUF (English phone / PC): https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF ๐ฅ๏ธ Live demo (answering on a CPU, no GPU): https://huggingface.co/spaces/FINAL-Bench/POCKET-35B-CPU ๐ Collection: https://huggingface.co/collections/FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6
View all activity
Organizations
SeaWolf-AI
's models
4
Sort:ย Recently updated
SeaWolf-AI/Darwin-36B-KR
Updated
about 1 month ago
SeaWolf-AI/Darwin-Qwen3.5-27B-x-Qwen3.5-27B-Claude-4-08162
28B
โข
Updated
Apr 12
โข
6
SeaWolf-AI/Darwin-Darwin-4B-Opus-x-gemma-4-E4B-it-The-D-08412
8B
โข
Updated
Apr 10
โข
2
โข
7
SeaWolf-AI/Darwin-gemma-4-E4B-it-x-Gemma-4-E4B-Claude-4-08292
Updated
Apr 8
โข
7