VIDRAFT_LAB's picture
🔄 In a Training Loop

VIDRAFT_LAB

SeaWolf-AI

AI & ML interests

Contact: arxivgpt@gmail.com

Recent Activity

repliedto their post about 10 hours ago
POCKET now speaks Gemma 4 — a 26B model that loads in every app, and runs on your PC with no GPU We're adding a Gemma-4 sibling to POCKET: POCKET-26B, built from Google's Gemma-4-26B-A4B (Apache-2.0). Our flagship POCKET-35B is a Qwen-family MoE and needs a recent llama.cpp; POCKET-26B trades a little size for the thing people kept asking for — it just loads, everywhere, today: Ollama, LM Studio, PocketPal, MLX, any stock llama.cpp. No fork, no bleeding-edge runtime, no CUDA, no cloud. It's a sparse Mixture-of-Experts (25.2B total, ~4B active per token), so the work per token stays small — a real 26B that generates on a CPU with no graphics card. Two things make it stand out: 1) Universal compatibility. Gemma 4 is a standard, widely-supported architecture, so POCKET-26B runs on the tools you already have — no waiting for your app to add a new model type. 2) Quality that survives compression. Measured GPQA-Diamond (198 q, greedy): • Full base: 67.7% • POCKET-26B Q4_K_M (17 GB): 67.7% — lossless • POCKET-26B Q2_K (11 GB): 67.2% — near-lossless, at 11 GB Live, on a CPU-only box (our demo Space — POCKET-26B vs Bonsai-27B, same machine, same stock llama.cpp): POCKET-26B ≈ 19 tok/s vs Bonsai ≈ 6 tok/s → about 3× faster generation, no GPU. (Honest notes: shared CPU box, sequential race; a dedicated machine is faster.) Where it fits in the family: • POCKET-35B (Qwen MoE) — bigger, top-tier, needs a recent llama.cpp. • POCKET-26B (Gemma 4) — loads in any app, quality-robust when compressed. The demo runs the Q4_K_M build; Q2_K (11 GB) is the smallest footprint. For a true ≤8 GB phone, the 5 GB POCKET-KR (Qwen) is still the pick. Try it and grab it: 🖥️ Live demo (Gemma4-based, answering on a CPU, no GPU): https://huggingface.co/spaces/FINAL-Bench/POCKET-26B-CPU 📦 POCKET-26B-GGUF (Q4_K_M 17 GB · Q2_K 11 GB): https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF 📚 POCKET collection: https://huggingface.co/collections/FINAL-Bench/pocket-models
reacted to theirpost with 👍 1 day ago
🖼️ POCKET-Image — the POCKET series goes visual: character-perfect text in any language, on-device A new model in VIDRAFT's POCKET family. POCKET put 35B-class models on phones and no-GPU PCs. POCKET-Image carries the same "big capability, small hardware" idea into image generation — and fixes the one thing nearly every image model gets wrong: text. Type "안녕하세요" into a typical model and you get "안ㅐ기." Hangul alone composes 11,172 syllable blocks; Arabic connects its letters; Thai stacks marks. Diffusion models draw scripts as shapes, so they smear. POCKET-Image renders every glyph exactly — 한국어 · 中文 · 日本語 · العربية (RTL) · ไทย · Latin and more — onto any scene you describe. What it is: • 100% accurate text, any language — where global models produce gibberish • Any background from a prompt — text is optional (empty → a pure image) • No GPU, no NPU — runs on plain CPU + RAM via the POCKET-Core engine • Measured footprint: 8.6 GB (RTX 3050/4060) · 4.5 GB (offloaded, 6 GB cards) · 13.4 GB (MacBook, 16 GB+) • Windows · macOS · Linux · fully local, no cloud Built on the open, commercial-friendly Z-Image (Apache-2.0) foundation. Honest note: the text is the guaranteed-correct part — the surrounding scene is ordinary generation, so a busy foreground can crowd the letters. We say so; clean backgrounds stay razor-sharp. 🎨 Studio — generate right here, any language: https://huggingface.co/spaces/FINAL-Bench/POCKET-Image-Studio 🧩 Model card: https://huggingface.co/FINAL-Bench/POCKET-Image-Zimage 📚 The POCKET collection: https://huggingface.co/collections/FINAL-Bench/pocket-models
posted an update 1 day ago
🖼️ POCKET-Image — the POCKET series goes visual: character-perfect text in any language, on-device A new model in VIDRAFT's POCKET family. POCKET put 35B-class models on phones and no-GPU PCs. POCKET-Image carries the same "big capability, small hardware" idea into image generation — and fixes the one thing nearly every image model gets wrong: text. Type "안녕하세요" into a typical model and you get "안ㅐ기." Hangul alone composes 11,172 syllable blocks; Arabic connects its letters; Thai stacks marks. Diffusion models draw scripts as shapes, so they smear. POCKET-Image renders every glyph exactly — 한국어 · 中文 · 日本語 · العربية (RTL) · ไทย · Latin and more — onto any scene you describe. What it is: • 100% accurate text, any language — where global models produce gibberish • Any background from a prompt — text is optional (empty → a pure image) • No GPU, no NPU — runs on plain CPU + RAM via the POCKET-Core engine • Measured footprint: 8.6 GB (RTX 3050/4060) · 4.5 GB (offloaded, 6 GB cards) · 13.4 GB (MacBook, 16 GB+) • Windows · macOS · Linux · fully local, no cloud Built on the open, commercial-friendly Z-Image (Apache-2.0) foundation. Honest note: the text is the guaranteed-correct part — the surrounding scene is ordinary generation, so a busy foreground can crowd the letters. We say so; clean backgrounds stay razor-sharp. 🎨 Studio — generate right here, any language: https://huggingface.co/spaces/FINAL-Bench/POCKET-Image-Studio 🧩 Model card: https://huggingface.co/FINAL-Bench/POCKET-Image-Zimage 📚 The POCKET collection: https://huggingface.co/collections/FINAL-Bench/pocket-models
View all activity

Organizations

FINAL_Bench's profile picture Gemma Challenge's profile picture PPALLI's profile picture