-
nota-ai/Solar-Open2-250B-Nota-INT4
Text Generation • 41B • Updated • 69 • 1 -
nota-ai/Solar-Open2-250B-Nota-NVFP4
Text Generation • 145B • Updated • 266 • 9 -
nota-ai/Solar-Open2-250B-Nota-INT4-GlobalPruned
Text Generation • 35B • Updated • 32 • 3 -
nota-ai/Solar-Open2-250B-Nota-NVFP4-GlobalPruned
Text Generation • 117B • Updated • 24 • 3
AI & ML interests
Hardware-aware AI Model Optimization
Recent Activity
Nota AI bridges the gap between high-performance AI models and edge devices.
From our automated optimization platform to bespoke AI solutions, we ensure your AI functions efficiently—everywhere it is needed.
🌟 Spotlight
Sovereign AI Foundation Model Project
Nota AI participates in the Sovereign AI Foundation Model Project, a key initiative by the South Korean government (NIPA) to develop global-tier foundation models. As a core optimization partner, we focus on compressing massive LLMs for practical deployment.
🔥 New Release: Solar-Open2-250B Series
Our latest compression of Upstage's flagship Solar-Open2-250B—a 250B-parameter Mixture-of-Experts model—combining routing-aware INT4/NVFP4 quantization with global expert pruning. The result: a frontier-scale MoE that serves on a single B200 (NVFP4) or two H100s (INT4), at near-lossless accuracy. Built on our proprietary DREAM-MoE and SRA-MoE algorithms, which preserve expert-routing decisions under aggressive low-bit quantization.
🏆 Solar-Open2-250B-Nota-INT4-GlobalPruned
The most deployable variant—global expert pruning + INT4 brings the full 250B MoE down to 117.8 GB, running on just 2× H100.
- Global Expert-Sensitivity Pruning: Removes experts using a model-wide importance score (non-uniform, per-layer)—scoring 83.39 average on high-difficulty reasoning benchmarks vs. 78.79 for uniform pruning.
- Near-Lossless: Only −0.6 pt vs. the unpruned INT4 baseline (83.99), ~99% accuracy retention.
- Hardware Efficiency: 117.8 GB weight footprint—serves 131K context on 2× H100 (tensor-parallel), down from the 4× H100 the unpruned INT4 requires.
- Backend Ready: vLLM (W4A16, AutoRound/GPTQ-compatible).
🏆 Solar-Open2-250B-Nota-NVFP4-GlobalPruned
NVFP4 quantization + global expert pruning, tuned for NVIDIA Blackwell—fits the pruned 250B MoE onto a single B200 with 131K context.
- Requires Blackwell (B200/GB200); leverages the FP4 tensor cores unavailable on Hopper, Ada, and Ampere.
Full-Fidelity Quantized Variants (no pruning)
- Solar-Open2-250B-Nota-INT4 — W4A16 quantization, 500.6 GB → 142.9 GB (−71.5%), 81.17 average vs. 81.57 BF16 (−0.5%). vLLM-ready, tensor-parallel across 4+ GPUs.
- Solar-Open2-250B-Nota-NVFP4 — W4A4 for Blackwell, 500.6 GB → 153.3 GB (−69.4%), 99.7% accuracy retention (81.35 vs. 81.57 BF16 average).
🚀 Our Core Business
🛠️ AI Platform: NetsPresso"We make AI lighter, faster, and ready for deployment." NetsPresso is our proprietary platform that accelerates model optimization, enabling you to secure on-device latency and accuracy without deep hardware expertise.
👉 Ready to optimize? Try NetsPresso Now | Request Private Demo |
🌍 AI Solutions"We provide end-to-end AI solutions powered by our core optimization technology." 1. Nota Vision AgentPowered by Vision Language Models (VLM), this agent goes beyond simple detection to understand complex situational contexts. 2. Edge AI Solutions
|
📚 Tech BlogGain insights into our engineering philosophy. We share deep dives into model compression methodologies and NPU acceleration techniques to help you stay ahead. |
🔗 Connect with Us
|
-
nota-ai/Solar-Open-100B-NotaMoEQuant-NVFP4
Text Generation • 59B • Updated • 14 • 4 -
nota-ai/Solar-Open-100B-Nota-FP8
Text Generation • 103B • Updated • 25 • 31 -
nota-ai/Solar-Open-100B-NotaMoEQuant-Int4
Text Generation • 2B • Updated • 254 • 47 -
nota-ai/Qwen3-30B-A3B-NotaMoEQuant-Int4
Text Generation • 0.6B • Updated • 10 • 8
-
nota-ai/Solar-Open2-250B-Nota-INT4
Text Generation • 41B • Updated • 69 • 1 -
nota-ai/Solar-Open2-250B-Nota-NVFP4
Text Generation • 145B • Updated • 266 • 9 -
nota-ai/Solar-Open2-250B-Nota-INT4-GlobalPruned
Text Generation • 35B • Updated • 32 • 3 -
nota-ai/Solar-Open2-250B-Nota-NVFP4-GlobalPruned
Text Generation • 117B • Updated • 24 • 3
-
nota-ai/Solar-Open-100B-NotaMoEQuant-NVFP4
Text Generation • 59B • Updated • 14 • 4 -
nota-ai/Solar-Open-100B-Nota-FP8
Text Generation • 103B • Updated • 25 • 31 -
nota-ai/Solar-Open-100B-NotaMoEQuant-Int4
Text Generation • 2B • Updated • 254 • 47 -
nota-ai/Qwen3-30B-A3B-NotaMoEQuant-Int4
Text Generation • 0.6B • Updated • 10 • 8