AI & ML interests

Hardware-aware AI Model Optimization

Recent Activity

Organization Card
Nota AI Banner

Nota AI bridges the gap between high-performance AI models and edge devices.
From our automated optimization platform to bespoke AI solutions, we ensure your AI functions efficiently—everywhere it is needed.

Website LinkedIn NetsPresso


🌟 Spotlight

Sovereign AI Foundation Model Project

Nota AI participates in the Sovereign AI Foundation Model Project, a key initiative by the South Korean government (NIPA) to develop global-tier foundation models. As a core optimization partner, we focus on compressing massive LLMs for practical deployment.

🔥 New Release: Solar-Open2-250B Series

Our latest compression of Upstage's flagship Solar-Open2-250B—a 250B-parameter Mixture-of-Experts model—combining routing-aware INT4/NVFP4 quantization with global expert pruning. The result: a frontier-scale MoE that serves on a single B200 (NVFP4) or two H100s (INT4), at near-lossless accuracy. Built on our proprietary DREAM-MoE and SRA-MoE algorithms, which preserve expert-routing decisions under aggressive low-bit quantization.

🏆 Solar-Open2-250B-Nota-INT4-GlobalPruned

The most deployable variant—global expert pruning + INT4 brings the full 250B MoE down to 117.8 GB, running on just 2× H100.

  • Global Expert-Sensitivity Pruning: Removes experts using a model-wide importance score (non-uniform, per-layer)—scoring 83.39 average on high-difficulty reasoning benchmarks vs. 78.79 for uniform pruning.
  • Near-Lossless: Only −0.6 pt vs. the unpruned INT4 baseline (83.99), ~99% accuracy retention.
  • Hardware Efficiency: 117.8 GB weight footprint—serves 131K context on 2× H100 (tensor-parallel), down from the 4× H100 the unpruned INT4 requires.
  • Backend Ready: vLLM (W4A16, AutoRound/GPTQ-compatible).

🏆 Solar-Open2-250B-Nota-NVFP4-GlobalPruned

NVFP4 quantization + global expert pruning, tuned for NVIDIA Blackwell—fits the pruned 250B MoE onto a single B200 with 131K context.

  • Requires Blackwell (B200/GB200); leverages the FP4 tensor cores unavailable on Hopper, Ada, and Ampere.

Full-Fidelity Quantized Variants (no pruning)

  • Solar-Open2-250B-Nota-INT4 — W4A16 quantization, 500.6 GB → 142.9 GB (−71.5%), 81.17 average vs. 81.57 BF16 (−0.5%). vLLM-ready, tensor-parallel across 4+ GPUs.
  • Solar-Open2-250B-Nota-NVFP4 — W4A4 for Blackwell, 500.6 GB → 153.3 GB (−69.4%), 99.7% accuracy retention (81.35 vs. 81.57 BF16 average).

🚀 Our Core Business

🛠️ AI Platform: NetsPresso

"We make AI lighter, faster, and ready for deployment."

NetsPresso is our proprietary platform that accelerates model optimization, enabling you to secure on-device latency and accuracy without deep hardware expertise.

  • Develop & Compress: Create lightweight models effortlessly using our Model Zoo and advanced Compressor (Structured Pruning).
  • Optimize & Convert: Maximize speed on verified hardware (NVIDIA, Arm, Qualcomm, etc.) with Graph Optimization and Graph Quantization.
  • Test on Real Devices: Validate performance instantly on actual devices via our Device Farm to eliminate deployment failures.

👉 Ready to optimize? Try NetsPresso Now | Request Private Demo

🌍 AI Solutions

"We provide end-to-end AI solutions powered by our core optimization technology."

1. Nota Vision Agent

Powered by Vision Language Models (VLM), this agent goes beyond simple detection to understand complex situational contexts.
It interprets video feeds through natural language prompts, delivering real-time insights locally without cloud dependency.

2. Edge AI Solutions

📚 Tech Blog

Gain insights into our engineering philosophy. We share deep dives into model compression methodologies and NPU acceleration techniques to help you stay ahead.

👉 Read our Tech Blog

🔗 Connect with Us


© 2026 Nota Inc. All rights reserved.

datasets 0

None public yet