Menu

Post image 1
Post image 2
Post image 3
Post image 4
Post image 5
Post image 6
1 / 6
121

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 · Hugging Face

Hacker News·about 1 month ago
#y9ZyDjUo
Reading 0:00
15s threshold

Model Summary Total Parameters 30B (3B active) Architecture MoE - Mamba-2 + MoE + Attention hybrid Context Length Up to 1M tokens Single-GPU Deployment 1× DGX Spark (GB10) or 1× H100 Supported Hardware NVIDIA Blackwell (DGX Spark / GB10, GB200, GeForce RTX 5090); NVIDIA Hopper (H100, H200); NVIDIA Ampere via W4A16 Supported Languages English (and coding languages), Spanish, French, German, Italian, Japanese Speculative Decoding DSpark for low-concurrency Data Centre and DGX Spark Workflows — Read more below , also provided are MTP (Multi-Token Prediction) and DFlash Recommended Sampling Temperature 1.0, Top_P 0.95 Best For Long-running autonomous agents, sub-agent workhorse deployments, and efficient local inference on personal hardware License OpenMDW License Agreement, version 1.1 Release Date August 11, 2026 Model Overview Model Developer: NVIDIA Corporation Model Dates: December 2025 - May 2026 Data Freshness: The pre-training data has a cutoff date of September 2025.…

Continue reading — create a free account

Join HashtagPLUS to read full articles, follow hashtags, vote, and join the conversation.

Read More