Laguna S 2.1 · Total Parameters
118B
Total parameters with only 8B active per token via MoE routing
Laguna M.1 · Total Parameters
225B
23B activated parameters — trained on 6,144 NVIDIA H200 GPUs
Laguna XS.2 · Total Parameters
33B
Only 3B activated per token; runs on a Mac with 36 GB RAM via Ollama
Laguna S 2.1 · Max Context
1M
1,048,576-token context window; default 256K on OpenRouter
Training Data Scale
30T+
Both M.1 and XS.2 pre-trained on 30+ trillion tokens from scratch
S 2.1 Training Speed
<9wk
Under 9 weeks from training start to public release; <4 weeks on 4,000 H200s
Company & Infrastructure Numbers
$626M
Total Funding Raised
$3B
Series B Valuation (Oct 2024)
$12B
Target Series C Valuation (2025)
2GW
Project Horizon AI Campus (Texas)
40K+
GB300 NVL72 GPUs (CoreWeave deal)
60
Applied Research Team Members
~$50M
Estimated Annual Revenue
2023
Founded (San Francisco)
15%
Muon Optimizer Edge over AdamW — Poolside's distributed Muon implementation achieves the same training loss in ~15% fewer steps than the standard AdamW optimizer, while requiring only one state per parameter (vs. two for AdamW), cutting checkpoint memory.
6,144
GPUs for Laguna M.1 Pre-training — Laguna M.1 was trained using 6,144 interconnected NVIDIA Hopper (H200) GPUs with NVLink; Laguna XS.2 used 2,048 H200 GPUs.
4,000
H200 GPUs for Laguna S 2.1 — Pre-training for the S 2.1 model began May 22, 2026 on 4,000 NVIDIA H200 GPUs, completing the entire cycle in under 4 weeks.
<1%
Muon Optimizer Overhead — During Laguna M.1 pre-training, the combined Muon optimizer overhead (including Newton–Schulz iterations) was reduced to under 1% of total training step time.
13%
Synthetic Data in Training Mix — Approximately 13% of Laguna XS.2's final training mix consists of synthetic data, with the broader Laguna family having used 4.4T+ synthetic tokens in total.
~60
AutoMixer Proxy Models — Instead of hand-tuned data mixing ratios, Poolside trained ~60 proxy models on different data compositions to automatically optimize training data mix via learned regression.
6.8%
Laguna S 2.1 Parameter Activation Rate — Laguna S 2.1 activates only ~6.8% of its 118B parameters on any given token, enabling frontier-class performance on hardware as compact as a single DGX Spark.