Laguna S 2.1
118B open-weight coding MoE model designed for agentic software engineering
Laguna S 2.1 is a 118-billion-parameter Mixture-of-Experts (MoE) model with 8 billion activated parameters per token, developed by Poolside for long-horizon agentic coding tasks. It features a 1,048,576-token context window, native reasoning/thinking capabilities, and a mixed attention architecture combining global and sliding-window attention. The model is open-weight under the OpenMDW-1.1 license and can be self-hosted on a single NVIDIA DGX Spark. On benchmark tasks like Terminal-Bench 2.1 (70.2%), SWE-Bench Pro (59.4%), and DeepSWE (40.4%), it matches or exceeds models several times its size, including DeepSeek-V4-Pro-Max.
Who it's for
Pricing · freemium
checked today| Plan | Price | Includes |
|---|---|---|
| Free (Limited Context) | Free | 262K token context window · 32K max output tokens · Free usage with model improvement clause |
| Paid via OpenRouter (1M Context) | Free | $0.10 per million input tokens · $0.20 per million output tokens · 1M token context window · 131K max output tokens |
| Featherless Flat-Rate | $10 /mo | Flat-rate pricing from $10/month · OpenAI-compatible API · 250K context access |
AI-researched pricing — verify on the official site before subscribing.
Use it for
- — Agentic coding and automation
- — Long-horizon software engineering tasks
- — Complex codebase analysis and modification
- — Terminal-based coding agents
- — Multi-step problem solving with tool integration
- — Self-hosted deployment for compliance and sovereignty
- — Building coding assistants and copilots
Get the most out of it
- 01Use pool, Poolside's native agent harness, for tighter integration with the model's 1M context window and native thinking capabilities—it offers superior ergonomics compared to third-party wrappers
- 02Enable thinking/reasoning mode (enabled by default) to improve benchmark performance significantly; Terminal-Bench improves from 60.4% to 70.2% and DeepSWE from 16.5% to 40.4%
- 03Self-host for long-running agentic work: the 236GB BF16 checkpoint fits on a single DGX Spark, allowing you to move high-volume coding tasks off metered APIs onto controlled hardware
- 04Leverage the 1M token context window for analyzing entire codebases and long task histories without chunking, enabling better reasoning and fewer round-trips
- 05For consumer hardware deployment, use the smaller Laguna XS 2.1 variant (33B, 3B active parameters) which runs comfortably on 24GB VRAM machines