North Mini Code runs on a single H100, but generates 3x the output tokens

Cohere open-sourced North Mini Code on Tuesday, a 30 billion parameter mixture-of-experts coding agent that activates 3 billion parameters per token and runs on a single H100 under an Apache 2.0 license. The release is positioned as a sovereignty-first alternative to managed frontier coding models, and that positioning holds on cost control and data residency. The harder question is capability and verbosity: independent analysis places the model 18th of 127 open-weight models on the Artificial Analysis Intelligence Index and 8th on output speed, while measuring output token volume at three times the class median. The verbosity number reshapes the deployment economics the single-H100 framing implies, and the cost math is workload-dependent in ways the launch announcement does not address.

The architecture is a 30 billion parameter sparse mixture-of-experts model with 128 experts, eight of which activate per token. Compute at inference is closer to a 3 billion parameter model despite the 30 billion total. The model supports a 256,000 token context window with a 64,000 token maximum generation length, available on Hugging Face under Apache 2.0. Cohere co-founder Nick Frosst demoed it running on a Mac Studio via MLX at around 20 gigabytes of RAM, the same machine he uses for his own coding work. That single-machine footprint is the load-bearing claim of the launch, and the source does not benchmark it under concurrent multi-user load, which is the production condition most teams will run it in.

Cohere's training approach is more interesting than the architecture suggests. The model was trained on more than 70,000 verifiable tasks across approximately 5,000 repositories, deduplicated against SWE-Bench, through two stages of supervised fine-tuning followed by reinforcement learning with verifiable rewards. Rather than optimizing against a single agent scaffold, Cohere trained across three: SWE-Agent with a rich CLI, Mini-SWE-Agent with a single bash tool, and OpenCode with structured JSON outputs. The team reports a 10 percentage point gain on OpenCode from the multi-harness approach while maintaining SWE-Agent performance. Multi-harness training is a real engineering choice that addresses tool-format brittleness, but the source does not compare it against the single-harness baseline on the same evaluation set, so the 10 point gain reads as evidence the approach works, not evidence it dominates simpler training.

The competitive picture centers on Mistral Devstral Small 2, a 24 billion parameter dense model. According to Cohere's internal testing under identical hardware, North Mini Code achieves 2.8x higher output throughput and a 30 percent inter-token latency advantage over Devstral Small 2. Cohere's Hugging Face technical post also claims the model outperforms open-source models up to four times its parameter count, including 120 billion parameter models, on its reported benchmarks. The vendor self-comparison is the floor of the case Cohere is making, and the ceiling is the 4x-parameter-count claim. Both need independent reproduction before they are useful for procurement decisions.

Independent measurement tells a more cautious story. Artificial Analysis ranks the model 8th of 127 comparable open-weight models on output speed at 210 tokens per second, with a time-to-first-token of 0.25 seconds against a class median of 1.95 seconds. On the Artificial Analysis Intelligence Index, the model places 18th of 127. The intelligence ranking is the gap the source does not frame as a gap: North Mini Code is fast, but it is not a frontier coding model by independent measure. The model also generated 75 million output tokens to complete the Intelligence Index against a class median of 25 million, the 3x verbosity factor that the launch framing does not lead with. In high-volume agentic pipelines where every orchestration step pays a per-token bill, that verbosity compounds into inference cost and latency in ways the single-H100 headline does not capture.

Frosst's launch framing makes the positioning explicit. He told the launch video audience that "suddenly people are thinking like hey, am I getting enough economic value out of the tokens from a model?" and positioned local deployment as the answer. On X, he contrasted the model with Anthropic's Claude Fable 5, which the source calls "the most capable publicly available managed coding model" and which runs at $50 per million output tokens. Frosst wrote that North Mini Code is "small, cost effective, apache 2.0, and locally deployable... vs large, expensive, proprietary and hegemonic." The dichotomy the founder is drawing is between deployment control and managed capability, and the 18th-of-127 intelligence ranking is what that dichotomy costs in raw model quality.

Engineering teams now face a real three-way comparison rather than a two-way one. Against GitHub Copilot, Cursor, and Claude Fable 5 (all managed, no on-premises option), North Mini Code offers residency and licensing control. Against Mistral Devstral Small 2, it claims throughput and latency advantages that need independent reproduction. Against Claude Fable 5 at $50 per million output tokens, it competes on cost only if the 3x output verbosity does not erase the local-deployment savings in real workloads. Cohere's choice to benchmark the model on Terminal-Bench v2, which tests agents in real terminal environments rather than synthetic code generation tasks, suggests the team knows that leaderboard scores on coding benchmarks do not decide production agentic deployment, and that evaluation matters more for the use cases North Mini Code is positioned for.

The 18th-of-127 intelligence ranking, the 3x output token volume, and the absence of independent reproduction for the vendor's throughput and latency claims together mean the deployment case is determined by workload volume, tool-failure rate, and concurrency profile, none of which the source evaluates. The single-H100 footprint is the cleanest part of the story, and the Apache 2.0 license is the strongest signal Cohere has sent about how it expects this model to be adopted. The economics depend on whether verbosity holds at three times comparable models under production agentic traffic, a measurement the source does not provide and that determines whether the sovereignty framing or the frontier-capability framing is the more accurate one for any given team.

Subscribe to AI Enthusiast Log

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe