LIFEHUBBER
Choose theme

AI Resources

Cohere North Mini Code

Hugging Face likes: 577 Hugging Face downloads, last 30 days: 9.2K Declared license: Apache-2.0: Apache-2.0 Last modified June 15, 2026: Modified 3mo ago
Stats from Hugging Face

Cohere North Mini Code is a Cohere Labs coding model release aimed at code generation, agentic software engineering, and terminal-based tasks.

Cohere and Hugging Face describe North Mini Code as a 30B-total, 3B-active mixture-of-experts model with Apache 2.0 Hugging Face weights, BF16 and FP8 variants, and try paths through OpenCode, Cohere API, and related Cohere surfaces. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.

What it is

A coding-focused MoE model

North Mini Code is presented as a 30B-total, 3B-active sparse mixture-of-experts model for coding work, software-engineering tasks, and agent-style terminal workflows.

Why it stands out

Built around coding-agent harnesses

The release materials frame the model around code generation, agentic software engineering, terminal tasks, OpenCode, SWE-Bench-style harnesses, and Terminal-Bench-style evaluation rather than only chat completions.

Availability

Weights and hosted try paths

The official materials list BF16 and FP8 Hugging Face model pages, Apache 2.0 licensing, OpenCode access before downloading, and Cohere API or hosted Cohere deployment paths to inspect separately.

Why it matters

What makes it useful

A coding model still needs a harness to read files, run commands and return tool results. North Mini Code gives builders a model choice for that workflow, with downloadable weights for their own server or a hosted route for trying the same kind of task.

Notable points

What stands out

The model has 30B total parameters but activates 3B per token; the download and serving requirements still depend on the whole checkpoint. Cohere lists a 256K context limit and 64K maximum generation, which are separate limits on the combined conversation and the reply. Published benchmark results describe Cohere’s tested harnesses rather than every deployment. Cohere also evaluates complex code generation with SciCode and LiveCodeBench, separately from the tool-using agent harnesses.

Before using

What to review

The Hugging Face model cards identify Apache 2.0. Review the current terms at the main official source to decide whether they suit your intended use, alongside setup notes and current access limits.

Which path is actually being used: BF16 weights, FP8 weights, local runtime, OpenCode, Cohere API, Model Vault, or another provider route.

Hardware and runtime requirements, especially GPU memory, SGLang or vLLM version notes, tool-call parsing, and whether the FP8 checkpoint fits the planned serving stack.

Coding-agent permissions before connecting the model to terminals, repositories, package managers, browsers, credentials, private code, or production systems.

Generated-code review, tests, dependency checks, security review, and privacy settings before using model output in real software work.

Cohere-reported benchmark and provider claims as source claims to inspect, not as a LifeHubber performance judgment.

The FP8 model card says that checkpoint is designed for vLLM and is not compatible with Transformers. For Transformers, the publisher points to the BF16 checkpoint.

Reader fit

Who may find it relevant

Readers comparing public coding models that can sit underneath software agents and terminal-based coding workflows.

Builders looking at OpenCode-style harnesses, local or hosted inference paths, and model choices for agentic software engineering experiments.

Less relevant for readers who only want a general consumer chatbot, a no-setup coding assistant, or a small model for everyday laptop use.

Editorial note

Why LifeHubber lists it

For builders whose terminal agent must carry command evidence into later turns, North Mini Code connects each result to the tool call that produced it while retaining structured reasoning. After a terminal tool runs, return its stdout and return_code in a dictionary with the matching tool_call_id, while preserving the assistant’s structured reasoning and tool call in the conversation. The model card demonstrates that exchange. It lets the next turn continue from the actual command result rather than a disconnected pasted log.

Source links

Source materials

Reader note

Before relying on this entry

LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.

What to explore next

Compare other coding models before choosing a route.

North Mini Code keeps active compute relatively small for its total size. Continue with a compact software-engineering model family or a larger multimodal coding model built for longer agent workflows.

Advertisements

Advertisements

For project maintainers

Listed here? You can use the badge.

If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.

See what’s moving