Choose theme
AI Resources
Cohere North Mini Code
Cohere North Mini Code is a Cohere Labs coding model release aimed at code generation, agentic software engineering, and terminal-based tasks.
Cohere and Hugging Face describe North Mini Code as a 30B-total, 3B-active mixture-of-experts model with Apache 2.0 Hugging Face weights, BF16 and FP8 variants, and try paths through OpenCode, Cohere API, and related Cohere surfaces. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
A coding-focused MoE model
North Mini Code is presented as a 30B-total, 3B-active sparse mixture-of-experts model for coding work, software-engineering tasks, and agent-style terminal workflows.
Why it stands out
Built around coding-agent harnesses
The release materials frame the model around code generation, agentic software engineering, terminal tasks, OpenCode, SWE-Bench-style harnesses, and Terminal-Bench-style evaluation rather than only chat completions.
Availability
Weights and hosted try paths
The official materials list BF16 and FP8 Hugging Face model pages, Apache 2.0 licensing, OpenCode access before downloading, and Cohere API or hosted Cohere deployment paths to inspect separately.
Why it matters
What makes it useful
A coding model still needs a harness to read files, run commands and return tool results. North Mini Code gives builders a model choice for that workflow, with downloadable weights for their own server or a hosted route for trying the same kind of task.
What to know
Where it fits
The model card’s self-served OpenCode path puts North Mini Code behind a local vLLM endpoint. OpenCode connects to that endpoint with interleaved reasoning enabled; vLLM uses Cohere’s tool and reasoning parsers.
Notable points
What stands out
The model has 30B total parameters but activates 3B per token; the download and serving requirements still depend on the whole checkpoint. Cohere lists a 256K context limit and 64K maximum generation, which are separate limits on the combined conversation and the reply. Published benchmark results describe Cohere’s tested harnesses rather than every deployment. Cohere also evaluates complex code generation with SciCode and LiveCodeBench, separately from the tool-using agent harnesses.
Before using
What to review
The Hugging Face model cards identify Apache 2.0. Review the current terms at the main official source to decide whether they suit your intended use, alongside setup notes and current access limits.
Which path is actually being used: BF16 weights, FP8 weights, local runtime, OpenCode, Cohere API, Model Vault, or another provider route.
Hardware and runtime requirements, especially GPU memory, SGLang or vLLM version notes, tool-call parsing, and whether the FP8 checkpoint fits the planned serving stack.
Coding-agent permissions before connecting the model to terminals, repositories, package managers, browsers, credentials, private code, or production systems.
Generated-code review, tests, dependency checks, security review, and privacy settings before using model output in real software work.
Cohere-reported benchmark and provider claims as source claims to inspect, not as a LifeHubber performance judgment.
The FP8 model card says that checkpoint is designed for vLLM and is not compatible with Transformers. For Transformers, the publisher points to the BF16 checkpoint.
Reader fit
Who may find it relevant
Readers comparing public coding models that can sit underneath software agents and terminal-based coding workflows.
Builders looking at OpenCode-style harnesses, local or hosted inference paths, and model choices for agentic software engineering experiments.
Less relevant for readers who only want a general consumer chatbot, a no-setup coding assistant, or a small model for everyday laptop use.
Editorial note
Why LifeHubber lists it
For builders whose terminal agent must carry command evidence into later turns, North Mini Code connects each result to the tool call that produced it while retaining structured reasoning. After a terminal tool runs, return its stdout and return_code in a dictionary with the matching tool_call_id, while preserving the assistant’s structured reasoning and tool call in the conversation. The model card demonstrates that exchange. It lets the next turn continue from the actual command result rather than a disconnected pasted log.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
What to explore next
Compare other coding models before choosing a route.
North Mini Code keeps active compute relatively small for its total size. Continue with a compact software-engineering model family or a larger multimodal coding model built for longer agent workflows.
More in AI Models
Keep browsing this category
Explore more AI model resources.
Gemma 4
google/gemma-4
A Google DeepMind Gemma 4 model family collection with public checkpoints including Gemma 4 12B, a dense multimodal model Google describes around local agentic workflows, native audio input, and encoder-free vision/audio handling.
MiniMax H3 Integrations
MiniMax-AI/awesome-minimax-h3-integration
A community-maintained MiniMax H3 integration index that maps checkpoints, hardware and VRAM starting points, runtimes, ComfyUI nodes, prompting tools, acceleration routes, and deployment options.
MiniCPM5-2B
OpenBMB/MiniCPM5-2B
An OpenBMB 2.52B-parameter language model for local assistants, coding agents, tool use, reasoning, and long-context workflows, with 131K context plus BF16, GGUF, MLX, GPTQ, and DSpark variants.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.