Every model the fleet can reach, by tier. Fit is 0 to 3 per job: 3 = first choice, 0 = not for this. Sizes are vendor-published or marked undisclosed; nothing is guessed. Evidence lines say m: when we measured it here and g: when it is general knowledge of the model.
Fit:3 first choice2 good1 can0 not for this
| Model | Tier · plan | Size · context | Reasoning | Bulk HTML | Bulk code | Ultra-effort code | Code review | SEO | Social | Marketing copy | Google / Meta ads | Research | Vision | Speed | Role · lane | Pros · cons · evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
Fable 5.1
claude-fable-5-1
Anthropic · anthropic
|
Top tier Claude Max 20x (jesse) |
undisclosed 1M ctx |
extended thinking, always on | 1 | 2 | 3 | 3 | 2 | 1 | 2 | 2 | 2 | 2 | 0 | Signs off every ship and every session wrap. Multi-hour marathons. Agent model:fable + [paid-subagent-OK] Ultra-effort code · Code review |
+ Judgment. Long-horizon autonomy. Catches what every cheaper lane passed (jmfield 09-10 crash). - Slowest, ~2x limit drain. Same family as Opus, so alone it is an echo, never a cross-arch check. m: 07-01 bench, zero edge on short tasks; edge = autonomy over hours. Reviews every deploy since 08-2026. |
|
GPT-6 Astra
gpt-6-astra
OpenAI · openai
|
Top tier ChatGPT Plus base ($0 per call, scarce weekly quota) |
undisclosed undisclosed ctx |
reasoning model, effort selectable | 1 | 2 | 3 | 3 | 1 | 0 | 1 | 2 | 1 | 1 | 0 | FINAL cross-arch pass after Fable on the ship gate. Manual only. review-now.py --lane gpt6_astra (Codex on ChatGPT Plus) Ultra-effort code · Code review |
+ Cross-arch to every Claude-authored diff. Terse, line-anchored. Free via Codex since 09-16. - Burns Plus quota fast, so once per ship. Code/logic/security only; a retrieval-free lane fabricated facts on blog audits. Never on a loop. m: 09-07 audit-script bench 4/7 defects, tied with terra. m: 09-16 caught two card errors Fable's pass missed. |
|
Opus 5
claude-opus-5
Anthropic · anthropic
|
Subscription Claude Max |
undisclosed 1M ctx |
extended thinking, effort levels | 1 | 2 | 3 | 2 | 2 | 1 | 2 | 3 | 2 | 2 | 1 | Plan, decompose, integrate, decide. Never the bulk author. the session lead Ultra-effort code · Google / Meta ads |
+ Best judgment per token short of Fable. Drives every skill and hook. - Most expensive tokens in the system; a 400-line lead answer is a routing error. g: architecture and security routed here by rule. Meta/Google ads reasoning lives in the paid-ads skill it drives. |
|
Sonnet 5
claude-sonnet-5
Anthropic · anthropic
|
Subscription Claude Max |
undisclosed 1M ctx |
extended thinking optional | 2 | 3 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | 2 | Execution subagents, ~80% of any build. subagents via the Agent tool Bulk code |
+ Fast, cheap on the plan, strong all-round coder. - Anthropic family: cannot verify Opus or Fable work. g: the default worker for UI, standard features, prompt tuning. |
|
Haiku 4.5
claude-haiku-4-5-20251001
Anthropic · anthropic
|
Subscription Claude Max |
undisclosed 200K ctx |
none | 1 | 1 | 0 | 0 | 1 | 1 | 1 | 0 | 1 | 1 | 3 | Cheap probes, classification, one-line checks. --model haiku for probes and classification Speed |
+ Fastest Claude. - Not a reviewer, not a builder. m: used for the 09-16 token-identity probe. |
|
GPT-5.6 Sol
gpt-5.6-sol
OpenAI · openai
|
Subscription ChatGPT Plus |
undisclosed undisclosed ctx |
reasoning model | 1 | 2 | 2 | 2 | 1 | 0 | 1 | 1 | 1 | 1 | 1 | Review-tier lane, manual. --lane codex_sol (manual) |
+ Third GPT voice at $0. - Unbenched. Shares the weekly Plus allowance with Astra. m: 09-08 probed working, 0 billed requests. Not benched on defects. |
|
GPT-5.6 Terra
gpt-5.6-terra
OpenAI · openai
|
Subscription ChatGPT Plus |
undisclosed undisclosed ctx |
reasoning model | 1 | 3 | 2 | 3 | 1 | 0 | 1 | 1 | 1 | 1 | 1 | Default OpenAI-family reviewer on the per-turn hook and gated reviews. Second opinion on stuck bugs. codex_terra, panel seat 2 (automatic hook) · /codex:rescue · codex exec Bulk code · Code review |
+ Genuine cross-arch to Claude code at $0. Agentic repo sessions via the codex plugin. - Same family as gpt-oss, so not cross-arch to Groq drafts. 134s cold spawn if allowed to run git. m: 09-07 3/5 planted defects, only lane to catch an unbounded discount. m: 19.1s median with no-exec instruction. |
|
GPT-5.6 Luna
gpt-5.6-luna
OpenAI · openai
|
Subscription ChatGPT Plus |
undisclosed undisclosed ctx |
reasoning model | 1 | 2 | 2 | 2 | 1 | 0 | 1 | 1 | 1 | 1 | 1 | Second GPT look when terra passes and you want another. --lane codex_luna (manual) |
+ Different blind spots from terra on the same family. - Off the panel since 09-16: two openai seats is an echo. m: 09-07 3/5, caught what terra missed and missed what terra caught. |
|
Kimi K3
kimi-k3
Moonshot · moonshot
|
Ollama Pro Ollama Pro (extra-high-usage model per Ollama) |
undisclosed 256K ctx |
reasoning_effort required (medium/high), max_tokens 48000 (24000 returns empty) | 0 | 1 | 2 | 3 | 1 | 0 | 1 | 1 | 1 | 0 | 1 | Reviewer only. Optional first lane of the ship chain. --lane kimi_k3 · ship gate Code review |
+ Finds real bugs, cheap, cross-arch to Claude. - Empty reply when reasoning eats the budget. Never bulk. Same family as K2.7. m: 83% completion over 841 calls, best of the Ollama lanes. m: 3-17s at effort medium. |
|
Gemma 4 31B
gemma4:31b
Google · google
|
Ollama Pro Ollama Pro |
31B 128K ctx |
reasoning_effort medium | 1 | 1 | 1 | 3 | 1 | 1 | 1 | 0 | 1 | 1 | 3 | First reviewer on every turn. Only Google family on the plan. ollama_gemma, panel SEAT 1 (automatic hook) Code review · Speed |
+ Fastest reviewer we have, finds every planted defect, costs the allowance almost nothing. - Tersest answers. Unproven on multi-file traps. Seated 09-16, short history. m: 09-16 gauntlet 9/9 plants, 0 false flags, 4.5s median at 0.8/5.6/13.5KB. m: 08-29 re-bench 3/3, 0 decoys. |
|
GLM 5.3 Flash
glm-5.3-flash
Zhipu · zhipu
|
Ollama Pro Ollama Pro |
undisclosed 128K ctx |
reasoning_effort high | 1 | 2 | 2 | 3 | 1 | 0 | 1 | 0 | 1 | 1 | 2 | Gated reviewer. Agentic coding over long sessions. ollama_glm, panel seat 3 Code review |
+ Richest review text of the GLM line; names the convention conflict and gives a repro. - Verbose; strip its thinking block in loops. m: 09-16 gauntlet 9/9, 0 false flags, 10-16s. m: 46% completion on the panel window (timeouts on big diffs at the old cap). |
|
MiniMax M3
minimax-m3:cloud
MiniMax · msa
|
Ollama Pro Ollama Pro |
undisclosed 1M ctx |
reasoning_effort high | 1 | 2 | 2 | 2 | 2 | 1 | 1 | 1 | 2 | 3 | 1 | Gated reviewer. Vision. Prompts over 10k chars. Fact audits. ollama_minimax / ollama_minimax_vision, panel seat 4 Vision |
+ The only Ollama route that takes images or >10k-char prompts. Best fact auditor. - Too slow for the per-turn hook on real diffs. m: 09-16 gauntlet 9/9, 0 false flags, but 48s on 13.5KB. m: 7/7 on a published-content fact audit. m: 24% completion on the old per-turn slot. |
|
Mistral Large 3
mistral-large-3:675b
Mistral · mistral
|
Ollama Pro Ollama Pro |
675B 256K ctx |
reasoning_effort high | 1 | 2 | 2 | 2 | 1 | 1 | 2 | 1 | 2 | 1 | 1 | Gated reviewer. Large documents. Pass-2 code review. ollama_mistral / ollama_code, panel seat 5 |
+ European family, cross-arch to everything else on the panel. - Occasional rubber-stamp miss. m: 66% completion on the panel window. m: 5/6 on the canary (one bare NO DEFECTS). |
|
Qwen 3.5 397B
qwen3.5:397b
Alibaba · alibaba
|
Ollama Pro Ollama Pro |
397B 256K ctx |
reasoning_effort medium, max_tokens 24000+ | 1 | 2 | 2 | 2 | 2 | 2 | 2 | 1 | 2 | 2 | 1 | Gated reviewer, domain verify seat (laravel, android, ios, astro, voip, infra). Writing, analysis, vision. ollama_qwen (panel seat 6, domain verify) / ollama_chat |
+ Only Alibaba family on the panel. Open multimodal. Strong writer. - Slow; reasoning eats small budgets. m: 09-16 live probe: empty at max_tokens 200, OK at 24000. m: ~37s on the 08-27 diff. |
|
Kimi K2.7 Code
kimi-k2.7-code
Moonshot · moonshot
|
Ollama Pro Ollama Pro |
1T MoE, 32B active 256K ctx |
think:false for bulk | 2 | 3 | 2 | 0 | 1 | 0 | 1 | 0 | 1 | 1 | 2 | Backend / SWE first choice. Bulk HTML third behind gpt-oss and deepseek flash. ollama_devstral (SWE) / ollama_kimi (bulk) Bulk code |
+ Agentic coder, SWE-tuned. - 10k-char prompt cap. Moonshot family: never a verifier for K3 or itself. m: 09-11 bulk HTML bench, behind gpt-oss 8.8s and deepseek flash 12.5s. Routed as SWE first choice by rule. |
|
DeepSeek V4.1 Flash
deepseek-v4.1-flash
DeepSeek · deepseek
|
Ollama Pro Ollama Pro |
158B 128K ctx |
thinking optional | 3 | 2 | 2 | 1 | 1 | 1 | 1 | 0 | 1 | 2 | 2 | Bulk HTML fallback #2. Deep reasoning and agentic work. Vision. ollama_deepseek / ollama_deepseek3 Bulk HTML |
+ Fast bulk lane on the flat plan when Groq is bursty. - DeepSeek family invents defects on review (v4-pro history). m: 09-11 bulk HTML 12.5s, 8/8. Vision/tools/thinking confirmed; swapped in 09-13. |
|
DeepSeek V4 Pro
deepseek-v4-pro
DeepSeek · deepseek
|
Ollama Pro Ollama Pro |
1.6T MoE, 49B active 1M ctx |
thinking off by default | 1 | 2 | 2 | 0 | 1 | 0 | 1 | 0 | 1 | 0 | 1 | Backend / SWE fallback. Draft lane, not a reviewer. ollama_deepseek_pro |
+ Huge model, strong on discrete tasks. - Stale April build on Ollama. Never a review seat. m: empty at 24k and 48k tokens on an 11.5KB diff; invented defects. m: 07-01 tied Fable on short tasks. |
|
Nemotron 3 Super
nemotron-3-super
NVIDIA · nvidia
|
Ollama Pro Ollama Pro |
120B 128K ctx |
yes | 1 | 1 | 1 | 0 | 2 | 2 | 2 | 1 | 2 | 0 | 2 | Content and research, mid lane, fast. ollama_nemotron |
+ Fast content lane, NVIDIA family adds a fifth voice for content checks. - Unbenched on review; do not seat without a gauntlet. g: agentic reasoning tuned. Holds no review seat. |
|
Nemotron 3 Ultra
nemotron-3-ultra
NVIDIA · nvidia
|
Ollama Pro Ollama Pro |
undisclosed 128K ctx |
yes | 1 | 2 | 2 | 1 | 1 | 1 | 2 | 1 | 2 | 0 | 1 | High-throughput reasoning, long-running agents. Underused. ollama_nemotron_ultra |
+ Reasoning tier on the flat plan. - No measured role yet. g: reasoning tier. Fable 09-16: bench before seating. |
|
Nemotron 3 Nano 30B
nemotron-3-nano:30b
NVIDIA · nvidia
|
Ollama Pro Ollama Pro |
30B 128K ctx |
light | 1 | 1 | 0 | 0 | 1 | 1 | 1 | 0 | 1 | 0 | 3 | Fast cheap lane: classification, routing, quick edits. ollama_nemotron_nano / ollama_code_fast Speed |
+ Cheapest Ollama call. - Not for anything that needs to be right first time. g: replaced qwen3-coder-next 07-15 as the fast lane. |
|
GPT OSS 120B
openai/gpt-oss-120b
OpenAI (open weights) · openai
|
Free tier Groq free tier |
117B MoE, 5.1B active 128K ctx |
low/medium/high | 3 | 2 | 1 | 1 | 1 | 1 | 2 | 0 | 1 | 0 | 3 | BULK HTML/CSS author, first call for a page from scratch. fast_gpt_oss (Groq) Bulk HTML · Speed |
+ Fastest usable HTML author, free. Copy is decent. - Bursty free tier. OpenAI family: not cross-arch to Codex or Astra. Sub-3% of review defects in the 09-07 bench. m: 09-11 bench 8.8s, 8/8 on the HTML fixture. m: raw Groq 403s without a User-Agent. Params from OpenAI's gpt-oss model card. |
|
GPT OSS 20B
openai/gpt-oss-20b
OpenAI (open weights) · openai
|
Free tier Groq free tier |
21B MoE, 3.6B active 128K ctx |
low/medium/high | 2 | 1 | 0 | 0 | 1 | 1 | 1 | 0 | 0 | 0 | 3 | Tiny HTML fragments, 4.7s. fast_llama4 (repointed to gpt-oss-20b) / fast_ask model=openai/gpt-oss-20b Speed |
+ Fastest of all. - Small; 120B is the same price. m: 4.7s on the 09-11 fixture. |
|
Qwen 3.6 27B
qwen/qwen3.6-27b
Alibaba (open weights) · alibaba
|
Free tier Groq free tier |
27B 128K ctx |
thinking optional | 3 | 2 | 1 | 1 | 1 | 1 | 2 | 0 | 1 | 0 | 3 | Second bulk HTML voice on Groq, Alibaba family. fast_qwen / fast_kimi (Groq) Bulk HTML · Speed |
+ Cross-arch to gpt-oss drafts at the same price: free. - Slower than gpt-oss-120b; qwen3.8 is org-blocked on Groq until enabled. m: 09-11 bench 13.8s, 8/8 on the HTML fixture. |
|
Whisper large-v3 turbo
whisper-large-v3-turbo
OpenAI (open weights) · openai
|
Free tier Groq free tier |
809M audio ctx |
n/a | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 1 | 0 | 3 | Speech to text. fast_transcribe (Groq) Speed |
+ Free, fast transcription. - Audio only. g: the /watch skill's caption fallback. |
|
Gemini 3.8 Flash
gemini-3.8-flash
Google · google
|
Free tier Gemini free tier. Pro (gemini-3.1-pro-preview) is gated: env has it enabled on a prepaid key, doctrine says off |
undisclosed 1M ctx |
thinking optional | 1 | 1 | 1 | 0 | 3 | 2 | 3 | 1 | 3 | 2 | 2 | Research with retrieval. Blog and page copy. Blog Gate B (fact audit with retrieval). mcp__gemini-research__research / write_blog_post / write_page_content SEO · Marketing copy · Research |
+ Grounded research, 1M context, long-form copy. - Pro tier bills when called. Never a code reviewer. m: the only lane with retrieval, which is why it holds Blog Gate B and astra does not. |
|
Xiaomi MiMo
mimo
Xiaomi · xiaomi
|
Metered metered; MCP not registered in ~/.claude.json as of 09-16 (402 on 09-13) |
undisclosed undisclosed ctx |
reasoning + TTS | 0 | 1 | 2 | 1 | 1 | 1 | 1 | 0 | 1 | 0 | 1 | Reasoning second opinion and TTS. mcp__mimo__* |
+ Distinct family; TTS. - Metered, small balance. m: 09-13 returned 402 during the lane outage. |
|
Perplexity Sonar
sonar (search API)
Perplexity · perplexity
|
Metered metered, $10 credit (vault Servers/perplexity.md) |
undisclosed n/a ctx |
n/a (search) | 0 | 0 | 0 | 0 | 2 | 0 | 0 | 0 | 1 | 0 | 2 | AI-visibility probe only: does the fleet appear in Perplexity answers. NOT a research lane. perplexity-fleet-check skill · audit-ai-crawler-access.sh |
+ Cheapest way to see what Perplexity says about a client. - Metered. Research goes to Gemini Flash (retrieval, free). Never wire the agent MCP again. m: 08-26 the perplexity MCP billed 732K tokens on OpenAI models via Perplexity's Response API; removed, guard-perplexity-sonar-only.py refuses it. Sonar chat-completions deprecated 2026-09-27. |
|
Veo 3.1 Lite
veo-3.1-lite-generate-preview
Google · google
|
Metered metered |
n/a n/a ctx |
n/a | 0 | 0 | 0 | 0 | 0 | 3 | 0 | 2 | 0 | 3 | 1 | Text and image to video, $0.40 per 8s clip. gemini-veo.sh Social · Vision |
+ Cheapest video lane. - Metered; fal seedance is the fallback. m: about 125 clips per $50. |
|
GLM 5.3 (full)
glm-5.3
Zhipu · zhipu
|
On the plan, not wired Ollama Pro |
undisclosed 128K ctx |
yes | 1 | 2 | 2 | 1 | 1 | 0 | 1 | 0 | 1 | 1 | 0 | SKIP: 6/7 runs hit the 120s timeout on the review diff. none |
+ Bigger sibling of the seated flash. - Unseatable on latency. m: 08-2026 timeouts. |
|
GLM 5.2 / 5.1
glm-5.2, glm-5.1
Zhipu · zhipu
|
On the plan, not wired Ollama Pro |
undisclosed 128K ctx |
yes | 1 | 2 | 1 | 2 | 1 | 0 | 1 | 0 | 1 | 1 | 2 | SKIP: 5.3-flash at effort high is as fast with richer text. none |
+ Held the seat before flash. - No gain over flash. m: 5.2 was the seat at 8-11s; 5.3-flash 8.7s at high. |
|
GPT OSS 120B / 20B (Ollama copy)
gpt-oss:120b, gpt-oss:20b
OpenAI (open weights) · openai
|
On the plan, not wired Ollama Pro |
117B / 21B MoE 128K ctx |
low/medium/high | 3 | 2 | 1 | 1 | 1 | 1 | 2 | 0 | 1 | 0 | 2 | SKIP: Groq serves the same weights free and fast (Greg: 'groq is pennies'). Would compete with reviewers for the Ollama allowance. none Bulk HTML |
+ Provider redundancy for the bulk lane if Groq dies. - Same weights, same blind spots. m: 09-16 answers on Ollama. Fable: wire as fallback #1; Astra: bench first, correlated errors with Groq. |
|
Kimi K2.6
kimi-k2.6
Moonshot · moonshot
|
On the plan, not wired Ollama Pro |
1T MoE, 32B active 256K ctx |
yes | 2 | 2 | 1 | 0 | 1 | 0 | 1 | 0 | 1 | 1 | 2 | SKIP: older sibling of K2.7-code, no role K2.7 or K3 lack. none |
+ General (not code-tuned) Kimi. - Same family as two wired lanes. g: prior generation. |
|
MiniMax M2.7
minimax-m2.7
MiniMax · msa
|
On the plan, not wired Ollama Pro |
undisclosed 1M ctx |
lighter than M3 | 1 | 2 | 1 | 2 | 1 | 1 | 1 | 0 | 1 | 2 | 2 | BENCH FIRST: candidate large-diff backup if M3 keeps timing out at 48s. none |
+ Might finish where M3 stalls. - Unmeasured. g: prior generation, lighter reasoning. Not measured. |
|
DeepSeek V4 Flash :0731 / V4 Pro :0813
deepseek-v4-flash:0731, deepseek-v4-pro:0813
DeepSeek · deepseek
|
On the plan, not wired Ollama Pro |
158B / 1.6T 128K / 1M ctx |
thinking optional | 2 | 2 | 2 | 0 | 1 | 0 | 1 | 0 | 1 | 1 | 2 | SKIP: pinned older builds. v4.1-flash supersedes :0731; bare deepseek-v4-pro resolves (verified 09-16). none (bare tags are the wired ones) |
+ Reproducible pins. - Nothing the bare tags lack. m: both answer on 09-16. |