Model roster

Every model the fleet can reach, by tier. Fit is 0 to 3 per job: 3 = first choice, 0 = not for this. Sizes are vendor-published or marked undisclosed; nothing is guessed. Evidence lines say m: when we measured it here and g: when it is general knowledge of the model.

Fit:3 first choice2 good1 can0 not for this

ModelTier · planSize · contextReasoningBulk HTMLBulk codeUltra-effort codeCode reviewSEOSocialMarketing copyGoogle / Meta adsResearchVisionSpeedRole · lanePros · cons · evidence
Fable 5.1
claude-fable-5-1
Anthropic · anthropic
Top tier
Claude Max 20x (jesse)
undisclosed
1M ctx
extended thinking, always on 12332122220
Signs off every ship and every session wrap. Multi-hour marathons.
Agent model:fable + [paid-subagent-OK]
Ultra-effort code · Code review
+ Judgment. Long-horizon autonomy. Catches what every cheaper lane passed (jmfield 09-10 crash).
- Slowest, ~2x limit drain. Same family as Opus, so alone it is an echo, never a cross-arch check.
m: 07-01 bench, zero edge on short tasks; edge = autonomy over hours. Reviews every deploy since 08-2026.
GPT-6 Astra
gpt-6-astra
OpenAI · openai
Top tier
ChatGPT Plus base ($0 per call, scarce weekly quota)
undisclosed
undisclosed ctx
reasoning model, effort selectable 12331012110
FINAL cross-arch pass after Fable on the ship gate. Manual only.
review-now.py --lane gpt6_astra (Codex on ChatGPT Plus)
Ultra-effort code · Code review
+ Cross-arch to every Claude-authored diff. Terse, line-anchored. Free via Codex since 09-16.
- Burns Plus quota fast, so once per ship. Code/logic/security only; a retrieval-free lane fabricated facts on blog audits. Never on a loop.
m: 09-07 audit-script bench 4/7 defects, tied with terra. m: 09-16 caught two card errors Fable's pass missed.
Opus 5
claude-opus-5
Anthropic · anthropic
Subscription
Claude Max
undisclosed
1M ctx
extended thinking, effort levels 12322123221
Plan, decompose, integrate, decide. Never the bulk author.
the session lead
Ultra-effort code · Google / Meta ads
+ Best judgment per token short of Fable. Drives every skill and hook.
- Most expensive tokens in the system; a 400-line lead answer is a routing error.
g: architecture and security routed here by rule. Meta/Google ads reasoning lives in the paid-ads skill it drives.
Sonnet 5
claude-sonnet-5
Anthropic · anthropic
Subscription
Claude Max
undisclosed
1M ctx
extended thinking optional 23222222222
Execution subagents, ~80% of any build.
subagents via the Agent tool
Bulk code
+ Fast, cheap on the plan, strong all-round coder.
- Anthropic family: cannot verify Opus or Fable work.
g: the default worker for UI, standard features, prompt tuning.
Haiku 4.5
claude-haiku-4-5-20251001
Anthropic · anthropic
Subscription
Claude Max
undisclosed
200K ctx
none 11001110113
Cheap probes, classification, one-line checks.
--model haiku for probes and classification
Speed
+ Fastest Claude.
- Not a reviewer, not a builder.
m: used for the 09-16 token-identity probe.
GPT-5.6 Sol
gpt-5.6-sol
OpenAI · openai
Subscription
ChatGPT Plus
undisclosed
undisclosed ctx
reasoning model 12221011111
Review-tier lane, manual.
--lane codex_sol (manual)
+ Third GPT voice at $0.
- Unbenched. Shares the weekly Plus allowance with Astra.
m: 09-08 probed working, 0 billed requests. Not benched on defects.
GPT-5.6 Terra
gpt-5.6-terra
OpenAI · openai
Subscription
ChatGPT Plus
undisclosed
undisclosed ctx
reasoning model 13231011111
Default OpenAI-family reviewer on the per-turn hook and gated reviews. Second opinion on stuck bugs.
codex_terra, panel seat 2 (automatic hook) · /codex:rescue · codex exec
Bulk code · Code review
+ Genuine cross-arch to Claude code at $0. Agentic repo sessions via the codex plugin.
- Same family as gpt-oss, so not cross-arch to Groq drafts. 134s cold spawn if allowed to run git.
m: 09-07 3/5 planted defects, only lane to catch an unbounded discount. m: 19.1s median with no-exec instruction.
GPT-5.6 Luna
gpt-5.6-luna
OpenAI · openai
Subscription
ChatGPT Plus
undisclosed
undisclosed ctx
reasoning model 12221011111
Second GPT look when terra passes and you want another.
--lane codex_luna (manual)
+ Different blind spots from terra on the same family.
- Off the panel since 09-16: two openai seats is an echo.
m: 09-07 3/5, caught what terra missed and missed what terra caught.
Kimi K3
kimi-k3
Moonshot · moonshot
Ollama Pro
Ollama Pro (extra-high-usage model per Ollama)
undisclosed
256K ctx
reasoning_effort required (medium/high), max_tokens 48000 (24000 returns empty) 01231011101
Reviewer only. Optional first lane of the ship chain.
--lane kimi_k3 · ship gate
Code review
+ Finds real bugs, cheap, cross-arch to Claude.
- Empty reply when reasoning eats the budget. Never bulk. Same family as K2.7.
m: 83% completion over 841 calls, best of the Ollama lanes. m: 3-17s at effort medium.
Gemma 4 31B
gemma4:31b
Google · google
Ollama Pro
Ollama Pro
31B
128K ctx
reasoning_effort medium 11131110113
First reviewer on every turn. Only Google family on the plan.
ollama_gemma, panel SEAT 1 (automatic hook)
Code review · Speed
+ Fastest reviewer we have, finds every planted defect, costs the allowance almost nothing.
- Tersest answers. Unproven on multi-file traps. Seated 09-16, short history.
m: 09-16 gauntlet 9/9 plants, 0 false flags, 4.5s median at 0.8/5.6/13.5KB. m: 08-29 re-bench 3/3, 0 decoys.
GLM 5.3 Flash
glm-5.3-flash
Zhipu · zhipu
Ollama Pro
Ollama Pro
undisclosed
128K ctx
reasoning_effort high 12231010112
Gated reviewer. Agentic coding over long sessions.
ollama_glm, panel seat 3
Code review
+ Richest review text of the GLM line; names the convention conflict and gives a repro.
- Verbose; strip its thinking block in loops.
m: 09-16 gauntlet 9/9, 0 false flags, 10-16s. m: 46% completion on the panel window (timeouts on big diffs at the old cap).
MiniMax M3
minimax-m3:cloud
MiniMax · msa
Ollama Pro
Ollama Pro
undisclosed
1M ctx
reasoning_effort high 12222111231
Gated reviewer. Vision. Prompts over 10k chars. Fact audits.
ollama_minimax / ollama_minimax_vision, panel seat 4
Vision
+ The only Ollama route that takes images or >10k-char prompts. Best fact auditor.
- Too slow for the per-turn hook on real diffs.
m: 09-16 gauntlet 9/9, 0 false flags, but 48s on 13.5KB. m: 7/7 on a published-content fact audit. m: 24% completion on the old per-turn slot.
Mistral Large 3
mistral-large-3:675b
Mistral · mistral
Ollama Pro
Ollama Pro
675B
256K ctx
reasoning_effort high 12221121211
Gated reviewer. Large documents. Pass-2 code review.
ollama_mistral / ollama_code, panel seat 5
+ European family, cross-arch to everything else on the panel.
- Occasional rubber-stamp miss.
m: 66% completion on the panel window. m: 5/6 on the canary (one bare NO DEFECTS).
Qwen 3.5 397B
qwen3.5:397b
Alibaba · alibaba
Ollama Pro
Ollama Pro
397B
256K ctx
reasoning_effort medium, max_tokens 24000+ 12222221221
Gated reviewer, domain verify seat (laravel, android, ios, astro, voip, infra). Writing, analysis, vision.
ollama_qwen (panel seat 6, domain verify) / ollama_chat
+ Only Alibaba family on the panel. Open multimodal. Strong writer.
- Slow; reasoning eats small budgets.
m: 09-16 live probe: empty at max_tokens 200, OK at 24000. m: ~37s on the 08-27 diff.
Kimi K2.7 Code
kimi-k2.7-code
Moonshot · moonshot
Ollama Pro
Ollama Pro
1T MoE, 32B active
256K ctx
think:false for bulk 23201010112
Backend / SWE first choice. Bulk HTML third behind gpt-oss and deepseek flash.
ollama_devstral (SWE) / ollama_kimi (bulk)
Bulk code
+ Agentic coder, SWE-tuned.
- 10k-char prompt cap. Moonshot family: never a verifier for K3 or itself.
m: 09-11 bulk HTML bench, behind gpt-oss 8.8s and deepseek flash 12.5s. Routed as SWE first choice by rule.
DeepSeek V4.1 Flash
deepseek-v4.1-flash
DeepSeek · deepseek
Ollama Pro
Ollama Pro
158B
128K ctx
thinking optional 32211110122
Bulk HTML fallback #2. Deep reasoning and agentic work. Vision.
ollama_deepseek / ollama_deepseek3
Bulk HTML
+ Fast bulk lane on the flat plan when Groq is bursty.
- DeepSeek family invents defects on review (v4-pro history).
m: 09-11 bulk HTML 12.5s, 8/8. Vision/tools/thinking confirmed; swapped in 09-13.
DeepSeek V4 Pro
deepseek-v4-pro
DeepSeek · deepseek
Ollama Pro
Ollama Pro
1.6T MoE, 49B active
1M ctx
thinking off by default 12201010101
Backend / SWE fallback. Draft lane, not a reviewer.
ollama_deepseek_pro
+ Huge model, strong on discrete tasks.
- Stale April build on Ollama. Never a review seat.
m: empty at 24k and 48k tokens on an 11.5KB diff; invented defects. m: 07-01 tied Fable on short tasks.
Nemotron 3 Super
nemotron-3-super
NVIDIA · nvidia
Ollama Pro
Ollama Pro
120B
128K ctx
yes 11102221202
Content and research, mid lane, fast.
ollama_nemotron
+ Fast content lane, NVIDIA family adds a fifth voice for content checks.
- Unbenched on review; do not seat without a gauntlet.
g: agentic reasoning tuned. Holds no review seat.
Nemotron 3 Ultra
nemotron-3-ultra
NVIDIA · nvidia
Ollama Pro
Ollama Pro
undisclosed
128K ctx
yes 12211121201
High-throughput reasoning, long-running agents. Underused.
ollama_nemotron_ultra
+ Reasoning tier on the flat plan.
- No measured role yet.
g: reasoning tier. Fable 09-16: bench before seating.
Nemotron 3 Nano 30B
nemotron-3-nano:30b
NVIDIA · nvidia
Ollama Pro
Ollama Pro
30B
128K ctx
light 11001110103
Fast cheap lane: classification, routing, quick edits.
ollama_nemotron_nano / ollama_code_fast
Speed
+ Cheapest Ollama call.
- Not for anything that needs to be right first time.
g: replaced qwen3-coder-next 07-15 as the fast lane.
GPT OSS 120B
openai/gpt-oss-120b
OpenAI (open weights) · openai
Free tier
Groq free tier
117B MoE, 5.1B active
128K ctx
low/medium/high 32111120103
BULK HTML/CSS author, first call for a page from scratch.
fast_gpt_oss (Groq)
Bulk HTML · Speed
+ Fastest usable HTML author, free. Copy is decent.
- Bursty free tier. OpenAI family: not cross-arch to Codex or Astra. Sub-3% of review defects in the 09-07 bench.
m: 09-11 bench 8.8s, 8/8 on the HTML fixture. m: raw Groq 403s without a User-Agent. Params from OpenAI's gpt-oss model card.
GPT OSS 20B
openai/gpt-oss-20b
OpenAI (open weights) · openai
Free tier
Groq free tier
21B MoE, 3.6B active
128K ctx
low/medium/high 21001110003
Tiny HTML fragments, 4.7s.
fast_llama4 (repointed to gpt-oss-20b) / fast_ask model=openai/gpt-oss-20b
Speed
+ Fastest of all.
- Small; 120B is the same price.
m: 4.7s on the 09-11 fixture.
Qwen 3.6 27B
qwen/qwen3.6-27b
Alibaba (open weights) · alibaba
Free tier
Groq free tier
27B
128K ctx
thinking optional 32111120103
Second bulk HTML voice on Groq, Alibaba family.
fast_qwen / fast_kimi (Groq)
Bulk HTML · Speed
+ Cross-arch to gpt-oss drafts at the same price: free.
- Slower than gpt-oss-120b; qwen3.8 is org-blocked on Groq until enabled.
m: 09-11 bench 13.8s, 8/8 on the HTML fixture.
Whisper large-v3 turbo
whisper-large-v3-turbo
OpenAI (open weights) · openai
Free tier
Groq free tier
809M
audio ctx
n/a 00000100103
Speech to text.
fast_transcribe (Groq)
Speed
+ Free, fast transcription.
- Audio only.
g: the /watch skill's caption fallback.
Gemini 3.8 Flash
gemini-3.8-flash
Google · google
Free tier
Gemini free tier. Pro (gemini-3.1-pro-preview) is gated: env has it enabled on a prepaid key, doctrine says off
undisclosed
1M ctx
thinking optional 11103231322
Research with retrieval. Blog and page copy. Blog Gate B (fact audit with retrieval).
mcp__gemini-research__research / write_blog_post / write_page_content
SEO · Marketing copy · Research
+ Grounded research, 1M context, long-form copy.
- Pro tier bills when called. Never a code reviewer.
m: the only lane with retrieval, which is why it holds Blog Gate B and astra does not.
Xiaomi MiMo
mimo
Xiaomi · xiaomi
Metered
metered; MCP not registered in ~/.claude.json as of 09-16 (402 on 09-13)
undisclosed
undisclosed ctx
reasoning + TTS 01211110101
Reasoning second opinion and TTS.
mcp__mimo__*
+ Distinct family; TTS.
- Metered, small balance.
m: 09-13 returned 402 during the lane outage.
Perplexity Sonar
sonar (search API)
Perplexity · perplexity
Metered
metered, $10 credit (vault Servers/perplexity.md)
undisclosed
n/a ctx
n/a (search) 00002000102
AI-visibility probe only: does the fleet appear in Perplexity answers. NOT a research lane.
perplexity-fleet-check skill · audit-ai-crawler-access.sh
+ Cheapest way to see what Perplexity says about a client.
- Metered. Research goes to Gemini Flash (retrieval, free). Never wire the agent MCP again.
m: 08-26 the perplexity MCP billed 732K tokens on OpenAI models via Perplexity's Response API; removed, guard-perplexity-sonar-only.py refuses it. Sonar chat-completions deprecated 2026-09-27.
Veo 3.1 Lite
veo-3.1-lite-generate-preview
Google · google
Metered
metered
n/a
n/a ctx
n/a 00000302031
Text and image to video, $0.40 per 8s clip.
gemini-veo.sh
Social · Vision
+ Cheapest video lane.
- Metered; fal seedance is the fallback.
m: about 125 clips per $50.
GLM 5.3 (full)
glm-5.3
Zhipu · zhipu
On the plan, not wired
Ollama Pro
undisclosed
128K ctx
yes 12211010110
SKIP: 6/7 runs hit the 120s timeout on the review diff.
none
+ Bigger sibling of the seated flash.
- Unseatable on latency.
m: 08-2026 timeouts.
GLM 5.2 / 5.1
glm-5.2, glm-5.1
Zhipu · zhipu
On the plan, not wired
Ollama Pro
undisclosed
128K ctx
yes 12121010112
SKIP: 5.3-flash at effort high is as fast with richer text.
none
+ Held the seat before flash.
- No gain over flash.
m: 5.2 was the seat at 8-11s; 5.3-flash 8.7s at high.
GPT OSS 120B / 20B (Ollama copy)
gpt-oss:120b, gpt-oss:20b
OpenAI (open weights) · openai
On the plan, not wired
Ollama Pro
117B / 21B MoE
128K ctx
low/medium/high 32111120102
SKIP: Groq serves the same weights free and fast (Greg: 'groq is pennies'). Would compete with reviewers for the Ollama allowance.
none
Bulk HTML
+ Provider redundancy for the bulk lane if Groq dies.
- Same weights, same blind spots.
m: 09-16 answers on Ollama. Fable: wire as fallback #1; Astra: bench first, correlated errors with Groq.
Kimi K2.6
kimi-k2.6
Moonshot · moonshot
On the plan, not wired
Ollama Pro
1T MoE, 32B active
256K ctx
yes 22101010112
SKIP: older sibling of K2.7-code, no role K2.7 or K3 lack.
none
+ General (not code-tuned) Kimi.
- Same family as two wired lanes.
g: prior generation.
MiniMax M2.7
minimax-m2.7
MiniMax · msa
On the plan, not wired
Ollama Pro
undisclosed
1M ctx
lighter than M3 12121110122
BENCH FIRST: candidate large-diff backup if M3 keeps timing out at 48s.
none
+ Might finish where M3 stalls.
- Unmeasured.
g: prior generation, lighter reasoning. Not measured.
DeepSeek V4 Flash :0731 / V4 Pro :0813
deepseek-v4-flash:0731, deepseek-v4-pro:0813
DeepSeek · deepseek
On the plan, not wired
Ollama Pro
158B / 1.6T
128K / 1M ctx
thinking optional 22201010112
SKIP: pinned older builds. v4.1-flash supersedes :0731; bare deepseek-v4-pro resolves (verified 09-16).
none (bare tags are the wired ones)
+ Reproducible pins.
- Nothing the bare tags lack.
m: both answer on 09-16.