System architecture v30 · 2026-09-16

Greg’s working
AI workflow

Every model, skill, agent and gate actually wired on this machine, read from ~/.claude rather than from a wishlist. Cross-model review fires at session end, and no two seats on a panel share a family.

6Review seats
153Skills
40Agents
15MCP servers
12Ollama Pro models

Full model roster33 models: tier, size, reasoning, what each is good at, pros and cons

What changed in v30

v30, 2026-09-16: TWO ACCOUNTS, ASTRA ON PLUS, A MEASURED PANEL. Claude Code now has a real profile switcher: claude = jesse Max 20x (own config dir ~/.claude-jesse, all MCPs symlinked), claude-jmf = admin@jmfield.com Max 5x; Orca panes stay jmf. Codex CLI 0.154 removed mcp-server, so mcp__codex__codex is dead: Codex work rides the codex plugin (/codex:rescue) or codex exec, and luna / terra / sol / gpt-6-astra all answer on the Plus plan at $0. Astra moved off the funded key and is now the FINAL cross-arch pass AFTER Fable on the ship gate (Fable on the 20x plan is the abundant one, Astra on Plus base is scarce, so it runs last, once). A 27-call gauntlet (3 diff sizes x 3 plants + decoy) reordered the panel: gemma4 9/9 at 4.5s takes seat 1, terra 2, glm 3, minimax 4 (9/9 but 48s on 13.5KB), mistral 5, qwen3.5 6 (new: Alibaba family, domain verify seat). luna off the panel (two OpenAI seats is an echo), Fable and Opus have cards, a full model roster page went up.

v29: BULK HTML MOVED TO GROQ. Eight lanes benched on one real landing-page spec (four Groq, four Ollama) with eight machine gates: Groq gpt-oss-20b 4.7s and gpt-oss-120b 8.8s, both 8/8, free tier; deepseek-v4.1-flash 12.5s on the flat Ollama plan; Kimi K2.7 passes only with thinking forced off and once returned nothing; glm-5.3-flash hit the token cap. Groq retired every Llama model plus Kimi K2 and qwen3-32b, so groq-mcp was repointed. Because gpt-oss is OpenAI-family, codex_terra is no longer cross-arch to HTML drafts; kimi_k3 on the ship gate is. kimi_k3 left the AUTOMATIC panel earlier (too slow for a 30s slice) and stays the manual cross-arch lane Fable consults; the panel table below reflects that. Cerebras key added, 402 until billing clears.

v28: THE GATED REVIEW CHAIN IS WIRED: the Ollama Pro lanes draft, then kimi_k3 → ollama_glm → ollama_mistral → ollama_minimax → codex_terra review, then Fable signs off. Five families, no per-call spend, same panel at internal, client and prod. The automatic hook runs seats 1 and 2; the rest fire on gated reviews and the ship gate. gpt-6-astra is manual only and never sits on an automatic panel. Moonshot, OpenCode paid, Gemini Pro and DeepSeek-direct are all unwired. Plan tiers differ: Claude Max 20x, Ollama basic Pro, ChatGPT basic Plus.

01 · Request lifecycle

How one prompt moves through the stack

Input lands on the planner, the lead agent executes and integrates, and everything below is the routing table it draws from.

INPUT
User Request
O
Opus 5
PLANNER

Reasoning, strategy, architecture, plans

nativeFable 5 marathon lane
S
Sonnet 5
LEAD AGENT

Primary executor · Reviews all outputs · Integrates · Deploys

native delegates via MCP
RESEARCH / REASONING
Gemini 3.8 Flashgemini-3.8-flash

Research, blog posts, FAQs, comparisons, case studies

researchcomparefind_best
CODE / BACKEND
GPT OSS 20Bvia Groq · 4.7s, the fastest lane on the stackFREE

Small sibling of the bulk author, benched 4.7s for a full landing page on 2026-09-11. The fast_llama4 tool now serves it: Groq retired Llama 4 Scout, Llama 3.3, Kimi K2 and qwen3-32b in 2026-09.

fast_llama4
GPT OSS 120Bvia Groq · BULK HTML/CSS author, 8.8s, benched 8/8 on 2026-09-11FREEBULK HTML

First call for HTML/CSS from scratch (RULE #0.5): fastest lane, free tier, bursty. Fallback deepseek-v4.1-flash on Ollama. OpenAI family, so NOT cross-arch to Codex or Astra output; a wired free fallback reviewer for Claude or moonshot drafts only, unseated from panels because it cannot co-sit with the Codex seats.

fast_gpt_oss
Qwen 3.6 27Bvia Groq · 13.8s, second bulk HTML voiceFREE

Alibaba family on Groq, so a cross-arch second draft against gpt-oss at the same price. Benched 13.8s, 8/8 on 2026-09-11. fast_code and fast_ask default to gpt-oss-120b.

fast_qwenfast_kimi
DeepSeek V4 Pro1.6T / 49B active (Ollama copy)OLLAMA PRO · STALE BUILD

Draft lane, not a reviewer. Ollama's copy is a stale April build. The metered direct API is deleted and its MCP unregistered.

ollama_deepseek_pro
DeepSeek V4.1 Flash158BOLLAMA PRO

Bulk HTML fallback behind Groq gpt-oss (12.5s, 8/8 on 2026-09-11). Deep reasoning and agentic work. Use deepseek-v4.1-flash -- the old deepseek-v4-flash:0731 pin was superseded 2026-09; the bare v4.1 name is the current live build (confirmed 2026-09-11, vision/tools/thinking/cloud).

ollama_deepseekollama_deepseek3
Gemma 4 31Bgemma4:31b · review panel SEAT 1 · 4.5s median on 12KBOLLAMA PROSEATED 2026-09-16

The only Google-family lane on the plan. Re-bench 2026-08-29: 3/3 planted defects, 0 decoys flagged, 0 errors, fastest reviewer we have. Seat 1: one of the two automatic-hook lanes. Gauntlet 2026-09-16: 9/9 plants, 0 false flags, 4.5s median at every size. Terse; unproven on multi-file traps.

ollama_gemma
GLM 5.3 Flashglm-5.3-flash · review panel seat 3OLLAMA PRO

Seat 3, gated reviews only; the automatic hook takes seats 1-2 (gemma, terra). Gauntlet 2026-09-16: 9/9 plants, 0 false flags, 10-16s. Agentic coding over long sessions. Verbose, strip its thinking block in loops.

ollama_glm
MiniMax M3OLLAMA PRO

Seat 4, gated reviews. Gauntlet 2026-09-16: 9/9 plants, 0 false flags, but 48s on a 12KB diff, too slow for the per-turn hook. MSA architecture, so a genuine cross-arch second opinion. 1M context and vision: the only Ollama route that takes images or prompts over 10k chars.

ollama_minimaxollama_minimax_vision
Kimi K3kimi-k3 · via OLLAMA · ship-gate + manual review lane, off the automatic panelOLLAMA PRO SUB

Unseated from the automatic panel 2026-09-08: ship-gate lane beside Fable and --lane kimi_k3 only. Reviewer only, never bulk (K2.7-code is the bulk Kimi). Moonshot family, served by Ollama Pro. Needs reasoning_effort: medium and max_tokens: 48000, or it returns empty.

kimi_k3
Fable 5.1claude-fable-5-1 · Agent model:fable · SIGNS OFF every shipCLAUDE MAX 20xTOP TIER

The judge. Reviews every session wrap and every deploy; nothing ships without Fable AND a cross-arch stamp. Multi-hour marathons are its one real edge. Needs [paid-subagent-OK]. Never on a loop or cron.

fable_5
Opus 5claude-opus-5 · lead / planner / integratorCLAUDE MAX

The session lead: spec, decompose, integrate, decide. Never the bulk author (RULE #0.5). Deep reasoning, architecture and security-adjacent work route here, not to Fable.

opus_5
GPT-6 Astravia Codex on ChatGPT Plus · FINAL review AFTER Fable · ship gate + manual --laneCHATGPT PLUS SUBOFF EVERY PANEL

Flipped to Codex 2026-09-16: no API key, no per-call cost. Heaviest OpenAI lane, so it is the FINAL cross-arch pass, run once after Fable on a clean diff (Fable rides the 20x plan, Astra the Plus base quota), the ship chain is k3 (optional) -> Fable -> Astra; reach it with --lane gpt6_astra. Never an automatic panel seat. Code, architecture and security. Never facts.

gpt6_astra
GPT-5.6 Solvia Codex on ChatGPT Plus · review tier onlyCHATGPT PLUS SUBOFF EVERY PANEL

Seated 2026-09-13 as a review-tier lane, never an automatic panel seat: the shared weekly Plus allowance is the budget. Reach it with --lane codex_sol.

codex_sol
GPT-5.6 Terravia Codex on ChatGPT Plus · review panel seat 2CHATGPT PLUS SUB

Seat 2, the default OpenAI-family reviewer and a genuine cross-architecture check against Claude-authored code (astra and sol are the manual review-tier siblings). Runs through Codex on the ChatGPT Plus sub: no API key, no per-call cost. Fires on the automatic panel (seat 2) and gated reviews; the ship chain itself is k3 (optional) -> Fable -> Astra.

codex_terra
GPT-5.6 Lunavia Codex on ChatGPT Plus · manual --lane onlyCHATGPT PLUS SUBOFF EVERY PANEL

Second OpenAI-family pass, off the panel since 2026-09-16 (terra holds the openai seat; two seats of one family is an echo). Bench 2026-09-07: terra 3/5, luna 3/5, each catching what the other missed, so run luna when terra passes and you want a second GPT look. --lane codex_luna.

codex_luna
Kimi K2.7 CodeOLLAMA PRO

Backend / SWE first choice as ollama_devstral. Bulk HTML only with thinking OFF, and third behind Groq gpt-oss and deepseek-v4.1-flash. Moonshot family: never a verifier for kimi_k3 or its own drafts. ollama_kimi is the same weights.

ollama_kimiollama_devstral
Mistral Large 3 675BOLLAMA PRO

Seat 5 of six on the gated panel. Complex reasoning, large documents, vision. Also serves ollama_code.

ollama_mistralollama_code
CONTENT / UTILITY
Nemotron 3 SuperNVIDIA 120BOLLAMA PRO

Agentic reasoning, content, mid lane (fast)

ollama_nemotron
Nemotron 3 UltraNVIDIAOLLAMA PRO

High-throughput reasoning + long-running agents. Content and research lane; holds no review seat.

ollama_nemotron_ultra
Nemotron 3 Nano 30BNVIDIAOLLAMA PRO

Fast/cheap agentic lane, quick tasks, classification, routing, lightweight reasoning.

ollama_nemotron_nanoollama_code_fast
Qwen3.5 397Bqwen3.5:397b · review panel seat 6 · domain verify seatOLLAMA PRO

Open multimodal, Alibaba family (the panel's only one). Reasoning, writing, analysis, vision. Seated 2026-09-16 as gated-panel seat 6 and the domain verify lane (laravel, android, ios, astro, voip, infra), replacing the disabled opencode qwen3.6. Off the automatic hook.

ollama_chatollama_qwen
Veo 3.1 · Geminigemini-veo.shVIDEO

Text and image to video. Lite $0.40/8s is the default, about 125 clips per $50. Cheapest video lane.

veo-3.1-lite-generate-previewgemini-veo.shfal-ai/seedance (fallback)
Findings flow to Fable, then Astra takes the final cross-arch pass; Opus 5 / Sonnet 5 integrate and deploy
// MCP_SERVERS, 15 ACTIVE
chrome-devtools RULE

Render before judging. DOM + accessibility tree + screenshot, not curl. Default for any UX, design audit, redesign, or rendering task per CLAUDE.md.

new_page · take_snapshot · take_screenshot
dataforseo SEO

Live SERP, keyword volume, backlinks, on-page Lighthouse, business listings, AI visibility (LLM mentions). 80+ tools, the canonical SEO data source.

serp · keywords · backlinks · ai_opt
stitch-design UI

Google Stitch, UI generation from prompts. Design systems, screen variants, project scaffolding. Convert output to semantic CSS before shipping.

create_design_system · generate_screen · variants
fal-image

fal.ai image generation + edit endpoints, fallback when cc-nano-banana (Gemini) is unavailable or for specific fal-only models.

generate_image · edit_image

The 15 user-scope servers in ~/.claude.json: chrome-devtools · claude-design · dataforseo · fal-image · figma · gemini-research · google-ads · groq-fast · magic · mobbin · ollama-pro · opencode · refero · stitch-design · ubersuggest. Project-scoped adds onepagelove; claude.ai connectors (Gmail, Drive, Calendar, Dropbox) and plugin servers (playwright, context7) are separate again. Codex is no longer an MCP server: Codex 0.154 removed mcp-server, so it is reached through the codex plugin and codex exec. Disabled and deliberately not listed above: mirage, xiaomi-mimo, glm-free, mistral-code. Unregistered 2026-09-07: kimi-direct (Moonshot API retired) and deepseek (metered direct API; Ollama Pro serves the same model as ollama_deepseek_pro). Both keys vaulted.

// BUILD_PIPELINES
ASTRO SSR FRONT-END BUILD
STEP 1
Design Input

Screenshot, brief, or reference file

DESIGN.md awesome-design-md
STEP 2
Skills Layer

Claude Code skills invoked by Opus

/stitch /astro-ssr /taste /design-system /codex:rescue /memory-recall /skill-auditor
STEP 3
Subagents

Specialized agents spawned per task

ollama_kimi ollama_devstral codex:codex-rescue
STEP 4
Models Execute

Fast code gen + design research

ollama_code_fast Gemini 3.8 Flash GPT OSS 20B
STEP 5
Review + Deploy

Code review then ship

/codex:rescue rsync deploy
LIBRARIES: awesome-design-md awesome-agent-skills awesome-claude-code-subagents codex-plugin-cc
ANDROID APP BUILD , SIDELOADED , 4 APPS ON THIS STACK
STEP 1
Spec, by the lead

5-part contract before any worker spawns. No spec, no spawn.

Opus /android-app-build
STEP 2
Author, Ollama Pro sub

Kotlin + Compose screens. The lead never bulk-authors.

ollama_devstral ollama_deepseek_pro recraft-v3 (art)
STEP 3
Cross-arch review

Never review a DeepSeek draft with DeepSeek. Same family reads as agreement.

fast_gpt_oss codex exec gpt-5.6-terra
STEP 4
Gates, then a real device

Invariant tests, then hardware. A subagent cannot see a screen.

gradlew test guarded adb screenrecord
STEP 5
Sign, publish, poll

Signer compared to the last shipped APK. Manifest and binary propagate independently.

apksigner verify CF Pages Fable wrap review
HARD RULE: a check that cannot fail is not a check never a raw adb tap both install paths: fresh and upgrade data invariants are tests, not eyeballs
// QUALITY_GATES
PRE-DEPLOY GATES, MANDATORY BEFORE EVERY SHIP
SHIP GATE, BLOCKING

Enforced by fable-ship-gate.py. A deploy or push without a stamp is refused, not warned.

python3 scripts/review-now.py <path> --lane kimi_k3 cross-arch lane. Empty reply is NOT REVIEWED, never a pass Fable sign-off Agent tool, model: fable. Freeze and md5 the tree first python3 scripts/fable-stamp.py --repo … --verdict pass writes the stamp the gate looks for
POST-DEPLOY AUDITS, EXIT 1 ON FAIL

Run after the CDN purge, never before: an audit against a stale edge measures the previous build.

audit-seo-files.sh audit-design.sh audit-cwv.sh audit-emdash-live.sh audit-wall-of-text.py audit-clipped-content.mjs --mobile --open-menus audit-ai-crawler-access.sh audit-live-security.sh snippet-eligibility-sweep.sh post-deploy-sweep.py seo-compliance-check.py <url> per-PAGE gate, run on every page we touch: title, meta, one h1, canonical (absolute and path-matched), indexability, lang, viewport, img alt, dashes in copy AND alt text, JSON-LD parses. Open Graph is advisory, not gated. --selftest proves all 27 mutations can fail
CWV comes from the PageSpeed Insights API, never headless Lighthouse.
PER-STACK, BEFORE THE SHIP GATE
Astro SSR npm run build:css if Tailwind changed, Playwright desktop + mobile, commit before deploy Laravel php artisan test, phpstan --level=5, check N+1, config:clear Any repo gitleaks runs on every commit via global core.hooksPath. A blocked commit is a real secret, not a nuisance Client-facing pages taste + impeccable + humanizer, then the web-ship-gate stamp
// SAFETY_RAILS
ANTI-MISTAKE SYSTEM, RULES THAT PREVENT DATA LOSS
Rules that block destructive operations.
DEPLOY SAFETY
rsync Rules
NEVER use rsync --delete, destroys server-only files
MD5 drift-check before every deploy, compare local vs server checksums first
All deploy scripts use ./deploy.sh with built-in drift protection
Server is source of truth, pull before push
deploy.sh drift-check md5sum compare
DEPLOYMENT
Deploy Targets
Deploy targets configured per-project in each repo's CLAUDE.md
Each site has a deploy.sh script with correct server + path
EDIT RULES
File Safety
Never edit compiled HTML on the server, all sites are Astro SSR, edit .astro source
Check for astro.config.mjs FIRST before touching any .html file
Playwright screenshot required after every frontend edit
Multi-device sync: push to backup remote, pull on second machine
/astro-ssr skill playwright hook
SERVER STATE
Never reset unfamiliar state
git reset --hard HEAD on a shared-CMS server destroys uncommitted live edits.
Audit dirty state + stashes before any destructive git op on a server with active users.
SSH HYGIENE
ControlMaster for rsync chains
Add ControlMaster auto + ControlPersist 10m to ~/.ssh/config for contabo.
Multiplexed sessions avoid rate-limit lockouts during rapid deploy chains.
BOT-WALL POLICY
Escalation ladder LOCKED
Patchright = top tier. CloakBrowser rejected.
LinkedIn / Trustpilot / G2 / Google SERP scraping ABANDONED, use official APIs.
Public marketing pages only. Never logged-in or personal data (CFAA / DMCA §1201 / EU sui generis floor).
SEO AUDIT WORKFLOW, DO NOT SKIP
STEP 1, FIRST PASS (NO PER-CALL COST)
Technical audit → seo-audit skill + openseo (251 rules) or gemini-research
AI citability → geo-optimizer (47 criteria), llms.txt, AI bot rules, passage optimization
E-E-A-T + citations → seo-geo skill
Never spawn general-purpose agents for SEO research, they default to Sonnet and burn tokens
STEP 2, SERP LANDSCAPE ONLY
seo-specialist agent = SERP landscaping ONLY, who's ranking, competitors, review counts, directory gaps
seo-specialist runs on Haiku and hallucinates technical facts, never trust its page counts, schema claims, or indexed URL lists
STEP 3, ACT
Cross-reference hub scores + Gemini findings + SERP data before touching any file
seo-* skill suite for implementation (24 skills, schema, sitemap, technical, content, geo, local, maps, cluster, drift, programmatic, competitor-pages …)
Playwright screenshot + deploy + IndexNow ping after every SEO change
AUDIT GAPS, BACKLOG

A client POC, SEO audit never ran. Schedule seo-audit + seo-geo + seo-local once content is final.

NEW, LATE MAY 2026
Ship Workflow (mandatory)

Every deployed page runs: taste → humanizer → frontend-design → impeccable → SEO suite. SEO skills are not optional, not batch-only, per-page on every ship.

Factcheck Gate (step 2.5)

blog-factcheck verifies every claim against cited sources via WebFetch. BLOCKS ship on NOT-FOUND, unverified entity/date/quote, rubric < 90, or P0 fabricated stat. Pair with blog-factcheck-fix, Ollama Pro + subagents rewrite, Claude orchestrates (10-25× cheaper than Claude rewriting).

Browser-verify Alpine (hard rule)

For Alpine-heavy templates (phone.html, contacts/detail.html), open as affected user with DevTools open BEFORE commit. Server-side smoke tests miss silent x-text binding failures.

UPDATED 2026-09-07
kimi_k3 optional ship-gate / manual reviewer

Kimi K3 is a ship-gate / manual review lane (off the automatic panel since 2026-09-08; gemma4 holds seat 1 since 2026-09-16), not DeepSeek. Parity rule: Opus / Sonnet / DeepSeek / GLM / Gemini-Pro tie on correctness, spend paid verify only for cross-arch echo-breaking + code craft, never correctness.

chrome-devtools rule

Render before judging. DOM + a11y + screenshot, not curl. Default for any UX, render, audit, or redesign task.

DAILY-USE SKILLS, HIGHEST LEVERAGE
humanizer

Remove AI-writing tells. Detects/fixes inflated symbolism, em-dash overuse, rule-of-three, AI vocab. Run on every Gemini/GPT-drafted copy before publish.

frontend-design + impeccable

Distinct production-grade UI. impeccable flags AI-slop tells (side-stripe borders, gradient text, glassmorphism, hero-metric clichés). Pair for design audit + polish.

taste-skill

Senior UI/UX engineer. Editorial typography, gapless bento grids, strict GSAP scroll triggers, massive section spacing. Default for any web dev (per CLAUDE.md taste-skill rule).

SEO suite, 24 skills

seo-audit · seo-technical · seo-content · seo-geo · seo-google · seo-local · seo-maps · seo-schema · seo-sitemap · seo-cluster · seo-drift · seo-firecrawl · seo-rotation · seo-sxo · seo-image-gen · seo-programmatic · seo-competitor-pages · seo-ecommerce · seo-plan · seo-dataforseo · seo-page · seo-backlinks

cc-nano-banana

Required for ALL image generation. Nano Banana (Gemini CLI) for blog images, thumbnails, icons, diagrams, illustrations, photos. Free local fallback (2026-06-07): sd.cpp (Metal + SDXL, local-image.sh) for bulk/uncensored stills, keeps fal=video, recraft=text-in-image.

notebooklm

Query + manage Google NotebookLM via CLI. Vault-to-Master sync after every session.

visual-review + playwright

Mandatory after any frontend edit, screenshot via Playwright, Read the PNG before calling work done. Per CLAUDE.md rule #4.

verify-my-work + verification-before-completion

Run after ANY code edit, deployment, or fix. Requires running verification commands and confirming output before claiming completion.

handoff

Write handoff doc so fresh agent continues in clean context window. Use for scope-creep split-outs, ~120k dumb-zone compression, planner→prototype→planner round-trip, or cross-CLI pass (Codex, Copilot, OpenCode). Triggers: "handoff", "spawn agent for", "fresh session".

ponytail + caveman

ponytail least-code ladder (YAGNI → stdlib → native → one line) cuts the code; caveman cuts the prose. Run both. A 10-run bench in June 2026 measured ponytail at −28% LOC on realistic builds, same correctness, faster. Levels: light/full/ultra. Installed 2026-06-19.

NEW TOOLS, APRIL 2026
/design-system

71 DESIGN.md files from Stripe, Figma, Apple, PlayStation, WIRED, VoltAgent/awesome-design-md

/codex:rescue

Codex CLI second opinion on ChatGPT Plus (gpt-5.6-terra default; gpt-5.4-mini is not callable on this plan), invoke when stuck >2 attempts or need architecture review

/skill-auditor

Weekly audit: broken symlinks, GitHub freshness, usage stats, duplicate detection

/start + /wrap

Session lifecycle, /start reads Memory+Obsidian+NotebookLM, /wrap commits+pushes+updates Obsidian

142 Active Skills

142 active skills, 151 stashed in skills-disabled/ to keep the session baseline small. Active set: humanizer, frontend-design, impeccable, taste-skill, full 24-skill SEO suite, Astro suite, cc-nano-banana, notebooklm, visual-review, verify-my-work. Plus plugin marketplaces (caveman, ponytail, obra-superpowers, interface-design, codex, memsearch).

Seedance 1.0 (fal, fallback)

ByteDance text-to-video via fal.ai, Pro (1080p) + Lite (720p). Live now.

SEO INTELLIGENCE TOOLS, INSTALLED GLOBALLY
seo-audit-skill CLI

251 audit rules across 20 categories, @seomator engine. CLI: seo-audit skill + openseo

technical content schema
geo-optimizer PYTHON

47 AI citability criteria, llms.txt, AI bot rules, passage-level optimization for ChatGPT/Perplexity/Gemini

GEO llms.txt AI search
NEW, APRIL 28, 2026
APRIL 23, 2026
MemSearch

Semantic vector search across 1,238 curated memory files. ONNX bge-m3 embeddings (local, free). Auto-captures sessions via hooks. Backed by MEMORY.md multi-index (root + 5 sub-indexes: active / infra / apis-skills / feedback / hot).

/memory-recall 4 hooks
Memory Scripts

memory-audit.py (orphans / empties / oversized) + reorganize-memory.sh (split >50-line files, fix index) + sync-memsearch-ifchanged.sh (AUTO re-index, do NOT manual sync-memsearch.sh) + weekly agent-os hygiene (Layer 5, report-only, Sat)

Level 2 audit
Karpathy Guidelines

4 behavioral principles from Andrej Karpathy: Think Before Coding, Simplicity First, Surgical Changes, Goal-Driven Execution. + RULE #0.7 (Boris Cherny, 2026-06-07): Loops not Prompts · Pre-compute > Inference · Tokens over Headcount.

cc-habits.md rules skill
NotebookLM Tier 3

Semantic across vault + master notebook. Master notebook. CLI authenticated (Mac + Fedora). Vault → project notebook → Master after every session.

notebooklm skill google account
Obsidian Vault Tier 2

the Obsidian vault, Projects, Servers, Sessions, NotebookLM. Per-project file pattern. Check FIRST before asking.

RULE #0
PLANNED, IN PROGRESS
ECharts Dashboard

Interactive reporting dashboard for CrawlHound, grade history, scan trends, site-wide metrics visualization

PDF Reports (ReportLab)

Downloadable PDF audit reports for CrawlHound, MrBotsworth, and gjapp, branded, shareable

CI/CD, Gitea Actions

Automated build + test + deploy pipeline on git.myseodesk.com, lint, SEO checks, rsync deploy on push

// HOOK_LAYER_ENFORCEMENT, v1 (2026-05-05)
DRIFT SIGNALS FROM CLAUDE USAGE REPORT, WHAT THE HOOKS FIX
72% DRIFT

Subagent-heavy sessions. Top spawns: backend-developer, ollama_kimi, code-reviewer, all have sub-tier MCP twins.

target after hooks: < 30%
76% DRIFT

Usage at >150k context. Long mixed-topic sessions, no /clear or /compact between unrelated tasks.

target after hooks: < 30%
4% DRIFT

Skill invocation share. The 100+ installed skills (visual-review, redesign, audits) barely fire, same work runs through paid subagents instead.

target after hooks: > 15%
THE HOOKS THAT MATTER, INSTALLED IN ~/.claude/hooks/
agent-spawn-gate.py PreToolUse

Fires on every Agent tool call. Drift list (6 named agents): warns ≤65% context used, blocks >65%. code-reviewer + general-purpose in ALWAYS_BLOCK. Whitelist (Explore/Plan/gsd/audit/seo/feature-dev) silent. Override: [paid-subagent-OK] token logs + allows.

delegation-table-injector.py UserPromptSubmit

Three triggers: session start, topic-switch keywords, post-/compact detection (transcript msg-count drop >50%). Injects ~250 tokens of routing rules from ~/.claude/refs/delegation-cheatsheet.md.

topic-switch-nudge.py UserPromptSubmit

Detects topic-switch keywords (“okay new task”, “switching to”, “moving on”) and appends a /clear reminder. 10-message debounce so it doesn’t fire on every consecutive switch.

skill-trigger-injector.py UserPromptSubmit

Highest-leverage hook. 46 keyword triggers in ~/.claude/rules/skill-triggers.json (regex, priority-ranked). On match, injects an EXTREMELY_IMPORTANT system-reminder naming the skill, same mechanism as using-superpowers.

chrome-devtools-reminder.py UserPromptSubmit NEW

Enforces the “render before judging” rule. When a prompt looks like UX / design / audit / redesign work, injects a reminder to use mcp__chrome-devtools__new_page + take_snapshot + take_screenshot from the main session, not curl, not WebFetch, not subagent.

SIX DRIFT SUBAGENTS, ALL DISABLED ON DISK, WITH THE SUB-TIER ROUTE THAT REPLACED THEM
drift agent sub-tier mcp route when in-agent is justified
backend-developerollama_devstral, ollama_deepseek_pro5+ files coordinated, or hard arch
ollama_kimifast_gpt_oss (HTML), ollama_code (review)multi-framework full-stack work
fastapi-developerollama_deepseek_pro, ollama_devstralcomplex async patterns + tests
code-reviewerHARD-BLOCKED. Chain: kimi_k3 (DEFAULT) → ollama_glm → ollama_mistral → ollama_minimax → codex_terra[paid-subagent-OK] override only
security-auditorapi-keys skill; security-auditor is a DISABLED agentwhole-system audit (5+ services)
database-administratorsql-pro is a DISABLED agent. Use the mysql/postgres MCP directlyHA/replication infra, not just queries
VERIFICATION CHAIN, MULTI-EYE REVIEW PATTERN

Apply when work is non-trivial: >50 LOC change, security-sensitive, or first-time deploy. Six distinct families, none metered. Gated panel order (from model-seats.yml): gemma4 (google), terra (openai), glm-5.3-flash (zhipu), minimax-m3 (msa), mistral-large-3 (mistral), qwen3.5 (alibaba). The automatic Stop hook runs seats 1-2 only; gemma finds every planted defect in 4.5s on a 13.5KB diff (gauntlet 2026-09-16), which is why it holds the per-turn seat. kimi_k3 is off the panel since c31c794: it is the optional first lane of the ship chain. The ship chain is kimi_k3 (optional) → Fable → gpt-6-astra FINAL.

1. DRAFT
Groq gpt-oss (HTML) / ollama_devstral / ollama_deepseek_pro
2. SEAT 1 (auto hook)
ollama_gemma (google), 4.5s
3. SEAT 2 (auto hook), OPENAI
codex_terra via ChatGPT Plus
4. SEATS 3-6, GATED
ollama_glm · ollama_minimax · ollama_mistral · ollama_qwen
5. SHIP CHAIN
kimi_k3 (optional) → Fable reviews and fixes
6. FINAL PASS
gpt-6-astra, once, on the clean diff. Then Sonnet integrates
// SUBAGENTS, 40 ACTIVE
SPECIALIZED SUBAGENTS, SPAWNED BY OPUS VIA AGENT TOOL

The 40 that live in ~/.claude/agents. Anything under agents-disabled/ is left out, because it does not run.

SEO (19)
seo-technicalseo-contentseo-geoseo-localseo-mapsseo-schemaseo-sitemapseo-clusterseo-driftseo-ecommerceseo-flowseo-sxoseo-googleseo-dataforseoseo-backlinksseo-performanceseo-visualseo-image-genseo-specialist
PAID ADS (9)
audit-googleaudit-metaaudit-budgetaudit-creativeaudit-trackingaudit-compliancecreative-strategistcopy-writervisual-designer
CODE QUALITY (6)
ecc-comment-analyzerecc-harness-optimizerecc-opensource-sanitizerecc-pr-test-analyzerecc-refactor-cleanerecc-silent-failure-hunter
SALES (5)
sales-companysales-competitivesales-contactssales-opportunitysales-strategy
ADS PIPELINE (1)
format-adapter
// TOKEN_ESTIMATION
ROUGH COST PER PLAN TYPE, OPUS 5 ($5/M in · $25/M out)
Estimates assume Opus as planner + Sonnet subagents. Actual cost varies by context size and iterations.
Plan Type Lead Model ~Tokens ~Cost Notes
Quick Ask Ollama Pro (Nemotron/Chat) 1-3K $0 (Pro) ollama_nemotron / ollama_chat, Tier 1
Bug Fix Sonnet 5 5-15K ~$0.05-0.15 Read file → diagnose → patch → verify
Single Feature Opus 5 + Sonnet 20-60K ~$0.17-0.70 Plan → subagent execute → /codex:rescue
Static Page Build Opus + Groq gpt-oss 30-80K ~$0.17-0.50 DESIGN.md → /stitch → Playwright check → deploy
Astro SSR Feature Opus + ollama_devstral / ollama_code 50-120K ~$0.35-1 /astro-ssr skill + full pre-deploy gate
Laravel API Opus + ollama_devstral, verify codex_terra 60-150K ~$0.50-1.30 Models + migrations + tests + PHPStan + review
Full Site (5-10 pages) Opus + mix 150-400K ~$1.30-4 Full 6-stage pipeline: PLAN→DESIGN→BUILD→SEO→CRO→QA
SEO Audit Gemini 3.8 Flash 10-30K ~$0.10-0.30 Schema + compliance + content + schema, mostly Gemini
MINIMIZE COST
  • → Use Ollama Pro (Tier 1) for drafts, analysis, code review
  • → Use Groq gpt-oss-120b / 20b (free, 5-9s) for bulk HTML; deepseek-v4.1-flash on Ollama as the batch fallback
  • → Use Gemini 3.8 Flash for research/content
  • → NVIDIA NIM is archived. Groq is no longer a fallback: gpt-oss is the bulk HTML seat; its Llama, Kimi K2 and qwen3-32b lanes were retired by Groq in 2026-09
  • → /compact regularly in long sessions
CONTEXT WATCH
  • → Each file read adds ~1-5K tokens
  • → Large HTML files: 10-30K tokens each
  • → Long conversations drift, use /clear or new session
  • → Subagents get fresh context, use them for big files
  • → CLAUDE.md loaded every session (~5K tokens)
WHEN TO UPGRADE
  • → Stuck after 2 attempts → /codex:rescue
  • → Complex architecture → Opus (not Sonnet)
  • → Critical prod deploy → the automatic panel ollama_minimax → codex_terra → ollama_glm → ollama_mistral, then the ship gate: kimi_k3 (cross-arch, manual lane) and Fable sign off. No metered lane.
  • → Fresh frontend design take → Groq gpt-oss or ollama_glm; ollama_kimi only with thinking off. gpt-5.5 is DISABLED.
  • → No-per-call cross-arch verify → fast_gpt_oss (free tier) or Codex CLI on Plus
  • → Sub-tier model fails tool calls → switch to Sonnet
// REFERENCE_TABLE
Role Model MCP Server Use Case Cost
Planner Opus 5 native Reasoning, architecture, strategy paid
Lead Sonnet 5 native Executes tasks, reviews, deploys paid
Research Gemini 3.8 Flash gemini-research Research, comparisons, blog posts, FAQs free
Fast 20B GPT OSS 20B replaced Llama 3.3 70B, retired by Groq fast_ask model=openai/gpt-oss-20b Fastest lane on the stack: 4.7s for a full landing page, 8 of 8 gates FREE
Qwen Qwen 3.6 27B replaced Llama 4 Scout, retired by Groq fast_qwen / fast_kimi (both repointed) 131k context, 13.8s on the same landing spec, 8 of 8. qwen3.8-27b needs the org toggle FREE
Cross-arch GPT OSS 120B fast_gpt_oss Bulk HTML author since v29 (see HTML row). Off every panel: as OpenAI-family it cannot cross-check its own drafts or terra's FREE
HTML GPT-OSS 120B / 20B bulk seat since v29 fast_gpt_oss (Groq) Bulk HTML. Benched 09-11 on a real landing spec: 20B in 4.7s, 120B in 8.8s, both 8 of 8 gates, free tier. Fallback deepseek-v4.1-flash (Ollama). Kimi K2.7 only with thinking off. FREE
Content Nemotron 3 Super 120B ollama_nemotron Agentic reasoning, content PRO
Ollama Pro Mistral Large 3 675B replaced qwen3-coder 480B ollama_code Review panel seat 5, cross-arch vs the google, openai, zhipu and msa seats above PRO
Ollama Pro Qwen3.5 397B ollama_chat Reasoning, writing, analysis, vision, thinking PRO
Ollama Pro Mistral Large 3 675B ollama_mistral Complex reasoning, large documents, vision PRO
Ollama Pro DeepSeek V4.1 Flash 158B ollama_deepseek Deep reasoning, thinking, agentic tasks PRO
Ollama Pro Kimi K2.7 Code replaced devstral-2 123B ollama_devstral SWE coding, multi-file editing (surviving specialist) PRO
Ollama Pro GLM 5.3 Flash ollama_glm Agentic coding, sustained over 100s of rounds PRO
// REVIEW_PANELS (generated from model-seats.yml)
StakesPanel (after egress policy)
INTERNALollama_gemma → codex_terra → ollama_glm → ollama_minimax → ollama_mistral → ollama_qwen
CLIENTollama_gemma → codex_terra → ollama_glm → ollama_minimax → ollama_mistral → ollama_qwen
PRODollama_gemma → codex_terra → ollama_glm → ollama_minimax → ollama_mistral → ollama_qwen
KeyFamilyModelEgress
codex_lunaopenaigpt-5.6-lunaok
codex_solopenaigpt-5.6-solok
codex_terraopenaigpt-5.6-terraok
fable_5anthropicclaude-fable-5-1ok
gpt6_astraopenaigpt-6-astraok
gpt_ossopenaiopenai/gpt-oss-120bok
kimi_k3moonshotkimi-k3ok
ollama_deepseek_prodeepseekdeepseek-v4-prook
ollama_devstralmoonshotkimi-k2.7-codeok
ollama_gemmagooglegemma4:31bok
ollama_glmzhipuglm-5.3-flashok
ollama_kimimoonshotkimi-k2.7-codeok
ollama_minimaxmsaminimax-m3:cloudok
ollama_mistralmistralmistral-large-3:675bok
ollama_qwenalibabaqwen3.5:397bok
opus_5anthropicclaude-opus-5ok
qwen36alibabaopencode/qwen3.6-plusok
0
Review seats
0
Ollama Pro
0
Hook commands
0
MCP Servers
0
Active Skills
0
Active Agents

+ 52 hook commands across 15 events (incl. agent-spawn-gate, stop-cross-review, fable-ship-gate, delegation-table, skill-trigger) · 1,238 memory files in MemSearch · NotebookLM master notebook · Obsidian vault per-project files

// GATES, DESIGN MCPS, RECENT SKILL WORK
THE GATE SYSTEM

Rules that live in code, not in prose, because prose loses. 52 hook commands across 15 events. These are the ones that BLOCK rather than remind:

fable-ship-gate no deploy or push without a Fable sign-off plus a cross-arch stamp agent-spawn-gate refuses paid subagents that have a sub-tier MCP twin; code-reviewer and general-purpose always blocked web-ship-gate a client page ships only with a design-gate stamp guardrail-deny · pretooluse-block-pipe-to-shell destructive commands and curl-piped-to-shell never run gitleaks pre-commit (global core.hooksPath) every commit in every repo, not per-repo opt-in seats.py refuses to load a config that puts a metered lane on an automatic panel
Plus advisory injectors: brief-first, spec-contract-nudge, skill-trigger, delegation-table, retrieval-first, topic-switch.
DESIGN REFERENCE MCPS

Anchor a build to screens humans actually designed, instead of to model defaults. The house rule is that a gate measures the ABSENCE of known slop; it cannot supply a point of view, so the reference comes first.

mobbin real product screens, flows and sections refero screens, flows and styles; live again as of 2026-09-05 onepagelove single-page inspiration and section patterns. Project-scoped in ~/.claude.json, so not among the 15 user-scope servers stitch-design · figma design systems, screen generation, Code Connect magic (21st.dev) component MCP, currently UNAUTHENTICATED. The site itself browses fine with chrome-devtools
Skill that drives them: ui-component-anchors, cross-linked from anti-slop-client-build.
SKILLS TOUCHED, LAST 14 DAYS (18 of 142)

Where the work actually went: the client-facing build and quality lane.

anti-slop-client-build ui-component-anchors landing-page-fast overlook-hero hallmark impeccable taste-skill humanizer visual-review web-page blog-master brand-dna client-pitch-kit audit-report paid-ads local-search-gap seo-image-gen watch
House rule: when a skill misfires, the lesson is appended to a Gotchas section in its SKILL.md so it does not recur.