Articles by Grizzly Peak Software
Jev vs. Kodiak: I Made Two "System One" Decision Models Play Zork
I made Kodiak, my open-weights decision model, and Jev, TypeSafe AI's closed one, play Zork by choosing from the game's valid commands: same harness, same seeds. In about 20,000 moves, not one was invalid. Jev played better, then walked into a grue while its own danger answer said "yes" at 0.95. Here's what that, a no-model baseline and a 35x question-wording effect taught me about System 1 / System 2 cascades.
Read ArticleI Built an Open-Source "System One" Decision Model on My DGX Spark in Three Days
Kodiak is an open-weights decision model: give it a document and typed questions, and it answers all of them in one pass, in 8 ms, with calibrated confidence and an honest "I can't tell." I built it on my DGX Spark with Claude Code in three days. Here's how it works, what it costs, a 3 a.m. plot twist, and where it stands against bigger models.
Read ArticleThe AI-Native If Statement: Guarding Agent Tool Calls with Jev
A weekend building a guard for agent tool calls with Jev: rules set the floor, calibrated confidence handles judgement, and the Claude Code hook that ended up guarding the agent that built it.
Read ArticleJev Was Trained to Be Honest, Not Helpful: What That Means for Engineers
Jev, TypeSafe's new decision model, is trained to be calibrated, not to please. Here's what that means, what it's good for, and how to use it from Claude Code without building an agent framework.
Read ArticleLoad Testing a Next.js BFF: Why We Moved from Postman to k6
I started load testing our Next.js backend-for-frontend in Postman. I finished in k6. Here's why I switched mid-project, what Claude did for me along the way, and the one refactor that turned a pile of response times into numbers leadership actually asks for.
Read ArticleI Ran a Fruit Fly's Entire Brain on My DGX Spark, Then Gave It a Body
I got Eon Systems' whole-brain fruit fly emulation running on an NVIDIA DGX Spark: 138,639 neurons wired from a real fly's connectome. It reproduced the paper's sugar-to-feeding result, needed a CUDA 13 fix to build the faithful GPU path, and ended up driving a simulated body that walks and turns on its own descending-neuron signals.
Read ArticleHow to Risk-Check a Book Before Publishing on KDP: The 45-Minute Pre-Publish Audit (2026)
Every KDP problem is cheapest to fix before you hit publish — and most account disasters were preventable in minutes. The pre-publish audit a 200+ title operation runs on every book: the instant-stop gate, the six risk domains, and why risk is two numbers, not one.
Read ArticleKDP Account Suspended or Under Review: The First 48 Hours (2026)
Written from a real enforcement file that includes two KDP account terminations — both reversed. How to triage which emergency you're in, what to do (and not do) in the first 48 hours, what appeals can and cannot achieve, and the post-restoration audit nobody tells you about.
Read ArticleKDP Duplicate Content and Similarity Flags: Why Original Books Get Blocked (2026)
You wrote it yourself and Amazon still flagged it as "similar to other books." Similarity enforcement isn't a plagiarism check — here are the six ways original books trip it, the keyword-field mistake that has terminated accounts, and the five-minute pre-publish checks that prevent all of it.
Read ArticleAmazon KDP Content Quality Notice: What It Actually Means and What to Do (2026)
The "content quality" email is short and generic on purpose. Here's how to read it, which of the four enforcement modes you're actually in, what to do in the first 24 hours, and the response mistakes that turn a fixable flag into a terminated account.
Read ArticlePostgres Until It Hurts: The Exact Point You Should Switch
This article reveals the concrete operational thresholds where Postgres becomes cost-prohibitive, backed by real query diagnostics and migration case studies that prove when to act.
Read ArticleFeature Flags: The Hidden Technical Debt You Can't Afford
I'll show why feature flags accumulate debt when not actively managed, then give the exact metrics and process to stop ignoring them. You'll build a sustainable flag strategy that prevents codebase bloat.
Read ArticleStaging Environments Lie: The Hidden Cost of False Confidence
You'll learn why staging environments deceive your team about production readiness, and gain practical validation techniques to catch failures before deployment. By the end, you'll have a checklist to replace trust with proof.
Read ArticleMonitoring Metrics That Lie: Why Your Alerts Don't Prevent Outages
Stop chasing uptime metrics that mask real failures. This article reveals three user-impact metrics that prevent outages, with code to instrument them without rewriting your stack.
Read ArticleWhat a Weekend of Benchmarking Taught Me About My DGX Spark
I benchmarked four local models on an NVIDIA DGX Spark across three context depths and a purpose-built agent reliability harness. The fastest model wasn't the best one, and the model I'd been running for weeks was quietly fabricating data.
Read ArticleRetry Logic Amplifies Outages: The Silent Killer You're Ignoring
This article exposes how standard retry strategies exacerbate system crashes, not prevent them. You'll learn to replace naive retries with circuit-breaking resilience patterns that keep your system running during chaos.
Read ArticleOn-call Rot: How 24/7 Duty Stalls Your Engineering Career
On-call cycles aren't just tedious: they actively devalue engineering skills. This article exposes how they sabotage promotion paths while offering sustainable alternatives.
Read ArticleMicroservices for Two: The Hidden Cost of Overengineering
This article debunks microservices for small teams with real migration pain points and concrete alternative patterns. You'll walk away with a single-process architecture blueprint that scales better than microservices.
Read ArticleCI Cache Lies in Production: Why Your Builds Are Flaky
Exposes how CI caches lie about artifact integrity, not just cache misses. Teaches concrete validation patterns to prevent real-world build failures before they break production.
Read ArticleLLM Self-Hosting Costs: When You Actually Save Money
Most engineers overestimate savings until they run the numbers. The real threshold is when your inference volume exceeds 150 requests/day.
Read Article