Articles by Grizzly Peak Software

Jev vs. Kodiak: I Made Two "System One" Decision Models Play Zork
Builder's Journal
Jev vs. Kodiak: I Made Two "System One" Decision Models Play Zork

I made Kodiak, my open-weights decision model, and Jev, TypeSafe AI's closed one, play Zork by choosing from the game's valid commands: same harness, same seeds. In about 20,000 moves, not one was invalid. Jev played better, then walked into a grue while its own danger answer said "yes" at 0.95. Here's what that, a no-model baseline and a 35x question-wording effect taught me about System 1 / System 2 cascades.

Read Article
I Built an Open-Source "System One" Decision Model on My DGX Spark in Three Days
Builder's Journal
I Built an Open-Source "System One" Decision Model on My DGX Spark in Three Days

Kodiak is an open-weights decision model: give it a document and typed questions, and it answers all of them in one pass, in 8 ms, with calibrated confidence and an honest "I can't tell." I built it on my DGX Spark with Claude Code in three days. Here's how it works, what it costs, a 3 a.m. plot twist, and where it stands against bigger models.

Read Article
The AI-Native If Statement: Guarding Agent Tool Calls with Jev
AI Agents
The AI-Native If Statement: Guarding Agent Tool Calls with Jev

A weekend building a guard for agent tool calls with Jev: rules set the floor, calibrated confidence handles judgement, and the Claude Code hook that ended up guarding the agent that built it.

Read Article
Jev Was Trained to Be Honest, Not Helpful: What That Means for Engineers
AI Assisted Development
Jev Was Trained to Be Honest, Not Helpful: What That Means for Engineers

Jev, TypeSafe's new decision model, is trained to be calibrated, not to please. Here's what that means, what it's good for, and how to use it from Claude Code without building an agent framework.

Read Article
Load Testing a Next.js BFF: Why We Moved from Postman to k6
Devops
Load Testing a Next.js BFF: Why We Moved from Postman to k6

I started load testing our Next.js backend-for-frontend in Postman. I finished in k6. Here's why I switched mid-project, what Claude did for me along the way, and the one refactor that turned a pile of response times into numbers leadership actually asks for.

Read Article
I Ran a Fruit Fly's Entire Brain on My DGX Spark, Then Gave It a Body
Builder's Journal
I Ran a Fruit Fly's Entire Brain on My DGX Spark, Then Gave It a Body

I got Eon Systems' whole-brain fruit fly emulation running on an NVIDIA DGX Spark: 138,639 neurons wired from a real fly's connectome. It reproduced the paper's sugar-to-feeding result, needed a CUDA 13 fix to build the faithful GPU path, and ended up driving a simulated body that walks and turns on its own descending-neuron signals.

Read Article
How to Risk-Check a Book Before Publishing on KDP: The 45-Minute Pre-Publish Audit (2026)
Self-Publishing
How to Risk-Check a Book Before Publishing on KDP: The 45-Minute Pre-Publish Audit (2026)

Every KDP problem is cheapest to fix before you hit publish — and most account disasters were preventable in minutes. The pre-publish audit a 200+ title operation runs on every book: the instant-stop gate, the six risk domains, and why risk is two numbers, not one.

Read Article
KDP Account Suspended or Under Review: The First 48 Hours (2026)
Self-Publishing
KDP Account Suspended or Under Review: The First 48 Hours (2026)

Written from a real enforcement file that includes two KDP account terminations — both reversed. How to triage which emergency you're in, what to do (and not do) in the first 48 hours, what appeals can and cannot achieve, and the post-restoration audit nobody tells you about.

Read Article
KDP Duplicate Content and Similarity Flags: Why Original Books Get Blocked (2026)
Self-Publishing
KDP Duplicate Content and Similarity Flags: Why Original Books Get Blocked (2026)

You wrote it yourself and Amazon still flagged it as "similar to other books." Similarity enforcement isn't a plagiarism check — here are the six ways original books trip it, the keyword-field mistake that has terminated accounts, and the five-minute pre-publish checks that prevent all of it.

Read Article
Amazon KDP Content Quality Notice: What It Actually Means and What to Do (2026)
Self-Publishing
Amazon KDP Content Quality Notice: What It Actually Means and What to Do (2026)

The "content quality" email is short and generic on purpose. Here's how to read it, which of the four enforcement modes you're actually in, what to do in the first 24 hours, and the response mistakes that turn a fixable flag into a terminated account.

Read Article
Postgres Until It Hurts: The Exact Point You Should Switch
Software Engineering
Postgres Until It Hurts: The Exact Point You Should Switch

This article reveals the concrete operational thresholds where Postgres becomes cost-prohibitive, backed by real query diagnostics and migration case studies that prove when to act.

Read Article
Feature Flags: The Hidden Technical Debt You Can't Afford
Modern Development Tools
Feature Flags: The Hidden Technical Debt You Can't Afford

I'll show why feature flags accumulate debt when not actively managed, then give the exact metrics and process to stop ignoring them. You'll build a sustainable flag strategy that prevents codebase bloat.

Read Article
Staging Environments Lie: The Hidden Cost of False Confidence
Modern Development Tools
Staging Environments Lie: The Hidden Cost of False Confidence

You'll learn why staging environments deceive your team about production readiness, and gain practical validation techniques to catch failures before deployment. By the end, you'll have a checklist to replace trust with proof.

Read Article
Monitoring Metrics That Lie: Why Your Alerts Don't Prevent Outages
Cloud & Infrastructure
Monitoring Metrics That Lie: Why Your Alerts Don't Prevent Outages

Stop chasing uptime metrics that mask real failures. This article reveals three user-impact metrics that prevent outages, with code to instrument them without rewriting your stack.

Read Article
What a Weekend of Benchmarking Taught Me About My DGX Spark
AI Agents
What a Weekend of Benchmarking Taught Me About My DGX Spark

I benchmarked four local models on an NVIDIA DGX Spark across three context depths and a purpose-built agent reliability harness. The fastest model wasn't the best one, and the model I'd been running for weeks was quietly fabricating data.

Read Article
Retry Logic Amplifies Outages: The Silent Killer You're Ignoring
Cloud & Infrastructure
Retry Logic Amplifies Outages: The Silent Killer You're Ignoring

This article exposes how standard retry strategies exacerbate system crashes, not prevent them. You'll learn to replace naive retries with circuit-breaking resilience patterns that keep your system running during chaos.

Read Article
On-call Rot: How 24/7 Duty Stalls Your Engineering Career
Career
On-call Rot: How 24/7 Duty Stalls Your Engineering Career

On-call cycles aren't just tedious: they actively devalue engineering skills. This article exposes how they sabotage promotion paths while offering sustainable alternatives.

Read Article
Microservices for Two: The Hidden Cost of Overengineering
Modern Development Tools
Microservices for Two: The Hidden Cost of Overengineering

This article debunks microservices for small teams with real migration pain points and concrete alternative patterns. You'll walk away with a single-process architecture blueprint that scales better than microservices.

Read Article
CI Cache Lies in Production: Why Your Builds Are Flaky
Modern Development Tools
CI Cache Lies in Production: Why Your Builds Are Flaky

Exposes how CI caches lie about artifact integrity, not just cache misses. Teaches concrete validation patterns to prevent real-world build failures before they break production.

Read Article
LLM Self-Hosting Costs: When You Actually Save Money
AI Integration & Development
LLM Self-Hosting Costs: When You Actually Save Money

Most engineers overestimate savings until they run the numbers. The real threshold is when your inference volume exceeds 150 requests/day.

Read Article