Tom Fuertes

Gauntlet AI Cohort 4

05 Feb 2026

Gauntlet AI Cohort 4 ships a graded project every week. I made it through week 9 and opted out of the final. The apps are offline now, but three of them stuck with me.

YesAInd: enforce limits in code

My first project, a synced multiplayer board built with the AI SDK. Prompts are suggestions. Haiku ignored "create ONLY 3 objects" until the server rejected anything over the cap, which took the eval pass rate from 30% to 97%.

Ghostfolio: sidecar over fork

Natural-language queries for Ghostfolio, a NestJS portfolio tracker. Forking someone else's monolith cost me days fixing things I didn't care about. A sidecar that only talks to the app's API was smaller, and swapping models was trivial because nothing reached into the host app.

LegacyLens: trust the evals, not the demo

RAG over the GnuCOBOL compiler source. The answers looked fine, and the first eval scored function recall at zero. Clicking around the app was never going to catch that. Running evals in Langfuse and fixing the worst failure each round did.

-Tom