Experience engineering, AI-native product and creative code. What runs across the top is the work, not the set dressing.
Most products know what they do. Few know how they should feel to use.
The separations
Projects told from start to finish: the problem, the decisions that solved it and what went wrong along the way.
- STORYPRINTS — Eight answers from a parent turned into a physical object in two minutes
- VAMPMAKER — A SaaS that prepares the role-playing session before the session starts
- BRANDAI — Brand identity that is portable across AI tools
- RBT.STUDIO — When the product is you
- AIMPLAS — A 3D tablet app that accompanies a real helmet and scooter
- D-GO — A website for D-Go, a direct-drive motor for cargo bikes
- THE SMART LOLLIPOP — A landing page for The Smart Lollipop, a children’s health device
- EONIA — The missing layer of biological decision between measuring and intervening
- EL ÚLTIMO ZARPE — A roguelike without meta-progression, where only understanding accumulates
- PHOTOGRAPHY · ART DIRECTION — A visual grammar, not a style — from the medium-format negative to the printed book
- NOTJUSTCODE — Vision, design and market reviewed inside the editor, against the repository
- SDD HARNESS — A 5-phase wizard so agents build what I had in my head
Eight answers from a parent turned into a physical object in two minutes · SaaS · B2C · EdTech
- Client
- Side project — RBT Studio’s own product
- Role
- Product, design and full-stack development (a team of one)
- Timeline
- March 2026 – ongoing
- Stack
- Next.js 15, React 18, TypeScript, Supabase, Claude, OpenRouter, Upstash Redis + BullMQ, pdfkit, Lemon Squeezy
The gap sat between the made-to-order book and the chat window
A shop-bought colouring book is generic by definition: the hero is never your daughter. The personalised alternatives are print-on-demand books — expensive, with weeks of waiting and a single print run. At the other end, an AI chat can write you a story, but it returns text in a window: it is not a book, it has no cover, it doesn’t print, it can’t be coloured in, it doesn’t stay on the shelf.
The problem: Three constraints that shaped the technical design
The final artefact is an A4 PDF, not a screen: everything generated has to survive a home printer. The audience is children aged 3 to 12, so content nobody has vetted cannot reach a printed page. And every book costs real money in model calls — the product’s economics had to be in the code from day one, not bolted on afterwards.
How it was approached
- The wizard: eight questions, not a prompt. The first product decision was not to expose a free-text field. A parent doesn’t want to write a prompt; they want to answer questions with their daughter next to them. Name, age, setting, companion, value, identity, length and narrative style — each answer is a card. That validated StoryConfig is the only contract that enters the pipeline: nothing downstream asks the browser anything again.
- A queue, and a cron that rescues it. Generations take minutes and cost money, so they go into a queue (BullMQ on Upstash Redis) with progress streamed to the browser. It is drained two ways: in-process right after responding, and by a cron every 5 minutes. The cron isn’t decorative redundancy — it is the only thing that lets a stuck queue fix itself. The endpoint fails closed if there is no CRON_SECRET, because every call spends money on models.
- The illustrations cannot contain text. These are colouring pages: a letter baked into the bitmap can’t be removed, and the page prints as is. The image model is called via chat completion, with no negative prompt, so everything travels inside the text — the instruction is repeated in three different ways and quoted dialogue is stripped from the scene before the model sees it, because quotation marks are the strongest signal for a model to draw letters. It is prompt-level mitigation, not a guarantee, and the documentation says so.
- The economics live in the code. Every image is paid for in credits, and the plan buys a monthly allowance that expires when the period closes: an allowance that accumulates is not an allowance. For a while the subscription webhook wrote the plan and nothing else, so a subscriber had zero credits and got the same generic drawings as the free plan. Today grantPlanCredits() runs on creation, on update and on every monthly payment — idempotent by description, with a unique index in Postgres that actually enforces it.
- Two deliberate business decisions. The free plan delivers a complete, printable book: holding back the PDF would leave the free plan with nothing to judge the product by — what the subscription removes is the watermark. And every user’s first story gets real AI illustration, whatever their plan; otherwise the only book a visitor judges the product by is rendered with six generic SVGs, that is, with none of the personalisation they are being asked to pay for. A sign-up costs one image call: it is cheap advertising.
wizard → PDF: Wizard · 8 steps → Validated StoryConfig → BullMQ queue · Redis → Narrative Engine · Claude → Content Safety → Illustration Engine · OpenRouter → Supabase Storage → SVG reader → A4 PDF on demand
Solution: A complete SaaS, from landing page to PDF, built and run by one person
An eight-step wizard with an optional photo upload of the child, with explicit consent and never written to the story’s config column. A narrative engine on Claude with age-appropriate scenes and lengths from 6 to 18 pages derived from the number of scenes. An illustration engine on OpenRouter, black line on white and consistent across scenes. A content filter applied both to what the model returns and to the product’s only free-text field. An SVG web reader and PDF export with pdfkit. Four plans read from a single limits table, a custom share card, an admin panel and subscriptions with Lemon Squeezy. The PDF is not in the pipeline: it is composed on demand from the saved story, so a generation never blocks waiting on layout.
Results
The product is built and deployed, and not yet validated in the market. There are no user, conversion or revenue metrics to report, and I’d rather not make them up. What can be stated: the full loop works in production, from the eight wizard answers to the downloadable PDF, with live payments and subscriptions. The business invariants are locked by tests, not by discipline — the limits table and the pricing page cannot diverge because there is a test that fails if they do. And the queue self-heals: a generation whose in-process trigger is interrupted is picked up by the cron within the next 5 minutes. What remains to be validated is what matters: whether a parent will pay €6.99 a month for this.
- 376 — tests across 30 suites, all green
- 113 — commits in five months of one person’s work
- 4 — plans read from a single limits table
"A flag no route reads is not a limit: it is a comment that looks like one."
What was learned
- A flag no route reads is not a limit. The plans table existed long before the routes queried it. The credits bug — a subscriber getting the same drawings as the free plan — was not a code error: it was configuration nobody executed. Now every limit has a route that reads it and a test that proves it.
- What can’t be removed afterwards has to be prevented beforehand. A letter baked into a bitmap that will be printed has no later fix, so the mitigation has to live in the prompt and in the scene preprocessing — and be documented as mitigation, not as a guarantee.
- Taking the PDF out of the pipeline is what made the system robust. Composing on demand from the saved story means a layout failure never blocks a generation that has already been paid for.
See StoryPrints in production
Product in production, pending market validation.
A SaaS that prepares the role-playing session before the session starts · SaaS · Niche B2C · TRPG
- Client
- Side project — RBT Studio
- Role
- Full Stack / AI Engineer — product, architecture and development
- Timeline
- May 2026 – ongoing (Release 1.0)
- Stack
- TypeScript, NestJS 11, Next.js 16, React 19, Prisma, PostgreSQL · Neon, Zustand, Tailwind 4, OpenRouter, Pollinations.ai
Running a chronicle is closer to producing a series than to playing a game
The Storyteller arrives on Friday with a five-session arc to sustain, a city with its political hierarchy, a dozen NPCs that must sound different from each other and character sheets that match the edition being played. Almost all that work happens before anyone rolls a die. VampMaker was born to absorb that preparation, and the first version did it well. The problem was what happened next.
The problem: Nothing generated survived the browser
A code audit in August 2026 put the diagnosis in one sentence. The store wrote only to localStorage, and the full CRUD of the NestJS backend — twenty-odd endpoints, with their PostgreSQL schema, relations and cascade deletes — was never invoked: saveCampaign() and loadCampaigns() were written, tested against the API and had not a single call from the interface. The result was a product that presented itself as a SaaS — with an account, tiers, a “Free Plan” in the sidebar — and lost the user’s campaign when they switched devices. The real brief wasn’t adding AI: it was turning a working demo into a billable product.
How it was approached
- The spec before the code. With a finding of that calibre the temptation is to open the editor and start wiring things up. I wrote the specification first: four SDD documents — SPEC, ARCHITECTURE, SCAFFOLD, AGENTS — that fix goals, non-goals and verifiable acceptance criteria before touching a line. Defining the non-goals turned out to be as useful as the goals: payment gateway, multi-user collaboration, native app and PDF export were left out of Release 1.0 in writing, and that closed the door on three months of scope drift.
- Eight agents with their dependency graph. On top of that spec I set up a harness of eight agents. The project is developed by one person, so the split isn’t about human parallelism: it is about every work session — my own or with an LLM — having a closed scope and a checkable definition of done.
- Six decisions documented as ADRs. The server becomes the source of truth and localStorage drops from store to read-through cache. String identifiers (cuid) end to end, because the frontend assumed number and as soon as Prisma emits cuids the user sees “Campaign not found” over what is really a typing problem — this ADR blocks all the others. Whatever doesn’t fit the schema goes into Json columns, not new tables. Synchronous write-through persistence per operation: one extra request is irrelevant next to the 10–30 seconds of a generation, and in exchange a failure is attributable to a specific action. A strict API contract with a single error envelope. And quota as a Guard that reserves beforehand and an Interceptor that confirms afterwards, so a failed generation doesn’t consume balance.
- The AI layer, treated as engineering rather than conversation. Connecting OpenRouter is twenty lines. The hard part was getting the model to always return something renderable, in the right language and without contradicting the canon of the chosen edition. Language: some of the English leaking in didn’t come from the model but from a template in the code itself, and the explicit rule now separates keys (untouched, consumed by the frontend) from values (in Spanish, with the game terminology translated). Schema: “return ONLY valid JSON” is not a specification, so the prompt now carries the explicit schema with its branching between editions. Layered canon: a universal core, a per-edition layer and the curses only of the clans that appear in the request — before, all fifteen were injected and “the Masquerade is absolute law” was asserted even in a Dark Ages campaign, where it isn’t declared until 1666.
8-agent harness: CostProbe → ContractMigrator → PersistenceBuilder → ResilienceBuilder → ExperienceBuilder → QuotaBuilder → DocWriter
Solution: A monorepo with three workspaces and a contract that breaks the build, not production
backend, frontend and shared, where the shared package holds the domain types both ends consume. A NestJS 11 backend with per-domain modules, Prisma on PostgreSQL in Neon, JWT in an httpOnly cookie and a global exception filter. A Next.js 16 frontend with React 19, App Router, Zustand and a tabbed campaign view: arc, sessions, NPCs, sheets and maps. A generation engine with six endpoints on deepseek-chat via OpenRouter for text and Pollinations.ai for images — the image prompt is deliberately kept in English, because diffusion models perform worse in Spanish: what gets translated is what the user sees, not what the generator consumes. Four supported editions (V5, V20, Dark Ages and Sabbat), each with its lore, factions and city offices.
Results
There are no usage metrics to show: the product is in its Release 1.0 phase and server-authoritative persistence is in progress. What there is, is accumulated judgement, which is what really travels from one project to the next. Release 1.0 closes server-authoritative persistence, non-silent errors, a single navigation path, Markdown export to bring the material to the table and a per-account quota limit.
- 6 ADRs — structural decisions, each with its discarded alternative
- ~190 — tokens saved per generation by no longer repeating the language rule six times
- 4 — supported editions, each with its lore and offices
"Before touching the prompt, find out who writes the text. I spent a while convinced the English came from the model. Part of it came from a template in the code itself."
What was learned
- A prompt without a schema is an API without a contract. “Return valid JSON” is not a specification. If the interface expects fifteen specific keys, those fifteen keys go in the prompt — and the ones the code consumes are not translated even if the rest of the content is.
- Context that contradicts the domain costs more than missing context. Injecting Camarilla rules into a medieval campaign isn’t just wasting tokens: it is actively asking the model to get it wrong. Splitting the canon by edition improved the output and made it cheaper.
- A hook with useState is not shared state. Every component that called useAuth had its own user and its own request. Since the sidebar lives in the root layout and doesn’t remount on navigation, login really worked — cookie issued, valid session — but the app kept rendering as if it hadn’t. A state architecture error disguised as an authentication bug.
- Writing the spec first turned “fix the app” into eight verifiable units. It’s the difference between a refactor with no known end and eight blocks with a definition of done.
Repository and demo yet to be published.
Brand identity that is portable across AI tools · SaaS · B2C · Freelancers & Creators
- Client
- Side project — RBT Studio
- Role
- Product design · Full-stack · Multi-agent architecture · Prompt engineering
- Timeline
- Own product
- Stack
- Next.js, Supabase, OpenRouter, Upstash Redis, Multi-agent, Markdown
The visual consistency problem in small projects isn’t about talent — it’s about memory. Every time you use an LLM to generate something for your brand, you start from scratch. BrandAI orchestrates five specialised agents to generate a complete design.md: palette, typography, verbal tone, spacing, tokens. A Markdown file that travels with you to any LLM, Cursor or pipeline.
The problem: LLMs don’t remember who you are
A complete design system in Figma or Notion is out of reach for small projects in time and money. But without brand context, every AI generation produces something inconsistent with the last. The result is a visual identity that fragments with every tool you use.
orchestration pipeline — 5 agents: Natural-language input → Orchestrator → Palette agent → Typography agent → Verbal tone agent → Redis · hash cache → Consistency validator → design.md assembly → Supabase · v+1
Solution: A portable artefact generated by specialised agents
Five agents in a pipeline — orchestrator, palette, typography, verbal tone, consistency validator — generate a complete design.md in 60–90 seconds. The validator detects inconsistencies between layers and re-runs the affected agents. The result travels to the IDE, to the chat, to the image generator’s prompt.
Product decisions
- my role: Product design · Full-stack · Multi-agent architecture · Prompt engineering
- users: Wellness founders — no design budget
Freelancers — consistency across clients
Content creators — multiple brands of their own - key product decision: Markdown as the output format. Not Figma, not Notion, not JSON. The most portable file in the digital ecosystem — it works in any editor, any LLM, any CI/CD.
- key AI decision: A different model per agent: haiku for palette and typography, sonnet for orchestrator and validator, gpt-4o-mini for verbal tone. ~30% cost reduction with a Redis cache keyed by semantic hash.
Results
- 60–90s — to generate a complete design.md
- ~30% — cost reduction via Redis cache
- 40% — return rate for evolution mode within 30 days
"The most important decision in the product wasn’t technical — it was about format. Markdown is the most portable artefact in the digital ecosystem."
What was learned
- The consistency validator is the most expensive agent and the most valuable. Without it, the orchestrator assembles sections that contradict each other — a cold palette with a “warm and earthy” tone. With it, average satisfaction rises by 0.8 points out of 5.
- Evolution mode retains more than generation mode. With design.md versioned in Supabase, the cost of switching tools is high. Every brand update that comes back to BrandAI is a natural retention loop.
- Different models per agent via OpenRouter was decisive. GPT-4o-mini outperformed Claude Haiku on verbal tone for English-language wellness brands. A one-line config change, no refactor.
When the product is you · Portfolio · Digital identity · Creative Dev
- Client
- Own product — rbt-studio.com
- Role
- Design, development and art direction
- Timeline
- Ongoing
- Stack
- React, Vite, Canvas API, GSAP, Generative art, Tailwind
A portfolio is not a CV with CSS. It is the most complete argument you can make about who you are as a professional. rbt.studio was designed from a visual arts foundation — a Fine Arts background that informs every technical decision. A particle field in the hero, a generative Lissajous canvas, a dark palette with gold and sage accents. The code is the design.
The problem: Positioning at the intersection of AI engineer and creative dev
A standard developer portfolio doesn’t communicate the creative dimension. A designer portfolio doesn’t communicate technical depth. Ricard’s profile lives exactly at that intersection — and the portfolio has to prove it without explaining it. If you have to explain that you’re creative, you’ve already lost.
layers of the visual system: Identity — dark / teal / cream → DM Sans + JetBrains Mono → Hero particle field — Canvas 2D → Lissajous canvas — parametric → Work cases — editorial grid → Stack (AI · Frontend · Systems)
Solution: The code itself as a portfolio piece
The generative canvas isn’t decoration — it’s the first proof that the author understands complex visual systems. The hero’s particle field responds to the cursor with attraction physics. The Lissajous canvas draws Bowditch figures in real time. A technical recruiter opens it in DevTools and sees the code. A design recruiter sees the result.
Product decisions
- my role: Identity design · Frontend · Generative art · Motion
- context: 10+ years of full-stack experience. A Fine Arts background (University of Barcelona). The portfolio has to show both worlds at once — technical and creative.
- key aesthetic decision: A dark/cream palette with teal as the primary accent. No generic brand colours. DM Sans for display — a sans that crosses the editorial with the technical.
- technical differentiator: A generative Lissajous canvas — a parametric figure that varies in real time. A particle field in the hero with attraction physics. All rendered in Canvas 2D with no external libraries.
Results
- 10+ — years of experience condensed into a single page
- 0 — external libraries for the generative effects
- 2 — profiles in one — AI engineer and creative technologist
"Technical depth from 10+ years of full stack engineering. Creative precision from a fine arts background. The intersection is where I work."
What was learned
- The portfolio is the argument, not the summary. It doesn’t list projects — it demonstrates a point of view. The particle field isn’t decoration, it’s the first sentence of the argument: this developer understands real-time visual systems.
- Canvas 2D with no external dependencies is a deliberate decision. Three.js would have been faster. Choosing to implement the particle physics and Lissajous curves from scratch demonstrates a mathematical understanding of the system.
- The dark/cream palette communicates before a word is read. Developer portfolios tend to be white with blue. The contrast of this visual system signals from the first second that the author has an aesthetic judgement of their own.
A 3D tablet app that accompanies a real helmet and scooter · Mobile app · Interactive 3D · Industrial R&D
- Client
- Aimplas — Plastics Technology Centre
- Role
- Mobile app development and 3D integration
- Timeline
- Client project
- Stack
- React Native, Expo.dev, Three.js, Headless CMS, API REST
Aimplas developed, together with Stimulo Design Studio, a physical helmet and scooter prototype that brings together several materials innovations. The problem with a prototype like that is that its advances can’t be seen: they are inside the material. The tablet app is the layer that makes them visible — a companion used in front of the real object, at trade fairs and demonstrations, to unfold what the eye can’t reach.
The problem: The innovation was inside the material, not in plain sight
A next-generation helmet and scooter look, at first glance, like a helmet and a scooter. The innovations are in composition, process and structure. They had to be explained to very different audiences — from a casual visitor to a materials engineer — without turning the app into a technical manual or an empty brochure.
content → tablet pipeline: Headless CMS → REST API → Expo / React Native app → Profile-based navigation → Interactive 3D model → Detail grid → Innovations list
Solution: Three reading formats over the same content
Navigation adapts to the type of user and offers the same content at three depths: interactive 3D elements to explore the object, a grid of detail images for a quick read, and a complete list of innovations for those who want the data. Each profile enters where it suits them and goes down to the level they want.
Product decisions
- my role: UX/UI · Frontend development · API integration with a headless CMS
- context: A project for Aimplas in collaboration with Stimulo Design Studio. The app complements a real physical helmet and scooter prototype.
- users: Trade-fair visitors — a quick, visual tour
Technical profiles — the full list of innovations
Sales team — a guided demo on the physical object - key technical decision: All content lives in an external CMS consumed via API. Updating texts, images or innovations doesn’t require rebuilding or reinstalling the app on the tablets.
Results
- 3 — content formats over the same information
- 0 — reinstalls needed to update content
- 1 — app accompanying the physical prototype in demonstrations
"The app doesn’t replace the prototype — it accompanies it. The object is on the table; the tablet shows what’s inside the material."
What was learned
- Decoupled content is a requirement, not a convenience. A trade-fair app gets updated the day before the event. With a headless CMS via API, the change is a content change; with hardcoded content, it’s a build-and-deploy cycle on every tablet.
- 3D is the way in, not the destination. The interactive model attracts and orients, but whoever is after the technical data needs a list. Sustaining both formats on the same content source avoids duplicating the editorial work.
- Expo reduced distribution friction across a fleet of tablets. Updating the app on several demo devices is a logistical task — not a technical one — and it’s best solved from the choice of stack.
A website for D-Go, a direct-drive motor for cargo bikes · Product website · 3D · E-mobility
- Client
- D-Go
- Role
- Web development and 3D technical direction
- Timeline
- Client project
- Stack
- Frontend development, 3D graphics, Scroll-driven animation, UX/UI design
D-Go is a premium electric motor for cargo bikes. Its value lies in the engineering: torque, integration, direct drive. Everything that makes it good is hard to tell. The website was developed with the Stimulo Design Studio marketing team to solve exactly that — presenting complex technical information clearly and visually, without diluting it.
The problem: A spec sheet doesn’t sell engineering
Presenting complex technical data clearly and visually is the central challenge of the project. A specifications table is precise and communicates nothing; sales copy communicates and loses the precision that justifies the price of a premium product.
scroll → technical narrative: 3D motor assets → Technical script by section → User scroll → Narrative sync → 3D render per segment → Technical data layers → Product CTA
Solution: 3D graphics synced with the scroll
3D graphics were created and several narratives were synced with the user’s scroll. By handing control of the tempo to the visitor, the technical content stops being an imposition and becomes exploration — with notably more positive engagement with information that, in a static format, gets abandoned.
Product decisions
- my role: UX/UI · Frontend development · End-to-end site development
- context: A project developed in collaboration with the Stimulo Design Studio marketing team. Public product site — d-go.eu
- key design decision: Syncing several narratives with the user’s scroll. The visitor controls the tempo of the technical story: they move on once they’ve understood, and stop when they want to look.
- communication challenge: A direct-drive motor is explained with unintuitive concepts. 3D graphics make it possible to show the mechanism instead of describing it.
Results
- 3D — custom graphics to explain the mechanism
- 1:1 — scroll and narrative in sync — the user sets the tempo
- d-go.eu — public product site in production
"When the user controls the tempo of the story, technical information stops being an obstacle and becomes the reason they stay."
What was learned
- Handing over control of the tempo changes the relationship with dense content. An autoplay video imposes a rhythm; scroll negotiates it. In a technical product, that difference decides whether the visitor reaches the end of the page.
- 3D shows the mechanism; text only describes it. A direct-drive motor is understood by seeing it work. The animation isn’t ornament — it is the explanation.
- Working with the marketing team from the start avoids rework. The narrative script and the technical implementation were designed together; the scroll structure is the structure of the message.
A landing page for The Smart Lollipop, a children’s health device · Product landing · Startup · 3D + GSAP
- Client
- The Smart Lollipop
- Role
- Web development · 3D · GSAP
- Timeline
- Client project
- Stack
- Frontend development, 3D graphics, GSAP, Custom templates, UX/UI
The Smart Lollipop is a startup with a product unlike anything before it — and therefore with no inherited visual language to draw on. The project was developed in close collaboration with Stimulo Design Agency, combining creativity and technology: helping a startup define how to present its product and communicate its value is always a stimulating challenge.
The problem: A startup without a language of its own yet
Before building the website something earlier had to be solved: how this product presents what it does and why it matters. With no direct category references, every visual and narrative decision defines the positioning — and a startup can’t afford a website that ages in six months.
scroll → emotional narrative: Defining the message → Script by section → Animated 3D graphics → GSAP · scroll sync → Custom templates → Content management
Solution: 3D narrative with the tempo in the user’s hands
Animated 3D graphics were integrated to create narratives synced with the browser scroll: the user sets the pace of the story and builds empathy with the product. In parallel, a system of custom templates lets the team update content easily without breaking the consistency of the design.
Product decisions
- my role: Communication · UX/UI · Frontend development · 3D and GSAP animation
- context: A project in collaboration with Stimulo Design Agency for an early-stage startup. Public site — thesmartlollipop.com
- key design decision: 3D narratives synced with the browser scroll. The user controls the pace of the story, which makes it easier to connect emotionally with the product.
- key product decision: Custom templates to manage content flexibly without compromising the visual and structural integrity of the design.
Results
- 3D — animated graphics synced with the scroll
- ∞ — content updates without touching the design
- 1 — visual language defined from scratch with the startup
"Helping a startup define how it presents its product and communicates its value is always a stimulating challenge."
What was learned
- In a startup, the website is designed after the message. Starting with the interface without having defined what is being communicated produces a pretty site that positions nothing.
- Custom templates are the balance between freedom and consistency. A completely free editor degrades the design within weeks; a locked one forces a call to the developer for every change. Templates confine freedom to the space where it breaks nothing.
- Empathy with the product is built with rhythm, not adjectives. Syncing the animation with the scroll lets the user discover at their own speed — and what is discovered convinces more than what is asserted.
The missing layer of biological decision between measuring and intervening · Own product · Health · Local-first + AI
- Client
- RBT Studio’s own product
- Role
- Product, system design and full-stack + AI development — sole author
- Timeline
- June 2026 → ongoing (MVP v1.0 + server increment v1.5)
- Stack
- TypeScript, React Native · Expo, Skia + Reanimated, NestJS, PostgreSQL · pgvector, LangGraph + Claude, Render · EAS
The high-performance user has more biological data than any previous generation — and still takes the same supplements every morning, however they are doing that day. The market solved measurement (Oura, WHOOP) and intervention (evidence-based supplementation), and left the middle empty: the decision. A system that recommends shifts the cognitive load onto the user; one that decides absorbs it. EONIA turns a 30-second check-in into a calculated biological state and a capsule architecture adapted to that state, on the device.
The problem: The gap is in the decision, not the data
Measurement and intervention are well solved by the market; between them there is nothing. The difference isn’t incremental but structural: the EONIA user doesn’t need to understand the HPA axis to benefit from it. Every screen answers a single question — what do I need to do right now — with no charts and no jargon up front.
How it was approached
- Specify before programming, and decide in writing. The project started with a complete document set (SPEC · ARCHITECTURE · SCAFFOLD · DESIGN · HANDOFF) before the first line of product code, and every structural decision was recorded as an ADR with its discarded alternative. It isn’t bureaucracy: in a product that touches health, traceability of why the system decides what it decides is part of the product.
- The decision engine is pure and deterministic. @eonia/engine has no I/O, no React, no network, no language model. A check-in goes in, a state and an architecture come out. It is the only place where decisions are made. A deliberate consequence: the same package runs on the device and on the server and produces exactly the same verdict — which means offline mode isn’t a degraded version, but the same system.
- The Orb is an entity, not a chart. The central visualisation (Skia + Reanimated, a sumi-e brush stroke) has to communicate the biological state pre-cognitively: in under a second, without reading or interpreting. A donut chart would have been cheaper and would have failed the only requirement that mattered. The three layers of information — what to do, why, and the physiological mechanism — are an architectural boundary enforced in the code, not a design convention.
- Where AI comes in. The narrative layer generates the personalised report with LangGraph and Claude, but under a hard constraint: the engine decides, the LLM only explains. The preparation node is pure code — normalisation, trend arithmetic, label resolution — and the model receives a grounded context it can’t step outside of. Three specialists analyse in parallel, a node synthesises and a safety gate decides between delivering or escalating. A hallucination in the normalisation node would have propagated to every branch; that’s why that node isn’t a model call.
- The clinical catalogue had to come out of the code. The formulas lived as constants in the engine, which meant updating a protocol was a deployment. They were moved to the database and a back-office console was built with three non-hierarchical roles: admin, specialist and provider. Editing the catalogue invalidates the clinical signature and bumps the version, and the report cache key includes that version. It is the only way a pharmacist — not a commit — can sign off a protocol.
- An incident that deserves to be in the case study. With the beta already in testers’ hands, every user got locked out after verifying their email. The cause: force_organization_selection enabled in the identity provider. Every session received a “choose organisation” task that a B2C user can never satisfy, the session stayed pending, the client counted it anyway and single-session mode rejected any new login. It was fixed on two fronts: the configuration and the client, because a production app can’t depend on a third party’s console being configured correctly.
check-in → protocol: Check-in · 5 dimensions → Individual baseline → @eonia/engine · pure → 6 biological states → 3-day hysteresis → 6 architectures · circadian → LangGraph · grounded narrative → Safety gate → Optional sync · Postgres
Solution: A closed loop of signal, decision and execution
A 5-dimension check-in in ~30 seconds → EONIA Score against your personal baseline → one of 6 biological states → one of 6 capsule architectures with circadian windows per compound. The architecture only changes after 3 consecutive coherent days: deliberate hysteresis against noise-driven oscillation. The clinical catalogue lives in the database with professional sign-off and versioning, not in the code.
Product decisions
- my role: Product, system design and full-stack + AI development — sole author
- key architecture: A pure, deterministic engine — no I/O, no network, no language model. The same package runs on device and server and produces the same verdict: offline mode is not a degraded version.
- key AI decision: The engine decides, the LLM only explains. The narrative layer receives a grounded context it can’t step outside of — it doesn’t derive a state, an architecture or a number of its own.
- status: MVP v1.0 deployed (Render + APK via EAS) and server increment v1.5 in production. Closed beta.
Results
No user metrics yet — the beta is closed and no figures have been published. What is verifiable: a complete, deployed MVP v1.0, with a dockerised backend on Render via Blueprint and an Android APK distributed to testers via EAS. Server increment v1.5 in production, with optional authenticated sync, narrative reports and a back-office console with role-based access control. And an editable, signed clinical catalogue: protocols are no longer code, and the clinical signature is a state of the data that only a professional can set.
- 35 — tests on the decision core, with multi-day integration
- 5 — workspaces with clear boundaries: engine, mobile, backend, admin, narrative
- 46 — commits since kickoff, with the SDD set kept in sync with the code
"Putting the LLM in the wrong place is easy and expensive. Restricting it to explaining a decision already made by verifiable code is what makes the system auditable — and what lets offline mode lose nothing essential."
What was learned
- The ADR that delivered the most value wasn’t technical. Reconciling an inherited multitenant backend with a B2C product didn’t solve a code problem: it settled whether the product was B2C or B2B2C, and stopped the backend from dragging the product along.
- Whatever a professional must be able to change can’t live in the code. Moving the clinical catalogue to a signed, versioned database turned an app into an operable system — a pharmacist signs off a protocol, not a commit.
- A third party’s configuration is production surface. One misset flag in the identity provider locked out 100% of testers after email verification. It was fixed in the configuration and in the client: a production app can’t depend on a third party’s console being right.
Demo and beta on request.
A roguelike without meta-progression, where only understanding accumulates · Own video game · Preparation roguelike
- Client
- Side project — my own game
- Role
- Game design, development and art direction
- Timeline
- July 2026 – ongoing
- Stack
- Godot 4.7, GDScript, JSON as the rules layer, HTML/SVG prototype as oracle, Art direction
The modern roguelike teaches you to memorise builds: you die, you earn permanent coins, you come back stronger. The player doesn’t solve the difficulty; unlocks erode it. El Último Zarpe (The Last Sail) does the opposite — no persistent coins, no upgrade tree. A port called Grisal, numbered days, a ship that sails with or without you. You prepare the voyage with limited money, contradictory information and a hold that can’t fit everything. You usually die. What you learn by dying is all you take with you.
The problem: A game that kills constantly has a very narrow margin for error
If a death feels arbitrary, the player doesn’t learn: they leave. On top of that, an uncomfortable measurement after the port revealed that the supposedly replayable game had only 6 possible puzzles in long mode — memorisable in an afternoon. Three difficulty parameters did absolutely nothing.
How it was approached
- The real problem wasn’t the loop, it was unfairness. A game that kills the player constantly has a very narrow margin for error: if a death feels arbitrary, the player doesn’t learn, they leave. The first design decision wasn’t about content but about a contract — never lie to the player about the rules; do lie about the facts. And “fair” here isn’t a statement of intent: every generated voyage is proven solvable before you are allowed to start. A verifier proves by brute force that at least one combination of purchases survives, with an economic margin. If it can’t find one, that voyage never reaches the player.
- Retrying is the same puzzle, not a new one. The most counterintuitive decision in the project: going back to Grisal repeats the same seed on purpose. The roguelike temptation is to give a new run after every death, but that turns failure into noise — if the next voyage is a different one, what you learned about this one is worth nothing. By repeating the seed, the second attempt is the same enemy with more light: the player returns to port knowing the water runs out on day 3, and this time spends their questions on something else.
- The prototype first, in HTML. Before touching Godot I built the entire game as an HTML page with SVG and JS — playable from start to finish — so the design could be discussed by playing it rather than imagining it. Once the mechanics were clear, that prototype became something more useful than documentation: the executable specification of the port. The Godot engine wasn’t considered correct until it generated exactly the same voyage as the prototype for the same seed.
- We measured that the game was boring, and fixed it with numbers. After the port, an uncomfortable measurement: in long mode only 6 puzzles were possible, and three difficulty parameters did nothing at all. The game that was in theory infinitely replayable was, in practice, memorisable in an afternoon. The answer was to widen the repertoire — from 6 to 12 hazards, from 13 to 18 items — and, above all, to take the rules out of the code and put them into data. Today difficulty lives in a JSON file that gets tuned between playtest runs without recompiling.
- The port map is a variable too. The port has eleven locations and not all of them open every run: the same seed that generates the voyage draws which ones are open today, so the optimal route stops being memorisable. Two safeguards stop the rotation from degenerating into frustration. There is an information floor — if the draw leaves the port too poor, extra locations open — and a harder rule, learned by breaking it: whatever switches a mechanic off doesn’t rotate. At first the alley, home to the only informant able to point out that a clue is false, was part of the draw. It cost variety (from 126 possible maps to 56) and it was worth it.
seed → crossing: Seed + difficulty + length → Voyage generation → Solvability verifier → Port location rotation → Information floor → Preparation · daily actions → Cascading crossing → Logbook · what killed you
Solution: Proven solvability and rules outside the code
A verifier proves by brute force that every generated voyage has at least one combination of purchases that survives with an economic margin; if it can’t find one, that voyage never reaches the player. The repertoire went from 6 to 12 hazards and from 13 to 18 items, and difficulty now lives in a JSON file that gets tuned between playtest runs without recompiling.
Product decisions
- my role: Game design, development and art direction
- the contract with the player: Never lie about the rules; do lie about the facts. Clues can be false and people lie in the tavern, but the system is fair — and that is proven, not promised.
- most counterintuitive decision: Going back to Grisal repeats the same seed on purpose. If the next voyage were a different one, what was learned about this one would be worth nothing. The second attempt is the same enemy with more light.
- status: Playable from start to finish. Self-contained Windows executable, verified in a clean directory. No store or date.
Results
The game is playable from start to finish and ships as a self-contained Windows executable, verified in a clean directory. Everything measurable today is about the system, not the players: there are no external playtests yet, no retention metrics, no store and no date. Any figure here speaks to the soundness of the system, not to anyone having liked it. That is the next step, not a result already achieved.
- 405/405 — combinations with exact parity against the prototype, across two oracles
- 792 — possible puzzles in long mode — there used to be 6
- 56 — distinct port maps through location rotation
"What rotates are voices, never mechanics. In ~44% of runs the player lost the ability to detect lies without anything telling them: exactly the kind of difficulty you can’t learn from."
What was learned
- Measure variety before you believe in it. The infinitely replayable game had six puzzles. It wasn’t a design hunch that caught it: it was counting. If a generative system isn’t instrumented, it’s assumed to work because it’s generative.
- A playable prototype is worth more as an oracle than as a document. Writing the game twice — first in HTML to play it, then in the engine — gave a definition of correct that admits no opinion: same result for the same seed, or the port is wrong.
- What’s missing, and I know it. The preparation days are still decorative, mandatory water turns one purchase into a non-decision, and there’s almost never a hazard without an honest clue — which is exactly what would give the logbook between loops its full meaning.
In development. Working Windows build; no store or date yet.
A visual grammar, not a style — from the medium-format negative to the printed book · Art direction · 2013–2021 · 3 books · 1 campaign
- Client
- Personal and commissioned work
- Role
- Art direction, photography, editing and layout of the three books
- Timeline
- 2013 – 2021
- Stack
- 35 mm, 120 · medium format, Digital, InDesign, Lightroom Book, Campaign art direction
A style is recognised by its finish: a colour, a grain, a filter. What’s here is something else — a way of solving the frame that works the same with a medium-format camera in a dark studio as with a digital one in full sun in a volcanic crater. Isolate a form against a background that stays silent. Let in a single saturated colour. Think in squares. And keep the accident only when it puts the image in order. The subject changes, the continent changes, the camera changes; the operation is always the same.
The problem: A style doesn’t survive a change of subject, medium and country
Eight years of work spread across studio, film, campaign and book risk reading as four different portfolios. The question wasn’t what my style is, but what decision gets made before knowing what is being photographed — because only that withstands a change of context.
How it was approached
- 01 · A form against a flat field. The most constant decision, and the most invisible: isolate an object and give it a background with no information. Sky, black, a white wall, the sheet of the sea. The background isn’t what was behind — it’s what was chosen so that there would be nothing behind. When the work reaches paper, the flat field stops being the sky and becomes the page margin itself: in Geometrías the image never bleeds, it floats inside the white, which does exactly the job the sky did.
- 02 · A single saturated accent. There are never two colours fighting. There is a desaturated field — grey, beige, black, the fog — and exactly one red thing inside it. Red always comes in alone, and it is always what organises the frame: it fixes where the eye starts and forces everything else to give up saturation. The size doesn’t matter: a hood that fills half the image or a two-centimetre embroidery work for the same reason.
- 03 · The square as a unit of thought. The square starts out as an imposition of the medium-format negative, but it doesn’t stay there: it survives scanning, layout and reaches the printer. Two of the three books are square — 21 × 21 and 17.5 × 17.5 — without any camera forcing it. The square forces composition by weight rather than direction: there is no long side to push the eye, so the subject has to hold itself up.
- 04 · Control and accident, on the same table. On one side, the studio: flash, black background, objects frozen in mid-air. On the other, film: fogging, double exposures, light leaks. What’s interesting isn’t that they coexist, but the criterion for keeping the accident — the error is kept when it reinforces rules 01 and 02; if it gets in their way, the frame is discarded. The double exposure wasn’t planned: it stays because the overlap erases the background and leaves the figures floating, that is, because the accident made rule 01 on its own.
- The system on paper — three books, three incompatible pages. The books are where the method is really put to the test, because a sequence forgives nothing. Geometrías: pure form, the image contained within a wide margin. USA: territory and road, the image bleeds and alternates a full-bleed page with a floating one so the reading has a pulse. Sillas: pairs butting against each other, no margin and no folio. It isn’t a template applied three times — it’s the same criterion producing three different answers because the subject is different. The format follows the subject.
from frame to page: Decide what stays out → A silent background → One saturated accent → Square frame → Shot · control or accident → Discard criterion → Sequence → Page edge according to subject
Solution: Four operations applied before the subject
The background isn’t what was behind: it’s what was chosen so there would be nothing behind. There are never two colours fighting — there is a desaturated field and exactly one red thing inside it. The square survives scanning, layout and printing. And the accident is kept only when it reinforces the first two rules; if it gets in their way, the frame is discarded.
Product decisions
- my role: Art direction, photography, editing and layout of the three books
- the four rules: 01 A form against a flat field
02 A single saturated accent
03 The square as a unit of thought
04 Control and accident on the same table - application: Tenerife, April 2019 — two days in the Teide crater. The location is chosen for the background, the clothes for the colour, and the shot is decided before the model sits down.
- the system on paper: Geometrías (21 × 21, an image that never bleeds), USA (24.4 × 21, full bleed alternating with floating) and Sillas (17.5 × 17.5, pairs butting together with no margin or folio).
Results
Eight years of work that hold together as one body rather than four separate portfolios, with three books laid out and a campaign produced from start to finish. The result isn’t a style recognisable by its finish, but a grammar that withstands the change of subject, medium and country — and that includes knowing when to suspend itself: in Sillas the first rule is deliberately switched off, because the background is what names the absent.
- 8 years — 2013–2021 · 35 mm, 120 and digital
- 3 books — three incompatible page solutions, and all three correct
- 4 rules — applied before knowing what is being photographed
"Sillas isn’t about chairs: it’s about absence. There rule 01 is suspended, and it is suspended for a reason that can be stated — the background can’t be flattened because the background is what names the absent."
What was learned
- The format follows the subject. A closed form asks for a closed frame; a territory without edges asks for a page without edges; an absence asks for the images to touch. The rule isn’t use margins or use bleed — it’s that the edge of the page says the same as the image.
- Almost all the work is in subtraction. The background pushed aside, the colour that doesn’t come in, the long side that gets cropped. Deciding before shooting what stays out of the frame.
- A grammar is useful when you also know how to break it and say why. Knowing when not to isolate is the same decision as isolating, made the other way round.
If you have a project where someone has to decide how something looks — and sustain that decision across a campaign, a website or a printed object.
Vision, design and market reviewed inside the editor, against the repository · Open-source tool · Product audit · CLI + prompt
- Client
- Side project — RBT Studio
- Role
- Product design + prompt engineering + CLI
- Timeline
- July 2026 – August 2026 · v0.5.1 working, pending publication on npm
- Stack
- Node.js ≥18.17 · zero dependencies, Claude Code · skills + slash commands, Markdown as the product layer, MIT
A software project has automated review of almost everything except the one thing that decides whether it survives. The linter reviews style, tests review behaviour, CI that it compiles. Nobody reviews whether the product makes sense. That gap belongs to the people who code and decide at the same time — indie hackers, technical founders, freelancers with their own product — and their real alternative isn’t another tool: it’s showing it to someone with judgement, or finding out through the silence after launch.
The problem: Asking an LLM to critique your work produces flattery formatted as a report
Lists of lukewarm observations — "improve the visual hierarchy", "could confuse the user" — that don’t force anything to change. An audit framework that doesn’t solve this is no use, however good its structure. And an audit that consumes the whole context window isn’t either: whoever runs it ends up with no room to fix what the report has just pointed out.
How it was approached
- The product is a prompt, not an application. The first architecture decision was to recognise where the value really was. notjustcode is about 200 lines of Markdown in templates/; the CLI is 372 lines dedicated to copying them into the project’s .claude/ folder. That split — 99% of the code serving 1% of the value — is uncomfortable to look at, but it’s the right one: the installer runs at install time, and the audit happens at audit time. Putting product logic in the CLI would mean measuring the repository at the wrong moment.
- The rule that turns an observation into a finding. Every finding has to be presented as a complete chain: finding → evidence in the repo → consequence → confidence. Without evidence you can point at, the finding doesn’t go in the report. The consequence has to name what a specific person does instead of what was expected — abandons, retries, writes to support, doesn’t come back — and phrases like “worsens the experience” are explicitly forbidden. Banning empty consequences isn’t a style rule: it’s the structural defence against the characteristic failure of this category of tool.
- Context cost is a product decision. An audit that consumes the entire context window is useless even if the report is good: whoever runs it ends up with no room to fix what the report has just pointed out. That’s why the reading scope is written as part of the product. The command reads in full whatever talks about the product — README, manifest, routes, UI copy, design tokens — reads only the first ~50 lines of components and hooks, and completely ignores node_modules, lockfiles, tests, migrations and binaries.
- Market analysis sits behind a flag. The competition and positioning layer requires web search. It costs the same in a small project as in a large one, but its relative weight doesn’t: it adds around +70% context in a small project and only +6% in a large one. That asymmetry is why it’s --market and not the default behaviour.
- Specifying v0.6.0 before writing it. Interview mode was designed with a complete SDD process before touching a line: a SPEC with 7 functional requirements and their acceptance criteria, an ARCHITECTURE with 7 ADRs, plus SCAFFOLD, AGENTS and HANDOFF. Since the feature isn’t code — it’s five sections of Markdown — no automated tests are possible, and that forced a systematic preference for the enumerable over the interpretive.
npx notjustcode: Install into .claude/ → Bounded reading scope → 01 · Vision → 02 · Design · Nielsen over routes → 03 · Market (--market) → Empty-consequence filter → 04 · 3–5 prioritised actions
Solution: Forbidden consequences and a reading scope written as product
A finding’s consequence has to name what a specific person does instead of what was expected — abandons, retries, writes to support, doesn’t come back. The command reads in full whatever talks about the product and only the first ~50 lines of components and hooks; it ignores node_modules, lockfiles, tests, migrations and binaries. Market analysis sits behind a flag because its relative weight is asymmetric depending on the size of the repo.
Product decisions
- my role: Product design + prompt engineering + CLI
- the uncomfortable architecture: The product is a prompt, not an application. ~200 lines of Markdown are the value; 372 lines of CLI exist to copy them. It’s the right split: the installer runs at install time, the audit at audit time.
- the rule that holds it all up: Every finding is a complete chain: finding → evidence in the repo → consequence → confidence. Without evidence you can point at it doesn’t go in the report, and "worsens the experience" is explicitly forbidden.
- status: v0.5.1 working end to end · pending publication on npm · v0.6.0 interview mode specified with a full SDD process.
Results
The project is finished as a product and not yet published on npm, so there are no usage metrics to report. What there is: v0.5.1 working end to end, with install, clean uninstall and --dry-run. Two complete reports published in the repository without editing out the uncomfortable conclusions — one on notjustcode itself and another on a SaaS with an interface — which are the main marketing asset, because they let you evaluate the product without installing anything. And the tool audited itself: five findings, four prioritised recommendations, and recommendation 3 is now the complete specification of v0.6.0. The report on itself includes a section on its own moat that concludes the product has no technical defence whatsoever — that the moat is authorship, distribution and the willingness to publish the criticism others would keep to themselves.
- ~90k — context tokens per audit, versus ~420k reading the whole repo
- 4 layers — vision, design, market and actions — each action references its finding
- 0 — dependencies · Node ≥18.17 · MIT
"The tool audited itself and the report became the roadmap. The most useful finding it produced was about itself: it planned far more easily than it shipped."
What was learned
- Credibility depends on one rule, not the whole prompt. Without the ban on empty consequences, the same framework produces a friendly, useless report. It wasn’t a whim: it’s the conclusion of having watched the version without it fail.
- A product can be so cheap to copy that protecting the code is the wrong strategy. Accepting that in writing in the README itself pays off more than faking a technical moat that doesn’t exist: the moat is authorship, distribution and publishing the criticism others would keep to themselves.
- The interpretive can’t be regression-tested. Since interview mode isn’t code but sections of Markdown, no tests are possible — and that forced a systematic preference for the enumerable over the interpretive.
MIT. Public repository; package pending publication on npm.
A 5-phase wizard so agents build what I had in my head · Internal tool · Specification-Driven Development
- Client
- Side project / internal tool — RBT Studio
- Role
- Skill design and development
- Timeline
- May 2026 · in active use
- Stack
- Claude Code · skill, Specification-Driven Development, Multi-agent architecture, TDD
Every time I started a new project with agents, the starting point was different: sometimes straight into code, sometimes a spec improvised in a message. The result was always the same — specs inconsistent from one project to the next, and a disconnect between what I had in my head and what the agents ended up building.
The problem: What existed was either too little or too much
The simplest planning skills had no real depth: a spec template with nothing behind it to ensure the resulting code met it. The most complete ones went to the other extreme — heavy processes designed for large teams that, on a personal or mid-sized client project, are pure friction: lots of ceremony, little agility.
How it was approached
- Search first, build later. Before writing anything I looked for existing solutions — other structured-planning skills — and none fitted. The simplest ones had no real depth: a spec template with nothing behind it to ensure the resulting code met that spec. The most complete ones went to the other extreme: heavy processes designed for large teams, which on a personal or mid-sized client project turned into pure friction. I needed something in between.
- Treat it as a Lean system, not project management in disguise. The key design decision. Instead of a single-document template, I structured it as a 5-phase wizard, each phase with a concrete responsibility and a clear output that feeds the next: SPEC (what and why), ARCHITECT (how and which decisions), SCAFFOLD (base structure), HARNESS AGENTS (coordination) and HANDOFF (delivery and context).
- TDD as the only non-negotiable discipline. The system’s pivot point is TDD: every implementation phase requires the tests to pass before it is accepted, which acts as a real anchor against drift — both mine and the agents’. It’s the difference between “I think this meets the spec” and “this meets the spec because the tests confirm it”. Adding more phases or more ceremony would have reintroduced the friction I wanted to avoid; removing the test discipline would have taken me back to square one.
5-phase wizard: SPEC → ARCHITECT → SCAFFOLD → HARNESS AGENTS → HANDOFF → TDD as the anchor in every phase
Solution: A Lean system with one non-negotiable discipline
Five sequential phases, each with a concrete responsibility and a clear output that feeds the next, generating the documentation artefacts before touching implementation code. Adding more phases would have reintroduced the friction; removing the test discipline would have taken me back to square one — good intentions without verification.
Product decisions
- my role: Skill design and development
- the 5 phases: SPEC what and why
ARCHITECT how and which decisions
SCAFFOLD base structure
HARNESS AGENTS coordination
HANDOFF delivery and context - the anchor: TDD. Every implementation phase requires the tests to pass before it is accepted — the difference between "I think this meets the spec" and "this meets the spec because the tests confirm it".
- status: Created in May 2026, in active use. The complete PRD for Job Search OS came out of this process.
Results
There’s no single metric that sums up the impact, but the change in how I work is clear. There is confidence that nothing gets lost between the initial idea and the finished project, because every phase leaves an artefact the next one can consult. There is a shared language between me and the agents: with an explicit spec and architecture before implementing, there’s no ambiguity left for each agent to interpret its own way. And there is less rework, because the TDD approach catches interpretation errors early rather than at the end. I’ve already used it to kick off real projects — the complete PRD for Job Search OS came out of this process, pivoting from an initial idea of a manual CRM to an event-driven Career Operating System with a far more solid architecture than starting straight from code would have produced.
- 5 phases — each with an artefact the next one can consult
- 1 anchor — TDD — the system’s only non-negotiable discipline
- In use — since May 2026, on personal and client projects
"The right structure isn’t in the minimal template or in the exhaustive process, but in a Lean system with well-defined phases and a single non-negotiable discipline that acts as the anchor."
What was learned
- Confidence that nothing gets lost between the initial idea and the finished project — every phase leaves a traceable artefact.
- A shared language between me and the agents. With an explicit spec and architecture before implementing, there’s no ambiguity left for each agent to interpret its own way.
- Less rework. The TDD approach catches interpretation errors early, not at the end — for example, pivoting from a manual CRM to an event-driven Career Operating System before writing any code.
Internal tool, in active use since May 2026.