NOTJUSTCODE — The product audit nobody calls
Vision, design and market reviewed inside the editor, against the repository · Open-source tool · Product audit · CLI + prompt
- Client
- Side project — RBT Studio
- Role
- Product design + prompt engineering + CLI
- Timeline
- July 2026 – August 2026 · v0.5.1 working, pending publication on npm
- Stack
- Node.js ≥18.17 · zero dependencies, Claude Code · skills + slash commands, Markdown as the product layer, MIT
A software project has automated review of almost everything except the one thing that decides whether it survives. The linter reviews style, tests review behaviour, CI that it compiles. Nobody reviews whether the product makes sense. That gap belongs to the people who code and decide at the same time — indie hackers, technical founders, freelancers with their own product — and their real alternative isn’t another tool: it’s showing it to someone with judgement, or finding out through the silence after launch.
The problem: Asking an LLM to critique your work produces flattery formatted as a report
Lists of lukewarm observations — "improve the visual hierarchy", "could confuse the user" — that don’t force anything to change. An audit framework that doesn’t solve this is no use, however good its structure. And an audit that consumes the whole context window isn’t either: whoever runs it ends up with no room to fix what the report has just pointed out.
How it was approached
- The product is a prompt, not an application. The first architecture decision was to recognise where the value really was. notjustcode is about 200 lines of Markdown in templates/; the CLI is 372 lines dedicated to copying them into the project’s .claude/ folder. That split — 99% of the code serving 1% of the value — is uncomfortable to look at, but it’s the right one: the installer runs at install time, and the audit happens at audit time. Putting product logic in the CLI would mean measuring the repository at the wrong moment.
- The rule that turns an observation into a finding. Every finding has to be presented as a complete chain: finding → evidence in the repo → consequence → confidence. Without evidence you can point at, the finding doesn’t go in the report. The consequence has to name what a specific person does instead of what was expected — abandons, retries, writes to support, doesn’t come back — and phrases like “worsens the experience” are explicitly forbidden. Banning empty consequences isn’t a style rule: it’s the structural defence against the characteristic failure of this category of tool.
- Context cost is a product decision. An audit that consumes the entire context window is useless even if the report is good: whoever runs it ends up with no room to fix what the report has just pointed out. That’s why the reading scope is written as part of the product. The command reads in full whatever talks about the product — README, manifest, routes, UI copy, design tokens — reads only the first ~50 lines of components and hooks, and completely ignores node_modules, lockfiles, tests, migrations and binaries.
- Market analysis sits behind a flag. The competition and positioning layer requires web search. It costs the same in a small project as in a large one, but its relative weight doesn’t: it adds around +70% context in a small project and only +6% in a large one. That asymmetry is why it’s --market and not the default behaviour.
- Specifying v0.6.0 before writing it. Interview mode was designed with a complete SDD process before touching a line: a SPEC with 7 functional requirements and their acceptance criteria, an ARCHITECTURE with 7 ADRs, plus SCAFFOLD, AGENTS and HANDOFF. Since the feature isn’t code — it’s five sections of Markdown — no automated tests are possible, and that forced a systematic preference for the enumerable over the interpretive.
npx notjustcode: Install into .claude/ → Bounded reading scope → 01 · Vision → 02 · Design · Nielsen over routes → 03 · Market (--market) → Empty-consequence filter → 04 · 3–5 prioritised actions
Solution: Forbidden consequences and a reading scope written as product
A finding’s consequence has to name what a specific person does instead of what was expected — abandons, retries, writes to support, doesn’t come back. The command reads in full whatever talks about the product and only the first ~50 lines of components and hooks; it ignores node_modules, lockfiles, tests, migrations and binaries. Market analysis sits behind a flag because its relative weight is asymmetric depending on the size of the repo.
Product decisions
- my role: Product design + prompt engineering + CLI
- the uncomfortable architecture: The product is a prompt, not an application. ~200 lines of Markdown are the value; 372 lines of CLI exist to copy them. It’s the right split: the installer runs at install time, the audit at audit time.
- the rule that holds it all up: Every finding is a complete chain: finding → evidence in the repo → consequence → confidence. Without evidence you can point at it doesn’t go in the report, and "worsens the experience" is explicitly forbidden.
- status: v0.5.1 working end to end · pending publication on npm · v0.6.0 interview mode specified with a full SDD process.
Results
The project is finished as a product and not yet published on npm, so there are no usage metrics to report. What there is: v0.5.1 working end to end, with install, clean uninstall and --dry-run. Two complete reports published in the repository without editing out the uncomfortable conclusions — one on notjustcode itself and another on a SaaS with an interface — which are the main marketing asset, because they let you evaluate the product without installing anything. And the tool audited itself: five findings, four prioritised recommendations, and recommendation 3 is now the complete specification of v0.6.0. The report on itself includes a section on its own moat that concludes the product has no technical defence whatsoever — that the moat is authorship, distribution and the willingness to publish the criticism others would keep to themselves.
- ~90k — context tokens per audit, versus ~420k reading the whole repo
- 4 layers — vision, design, market and actions — each action references its finding
- 0 — dependencies · Node ≥18.17 · MIT
"The tool audited itself and the report became the roadmap. The most useful finding it produced was about itself: it planned far more easily than it shipped."
What was learned
- Credibility depends on one rule, not the whole prompt. Without the ban on empty consequences, the same framework produces a friendly, useless report. It wasn’t a whim: it’s the conclusion of having watched the version without it fail.
- A product can be so cheap to copy that protecting the code is the wrong strategy. Accepting that in writing in the README itself pays off more than faking a technical moat that doesn’t exist: the moat is authorship, distribution and publishing the criticism others would keep to themselves.
- The interpretive can’t be regression-tested. Since interview mode isn’t code but sections of Markdown, no tests are possible — and that forced a systematic preference for the enumerable over the interpretive.
MIT. Public repository; package pending publication on npm.