AI-built, but robust
You stopped reading every line your AI writes. Kelson makes sure it still gets checked.
Nobody can keep up with what agents produce, so most AI code merges half-read. Asking another AI to review it helps, but a review is one step: you have to remember to run it, it only gives advice, and it forgets by tomorrow. Kelson runs the whole process around your agent, from spec to merge. The checks always run, the agent can’t skip or edit them, serious problems block the merge, and every mistake caught can become a permanent check.
“Delete old uploads” is ready. Uploads people delete will now be removed for good after 90 days, instead of kept forever. Every check passed, and the independent review found nothing serious.
The agent tried to delete 3 tests to get its change through. Sent back; fixed properly.
What you see. Not the code, unless you want it.
The problem
AI made writing code cheap. Checking it is now the job.
Coding agents can build whole features from a description, and they’re usually right. The trouble is the “usually”, and the fact that finding the exceptions now falls on you.
The agent marks its own homework
The same AI that wrote the code decides when it’s done. If the tests fail, nothing stops it changing the tests instead of the code. Researchers have caught leading models doing exactly that.
METR, 2025: frontier models patching the code that grades them.Passing the tests isn’t the same as good to merge
Tests only cover what someone thought to test. Plenty of AI changes pass and still aren’t something a careful developer would accept.
METR, 2026: about half of AI fixes that passed a standard benchmark’s tests would not have been merged by the project’s maintainers.Asking for a review isn’t a process
You can write rules for your agent, or ask another AI to review the pull request. Both help. But rules are suggestions an agent can drop in a long session, and a review is advice: it posts comments, and nothing makes anyone act on them. It only happens if you remember to ask.
And it forgets everything by tomorrow
You catch a problem in review this week. Next month the agent makes it again, and someone has to notice again. Each session starts fresh. What you learned lives in your head, not in the process.
What’s missing isn’t a smarter coding agent. It’s a build process around the agent: one it can’t skip, that checks its work properly, and that remembers what went wrong.
The solution
Lots of small safeguards, stacked.
No single check catches everything. Each layer has gaps. Stack enough layers that catch different things, and a mistake that slips through one gets stopped by the next. It’s how aviation and hospitals keep things safe, applied to code your AI writes.
You could set up every one of these yourself. Most developers who let AI write real code end up building some version of them by hand: rules files, review prompts, CI settings, locked branches. Kelson ships them together, tested, and outside the agent’s reach, so it can’t switch them off.
Think of it like a web framework. It doesn’t do anything you couldn’t write. It means you stop rewriting it for every project, and every improvement reaches all of them. Underneath sit the quieter layers too: a spending cap on every build, and builds that resume after a crash instead of starting over.
One feature, start to finish
You write what you want. You approve what you get.
Everything in between runs on its own, in your own GitHub, and only stops when a decision is genuinely yours.
- You
Describe the feature
What it should do, and how you’ll know it works. Kelson checks the description is clear enough to build from.
- Your AI agent
Builds it
Any coding agent, using your own account. It works from the locked copy of your spec.
- Kelson
Checks it
Automatic checks the agent can’t edit. Fail one and the work goes back to the agent with the reason.
- Another AI
Reviews it
Fresh eyes, no knowledge of how it was built. Anything serious goes back for another round.
- You
Say yes
A short summary in plain English. Tap to merge, or say not yet.
Rules decide the route from the change itself, never the agent. Small, low-risk changes take the light route: one review and a one-tap yes. Bigger ones get every detour they need. Dashed boxes run only when the change calls for them.
Example · what a check actually catches
The agent cut corners. It didn’t get through.
Asked to add a new rule for deleting old uploads, the agent got its change working, but three existing tests failed. So it deleted those three tests. In a normal setup, everything goes green and you’d only find out if you read the diff. Here the check caught it, the work went back, and the agent fixed the code instead.
- ✓Spec checked and locked
- ✓Agent built the feature: 14 files changed
- ✓No passwords or keys leaked into the code
- ✗Tests weakened: 9 checks in the upload tests went down to 6
- ↺Sent back to the agent with that reason
- ✓All 9 tests kept, and passing
- ✓Database change rehearsed on a throwaway copy
- ✓Independent review: nothing serious
- ●Summary sent to your phone
| Your agent, told to check its own work | A setup you built yourself | Your agent inside Kelson | |
|---|---|---|---|
| Who decides it’s done | The agent | The agent, if it follows your rules | Checks it can’t change, then you |
| Can it skip or edit the checks | Yes. Instructions are suggestions | Yes, unless you’ve locked everything down | No. Touching one stops the build and asks you |
| Remembers last month’s mistakes | No. Every session starts fresh | Only if you add a rule by hand | Yes. Caught mistakes can become permanent checks |
| Tried in a real environment | Rarely | If you’ve wired it up | Database rehearsal and before-and-after screenshots, built in |
| Runs while you’re away | Until it gets stuck or drifts | Depends what you built | Yes. It only stops when a decision is yours |
| Setting it up | Nothing to set up | Yours to build and maintain, per project | Ready-made, with your own rules on top |
It gets better every time
Every mistake your AI makes, Kelson learns from.
Every mistake that’s found, by the independent review, by you, or after it merged, is confirmed and written down: what went wrong, where, and which AI made it. That record is the lessons log. What happens next depends on the kind of mistake.
Could a program catch it?
It becomes a check. That’s the ratchet.
- Kelson drafts a new automatic check.
- It’s tested against the real mistake, so it’s proven to catch it.
- You approve it once.
- It runs on every build, before the review: first watching, then blocking once it’s proven not to fire on good changes.
Checks only get added, never quietly loosened. The ratchet only turns one way.
Needs judgement to spot?
It becomes guidance for the builder.
- Kelson looks for the mistakes AI builders keep making.
- They become a short list of things to avoid.
- The agent gets that list before it writes a line.
- The list is retested and pruned, so it stays short and current.
Fewer mistakes reach review, and builds need fewer rounds to get ready.
Checks stop mistakes. Guidance makes them rarer. The independent review still runs on every change either way, and it never relies on the list. In Kelson Cloud, the lessons log can learn from mistakes across many projects, not just yours.
What you get
The whole cycle, AI-first. Kelson is the builder at its centre.
Building software with AI is more than the build. The open-source core covers build, check and merge, and works on its own; context and planning plug in when you want them. Kelson Cloud runs all seven stages for you, set up and in one place.
- Set upAdd Kelson to a repository and set its rules.Open sourceCloud
- ContextWhat your project knows, kept short and current.Open source · optionalCloud
- PlanWhat to build next, as specs ready to build.Open source · optionalCloud
- BuildYour agent builds from a locked spec.Open sourceCloud
- Check & reviewChecks it can’t touch, tried for real, an independent AI.Open sourceCloud
- MergeYour yes, then merged.Open sourceCloud
- Get betterThe ratchet: mistakes become checks.Open sourceCloud
| Kelson open sourceCore builder, plus optional modules · run it yourself | Kelson CloudEverything, set up and run for you | |
|---|---|---|
| 1Set up | ||
| Installer | Adds Kelson to your repository in one pull request | Connect your repository, database and hosting. Kelson configures the rest |
| Your project’s rules | Written and maintained by you | Set up for you and kept up to date |
| 2Context | ||
| Project knowledge every build starts from | Optional module, run by you | Built in and managed for you |
| 3Plan | ||
| Roadmap, features and specs | Optional module, run by you | Built in, with approvals in one place |
| 4Build | ||
| Spec in, merged pull request out | ✓ | ✓ |
| Works with your agent and your own AI account | ✓ | ✓ |
| Runs in your own GitHub Actions | ✓ | ✓ |
| Spending cap and crash recovery on every build | ✓ | ✓ |
| 5Check & review | ||
| Tests the agent can’t touch, automatic checks | ✓ | ✓ |
| Database rehearsal and before-and-after screenshots | ✓ | ✓ |
| Independent AI review | ✓ | ✓ |
| 6Merge | ||
| Approvals | On GitHub, with a phone notification | One tap in the app, in plain English |
| Overnight queue and morning summary | A queue you start | Scheduled for you |
| Every build, cost and approval in one place | — | ✓ |
| Team: shared builds and approvals | — | ✓ |
| 7Get better | ||
| The ratchet: your mistakes become your checks | ✓ | ✓ |
| Lessons log | Your project’s own | Learns from mistakes across projects (opt in) |
| Check library | The ready-made checks | A maintained library, updated as tools and AI models change |
Kelson Cloud is in design. The core builder is being built now.
Is it for you
For developers already letting AI write real features.
A good fit if you
- Already use a coding agent for whole features, not just autocomplete
- Spend more time checking its work than you’d like
- Have started writing your own rules files and review steps, and are tired of maintaining them
- Want to step away while it builds, and come back to something you can trust
Not the right fit if you
- Want an app built from a single prompt without touching code. Kelson is for developers.
- Want a new coding agent. Kelson works with the one you already use.
- Need someone to guarantee the code. You still own your code, security and data.
Does Kelson write or review the code itself?
No. Your AI agent writes it and an AI reviewer reads it, possibly the same review tool you use today. Kelson runs the process around them: what gets built, which checks it faces, what happens when something fails, and when you’re asked.
Why not just ask Claude Code to review the pull request?
Do that, and Kelson will too. The difference is everything around the review. It runs on every change, without anyone remembering. The reviewer is a different AI that never saw the build. A serious finding blocks the merge and sends the work back automatically, instead of sitting as a comment. Automatic checks and a trial run on a real database catch what reading code can’t. And every confirmed finding can become a permanent check, so next month’s review starts where this one ended.
Where does it run?
In your own GitHub, using GitHub Actions. Your code and passwords stay in your repository. If the hosted service is ever down, your builds still run.
What does it cost?
The core is free and open source. You pay your AI provider directly for the agent’s work, with a cap on every build. The hosted version will be a subscription.
Building version one
Run it yourself, or let Kelson set it up.
Kelson is being built with its own process, and tested on a real product whose database problems only showed up once its code was live. Those problems became Kelson’s first database checks.