How I Built 2,400+ Free Tools With AI Agents (and the Bugs That Nearly Shipped)
There are more than 2,400 free, in-browser tools on this site — converters, calculators, generators, formatters. Most of them were written by AI coding agents, not by hand. That part is easy to brag about. The part worth writing down is what went wrong, because it taught me exactly where AI agents can and cannot be trusted.
The pipeline is simple. A generator script defines a batch of tools as plain data — a name, a slug, an icon, the formula, a couple of worked examples. An agent turns each entry into a self-contained .astro page, and a second pass inserts it into the tool registry and assigns it a category. Run it, and a dozen or a few hundred tools appear at once. For a while it felt like free output.
Then one batch of roughly 300 tools lied to me.
The batch that looked finished and wasn't
On review, more than a hundred of the "calculators" in that batch were the same blind formula wearing different names: take input A, multiply by B, maybe divide by 100, and print it under a label like Result or Divisor. The math was generic where it should have been specific, and a few tools computed the wrong thing outright — one "net revenue retention" calculator subtracted expansion revenue instead of adding it. Separately, around 140 meta descriptions had a stutter bug: "runs privately privately privately in your browser."
None of it looked broken. Each page rendered, each had a tidy explanation and a plausible FAQ, each passed a glance. That is the trap with AI-generated code: the failure mode isn't a red error, it's confident, well-documented wrongness. A human skim will wave it right through.
How I catch it now
The fix wasn't better prompts, it was cheaper suspicion. Before any batch is trusted, a triage pass scans for the tells — the exact boilerplate the agent falls back on when it's padding, generic output labels, and the doubled-word stutter:
// Triage an AI-written batch BEFORE trusting any of it.
// These three cheap checks caught most of the damage.
const DOUBLED_WORD = /\b(\w+)\s+\1\b/i; // "runs safely safely safely"
const GENERIC_LABELS = /^(Result|Output|Divisor|Input total)$/;
const TEMPLATE_FAQ = "the main output shown in the result panel";
for (const tool of batch) {
if (DOUBLED_WORD.test(tool.description)) flag(tool, "doubled word in copy");
if (GENERIC_LABELS.test(tool.outputLabel)) flag(tool, "blind-formula calculator");
if (tool.faqs.some(f => f.answer.includes(TEMPLATE_FAQ))) flag(tool, "boilerplate FAQ");
}
// Anything flagged gets rebuilt from a real spec — not patched in place. Verify the code that ships, not the plan
Triage only flags the obvious tells. The real check is deterministic: pull the actual function out of each shipped file, run it against known input-output pairs, and compare the numbers. A ciphers tool has to reproduce the canonical test vectors and round-trip; a calculator has to match values I worked out by hand. Every file also gets compiled in a routable temporary folder so a broken import can't sneak through, and anything with a canvas or a file upload gets driven in a real browser, because "it compiles" is not "it works."
The one rule underneath all of it: verify the code that ships, not the description of what it should do. The agent's summary is marketing. The function is the truth. When those two disagree — and across thousands of tools, they disagree more than you'd like — the function wins and the tool gets rebuilt from a real spec, not patched in place.
What building at this scale actually taught me
AI agents are extraordinary at volume and genuinely bad at knowing when they're wrong. They will produce a hundred tools an afternoon and hand you every one as if it were finished. The value you add as the human isn't typing — it's the deterministic gate the agent can't build for itself: the test that runs the real output and refuses to ship what doesn't match. Get that gate right and the volume is a superpower. Skip it and you're publishing confident nonsense at scale.
Frequently asked questions
Can AI agents really build thousands of working tools?
Yes, for volume — an agent can write hundreds of small, self-contained tools quickly. The catch is correctness: agents do not reliably verify their own output, so a meaningful share ship subtly wrong unless you add deterministic checks that run the actual code against known answers.
What is the most common way AI-generated code is wrong?
Confidently wrong logic wrapped in convincing copy. In one batch, over a hundred calculators were generic formulas with plausible labels, and a few computed the wrong thing entirely. The code looked finished and the explanation read well, which is exactly why a visual skim is not enough.
How do you verify AI-generated tools at scale?
Extract the shipped function, run it against known input-output pairs, and compare — do not trust the description. Compile every file in a routable temporary folder, and for anything with a canvas or file upload, drive it in a real browser. Verify the code that ships, not the plan the agent described.
Shipping real volume with AI agents?
I help teams build with AI coding agents and — more importantly — put the verification around them so what ships is actually correct. If you're generating at scale and want the guardrails right, let's talk.
AI consulting with Patrick Bushe