Skip to main content

You handed the task to an agent.
Why does it stop when you close your laptop?

We think cloud agents will build most of the world’s software. That only works if handing one a task feels like handing it to a teammate: describe it, walk away, and come back to a pull request with evidence.

Writing code is cheap, but trusting it is not.

A frontier model can write most of the changes you would ask for. The cost shows up afterwards, in an engineer’s attention.

Someone still has to watch the terminal, rerun the tests, open the page and click through it before anyone can say the change works. That attention is the scarce resource now, and a good agent should need less of it.

Agents need their own machine.

Real work means starting the app, running tests and driving a browser, often for several tasks at once. A laptop is a poor host for that. It sleeps when you close it, and two agents sharing one checkout get in each other’s way.

Every [code]smith task runs on its own cloud machine with your repository checked out, common languages and tools installed, and a browser for screenshots and recordings. You can run several tasks at once, and none of them use your laptop.

Your CI already knows how to install your dependencies, start your services and run your tests, so [code]smith checks its work there. It pushes to the pull request, waits for your GitHub Actions jobs, reads the logs when one fails, and pushes fixes until the checks pass. When those jobs run on Blacksmith, it pulls a job’s logs and resource usage directly.

Coming soon: [code]smith will build each task’s machine from your CI setup, so your dependencies, services and caches are ready before it runs its first command.

Claims need evidence.

A passing test suite says little about whether a page looks right or an animation got smoother. For that you need to see the change: screenshots of the page, a recording of the interaction, before and after numbers for a performance fix.

[code]smith puts that evidence in the pull request. It can run Proof, a separate session that exercises the change and posts its screenshots and recordings to the PR, next to the diff. Evidence is what lets you trust work you didn’t watch.

Work shouldn’t stop when you do.

Plenty of tasks stall on something slow, like a CI run or a week of data. [code]smith does the waiting. It checks back on CI, pushes fixes when a reviewer comments, and can schedule itself to pick the work up hours or days later.

When you come back, you find a pull request that is ready for review, with its CI results and evidence attached.

Work that outlives the conversation

See what it can do →Get started →