Paying for a machine that guesses

I pay for several AI models every month. Some days they save me an hour; on others I spend that hour proving them wrong.

· 2 min read · ai ·trust ·probabilistic-systems ·product-thinking

A paid probability machine produces a result ready for close inspection.

I pay for several AI models every month, then spend a fair amount of time checking what they give me.

There is something faintly absurd about that. The answer arrives in seconds, looks finished and may contain a command flag that has never existed.

I once called it tarot with an API after catching myself wanting to believe a particularly convenient answer. The machine sounded certain. I liked the answer. Neither fact made it true.

And yet I keep renewing the subscriptions.

Most of what I ask them to do is much less grand than the way AI is sold. I give one a messy log and ask where it would look first. I open an unfamiliar repository and ask it to sketch the shape of it before I start reading. The result is often good enough to get me moving, which is different from being right.

The model does not look embarrassed when I find that the file it quoted does not exist.

The value depends on what happens next. With code I have something solid to inspect. There is a diff, a test suite and eventually a deployment I can watch. If the suggestion is poor, I can throw it away. If it is almost right, correcting it may still be quicker than starting cold.

This is harder to judge than the usual productivity claims make it sound. A model can save twenty minutes and cost ten in checking. Fine. It can also send me into a dead end that takes an hour to unwind. Both experiences feel fast at the beginning.

I am comfortable with that bargain when the job is small and reversible. I am much less comfortable when the answer affects something expensive to undo, or when I cannot get back to the source. A confident paragraph is not much help if I have no independent way to test it or source of truth as a reference.

So the checking stays. I read the diff, open the documentation, run the tests and watch what happens after deployment.

Some days a model saves me an hour. On others, I spend all day proving it wrong. The bill does not distinguish between them, and neither does the confidence of the answer on screen.

← All writing