Skip to content

AI · 1 min read

The AI projects that actually ship have one thing in common

It is not the model, the framework or the budget. It is knowing what the system should do when it is unsure.

By Bilal Ahmed ·

We have scoped a lot of AI work in the last two years and turned down a fair amount of it. The projects that reach production and stay there share a design decision that has nothing to do with which model you pick.

Design the uncertain path first

Any classifier will be confident and wrong some percentage of the time. The question that decides whether a project survives contact with real users is: what happens then?

On a healthcare triage system we built, the answer was a calibrated confidence threshold and a human review queue. The model handles 81% of referrals on its own. The other 19% go to a person — not as a failure mode, but as the designed behaviour. That is what got it past the clinical governance board.

Measure against a labelled set, weekly

Vibes are not evaluation. Build a holdout set labelled by the people whose judgement the system is replicating, and run against it every week. When accuracy drifts — and it will, as your inputs change — you will know before your users tell you.

Log the overrides

Every time a human corrects the system, that is a labelled example arriving free. Capture it. Most teams do not, and then wonder why their model never improves after launch.

Bilal Ahmed

SoftCity Solutions

Talk to us about your project

Keep reading

More from the blog

Let's talk about what you are building

A 30-minute call, no deck and no obligation. We will tell you honestly whether we are the right fit — and who is, if we are not.

Or call +923019278056 · We reply to every enquiry within one working day.