Field note / AI & product
How to know if your AI feature is actually getting better
A good-looking answer is easy to mistake for a useful one. Here’s a practical way to test the difference.
Imagine you add an AI assistant to a product. It writes a landing page, sorts a support request, or suggests the next step in a workflow. The first few outputs look impressive. Then a real customer asks something messier.
That is the moment a demo stops being a test. The question is no longer “Can it produce an answer?” It is “Does it help with the work people actually came to do?”
I’ve been thinking about this as a product and design problem. If you cannot describe what a useful result looks like, you cannot tell whether the product improved. You can only tell that the output changed.
Test the job, not the performance of the demo.
Start with the work people bring you
Before changing prompts or switching models, collect a small set of tasks that reflects actual use. For a marketing assistant, that could include a clear offer, an incomplete brief, conflicting brand instructions, and a request that should be declined or sent back for clarification.
Real requests and support tickets are useful because they reveal the awkward cases a polished example misses. Add a few cases you write yourself, too. If every test comes from what people already ask, you may miss the valuable things they never thought the product could do.
Write down why each case matters before running it. “Our current model fails here” is not enough. The case should represent a real customer need, a meaningful risk, or a decision the product must get right.
Make the score explain itself
A single number is tempting. It is also easy to trust too soon. If the assistant writes copy, “good” is too vague to grade. Break it into claims a reviewer can check: Did it preserve the offer? Did it invent a testimonial? Is the call to action clear? Did it ask for missing information when the brief was incomplete?
Use a simple automatic check when the answer has a fixed shape, like a category or a valid data format. For open-ended work, use a clear rubric and review a sample of the judgments yourself. If a reasonable person disagrees with the score, the scoring rule may need work.
Run the same case more than once. If it passes on Monday and fails on Tuesday with the same setup, a tiny score change tells you very little. The test has to be steadier than the improvement you hope to measure.
“Is this landing page good?”
“Does it state the offer accurately, avoid invented proof, and give a clear next action?”
Keep a few examples out of sight
Once you can score the work, you can start improving it. This is where it gets easy to fool yourself. If you read every failed example and keep adjusting the prompt until those exact examples pass, you may have built a prompt that memorizes the test rather than handles new requests.
Set aside some cases before you make changes. Use one group to learn from and a separate group to check whether the improvement carries over. If the familiar cases get better and the unseen ones stay flat, treat that as a warning.
Change one thing, then look again
Pick a surface you can edit and undo: a prompt, a tool description, a model setting, or one rule in a workflow. Decide what you are trying to improve before changing it. Better accuracy? Faster responses? Lower cost at the same quality?
Make one meaningful change at a time. Compare the result with your starting point and with the normal variation in your test. Keep the change when the gain shows up in the unseen cases. Revert it when quality drops, or when only the familiar examples improve.
Sometimes the most useful finding is that the test is wrong. A confusing task, a grader that expects something never requested, or a broken run can all make the product look worse than it is. Fix the measurement before optimizing the product around it.
A small playbook for your next AI feature
- Name the job.Write one sentence about what the user needs to accomplish.
- Gather real cases.Include routine requests, difficult requests, and cases where the system should ask or stop.
- Define success.Use criteria someone else could apply without guessing what you meant.
- Check the checker.Read a handful of outputs and scores before trusting the total.
- Save unseen cases.Keep a separate set for testing whether changes generalize.
- Make one change.Compare quality, time, and cost. Keep only gains that hold up.
There is no magic score that makes a product useful. But a small, honest test can make the next decision clearer. That is enough to start.