How to Evaluate an AI Assistant for Writing, Research and Everyday Tasks
Test AI assistants on your own tasks and compare accuracy, instruction following, editing effort, privacy and cost.

Choosing an AI assistant on the strength of benchmark scores is a bit like choosing a car on its top speed. The numbers are real, they are measured carefully, and they have almost nothing to do with how the thing performs on the journey you actually make. Benchmarks measure performance on standardised problems; you have non-standard work.
A better approach takes about two hours and produces an answer specific to you.
Build a test set from your own work
Collect ten to fifteen tasks you genuinely did in the last month. Not invented examples — real ones, with the real messy context attached. A useful spread looks like this:
- Three writing tasks of different lengths and registers.
- Three research questions where you already know the correct answer, so you can catch a confident wrong one.
- Two summarisation tasks over documents you have read.
- Two format-conversion tasks — notes into a table, a transcript into an outline.
- Two tasks with fiddly constraints: a word limit, a required structure, a tone you have to hit.
Run the identical prompts through each candidate. Running them against several models in the same interface is easier than juggling accounts — a free AI chatbot running GPT, Claude and Gemini lets you paste one prompt and compare the outputs side by side, which removes the temptation to reword the prompt between attempts.
Score on four things, not on vibes
Instruction-following. Did it respect the constraints you set? A word limit, a required heading structure, “no bullet points”. This is where assistants differ most and where the difference matters most in daily use. Score it binary: followed or did not.
Verifiability. For the research questions, can you check the answer? An assistant that gives you a checkable claim with a named source is more useful than one that gives you a fluent paragraph you have to research from scratch. Count how many claims you could verify in under a minute.
Error mode. When it does not know, what does it do? Hedge, ask a clarifying question, or invent something plausible? The third behaviour is disqualifying for research work regardless of how good the other outputs are.
Edit distance. The honest measure for writing tasks: how much did you have to change before you would send it? Count it roughly — light edit, heavy edit, rewrote from scratch. This correlates with real time saved far better than any quality impression.
What not to weight
| Tempting signal | Why it misleads |
|---|---|
| Benchmark leaderboards | Measures exam-style tasks, not your work |
| Response length | Longer is easier to produce and harder to check |
| Confident tone | Uncorrelated with accuracy |
| Speed | Real; matters far less than one avoided rewrite |
| Context window size | Only matters if you actually feed it long documents |
Test the boring cases too
Most published comparisons focus on hard problems. Your usage will be dominated by easy ones: reformat this, shorten that, what is the word for when. Assistants differ less on these, but they differ on friction — how many turns it takes to get the thing you asked for. Time three routine tasks end to end on each candidate.
Check the practical constraints before you commit
- Data handling. What happens to what you paste in? Retention period, training use, deletion on request. For anything work-related this decides the question before quality does.
- Cost at your usage. Free tiers usually cap messages or restrict which model you reach. Work out what a normal week costs you.
- Export. Can you get your conversations out? Assistants accumulate useful context and some make it difficult to leave.
Re-run it twice a year
Model releases reshuffle the order regularly, and the assistant that won your test last spring may not win it now. Keep the test set in a document. Re-running fifteen saved prompts takes twenty minutes and is the only reliable way to know whether switching is worth the disruption.
The point of all this is not rigour for its own sake. It is that the gap between assistants on general benchmarks is narrow and the gap on your specific tasks is often wide — and only your own test set will show you which way it falls.=