How We Test AI Tools: Our Scoring Framework, Explained
The methodology behind every review on this site — the seven scoring criteria, the standard test prompts, how we handle vendor relationships, and what we refuse to score.

Most AI tool reviews are rewritten press releases. We built this page so you can judge whether ours are worth trusting — and so we can be held to it.

The seven criteria
Every tool we score is rated 1–10 on the same axes, weighted by category.
1. Output quality (30%). The core job, judged against a standard set of tasks. For writing tools that's a blog section, an email, and an edit pass. For image tools it's a product shot, a portrait, and a text-in-image poster. For coding tools it's a bug fix, a refactor, and a feature in an unfamiliar codebase.
2. Reliability (20%). How often it produces something usable on the first attempt. We record hit rate across fixed attempts, because a tool that's brilliant one time in five is a worse tool than one that's good every time.
3. Control (15%). Can you steer it? References, parameters, rules files, iterative refinement. Ceiling matters less than steerability.
4. Workflow fit (15%). Integrations, export formats, collaboration, and whether it forces you to change how you work.
5. Speed (5%). Time to first usable output, measured under normal load, not at 4am.
6. Value (10%). Cost per unit of real work, including the iterations you'll waste. Headline price is rarely the real price.
7. Trust (5%). Data retention, training policies, licensing clarity, security posture, and how honestly the vendor describes limitations.
The standard test set
Each category has fixed tasks we run against every tool, unchanged between reviews so scores stay comparable over time.
- Writing: an 800-word explainer from a brief, a cold email, an edit of a deliberately weak paragraph, and a voice-matching task using a provided sample.
- Image: photorealistic product on white, an environmental portrait, a poster containing a specific headline, and a consistency test across three images.
- Video: a camera move, a human walking, an object interaction, and a continuity test across two shots.
- Coding: a failing test to fix, a cross-file refactor, a new feature in a codebase the tool hasn't seen, and a security review of intentionally flawed code.
- Productivity: a two-week real-usage period on live work, not a demo.

How we actually run tests
- Same prompts, same day. Model quality varies; comparing a tool tested in January to one tested in July is meaningless.
- Paid tiers. We test what a serious user would buy, and note when free-tier quality differs.
- Fixed iteration budget. Three attempts per task. Anything beyond that counts against reliability.
- Two independent scorers, comparing only after both have finished.
- Real work, not demos. Every tool that scores above 7 has been used on live client or editorial work for at least two weeks.
What we won't do
- Score a tool we haven't used for a fortnight. News coverage of a launch is clearly labelled as coverage, not a review.
- Publish a number for an unreleased product.
- Take payment for a score. Ever, in any form.
- Use vendor-supplied benchmarks as evidence. We cite them as claims and test them where we can.
- Compare across categories. A 9 for an image generator and a 9 for a CRM assistant are not the same 9.

Public benchmarks: useful, limited
We read LMSYS Arena rankings and academic evaluations, and they inform which tools we prioritise. But leaderboard position correlates weakly with day-to-day usefulness, because benchmarks measure model capability while you experience product design, latency, integrations, and defaults.
A tool that routes to a slightly weaker model but never loses your work will beat the leaderboard winner in practice.
Disclosure
- We pay for our own subscriptions unless a vendor grants access, which we state in the review.
- Some links may be affiliate links. They never affect scores, and a tool's score is finalised before any commercial discussion.
- We update reviews when tools change materially, and note the revision date.
- If we get something wrong, we correct it in place and say what changed.

Tell us when we're wrong
If your experience with a tool contradicts our review, we want to hear it — especially if you use it in a context we didn't test. Several of our biggest score revisions came from reader reports, not retesting.
Start with a review: ChatGPT vs Claude, best AI coding assistants, or best AI image generators.
Keep reading

AI Tool Pricing Explained: Seats, Credits, Tokens and the Bills That Surprise You
How AI pricing models really work — per-seat, credit packs, token metering, and usage tiers — plus how to estimate cost before you commit and avoid the classic overage traps.

AI Tools for Small Business: A $100/Month Stack That Replaces Three Contractors
A practical, priced AI toolkit for small businesses — marketing, customer support, bookkeeping admin, and sales — with what to adopt first and what to skip.

The Best Free AI Tools in 2026 (No Trial, No Credit Card)
A tested list of genuinely free AI tools for writing, images, video, coding, research, and audio — including what each free tier actually limits.