Lumen AI logoLumen AI
AI Tool Reviews & Comparisons

How We Test AI Tools: Our Scoring Framework, Explained

The methodology behind every review on this site — the seven scoring criteria, the standard test prompts, how we handle vendor relationships, and what we refuse to score.

Lumen AI Editorial7 min readEdit this article
Reviewer comparing AI tool outputs side by side across two monitors with a scoring sheet

Most AI tool reviews are rewritten press releases. We built this page so you can judge whether ours are worth trusting — and so we can be held to it.

Desk with a scoring rubric printed beside a laptop running AI tool tests
Same prompts, same conditions, every time.

The seven criteria

Every tool we score is rated 1–10 on the same axes, weighted by category.

1. Output quality (30%). The core job, judged against a standard set of tasks. For writing tools that's a blog section, an email, and an edit pass. For image tools it's a product shot, a portrait, and a text-in-image poster. For coding tools it's a bug fix, a refactor, and a feature in an unfamiliar codebase.

2. Reliability (20%). How often it produces something usable on the first attempt. We record hit rate across fixed attempts, because a tool that's brilliant one time in five is a worse tool than one that's good every time.

3. Control (15%). Can you steer it? References, parameters, rules files, iterative refinement. Ceiling matters less than steerability.

4. Workflow fit (15%). Integrations, export formats, collaboration, and whether it forces you to change how you work.

5. Speed (5%). Time to first usable output, measured under normal load, not at 4am.

6. Value (10%). Cost per unit of real work, including the iterations you'll waste. Headline price is rarely the real price.

7. Trust (5%). Data retention, training policies, licensing clarity, security posture, and how honestly the vendor describes limitations.

The standard test set

Each category has fixed tasks we run against every tool, unchanged between reviews so scores stay comparable over time.

  • Writing: an 800-word explainer from a brief, a cold email, an edit of a deliberately weak paragraph, and a voice-matching task using a provided sample.
  • Image: photorealistic product on white, an environmental portrait, a poster containing a specific headline, and a consistency test across three images.
  • Video: a camera move, a human walking, an object interaction, and a continuity test across two shots.
  • Coding: a failing test to fix, a cross-file refactor, a new feature in a codebase the tool hasn't seen, and a security review of intentionally flawed code.
  • Productivity: a two-week real-usage period on live work, not a demo.
Editorial team debating tool scores around a shared screen
Two reviewers score independently before comparing.

How we actually run tests

  • Same prompts, same day. Model quality varies; comparing a tool tested in January to one tested in July is meaningless.
  • Paid tiers. We test what a serious user would buy, and note when free-tier quality differs.
  • Fixed iteration budget. Three attempts per task. Anything beyond that counts against reliability.
  • Two independent scorers, comparing only after both have finished.
  • Real work, not demos. Every tool that scores above 7 has been used on live client or editorial work for at least two weeks.

What we won't do

  • Score a tool we haven't used for a fortnight. News coverage of a launch is clearly labelled as coverage, not a review.
  • Publish a number for an unreleased product.
  • Take payment for a score. Ever, in any form.
  • Use vendor-supplied benchmarks as evidence. We cite them as claims and test them where we can.
  • Compare across categories. A 9 for an image generator and a 9 for a CRM assistant are not the same 9.
Abstract neural network representing benchmark evaluation of AI models
Public benchmarks inform our view; they don't decide it.

Public benchmarks: useful, limited

We read LMSYS Arena rankings and academic evaluations, and they inform which tools we prioritise. But leaderboard position correlates weakly with day-to-day usefulness, because benchmarks measure model capability while you experience product design, latency, integrations, and defaults.

A tool that routes to a slightly weaker model but never loses your work will beat the leaderboard winner in practice.

Disclosure

  • We pay for our own subscriptions unless a vendor grants access, which we state in the review.
  • Some links may be affiliate links. They never affect scores, and a tool's score is finalised before any commercial discussion.
  • We update reviews when tools change materially, and note the revision date.
  • If we get something wrong, we correct it in place and say what changed.
Shield representing editorial independence and disclosure standards
Disclosure is part of the review, not a footnote.

Tell us when we're wrong

If your experience with a tool contradicts our review, we want to hear it — especially if you use it in a context we didn't test. Several of our biggest score revisions came from reader reports, not retesting.

Start with a review: ChatGPT vs Claude, best AI coding assistants, or best AI image generators.

#Methodology#Reviews#Benchmarks#Transparency#Editorial